Circuit board, computing system, computer cabinet and computing method
By designing switching units, acceleration card units and monitoring units on the circuit board, connecting buses at different speeds to the control device, dynamically scheduling resources, the problem of high resource consumption when GPU increases computing power in the prior art is solved, and more efficient utilization of server computing power is achieved.
Patent Information
- Application Number
- CN202510623229.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the prior art, the addition of graphics processors (GPUs) to improve server computing power consumption is high, and the timeliness and efficiency of scheduling resources is low, which limits the development of artificial intelligence technology.
By designing the switching unit, acceleration card unit, expansion storage unit and monitoring unit on the circuit board, connecting to the control device with buses of different speeds, monitoring and transmitting the call frequency information of the acceleration card unit, dynamic scheduling strategies are realized to optimize resource utilization.
It reduces the time when the acceleration card unit calls model parameters, improves the resource utilization efficiency of server computing power and the timeliness of scheduling resources, and saves resource consumption.
Smart Images

Figure CN120144323A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more particularly to a circuit board, a computing system, a computer cabinet, and a computing method. Background Art
[0002] Nowadays, we are in an era of rapid development of artificial intelligence technology. The rapid development of large models has put forward higher requirements for the computing power of servers than ever before. However, in related technologies, the computing power is mainly increased by adding Graphics Processing Units (GPUs) and other means. This method consumes more resources, and the timeliness and efficiency of scheduling server resources are poor, which restricts the development of artificial intelligence technology. Summary of the Invention
[0003] In view of the above problems, this application provides a circuit board, a computing system, a computer cabinet, and a computing method.
[0004] According to the first aspect of this application, a circuit board is provided, including: a switching unit; an acceleration card unit connected to the switching unit, where the first memory of the acceleration card unit is used to store the first model parameters of the first expert model; an extended storage unit connected to the switching unit, where the extended storage unit is used to store the second model parameters of the second expert model; the call frequencies of the first expert model and the second expert model are different; a monitoring unit is connected to the acceleration card unit and connected to the first control device via multiple buses; the data transmission speeds of the multiple buses are different, so that the monitoring unit can transmit the call frequency information of the acceleration card unit calling different expert models to the first control device at different transmission speeds.
[0005] According to the second aspect of this application, a computing system is provided, including: multiple circuit boards as described above; and a first control device for collecting the call frequency information of the expert models configured in the multiple circuit boards and determining the scheduling strategy of the multiple circuit boards based on the received call frequency information.
[0006] According to the third aspect of this application, a computer cabinet is provided, including: multiple computing systems as described above; and a second control device for receiving the call frequency information from the first control devices of the multiple computing systems and determining the scheduling strategy of the multiple computing systems based on the received call frequency information.
[0007] According to the fourth aspect of the present application, a calculation method is provided, which is applied to a circuit board. The method includes: an acceleration card unit of the circuit board obtains first model parameters from the memory of the acceleration card unit of the circuit board, and / or obtains second model parameters from an extended storage unit of the circuit board via a switching unit of the circuit board; the acceleration card unit of the circuit board calls at least one of a first expert model and a second expert model based on the first model parameters and / or the second model parameters to perform a model processing task, and the call frequency of the first expert model is higher than that of the second expert model; a monitoring unit of the circuit board transmits call frequency information of the acceleration card unit of the circuit board calling different expert models to a first control device at different transmission speeds.
[0008] According to an embodiment of the present application, during operation, the acceleration card unit can obtain first model parameters from the first memory in a timely manner, or obtain second model parameters from an additionally extended extended storage unit via the switching unit. In this way, by expanding the memory accessible to the acceleration card unit, the present application avoids the acceleration card unit consuming a long time to call model parameters from other storage spaces, reduces the duration of the acceleration card unit calling model parameters, can thus give full play to the computing power of the acceleration card unit, and further can improve the computing power of the server without adding a circuit board, thereby saving resource consumption.
[0009] On this basis, the monitoring unit is connected to the acceleration card unit and can collect call frequency information of the acceleration card unit calling expert models with different call frequencies. At the same time, the monitoring unit can also be connected to the first control device via multiple buses and transmit different call frequency information to the first control device at different speeds via the multiple buses. In this way, for information that needs to be processed quickly (such as call frequency information of expert models with high-frequency calls), it can be transmitted to the first control device at a relatively fast speed, enabling the first control device to perform resource scheduling for the acceleration card unit in a timely manner, improving the timeliness of scheduling resources, and improving the efficiency of scheduling resources. Description of the Drawings
[0010] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objects, features, and advantages of the present application will become clearer. In the drawings:
[0011] Figure 1 A schematic diagram of a circuit board according to the first embodiment of the present application is shown.
[0012] Figure 2 A schematic diagram of a circuit board according to the second embodiment of the present application is shown.
[0013] Figure 3 A schematic diagram of a circuit board according to the third embodiment of the present application is shown.
[0014] Figure 4Shows a schematic diagram of a computing system according to an embodiment of the present application.
[0015] Figure 5 Shows a schematic diagram of a first circuit board and a second circuit board according to a first embodiment of the present application.
[0016] Figure 6 Shows a schematic diagram of a first circuit board and a second circuit board according to a second embodiment of the present application.
[0017] Figure 7 Shows a schematic diagram of the connection of units in different circuit boards according to an embodiment of the present application.
[0018] Figure 8 Shows a schematic diagram of the connection of a first control device and a circuit board according to an embodiment of the present application.
[0019] Figure 9 Shows a schematic diagram of a computer cabinet according to a first embodiment of the present application.
[0020] Figure 10 Shows a schematic diagram of a computer cabinet according to a second embodiment of the present application.
[0021] Figure 11 Shows a schematic diagram of the connection of a computer cabinet according to an embodiment of the present application.
[0022] Figure 12 Shows a schematic diagram of the connection of the circuit board of a first computing system and the circuit board of a second computing system according to an embodiment of the present application.
[0023] Figure 13 Shows a schematic diagram of a computing method according to an embodiment of the present application. Detailed implementation manners
[0024] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present application. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present application.
[0025] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0027] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning that those of ordinary skill in the art usually understand such expressions (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0028] Figure 1 A schematic diagram of a circuit board according to a first embodiment of the present application is shown.
[0029] As Figure 1 shown, the circuit board of this embodiment includes a switching unit SW, an acceleration card unit AI, a first memory AIS of the acceleration card unit, an extended storage unit ES, and a monitoring unit MON.
[0030] The acceleration card unit AI can be a GPU. The acceleration card unit AI can perform model processing tasks based on a Mixture of Experts (MoE) model. The acceleration card unit AI can be deployed with an expert model and can perform model processing tasks using the deployed expert model. The model processing task can refer to a task of data processing using an expert model.
[0031] The acceleration card unit AI is electrically connected to the first memory AIS of the acceleration card unit AI. The first memory AIS can store the first model parameters of the first expert model. The acceleration card unit AI can obtain the first model parameters from the first memory AIS. For example, the first memory AIS can be a Compression Attached Memory Module (CAMM). The CAMM can support multiple memory types. Specifically, the first memory AIS can adopt the form of a Graphics Double Data Rate 6 (GDDR6) memory chip, which has the advantages of high speed, high capacity, adjustable bandwidth, and low cost.
[0032] The acceleration card unit AI can also be electrically connected to the switching unit SW. For example, the switching unit SW can be a Peripheral Component Interconnect Express (PCIE) switch. The switching unit SW can be electrically connected to the extended storage unit ES. The extended storage unit ES can store the second model parameters of the second expert model. In this way, the acceleration card unit AI can obtain the second model parameters from the extended storage unit ES via the switching unit SW. It should be understood that the acceleration card unit AI can simultaneously read the first model parameters from the first memory AIS and read the second model parameters from the extended storage unit ES, or can separately read the first model parameters and the second model parameters at different time periods, and this application does not make any limitations in this regard. Moreover, in the embodiments of this application, in addition to the first memory AIS, the latency of the acceleration card unit AI accessing other memories (such as the memory of the extended storage unit) can be increased to avoid the acceleration card unit AI obtaining data from other memories and affecting the timeliness of obtaining data from the first memory AIS.
[0033] The acceleration card unit AI can also be electrically connected to the monitoring unit MON. For example, the monitoring unit MON can be implemented based on an embedded controller (Baseboard Management Controller, BMC), a field programmable gate array (Field Programmable Gate Array, FPGA), etc. The monitoring unit MON can collect relevant information about the acceleration card unit AI invoking the expert model, such as the working state information and invocation frequency information when the acceleration card unit AI invokes the expert model. Among them, the invocation frequency information can refer to the invocation frequency of the expert model by the acceleration card unit AI per unit time. This invocation frequency can be determined based on the frequency of access to the model parameters stored on the circuit board. In addition, the monitoring unit MON can also be electrically connected to the first control device CTD1 independent of the circuit board via multiple buses L1, …, Ln, and send the monitored relevant information via the multiple buses L1, …, Ln, where n is an integer greater than 1. And, since the data transmission speeds of the multiple buses L1, …, Ln are different, different buses with different transmission speeds can be selected to transmit the above relevant information. For example, the multiple buses L1, …, Ln can include a Universal Serial Bus (USB) bus, a Low-Voltage Differential Signaling (LVDS) tunnel protocol and interface (LTPI) bus, etc. In an embodiment of the present application, the invocation frequencies of the first expert model and the second expert model are different. The bus with a relatively high transmission speed among the multiple buses L1, …, Ln can be used to transmit the expert model with a relatively high invocation frequency among the first expert model and the second expert model.
[0034] Based on this, during the operation of the acceleration card unit AI, it can obtain the first model parameters from the first memory AIS in a timely manner, or obtain the second model parameters from the additionally extended extended storage unit ES via the switching unit SW. In this way, by expanding the memory accessible to the acceleration card unit AI in the present application, it is possible to avoid the acceleration card unit AI consuming too much time to call model parameters from other storage spaces, thereby reducing the duration of the acceleration card unit AI calling model parameters, fully leveraging the computing power of the acceleration card unit AI, and further improving the computing power of the server without increasing the circuit board, saving resource consumption.
[0035] Further, the monitoring unit MON is connected to the acceleration card unit AI, and can collect the call frequency information of the acceleration card unit AI for calling expert models with different call frequencies. At the same time, the monitoring unit MON can also be connected to the first control device CTD1 via multiple buses L1, …, Ln, and transmit different call frequency information to the first control device CTD1 at different speeds via the multiple buses L1, …, Ln. In this way, for information that needs to be processed quickly (such as the call frequency information of expert models with high-frequency calls), it can be transmitted to the first control device CTD1 at a relatively fast speed, which enables the first control device CTD1 to perform resource scheduling for the acceleration card unit AI in a timely manner, improving the timeliness of scheduling resources and the efficiency of scheduling resources.
[0036] In the embodiment of the present application, the ports of the switching unit SW can be configured to support different protocols based on firmware information. For example, a part of the ports of the switching unit SW can be configured as PCIE ports, and another part of the ports can be configured as Compute Express Link (CXL) ports. In this way, the switching unit SW can communicate with the acceleration card unit AI through PCIE signals and communicate with the extended storage unit ES through CXL signals.
[0037] In the embodiment of the present application, the circuit board may further include a control unit. The control unit may be a Central Processing Unit (CPU for short). The control unit is electrically connected to the switching unit SW and sends instructions to the switching unit SW. For example, the instructions sent by the control unit to the switching unit SW can be implemented based on PCIE signals, specifically PCIEx16 signals.
[0038] The control unit may be pre-configured with an initial scheduling policy. For example, the initial scheduling policy can be sent to the control unit by the above-mentioned first control device. The initial scheduling policy can be implemented based on a predetermined program. Based on the scheduling policy, the control unit can send scheduling instructions and other data to the switching unit SW. Under the control of the scheduling instructions, the switching unit SW can electrically connect other units to the control unit, so that the data from the control unit is sent to other units, thereby scheduling the other units. For example, the switching unit SW can send the control instructions from the control unit to the acceleration card unit AI to schedule the acceleration card unit AI to execute model processing tasks.
[0039] The switching unit SW can also send the second model parameters from the extended storage unit ES to the acceleration card unit AI under the control of a scheduling instruction. The acceleration card unit AI can also call the second expert model to execute a model processing task based on the second model parameters, and obtain the model processing result of the second expert model. Among them, the CXL link needs to have sufficient bandwidth to match the GPU requirements to avoid bandwidth becoming a bottleneck for accessing the resources of the extended storage unit ES.
[0040] Based on this, under the control of a control instruction, the acceleration card unit AI can obtain the first model parameters from the first memory AIS, or can obtain the second model parameters from the extended storage unit ES via the switching unit SW. After that, the acceleration card unit AI can call the first expert model to execute a model processing task based on the obtained first model parameters to obtain the model processing result of the first expert model; it can also call the second expert model to execute a model processing task based on the obtained second model parameters to obtain the model processing result of the second expert model; it can also call the first expert model and the second expert model simultaneously to execute a model processing task, so as to obtain the model processing result of the first expert model and the model processing result of the second expert model respectively. The call frequency of the first expert model is higher than that of the second expert model. In this way, during the working process, the acceleration card unit AI can preferentially read the first model parameters from the first memory AIS that can be accessed at a relatively high speed to execute the model processing task.
[0041] Based on this, the time taken for the acceleration card unit AI to read the first model parameters from the first memory AIS is shorter than the time taken to read the second model parameters from the extended storage unit ES via the switching unit SW. In this way, storing the first model parameters that need to be obtained frequently in the first memory AIS that the acceleration card unit AI can read relatively quickly facilitates the acceleration card unit AI to read the first expert parameters frequently, and storing the second model parameters with a low acquisition frequency in the extended storage unit with a relatively slow reading speed can make full use of the computing power of the acceleration card unit AI to perform high-frequency processing on the first expert parameters, reducing the waste of the computing power of the acceleration card unit AI, thereby reducing the resources consumed to increase the server computing power, improving the efficiency of the acceleration card unit AI in obtaining model parameters, and reducing the power consumption generated when the acceleration card unit AI obtains model parameters. Moreover, storing the second model parameters with a lower call frequency in the extended storage unit ES in advance can avoid temporarily scheduling the storage location of the second model parameters when the second model parameters are needed, thereby consuming additional scheduling time, and thus can read the model parameters of the expert model with a relatively low frequency from the extended storage unit ES in a timely manner. In this way, the efficiency of the acceleration card unit AI in executing the model processing task is improved.
[0042] Further, the monitoring unit MON can collect the different call frequency information of the acceleration card unit AI for different expert models, and transmit the different call frequency information to the first control device via multiple buses L1, …, Ln respectively. For example, the monitoring unit MON can collect the call frequency information of the first expert model when the acceleration card unit AI calls the first expert model, and can also collect the call frequency information of the second expert model when the acceleration card unit AI calls the second expert model.
[0043] For example, the multiple buses L1, …, Ln can include a first bus and a second bus. The monitoring unit MON and the first control device CTD1 can be electrically connected via the first bus and the second bus, and can send the call frequency information of the first expert model to the first control device via the first bus at a first transmission speed, and send the call frequency information of the second expert model to the second control device via the second bus at a second transmission speed. Among them, the first transmission speed is higher than the second transmission speed. In this way, for the first expert model with a relatively high call frequency, a bus with a relatively faster transmission speed is used for transmission, which is convenient for the first control device to obtain the call situation of the expert model with high-frequency calls by the acceleration card unit AI in a timely manner, so that the resource scheduling for the acceleration card unit AI can be carried out in a timely manner, improving the timeliness of scheduling resources and the efficiency of scheduling resources.
[0044] The embodiment of the present application is not limited thereto. The monitoring unit MON can also select multiple buses L1, …, Ln respectively based on the storage locations of the model parameters of different expert models, so as to transmit different call frequency information to the first control device CTD1 via multiple buses L1, …, Ln respectively. For example, the monitoring unit MON can select the first bus to transmit the first call frequency information of the model parameters stored in the first memory AIS when the acceleration card unit AI reads the model parameters from the first memory, and can also select the second bus to transmit the second call frequency information of the model parameters stored in the extended storage unit ES when the acceleration card unit AI reads the model parameters from the extended storage unit ES.
[0045] The first control device CTD1 may pre-store the mapping relationship between the first call frequency information, the second call frequency information, and the scheduling policy. For example, the form of this mapping relationship may be a mapping table or the like. When receiving the first call frequency information and the second call frequency information, it can use the mapping relationship stored in itself to determine the scheduling policy associated with the first call frequency information and the second call frequency information, and send the queried scheduling policy to the control unit. In this way, based on the association relationship between the storage location of the model parameters described above and the call frequency of the expert model, the call frequency information of the expert model to which the model parameters belong can be transmitted using the bus corresponding to the storage location for the specific storage location of the model parameters, avoiding the problem that in some cases, due to the change in the call frequency of the expert model, the speed of the bus for transmitting the call frequency information does not correspond to the call frequency information, thereby affecting the timeliness of transmitting the call frequency information of the expert model with high-frequency calls. The embodiments of the present application are not limited to this. The above mapping relationship may also be stored in the terminal device. The first control device CTD1 may also send the first call frequency information and the second call frequency information to the terminal device. The terminal device performs the above operations to obtain the scheduling policy, and sends the scheduling policy to the first control device CTD1, and then the first control device CTD1 sends the scheduling policy to the control unit.
[0046] In some other embodiments of the present application, each model processing task may include multiple model processing processes. In different model processing processes, the call frequencies of the same expert model by the acceleration card unit AI may be different. For example, in the first model processing process of the acceleration card unit AI, since the call frequency of the first expert model is higher than that of the second expert model, the first memory AIS may store the first model parameters, and the extended storage unit ES may store the second model parameters, so that the acceleration card unit AI can obtain the model parameters of the expert model with high-frequency calls in a timely manner. In this case, the monitoring unit MON may send different call frequency information to the first control device CTD1 via multiple buses L1,..., Ln, so that the first control device CTD1 determines the second model processing process after the first model processing process based on the received call frequency information. For example, the first control device CTD1 may pre-store the first mapping relationship between the call frequency information and the model processing process, and the second mapping relationship between the model processing process and the scheduling policy. For example, the forms of the first mapping relationship and the second mapping relationship may be mapping tables. In this way, the first control device CTD1 may use the first mapping relationship based on the received call frequency information to determine that the currently ongoing model processing process is the first model processing process, and determine the second model processing process after the first model processing process based on the sorting of multiple model processing processes in the first mapping relationship. Then, according to the second model processing process, the second mapping relationship may be used to determine the scheduling policy corresponding to the second model processing process. After that, when the first model processing process ends, the determined scheduling policy may be sent to the control unit. Among them, it may be determined that the first model processing process ends when the acceleration card unit AI stops calling the expert model. In another embodiment of the present application, the call frequency information may also be sent to the terminal device via the first control device CTD1, so that the terminal device determines the scheduling policy based on the call frequency information.
[0047] At least one of the first model parameter and the second model parameter is stored at a different location in the second model processing flow than in the first model processing flow. For example, compared with the first model processing flow, in the second model processing flow, the call frequencies of the first expert model and the second expert model can change. Thus, the model parameters stored in the first memory AIS can change. For example, it can change from storing the first model parameter to storing the second model parameter. Correspondingly, the model parameters stored in the extended storage unit ES can change from the second model parameter to the first model parameter. In this case, the storage locations of the first model parameter and the second model parameter change. For example, after receiving the scheduling policy from the first control device CTD1, the control unit can update its originally deployed initial scheduling policy to a new scheduling policy. Under the control of this scheduling policy, the control unit can send the first model parameter to the extended storage unit ES via the switching unit SW to write the first model parameter into the extended storage unit to change the storage location of the first model parameter, and / or send the second model parameter to the acceleration card unit AI via the switching unit SW to write the second model parameter into the first memory to change the storage location of the second model parameter. For example, after writing the first model parameter into the extended storage unit ES, the first model parameter can overwrite the second model parameter originally stored in the extended storage unit ES. After writing the second model parameter into the first memory, the second model parameter can overwrite the first model parameter originally stored in the extended storage unit ES. By changing the storage location of the model parameters, the transmission speeds of the call frequency information of the first expert model and the second expert model can be adjusted simultaneously, so that the transmission speeds of the call frequency information of the first expert model and the second expert model match the second model processing flow.
[0048] Specifically, the storage locations of the first model parameter and the second model parameter have changed. The first model parameter is stored in the extended storage unit ES as the model parameter of the first expert model with low-frequency calls, while the second model parameter is stored in the first memory AIS as the model parameter of the second expert model with high-frequency calls. Based on this, in the second model processing flow different from the first model processing flow, the first expert model becomes an expert model with low-frequency calls, and the second expert model changes to an expert model with high-frequency calls. Furthermore, the call frequency information of the second expert model can be transmitted at the first transmission speed, and the call frequency information of the first expert model can be transmitted at the second transmission speed.
[0049] The embodiments of the present application are not limited thereto. In the embodiments of the present application, the first model processing flow and the second model processing flow may also belong to different model processing tasks. For example, the calculation results obtained in the first model processing flow may not be used as the input information for the second model processing flow. Specifically, the second model processing flow may be the model processing flow of the model processing task executed by the AI of the acceleration card unit of other circuit boards, and so on.
[0050] On this basis, for the problem that in some solutions, the model deployed on a single GPU (such as a deep learning model architecture based on the self-attention mechanism (Transformer)) can only serially execute multiple model processing flows of the model processing task in a fixed pipeline during the execution of the same model processing task, making it difficult to schedule GPU-related resources in a timely manner, the above solution of the present application can decouple the serially executed model processing task and use multiple circuit boards to parallelly execute the model processing flows that were originally serially executed in a pipeline form. In the parallelly executed model processing flows, the monitoring units of the circuit boards parallelly executing the model processing flows can send the call frequency information corresponding to the model parameters at the corresponding transmission speed according to the storage location of the model parameters. For example, the monitoring unit transmits the call frequency information of the expert model with a relatively high call frequency at a relatively high speed. In this way, even when the AI of the acceleration card unit of any circuit board receives the scheduling strategy of the model processing flow of other circuit boards and parallelly executes the same model processing task with other circuit boards, the first control device CTD1 can also timely obtain the call frequency information of the expert models with high-frequency calls deployed on each of the multiple circuit boards parallelly processing the same model processing task based on the positions of the model parameters stored on these multiple circuit boards, and the call frequency information is timely sent to the first control device CTD1, so that the first control device CTD1 can timely perform resource scheduling for the multiple circuit boards according to the received call frequency information of the expert models with high-frequency calls, improving the timeliness of resource scheduling and the efficiency of resource scheduling.
[0051] Figure 2 Shows a schematic diagram of a circuit board according to the second embodiment of the present application.
[0052] Such as Figure 2As shown, the extended control unit CTRL may include an extended control subunit ES_1 and multiple extended storage subunits. For example, the multiple extended storage subunits may include a first extended storage subunit ES_21 and a second extended storage subunit ES_22. For example, the extended control subunit ES_1 may be a Memory Extension Controller (MXC). The extended storage subunits may be in the form of Double Data Rate Synchronous Dynamic Random Access Memory (DDR) chips. This memory is the physical form of a memory resource pool extended based on the CXL protocol, and its capacity can be maximized as much as possible according to the bandwidth. In addition to the acceleration card unit AI being able to interact with this storage subunit for data, the control unit CTRL can also interact with this storage subunit for data via the switching unit SW through the Global Interconnect Manager (GIM) of CXL 3.1 to access the memory resource function. In this way, two forms of acceleration can be performed based on the CPU as the general computing node and the GPU as the intelligent computing node, solving problems such as insufficient memory capacity or other related memory problems in both general computing and intelligent computing. The extended control subunit ES_1 can control at least one of the multiple extended storage subunits to store the data transmitted via the switching unit SW, and can also send the data stored in at least one of the multiple extended storage subunits to the switching unit SW for transmission to other units. By setting the extended control subunit ES_1 and the extended storage subunits on the circuit board, the extended control subunit ES_1 and the extended storage subunits can be used to store model parameters with low call frequencies. Thus, when the computing power of this circuit board is insufficient, the model parameters stored in the extended storage subunits can be sent to other circuit boards for calculation. In this way, the strategy of using memory to convert computing power adopted in this application can solve the problem of insufficient computing power, greatly improving the quality of the activated experts and the number of unactivated experts on a single circuit board. Moreover, in this application, a new architecture using the extended control subunit ES_1 and the extended storage subunits and combining with the CXL cache coherence protocol is adopted to expand the memory of the acceleration card unit AI. Furthermore, the extended control subunit ES_1 and the extended storage subunits can use multi-level caches and increase the cache capacity to improve the latency problem. It should be understood that Figure 2 the number of the extended control subunit ES_1 and the extended storage subunits in the above is only an example, and this application does not limit the number of the extended control subunit ES_1 and the extended storage subunits.
[0053] In this application, in addition to the first expert model and the second expert model, the acceleration card unit AI can also call a third expert model to execute model processing tasks. Among them, the call frequency of the first expert model is higher than that of the third expert model, and the call frequency of the third expert model is higher than that of the second expert model. The second memory CTRLS of the control unit CTRL can be used to store the third model parameters of the third expert model. The second memory CTRLS can adopt the form of ordinary DDR particles, and this form is adapted to the platform where the CPU is located. Moreover, in the case of using DDR, adopting a land grid array (LGA) socket will result in lower signal transmission loss and higher signal rate.
[0054] The acceleration card unit AI can also, under the control of a control instruction, obtain the third model parameters from the second memory CTRLS, and based on the third model parameters, call the third expert model to execute model processing tasks to obtain the model processing result of the third expert model. For example, the first expert model, the second expert model, and the third expert model can be divided based on a predetermined call frequency range. In the embodiments of this application, the first expert model, the second expert model, and the third expert model can be used to execute the same type of model processing tasks. For example, the expert models set on the same circuit board can be used to process the same type of minimum text units (tokens). For example, the first expert model, the second expert model, and the third expert model are all used to process punctuation marks, or all used to process words, and so on.
[0055] The second memory CTRLS of the control unit CTRL can store the first model parameters, the second model parameters, and the third model parameters. In the embodiments of this application, the control unit CTRL can, before the acceleration card unit AI executes model processing tasks, pre-schedule the resources of the first memory AIS of the acceleration card unit AI and the resources of the extended storage unit ES to store the model parameters. For example, the control unit CTRL can send a scheduling instruction, the first model parameters, and the second model parameters to the switching unit SW according to an initial scheduling strategy. The switching unit SW can, under the control of the scheduling instruction, send the first model parameters to the acceleration card unit AI to write the first model parameters into the first memory AIS of the acceleration card unit AI, or send the second model parameters to the extended storage unit ES to write the second model parameters into the extended storage unit ES.
[0056] Figure 3 FIG. shows a schematic diagram of a circuit board according to the third embodiment of this application.
[0057] As Figure 3 shown, in addition to the above control unit CTRL, switching unit SW, acceleration card unit AI, extended storage unit ES, and monitoring unit MON, the circuit board of this embodiment may further include a communication unit INT.
[0058] The communication unit INT can be a Network Interface Controller (NIC). The communication unit INT can be electrically connected to the switching unit SW, receive data from the control unit CTRL via the switching unit SW, and send the received data to other circuit boards based on the communication connection with other circuit boards.
[0059] The monitoring unit MON can be electrically connected to the first interface CON1, and the first interface CON1 can be a high-speed connector. The monitoring unit MON can output bus signals of multiple speeds via the first interface CON1. For example, the first interface CON1 can include a USB interface, LTPI, etc. The USB interface can be used to transmit USB signals. The LTPI interface can be used to transmit signals such as Inter-Integrated Circuit (I2C) signals, Universal Asynchronous Receiver / Transmitter (UART) signals, and General Purpose Input / Output (GPIO) signals.
[0060] The switching unit SW can be electrically connected to the second interface CON2, and the second interface CON2 can be a high-speed connector. The switching unit SW can be connected to the switching unit SW of other circuit boards via the second interface CON2.
[0061] The acceleration card unit AI can be electrically connected to the third interface CON3, and the third interface CON3 can be a high-speed connector. The acceleration card unit AI can be connected to the acceleration card unit AI of other circuit boards via the third interface CON3.
[0062] The communication unit INT can be electrically connected to the fourth interface CON4, and the fourth interface CON4 can be a high-speed connector. The communication unit INT can be connected to the communication unit INT of other circuit boards via the fourth interface CON4.
[0063] Specifically, after the circuit board is powered on and starts to work, the circuit board can not only be used as a computing core, but also as a GPU and a memory pool. When the acceleration card unit AI is working, in addition to obtaining the model parameters of the first memory AIS, the model parameters of the second memory CTRLS and the extended storage unit ES can also be obtained. For example, the acceleration card unit AI can call the Compute Unified Device Architecture (CUDA) to allocate an area on the first memory AIS, and can also call CUDA to allocate the same size of space on the second memory CTRLS, then initialize the data, and copy the initialized data from the first memory AIS to the second memory CTRLS. Moreover, the work of placing the memory can be automatically taken over by using the Uniform Memory Access (UMA) technology. In this way, only by calling the fixed Application Programming Interface (API) in CUDA, UMA can perform automatic processing.
[0064] However, in the case of multiple nodes (such as the above circuit boards) in the entire chassis and cabinet, since the processors of each node access the memory through a shared bus, when the number of processors increases, the bus will become a performance bottleneck, resulting in a slow memory access speed, affecting the overall system performance, and the memory that the control unit CTRL can carry is limited. Therefore, it is particularly important and convenient to call the memory of the extended storage unit ES. Through the CXL bus technology, the CXL resources of the control unit CTRL can be converted into a memory expansion resource pool, so that the acceleration card unit AI can easily call the extended memory resources. The prerequisite for the call is:
[0065] CXL interface integration: The model processing subunit AI_H supports the CXL.memory protocol, or can access the extended storage subunit through the extended control subunit ES_1. Among them, the CXL.memory protocol allows the control unit CTRL and the acceleration card unit AI to directly access or pool the memory of the extended storage subunit.
[0066] ② Physical connection: The CXL extended memory is connected to the acceleration card unit AI and the control unit CTRL through high-speed interconnection (such as a CXL link) to ensure low latency and high bandwidth.
[0067] ③ Virtualization support: Map the CXL extended memory to a unified virtual address space, and the acceleration card unit AI accesses the memory of the extended storage subunit through its built-in Memory Management Unit (MMU).
[0068] ④ Page table management: The AI driver of the accelerator card unit collaborates with the control unit CTRL to manage the page table of the memory of the extended storage subunit, and processes the address translation requests of the accelerator card unit AI through the Input-Output Memory Management Unit (IOMMU) or MMU.
[0069] ⑤ CXL.cache protocol: The accelerator card unit AI supports cache coherence, and uses CXL.cache to maintain the cache state between the accelerator card unit AI and the control unit CTRL to avoid data inconsistency.
[0070] ⑥ Hardware coherence domain: A coherence domain is built within the circuit board to enable the accelerator card unit AI and the control unit CTRL to transparently access the extended storage subunit without explicit refreshing.
[0071] ⑦ At the driver and operating system layer: The control unit CTRL identifies the access to the extended storage subunit as an allocable resource, provides an API for the CXL.mem module driver of the Linux system to manage, and can modify the preset driver of the accelerator card unit AI so that the accelerator card unit AI can support the memory allocation of the extended storage subunit, and enables the accelerator card unit AI to process DMA (Direct Memory Access) requests and address mapping.
[0072] After the memory of the extended storage subunit is called, the system still needs further optimization. Thus, based on the call frequencies of the detected first expert model, second expert model, and third expert model, the positions of the model parameters stored in the first memory AIS, extended storage unit ES, and second memory CTRLS can be adjusted among the first memory AIS, extended storage unit ES, and second memory CTRLS, placing the frequently accessed data in the first memory AIS and storing the data with low access frequency in the extended storage subunit. The data records such as call frequencies are obtained by the monitoring unit MON, and the monitoring unit MON can send the obtained call frequency information to the terminal device to update the scheduling policy of the control unit CTRL.
[0073] On a single circuit board, a small-scale hybrid expert model can be deployed based on the capacities of the first memory AIS, extended storage subunit, and second memory control unit CTRLS. The number of expert models that can be deployed on a single circuit board can be calculated by the following formula:
[0074] Maximum number of expert models = [(Capacity of the first memory AIS - Overhead of the 2GB circuit board) + (Total memory capacity that can be called by the acceleration card unit AI × Predetermined model parameter call factor)] / (Number of parameters of the expert model × Number of bytes processed by the acceleration card unit AI per unit time + Occupation value of activation value) × (50% - 70%) (1)
[0075] As can be seen from the above formula, 50 expert models can be deployed on a single circuit board. Compared with the 20 expert models in the related art, the number of model deployments is increased, which is convenient for improving the computing power of the circuit board of the present application.
[0076] Furthermore, multiple groups of expert models can be deployed on a single circuit board, and each group of expert models includes expert models of the same type. Specifically, the architecture and working process of MOE can be expressed by the following two formulas:
[0077] MOE (x)=∑K(Gi(x)Ei(x)) (2)
[0078] G(x)=TopK(Softmax(Wg(x)+€)): (3)
[0079] Among them, Ei(x) is the model processing result of the input Token passing through the i-th expert model. Gi(x) is the output weight of the i-th expert model (between 0 and 1). G(x) represents the total model processing result. TopK is used to select the top K largest values from the input data. Softmax represents the activation function. Wg represents the predetermined weight. x represents the input data. € represents the predetermined variable.
[0080] From the architecture and working process of MOE, it can be seen that in addition to having few activation parameters and fast inference speed, it also has strong scalability. By increasing or decreasing the number of expert models, the parameter scale of the expert models can be adjusted at will, so as to adapt to more different tasks.
[0081] Starting from the working principle of the MOE architecture, the present application further deploys 1 to 2 expert models belonging to the same type of model processing tasks for a single circuit board CD, which can further reduce the activation parameters of the single circuit board CD, so that the memory resources can be better allocated to the expert models deployed on the single circuit board CD, reducing the communication overhead caused by data interaction between multiple circuit boards CD during the operation of the computing system CS, and avoiding memory-related problems such as video memory fragmentation.
[0082] Compared with the excessive stacking of the number of GPUs and the increase in the power consumption of GPUs in the related art, the present application further refines the hardware architecture based on the MOE model. It not only increases the number of expert models in the MOE model to improve the computing ability, but also makes each computing power hardware node more refined. Instead of deploying the expert models in the MOE model on all GPUs of the entire machine cluster, the expert models will be deployed on the confirmed single computing power chip according to the characteristics of business processing. That is, a large feedforward neural network (FFN) in the Transformer model is replaced with multiple small expert models (also in the FFN structure). Furthermore, a sparse activation mechanism is adopted, and only 1 to 2 expert models of the same type are activated each time during inference, greatly reducing the number of computing parameters required during inference, thereby improving the inference efficiency of the entire MOE model. And when the stored parameter quantity is increased, the interpretability of model processing and the ability to obtain local information are improved. It should be noted that each of the expert models is also in the FFN structure. The expert models of the present application can be compressed, for example, the compression directions can be sparsity, distillation, quantization, etc.
[0083] Based on this, the number of experts in the present application can be adjusted according to system requirements. Since only the expert models of the same type are activated each time in the present application, it is necessary to improve the parallel ability of the expert models to handle dedicated transactions, so higher requirements are put forward for the memory capacity and bandwidth of the current acceleration card unit AI.
[0084] Figure 4 The schematic diagram of a computing system according to an embodiment of the present application is shown.
[0085] As Figure 4 shown, the computing system of this embodiment includes a plurality of circuit boards CD and a first control device CTD1. The circuit board CD here can be any of the above-mentioned circuit boards CD. A single computing system can be deployed in a single chassis.
[0086] The first control device CTD1 can collect the call frequency information of the expert models configured in the plurality of circuit boards CD and send the call frequency information to the terminal device. For example, the first control device CTD1 can receive the call frequency information from the monitoring unit MON of the circuit board CD and can determine the scheduling strategy of the circuit board based on the call frequency information. For example, the mapping relationship between the call frequency information and the scheduling strategy can be preset, so that the scheduling strategy corresponding to the call frequency information can be determined using this mapping relationship.
[0087] Optionally, the multiple circuit boards CD include a first circuit board CD1 and a second circuit board CD2. The model parameters stored in the first circuit board CD1 and the model parameters stored in the second circuit board CD2 are respectively used to perform different types of model processing tasks. For example, the model parameters stored in the first circuit board CD1 and the model parameters stored in the second circuit board CD2 can be used to process different types of data. Specifically, the types of the minimum units (tokens) processed by the model parameters stored in the first circuit board CD1 and the model parameters stored in the second circuit board CD2 are different. For example, the model parameters of the first circuit board CD1 can be used to process punctuation marks, while the model parameters of the second circuit board CD2 can be used to process words, etc.
[0088] In an embodiment of the present application, the model parameters stored in the first circuit board CD1 may be target model parameters for performing a target model processing task. The first circuit board CD1 is electrically connected to the second circuit board CD2 and can send the target model parameters to the second circuit board CD2. After receiving the target model parameters, the second circuit board CD2 can call the target expert model to perform the target type of model processing task based on the target model parameters, and obtain the target model processing result. Based on this, by scheduling the computing power resources of the second circuit board CD2 to assist in processing the target model parameters of the first circuit board CD1, problems such as the first circuit board CD1 malfunctioning due to excessive frequency of calling the target expert model can be avoided, thereby improving the reliability of the circuit board CD and increasing the redundancy of the computing system.
[0089] Figure 5 A schematic diagram of the first circuit board and the second circuit board according to the first embodiment of the present application is shown.
[0090] In Figure 5In the illustrated embodiment, the first circuit board CD1 includes a first control unit CTRL1 and an acceleration card unit AI1 of the first circuit board that are electrically connected via a switching unit SW1 of the first circuit board. The first control unit CTRL1 may send a first scheduling instruction, a first control instruction, and target model parameters to the switching unit SW1 of the first circuit board according to a first target scheduling policy. The switching unit SW1 of the first circuit board may, under the control of the first scheduling instruction, send the first control instruction and the target model parameters to the acceleration card unit AI1 of the first circuit board based on the first target scheduling policy. It should be understood that in this application, both the first circuit board CD1 and the second circuit board CD2 are the circuit boards described above. Therefore, both of them also have structures corresponding to the circuit boards described above. For example, the acceleration card unit AI1 of the first circuit board and the acceleration card unit AI2 of the second circuit board are both the acceleration card units described above. The same applies to other units in the first circuit board CD1 and the second circuit board CD2, and will not be elaborated here one by one. Similarly, the circuit boards described in each part below in this application are also similar to the circuit boards described above, and the units of the circuit boards described in each part below are also similar to those described above, and will not be elaborated in this application one by one.
[0091] The second circuit board CD2 includes an acceleration card unit AI2 of the second circuit board. The acceleration card unit AI1 of the first circuit board is electrically connected to the acceleration card unit AI2 of the second circuit board and may, under the control of the first control instruction, send the target model parameters to the acceleration card unit AI2 of the second circuit board to write the target model parameters into the memory of the acceleration card unit AI2 of the second circuit board. The memory here refers to the first memory of the above acceleration card unit.
[0092] Figure 6 A schematic diagram of the first circuit board CD1 and the second circuit board CD2 according to the second embodiment of the present application is shown.
[0093] In Figure 6 the illustrated embodiment, the first circuit board CD1 includes a first control unit CTRL1 and an acceleration card unit AI1 of the first circuit board that are electrically connected via a switching unit SW1 of the first circuit board. The first control unit CTRL1 may send a second scheduling instruction, a second control instruction, and target model parameters to the switching unit SW1 of the first circuit board according to a second target scheduling policy.
[0094] The second circuit board CD2 may include a switching unit SW2 of the second circuit board and an extended storage unit ES2 of the second circuit board. The switching unit SW1 of the first circuit board is electrically connected to the switching unit SW2 of the second circuit board, and may, under the control of the second scheduling instruction and based on the second target scheduling policy, send target model parameters to the switching unit SW of the second circuit board. The switching unit SW of the second circuit board is electrically connected to the extended storage unit ES of the second circuit board. Thus, via the switching unit SW2 of the second circuit board, the switching unit SW1 of the first circuit board may send the target model parameters to the extended storage unit ES of the second circuit board to write the target model parameters into the extended storage unit ES of the second circuit board.
[0095] The embodiments of the present application are not limited thereto. The first circuit board CD1 may further include a first communication unit, and the second circuit board CD2 may further include a second communication unit. The switching unit SW1 of the first circuit board may also be electrically connected to the first communication unit. Thus, the first communication unit may receive the target model parameters via the switching unit SW1 of the first circuit board. The first communication unit is communicatively connected to the second communication unit and may send the target model parameters to the second communication unit. The second communication unit is electrically connected to the switching unit SW2 of the second circuit board and may send the target model parameters to the extended storage unit ES2 of the second circuit board via the switching unit SW2 of the second circuit board.
[0096] Figure 7 A schematic diagram of the connection of units in different circuit boards according to an embodiment of the present application is shown.
[0097] As Figure 7 shown, the acceleration card units of different circuit boards in the computing system may be connected to each other. For example, the first acceleration card unit may be connected to the second acceleration card unit, and the first acceleration card unit and the second acceleration card unit may be connected to the Nth acceleration card unit, and so on. Thus, each acceleration card unit may be connected to other acceleration card units, facilitating data transmission between the acceleration card units. Herein, the first acceleration card unit may refer to the acceleration card unit of the first circuit board, and the same applies to other acceleration card units, which will not be elaborated herein.
[0098] Similarly, the switching units of different circuit boards in the computing system may be connected to each other. For example, the first switching unit may be connected to the second switching unit, and the first switching unit and the second switching unit may be connected to the Nth switching unit, and so on. Thus, each switching unit may be connected to other switching units, facilitating data transmission between the switching units. Herein, the first switching unit may refer to the switching unit of the first circuit board, and the same applies to other switching units, which will not be elaborated herein.
[0099] Similarly, the communication units of different circuit boards in a computing system can be connected to each other. For example, the first communication unit can be connected to the second communication unit, and the first communication unit and the second communication unit can be connected to the Nth communication unit, and so on. In this way, each communication unit can be connected to other communication units, facilitating data transmission between communication units. Among them, the first communication unit can refer to the communication unit of the first circuit board, and the same applies to other communication units, which will not be elaborated here. It should be noted that for the sake of illustration, in Figure 7 the first communication unit and the Nth communication unit are connected by a black line, but it should be understood that this is only for illustration, and in fact, the communication units can be communicatively connected.
[0100] Furthermore, in this application, the switching units of multiple circuit boards are connected to form a FABRIC (structure) link. Within a single computing system, this network architecture has good nanosecond-level low-latency performance. On this basis, the acceleration card units of multiple circuit boards are connected to achieve peer-to-peer communication (Peer-to-Peer Communication, P2P) between acceleration card units, thus forming a dual-ring network with the above FABRIC network architecture. In this way, the model parameters with high call frequencies can be transmitted between the connected acceleration card units, and the model parameters with high call frequencies can also be transmitted between the connected switching units. The network formed by the connected acceleration card units and the network formed by the connected switching units are redundant with each other, and the data volume related to the reduce function of the link can be reduced to half of the original, which can not only meet the general low-latency transmission requirements but also meet the special requirements of the communication library. In addition, the above FABRIC link can also be used to transmit the data stored in the extended storage unit and the data in the memory of the control unit. Since the network formed by the acceleration card units does not require a third-party unit for data transfer, the latency caused by the third-party unit for data transfer can be avoided. Therefore, it is used as a dedicated link for transmitting data with high-frequency calls and can also be regarded as an exclusive link for the acceleration card units to call each other's memory.
[0101] The data throughput of the network formed by the connected communication units is relatively large compared to the above dual-ring network, but the data needs to be demodulated during the data transmission process. In this way, it can be used to transmit data with call frequencies other than those with high-frequency calls.
[0102] On this basis, through the above connection methods between the switching units, acceleration card units, and communication units of different circuit boards, three balanced topologies can be formed, thereby reducing the data volume related to the reduce function of the link and facilitating the optimization of system resource allocation.
[0103] Optionally, the first control device may send the target call frequency information of the target expert model from the first circuit board to the terminal device. For example, the target call frequency information may be sent to the terminal device through an RJ45 (Registered Jack) network interface. After receiving the target call frequency information, the terminal device may display the target call frequency information through a visualization interface, so that relevant personnel can send a target scheduling policy to the first control device through the terminal device according to the displayed target call frequency information. In response to receiving the target scheduling policy corresponding to the target call frequency information from the terminal device, the first control device may send the target scheduling policy to the first circuit board to control the first circuit board to send the target model parameters to the second circuit board, thereby realizing the scheduling of the resources of the second circuit board. In the embodiments of the present application, the target scheduling policy may also be determined based on other information such as GPU load information, first memory occupancy information, second memory control unit occupancy information, power consumption information of the circuit board, etc. The method is similar to the above and will not be elaborated here.
[0104] For some exemplary server hardware architectures, due to their high complexity and great management difficulty, they will result in high operation and maintenance costs, high management difficulty, and high costs. Moreover, the number of nodes in such a server hardware architecture is relatively large, and the computing power of some of the nodes may not be fully utilized, resulting in a waste of resources and affecting the computing efficiency of the architecture. In addition, such a server hardware architecture is subject to hardware limitations. Since the hardware resources of a single server have an upper limit, it is difficult to further expand the hardware resources of the server after the hardware resources reach the upper limit, resulting in limited computing power of the server. Also, since a large amount of resources are deployed on the same server, the entire system may crash in the event of a failure of a certain server node. Furthermore, the above server hardware architecture has low efficiency in processing a large number of small memory read and write tasks. This application uses hardware-aware routing technology and dynamic supervision technology to determine status information such as the call frequency of circuit boards. Furthermore, in the case where the computing power of the acceleration card unit of any circuit board is insufficient, the model parameters stored on this circuit board can be sent to other circuit boards for processing. In this way, when any circuit board calls the first target expert model for executing a target model processing task of a target type, this circuit board can send the model parameters of the second target expert model for this target model processing task to other circuit boards. In this way, the above circuit board can execute a part of the data processing process of the target model processing task, while other circuit boards can call the second target expert model to execute the other part of the data processing process of this target model processing task, realizing the pipeline parallel processing of context tasks and solving the problem of low parallel computing ability of large models in related technologies. Moreover, the distribution of expert models between the above circuit board and other circuit boards can be dynamically adjusted through the above update and adjustment strategy, thereby enabling dynamic batching, solving the problem of insufficient computing power caused by insufficient memory of large models in related technologies, and balancing the differences in the computing times of different expert models, improving the computing efficiency.
[0105] In this way, by implementing scheduling across devices, systems, and even architectures, the limitations of the hardware architecture on the computing power scheduling of acceleration card units can be broken through, the resources of the acceleration card units can be fully utilized, resource waste can be reduced, costs can be lowered, the problems of insufficient large model parameters, few expert models, and low capabilities in related technologies can be solved, and the problem of resource waste caused by the fixed deployment method of current large models can be solved.
[0106] Optionally, multiple circuit boards are stacked vertically in sequence according to the number of calls of the configured expert models. The first circuit board is adjacent to the second circuit board, and the number of calls of the expert model of the first circuit board is higher than that of the expert model of the second circuit board. In this way, when the number of calls of the expert model of the first circuit board is high, by scheduling the resources of the second circuit board adjacent to the first circuit board to execute the model processing task based on the model parameters of the first circuit board, problems such as the first circuit board malfunctioning due to excessive frequency of calling the target expert model can be avoided, thereby improving the reliability of the circuit board. It should be understood that in the embodiments of the present application, the resources of other circuit boards called by the first circuit board are not limited to the circuit boards adjacent to the first circuit board, and may also include circuit boards with a lower number of calls of other expert models. Based on this, for a single computing system, in the circuit boards of the present application, a single acceleration card unit is connected to a single control unit, and special deployment of the expert model is performed on a single circuit board to implement a single circuit board as a small-scale basic model. On this basis, dynamic routing of model parameters between multiple circuit boards within a single computing system can form a medium-scale model. By deploying a medium-scale model in this way, the number of expert models, the resources that experts can schedule, and the data flow and gating network between expert models within the computing system can be optimized. Moreover, in the embodiments of the present application, a hierarchical communication strategy can be adopted to deploy groups of experts with frequent interactions on the same circuit board or within the same computing system, and use NVLink communication technology or PCIE FABRIC interconnection technology to improve the communication speed.
[0107] Optionally, the first control device is electrically connected to any one of the multiple circuit boards via multiple buses. The data transmission speeds of the multiple buses are different.
[0108] The first control device can collect the call frequency information of different types of expert models from any one of the circuit boards via the multiple buses. The type of the expert model represents the historical call frequency of the expert model. For example, the type of the expert model can be divided according to the historical call frequency and call frequency range of the expert. For example, the first expert model, the second expert model, and the third expert model described above are different types of expert models. In the embodiments of the present application, the call frequency information of the expert model with a high call frequency (such as the call frequency information of the first expert model) and the call frequency information of the expert model with a low call frequency (such as the call frequency information of the second expert model) can be collected from any one of the circuit boards via the multiple buses. The collected call frequencies of the expert models can be sent to the terminal device so as to schedule the circuit board to call the expert model to execute the model processing task in a timely manner according to the number of calls of different types of expert models.
[0109] Figure 8Shows a schematic connection diagram of a first control device and a circuit board according to an embodiment of the present application.
[0110] As Figure 8 shown, a plurality of buses include a first bus L1 and a second bus L2. The data transmission speed of the first bus L1 is higher than that of the second bus L2. The monitoring unit MON of the circuit board CD can collect the call frequency information of the expert model of the circuit board CD. The first control device CTD1 may include an acquisition unit ACQ. The acquisition unit ACQ can be electrically connected to the monitoring unit MON via the first bus L1 and the second bus L2, and collect the first call frequency information of the first type of expert model from the monitoring unit MON via the first bus L1, and collect the second call frequency information of the second type of expert model from the monitoring unit MON via the second bus L2. The first control device CTD1 may further include a device control unit ECT. The device control unit ECT can send the first call frequency information and the second call frequency information to the terminal device TE.
[0111] The acquisition unit ACQ may include a first acquisition subunit and a second acquisition subunit. The first bus L1 may be the above-mentioned USB bus. The monitoring unit MON can be electrically connected to the first acquisition subunit via the first bus L1. The first acquisition subunit may be a USB HUB (USB hub) unit. Through the USB bus, data can be quickly transmitted to the USB HUB unit for real-time processing. Then, the USB HUB unit can transmit the collected information to the device control unit ECT. The second bus L2 can be electrically connected to the second acquisition subunit. The second bus L2 may be the above-mentioned LTPI bus. The FPGA can collect the model parameters of the medium call frequency from the monitoring unit MON via the second bus L2. For the call frequency information of the model parameters with a low call frequency, the FPGA can collect it through I2C signals or CPIO signals.
[0112] Optionally, the first control device CTD1 further includes a system switching unit SSW. For example, the system switching unit may be implemented based on a PCIE Switch unit. The system switching unit SSW is electrically connected to the switching units SW of the respective circuit boards CD. The device control unit ECT is electrically connected to the system switching unit SSW for sending a target scheduling policy to the system switching unit SSW, so that the switching unit of the circuit board CD receives the target scheduling policy via the system switching unit, and then writes the target scheduling policy into the control unit via the switching unit. For example, the switching unit of the first circuit board receives the target scheduling policy via the system switching unit, and then writes the target scheduling policy into the first control unit via the switching unit of the first circuit board.
[0113] Figure 9Shows a schematic diagram of a computer cabinet according to a first embodiment of the present application.
[0114] As Figure 9 shown, the computer cabinet of this embodiment includes a plurality of computing systems CS and a second control device CTD2. The computing system CS here can be any of the above computing systems.
[0115] The second control device CTD2 can receive call frequency information from the first control device of the plurality of computing systems CS, and can determine a scheduling policy based on the received call frequency information. The specific method is similar to that described above and will not be elaborated here. For example, the computer cabinet can correspond to a server cabinet.
[0116] Figure 10 Shows a schematic diagram of a computer cabinet according to a second embodiment of the present application, Figure 11 Shows a schematic connection diagram of a computer cabinet according to an embodiment of the present application.
[0117] As Figure 10 and Figure 11 shown, the computer cabinet further includes a collection device CLD. The collection device CLD can be a Remote Power Distribution Unit (RPDU). The collection device CLD can be electrically connected to the circuit boards of the plurality of computing systems CS and collect the power consumption information of the circuit boards. In one embodiment of the present application, the collection device CLD can directly send the power consumption information to the terminal device. In another embodiment, the collection device CLD can be electrically connected to the second control device CTD2 and send the power consumption information to the terminal device via the second control device CTD2. After receiving the power consumption information, the terminal device can determine a scheduling policy based on the power consumption information and send the scheduling policy to the second control device CTD2. The second control device CTD2 can send the scheduling policy to the circuit boards in the computing system CS, thereby scheduling the resources of the circuit boards.
[0118] Optionally, the second control device CTD2 may be an industrial-grade CPU based on the Edge series. This CPU has characteristics such as wide temperature range and fast response. The second control device CTD2 may include a system control unit SCT, a first architecture switching unit SYC1, and a second architecture switching unit SYC2. The first architecture switching unit SYC1 may be implemented based on a PCIE Switch unit. The second architecture switching unit SYC2 may be implemented based on an Ethernet Switch (ETH SW) unit. The switching unit SW of the circuit boards of multiple computing systems CS may be connected to the first architecture switching unit SYC1 and perform data transmission via the first architecture switching unit SYC1. For example, the switching unit SW may also be connected to the first architecture switching unit SYC1 via the above-mentioned system switching unit. The communication unit of the circuit boards of multiple computing systems CS may be connected to the second architecture switching unit SYC2 and perform data transmission via the second architecture switching unit SYC2. In an embodiment of the present application, the system control unit SCT may receive the above-mentioned scheduling policy corresponding to the power consumption information from the terminal device and send it to the switching unit SW of the circuit board via the first architecture switching unit SYC1, so as to write the scheduling policy into the control unit of the circuit board via the switching unit SW. In another embodiment of the present application, the system control unit SCT may receive the above-mentioned scheduling policy corresponding to the power consumption information from the terminal device and send it to the communication unit INT of the circuit board via the second architecture switching unit SYC2, so as to write the scheduling policy into the control unit of the circuit board via the communication unit INT and the switching unit SW connected to the communication unit INT.
[0119] Optionally, the multiple computing systems include a first computing system and a second computing system. The first computing system includes the circuit board of the first computing system. The second computing system includes the circuit board of the second computing system. The acquisition device CLD is electrically connected to the circuit board of the first computing system and the circuit board of the second computing system, and may acquire the power consumption information of the circuit board of the first computing system and the circuit board of the second computing system. Then, the acquisition device CLD may send the power consumption information to the terminal device and send the scheduling policy corresponding to the power consumption information received from the terminal device to the circuit board of the first computing system. The circuit board of the first computing system may send the model parameters of the circuit board of the first computing system to the circuit board of the second computing system according to the scheduling policy corresponding to the power consumption information. The circuit board of the second computing system may call the expert model of the circuit board of the first computing system based on the model parameters of the circuit board of the first computing system to execute the model processing task and obtain the model processing result of the circuit board of the first computing system.
[0120] Multiple architectures can also be connected via the first architecture exchange unit SYC1 and the second architecture exchange unit SYC2 described above. In this way, physical conditions are provided for clustering multiple chassis within the same architecture and between multiple architectures through a network. On this basis, the RDMA technology can be used to enable mutual access between multiple first control devices, ensuring the efficiency of all-to-all bandwidth communication and reducing the latency of gradient synchronization among experts. Moreover, conditions are provided for networking multiple chassis within a cabinet through PCIE, which can be more applied to low-latency scenarios. The second control device CTD2 can also include a rack management controller (RMC). The RMC can play a role in supervising and managing the entire architecture. For example, the RMC can determine a computing system CS with high power consumption within the architecture. The computing system CS can be a system with a high memory usage frequency of the acceleration card unit and a high call frequency of the expert model. The RMC can send power consumption information to the terminal device, obtain the feedback scheduling policy, and add the scheduling policy to the image file loaded by the RMC. The RMC can send the power consumption information based on the RJ45 network interface. Priority scheduling can be performed on the resources of the computing system with high power consumption, and the remaining computing systems can be used to provide redundancy for the computing system with high power consumption to ensure the stable operation of the computing system with high power consumption. Within a cabinet, multiple server chassis can be placed, and multiple ports can be formed in the form of two types of switches, namely the second architecture exchange unit SYC1 and the second architecture exchange unit SYC2, to exchange resources of each chassis and each node. For the MOE model, local communication can be used for tuning at this time. The expert groups with frequent interactions are deployed within the same rack, and technologies such as NVLINK are used for tuning. An asynchronous update mechanism is adopted for the global parameters (such as the gating network) on the non-critical path to reduce communication congestion.
[0121] Figure 12 The connection schematic diagram of the circuit board of the first computing system and the circuit board of the second computing system according to an embodiment of the present application is shown.
[0122] As Figure 12 shown, the circuit board TCD1 of the first computing system includes a control unit TCTRL1 of the circuit board of the first computing system and a communication unit TINT1 of the circuit board of the first computing system that are electrically connected via an exchange unit TSW1 of the circuit board of the first computing system. The circuit board TCD2 of the second computing system includes a communication unit TINT2 of the circuit board of the second computing system and an extended storage unit TES2 of the circuit board of the second computing system that are electrically connected via an exchange unit TSW2 of the circuit board of the second computing system.
[0123] The control unit TCTRL1 of the circuit board of the first computing system can send a first target scheduling instruction, a first target control instruction, and the model parameters of the circuit board TCD1 of the first computing system to the switching unit TSW1 of the circuit board of the first computing system according to a scheduling strategy corresponding to the power consumption information.
[0124] In an embodiment of the present application, the switching unit TSW1 of the circuit board of the first computing system can, under the control of a second target scheduling instruction, based on a scheduling strategy corresponding to the power consumption information, send the model parameters of the circuit board TCD1 of the first computing system to the communication unit TINT1 of the circuit board of the first computing system, so as to send the model parameters of the circuit board TCD1 of the first computing system to the communication unit TINT2 of the circuit board of the second computing system via the communication unit TINT1 of the circuit board of the first computing system, and write the model parameters of the circuit board TCD1 of the first computing system received via the communication unit TINT2 of the circuit board of the second computing system into the extended storage unit TES2 of the circuit board of the second computing system.
[0125] The circuit board TCD2 of the second computing system may further include a switching unit TSW2 of the circuit board of the second computing system. The communication unit TINT2 of the circuit board of the second computing system can write the model parameters of the circuit board TCD1 of the first computing system into the extended storage unit TES2 of the circuit board of the second computing system via the switching unit TSW2 of the circuit board of the second computing system.
[0126] The embodiment of the present application is not limited thereto. The circuit board TCD2 of the second computing system may further include an acceleration card unit of the circuit board of the second computing system. The acceleration card unit of the circuit board of the second computing system can be electrically connected to the switching unit TSW2 of the circuit board of the second computing system. Thus, in another embodiment of the present application, the communication unit TINT2 of the circuit board of the second computing system can also write the model parameters of the circuit board TCD1 of the first computing system into the memory of the acceleration card unit of the circuit board of the second computing system.
[0127] In the embodiment of the present application, multiple computer cabinets can be deployed. The second control devices of the multiple computer cabinets can be connected to each other to perform data interaction between the multiple computer cabinets. The high-speed network card signal of the communication unit can not only travel through electricity within the cabinet, but also perform optical transmission in the scenario of long-distance cross-cabinet transmission. By adopting the network group clustering method, cross-chassis, cross-cabinet, and interconnection of hundreds, thousands, and tens of thousands of network cards can be realized among data nodes.
[0128] Figure 13 A schematic diagram of a computing method according to an embodiment of the present application is shown.
[0129] As Figure 13As shown, the calculation method of this embodiment can be applied to a circuit board, and the calculation method may include: operation S1310 to operation S1330.
[0130] In operation S1310, the acceleration card unit of the circuit board obtains the first model parameter from the memory of the acceleration card unit of the circuit board, and / or obtains the second model parameter from the extended storage unit of the circuit board via the switching unit of the circuit board.
[0131] In operation S1320, the acceleration card unit of the circuit board invokes at least one of the first expert model and the second expert model based on the first model parameter and / or the second model parameter to perform a model processing task, and obtains a model processing result;
[0132] In operation S1330, the monitoring unit of the circuit board transmits the call frequency information of the acceleration card unit of the circuit board for different expert models to the control device at different transmission speeds.
[0133] In the embodiment of the present application, operations S1310 to S1330 are similar to the operations performed by the circuit board described above.
[0134] For example, the monitoring unit of the circuit board transmits the call frequency information of the acceleration card unit of the circuit board for different expert models to the control device at different transmission speeds, which may include: the monitoring unit respectively selects a plurality of buses based on the storage locations of the model parameters of different expert models, and thus transmits different call frequency information to the control device via the plurality of buses, so that the control device determines the scheduling strategy of the circuit board based on the received call frequency information.
[0135] For example, in the first model processing flow of the acceleration card unit, the memory is used to store the first model parameter, and the extended storage unit is used to store the second model parameter. The control device determines the scheduling strategy of the circuit board based on the received call frequency information, which may include: the control device determines the second model processing flow after the first model processing flow based on the received call frequency information, and determines the scheduling strategy based on the second model processing flow, where the storage location of at least one of the first model parameter and the second model parameter in the second model processing flow is different from that in the first model processing flow.
[0136] For example, the calculation method may further include: the control unit of the circuit board receiving a scheduling policy from a control device; under the control of the scheduling policy, the control unit sending a first model parameter to an extended storage unit via a switching unit to write the first model parameter into the extended storage unit to change the storage location of the first model parameter; and / or, sending a second model parameter to an acceleration card unit via the switching unit to write the second model parameter into a memory to change the storage location of the second model parameter; wherein, after changing the storage location, the transmission speed of the call frequency information of each of the first expert model and the second expert model matches the second model processing flow.
[0137] For example, the multiple buses include a first bus and a second bus; transmitting different call frequency information to a control device via the multiple buses respectively may include: a monitoring unit sending the call frequency information of the model parameters stored in a memory to the control device via the first bus at a first transmission speed; and / or sending the call frequency information of the model parameters stored in an extended storage unit to the control device via the second bus at a second transmission speed; wherein, the first transmission speed is higher than the second transmission speed.
[0138] For example, when the acceleration card unit of the circuit board executes a model processing task based on one of the first model parameter and the second model parameter in the above calculation method: when the acceleration card unit of the circuit board executes a model processing task based on one of the first model parameter and the second model parameter: the acceleration card unit of the circuit board sends the other one of the first model parameter and the second model parameter to another circuit board to write the other one of the model parameters into the other circuit board, so that the other circuit board executes a model processing task based on the other one of the model parameters. It should be understood that the calculation method of the present application is not limited thereto and will not be elaborated herein.
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0140] Those skilled in the art will understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0141] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the various embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.
Claims
1. A circuit board, characterized in that: include: Exchange unit; An acceleration card unit, connected to the switching unit, wherein a first memory of the acceleration card unit is used to store a first model parameter of a first expert model; an extended storage unit connected to the exchange unit, the extended storage unit being used to store a second model parameter of a second expert model; a calling frequency of the first expert model and a calling frequency of the second expert model are different; A monitoring unit is connected to the acceleration card unit and is connected to the first control device via multiple buses; the data transmission speeds of the multiple buses are different, so that the monitoring unit can transmit the calling frequency information of the acceleration card unit calling different expert models to the first control device at different transmission speeds.
2. The circuit board according to claim 1, characterized in that: The acceleration card unit is used to: obtain the first model parameters from the first memory to call the first expert model to perform a model processing task; and / or obtain the second model parameters from the extended storage unit via the exchange unit to call the second expert model to perform a model processing task; the calling frequency of the first expert model is higher than the calling frequency of the second expert model.
3. The circuit board according to claim 2, characterized in that: The monitoring unit is further used to collect different calling frequency information of the acceleration card unit for different expert models, and transmit the different calling frequency information to the first control device via the multiple buses.
4. The circuit board according to claim 3, characterized in that: The monitoring unit is also used to select the multiple buses respectively based on the storage locations of the model parameters of the different expert models, so as to transmit the different call frequency information to the first control device via the multiple buses respectively, so that the first control device determines the scheduling strategy of the circuit board based on the received call frequency information.
5. The circuit board according to claim 4, characterized in that: In the first model processing flow of the acceleration card unit, the first memory is used to store the first model parameters, and the extended storage unit is used to store the second model parameters; The monitoring unit is further configured to send the different call frequency information to the first control device via the multiple buses, so that the first control device determines a second model processing flow after the first model processing flow based on the received call frequency information, and determines the scheduling strategy based on the second model processing flow; The storage location of at least one of the first model parameter and the second model parameter in the second model processing flow is different from the storage location in the first model processing flow.
6. The circuit board according to claim 5, characterized in that: Also includes: A control unit, connected to the switching unit, is used for: receiving the scheduling strategy from the first control device; Under the control of the scheduling strategy, sending the first model parameter to the extended storage unit via the switching unit to write the first model parameter into the extended storage unit to change the storage location of the first model parameter; and / or, sending the second model parameter to the acceleration card unit via the switching unit, so as to write the second model parameter into the first memory, so as to change the storage location of the second model parameter; After the storage location is changed, the transmission speed of the calling frequency information of the first expert model and the second expert model respectively matches the processing flow of the second model.
7. The circuit board according to claim 6, characterized in that: The control unit is also used for: After receiving the scheduling policy, the initial scheduling policy deployed by itself is updated to the scheduling policy.
8. The circuit board according to claim 7, characterized in that: The control unit is further used for: According to the initial scheduling strategy, sending an initial scheduling instruction, the first model parameter and the second model parameter to the switching unit; The switching unit is used to: Sending the first model parameter to the acceleration card unit to write the first model parameter into a first memory of the acceleration card unit; The second model parameters are sent to the extended storage unit to write the second model parameters into the extended storage unit.
9. The circuit board according to any one of claims 6 to 8, characterized in that: The second memory of the control unit is used to store the third model parameters of the third expert model; The acceleration card unit is also used for: Under the control of the control instruction received from the control unit via the switching unit, acquiring the third model parameter from the second memory; and Based on the third model parameters, the third expert model is called to perform a model processing task.
10. The circuit board according to claim 9, characterized in that: The calling frequency of the first expert model is higher than that of the third expert model, and the calling frequency of the third expert model is higher than that of the second expert model.
11. The circuit board according to any one of claims 3 to 8, characterized in that: The plurality of buses include a first bus and a second bus; the monitoring unit is connected to the first control device via the first bus and the second bus; The monitoring unit is also used for: Sending the calling frequency information of the model parameters stored in the first memory to the first control device via the first bus at a first transmission speed; and / or sending, via the second bus, the calling frequency information of the model parameters stored in the extended storage unit to the first control device at a second transmission speed; The first transmission speed is higher than the second transmission speed.
12. A computing system comprising: A plurality of circuit boards as claimed in any one of claims 1 to 11; as well as The first control device is used to collect the calling frequency information of the expert models configured in the multiple circuit boards, and determine the scheduling strategies of the multiple circuit boards based on the received calling frequency information.
13. The computing system according to claim 12, characterized in that: The multiple circuit boards include a first circuit board and a second circuit board, the first circuit board is used to store target model parameters; the model parameters stored in the first circuit board and the model parameters stored in the second circuit board are respectively used to perform different types of model processing tasks; The first circuit board is electrically connected to the second circuit board and is used to send the target model parameters to the second circuit board; The second circuit board is used to call the target expert model to execute the model processing task of the target type based on the target model parameters.
14. The computing system according to claim 13, characterized in that: The first control device is also used for: Sending target call frequency information of the target expert model from the first circuit board to a terminal device; In response to receiving a target scheduling policy corresponding to the target call frequency information from the terminal device, the target scheduling policy is sent to the first circuit board to control the first circuit board to send the target model parameters to the second circuit board.
15. The computing system according to claim 14, characterized in that: The first control device further comprises a device control unit and a system switching unit, wherein the system switching unit is electrically connected to the respective switching units of the plurality of circuit boards; Wherein, the device control unit is electrically connected to the system switching unit, and is used to send the target scheduling policy to the system switching unit, so that the switching unit of the first circuit board receives the target scheduling policy via the system switching unit, and thus writes the target scheduling policy into the first circuit board via the switching unit of the first circuit board.
16. The computing system according to any one of claims 14 to 15, characterized in that: The acceleration card unit of the first circuit board is used to send the target model parameters to the acceleration card unit of the second circuit board, so as to write the target model parameters into the first memory of the acceleration card unit of the second circuit board.
17. The computing system according to claim 16, characterized in that: The target scheduling strategy includes a first target scheduling strategy; The switching unit of the first circuit board is used to send the target model parameters to the acceleration card unit of the first circuit board based on the first target scheduling strategy.
18. The computing system according to any one of claims 14 to 15, characterized in that: The target scheduling strategy includes a second target scheduling strategy; The switching unit of the first circuit board is used to send the target model parameters to the switching unit of the second circuit board based on the second target scheduling strategy, so as to send the target model parameters to the extended storage unit of the second circuit board via the switching unit of the second circuit board, so as to write the target model parameters into the extended storage unit of the second circuit board.
19. The computing system according to any one of claims 13 to 15, characterized in that: The first circuit board is adjacent to the second circuit board, and the number of times the expert model of the first circuit board is called is higher than the number of times the expert model of the second circuit board is called.
20. The computing system according to any one of claims 13 to 15, characterized in that: The first control device comprises: a collection unit electrically connected to the monitoring unit of any circuit board via a first bus and a second bus, and configured to collect first call frequency information from the monitoring unit via the first bus, and collect second call frequency information from the monitoring unit via the second bus; A device control unit is used to determine a scheduling strategy for any circuit board based on the first call frequency information and the second call frequency information.
21. A computer cabinet, characterized in that: include: A plurality of computing systems as claimed in any one of claims 12 to 20; as well as The second control device is used to receive the call frequency information from the first control devices of the multiple computing systems, and determine the scheduling strategies of the multiple computing systems based on the received call frequency information.
22. The computer cabinet according to claim 21, characterized in that: The plurality of computing systems include a first computing system and a second computing system; The computer cabinet further includes an acquisition device electrically connected to the circuit board of the first computing system and the circuit board of the second computing system, for: Collecting power consumption information of a circuit board of the first computing system and a circuit board of the second computing system; Sending the power consumption information to a terminal device; sending a scheduling policy corresponding to the power consumption information received from the terminal device to a circuit board of the first computing system; The circuit board of the first computing system is used to send the model parameters of the circuit board of the first computing system to the circuit board of the second computing system according to the scheduling policy corresponding to the power consumption information; The circuit board of the second computing system is used to call the expert model of the circuit board of the first computing system to perform a model processing task based on the model parameters of the circuit board of the first computing system.
23. The computer cabinet according to claim 22, characterized in that: The circuit board of the first computing system includes a switching unit of the circuit board of the first computing system and a communication unit of the circuit board of the first computing system that are electrically connected, and the circuit board of the second computing system includes a communication unit of the circuit board of the second computing system and an extended storage unit of the circuit board of the second computing system that are electrically connected via the switching unit of the circuit board of the second computing system; The switching unit of the circuit board of the first computing system is used to send the model parameters of the circuit board of the first computing system to the communication unit of the circuit board of the first computing system based on the scheduling strategy corresponding to the power consumption information, so as to send the model parameters of the circuit board of the first computing system to the communication unit of the circuit board of the second computing system via the communication unit of the circuit board of the first computing system, so as to write the model parameters of the circuit board of the first computing system into the extended storage unit of the circuit board of the second computing system via the communication unit of the circuit board of the second computing system.
24. A calculation method, characterized in that: Applicable to a circuit board as claimed in any one of claims 1 to 11, The method comprises: The accelerator card unit of the circuit board obtains the first model parameter from the memory of the accelerator card unit of the circuit board, and / or obtains the second model parameter from the extended storage unit of the circuit board via the switching unit of the circuit board; The accelerator card unit of the circuit board calls at least one of the first expert model and the second expert model to perform a model processing task based on the first model parameter and / or the second model parameter, and the calling frequency of the first expert model is higher than the calling frequency of the second expert model; The monitoring unit of the circuit board transmits the calling frequency information of calling different expert models by the acceleration card unit of the circuit board to the control device at different transmission speeds.
25. The calculation method according to claim 24, characterized in that: The monitoring unit of the circuit board transmits the calling frequency information of calling different expert models by the acceleration card unit of the circuit board to the control device at different transmission speeds, including: The monitoring unit of the circuit board selects multiple buses based on the storage locations of the model parameters of the different expert models, thereby transmitting different call frequency information to the control device via the multiple buses, so that the control device determines the scheduling strategy of the circuit board based on the received call frequency information.
26. The calculation method according to claim 25, characterized in that: In the first model processing flow of the acceleration card unit, the memory is used to store the first model parameters, and the extended storage unit is used to store the second model parameters; The control device determines the scheduling strategy of the circuit board based on the received call frequency information, including: The control device determines, based on the received call frequency information, a second model processing flow following the first model processing flow, and determines the scheduling strategy based on the second model processing flow; The storage location of at least one of the first model parameter and the second model parameter in the second model processing flow is different from the storage location in the first model processing flow.
27. The calculation method according to claim 26, characterized in that: The method further comprises: The control unit of the circuit board receives the scheduling strategy from the control device; The control unit of the circuit board, under the control of the scheduling strategy, sends the first model parameter to the extended storage unit via the switching unit to write the first model parameter into the extended storage unit to change the storage location of the first model parameter; and / or sends the second model parameter to the accelerator card unit via the switching unit to write the second model parameter into the memory to change the storage location of the second model parameter; After the storage location is changed, the transmission speed of the calling frequency information of the first expert model and the second expert model respectively matches the processing flow of the second model.
28. The calculation method according to any one of claims 25 to 27, characterized in that: The plurality of buses include a first bus and a second bus; The transmitting the different call frequency information to the control device via multiple buses respectively includes: The monitoring unit of the circuit board sends the call frequency information of the model parameters stored in the memory to the control device via the first bus at a first transmission speed; and / or sends the call frequency information of the model parameters stored in the extended storage unit to the control device via the second bus at a second transmission speed; The first transmission speed is higher than the second transmission speed.
29. The calculation method according to any one of claims 25 to 27, characterized in that: The method also includes, when the acceleration card unit of the circuit board performs a model processing task based on one of the first model parameter and the second model parameter: the acceleration card unit of the circuit board sends the other model parameter of the first model parameter and the second model parameter to other circuit boards to write the other model parameter to the other circuit boards, so that the other circuit boards perform the model processing task based on the other model parameter; wherein the number of calls to the expert model of the circuit board is higher than the number of calls to the expert models of the other circuit boards.
Citation Information
Patent Citations
Scheduler, method of operating scheduler, and accelerator apparatus including scheduler
CN113760531A
Large model-oriented data storage optimization method
CN117931077A
Single GPU (Graphics Processing Unit) chip architecture system based on multiple types of extended memories
CN118247120A
Model training method, computing device and system
CN118278540A
Model operation method and device and electronic equipment
CN119322645A