Circuit board, computing system, computer cabinet, and computing method

By setting up switching units, acceleration card units, expansion storage units and monitoring units on the circuit board, optimizing model parameter storage and transmission strategies, the problems of high resource consumption and low scheduling efficiency in the existing technology are solved, and the server computing power is improved and the efficient utilization of resources is achieved.

CN120144323BActive Publication Date: 2025-08-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510623229.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-05
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the prior art, when increasing computing power by adding graphics processors (GPUs), the resource consumption is large, and the timeliness and efficiency of scheduling server resources are low, which restricts the development of artificial intelligence technology.

Method used

By setting up a switching unit, an acceleration card unit, an expansion storage unit and a monitoring unit on the circuit board, using the combination of different memory and buses, the efficient model parameter acquisition and scheduling strategy optimization of the acceleration card unit is realized, including storing high-frequency model parameters in the first memory, storing low-frequency model parameters in the extended memory unit, and transmitting call frequency information at different speeds through multiple buses.

Benefits of technology

Improve server computing power without adding circuit boards, save resource consumption, improve timeliness and efficiency of scheduling resources, and ensure the rapid processing of high-frequency model parameters and timely scheduling of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144323B_ABST
    Figure CN120144323B_ABST
Patent Text Reader

Abstract

The present application provides a circuit board, a computing system, a computing cabinet, and a computing method, which can be applied to the field of artificial intelligence technology. The circuit board includes: a switching unit; an accelerator card unit connected to the switching unit, wherein the first memory of the accelerator card unit is used to store the first model parameters of the first expert model; an extended storage unit connected to the switching unit, wherein the extended storage unit is used to store the second model parameters of the second expert model; the calling frequency of the first expert model is different from the calling frequency of the second expert model; the monitoring unit is connected to the accelerator card unit and connected to the first control device via multiple buses; the data transmission speeds of the multiple buses are different. The circuit board can give full play to the computing power of the accelerator card unit, thereby increasing the computing power of the server without adding circuit boards, saving resource consumption, and improving the timeliness and efficiency of scheduling resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a circuit board, a computing system, a computing cabinet, and a computing method. Background Art

[0002] Artificial intelligence (AI) technology is rapidly developing. The rapid development of large models places higher demands on server computing power than ever before. However, related technologies primarily increase computing power by adding graphics processing units (GPUs). This approach consumes a large amount of resources, resulting in poor timeliness and inefficiency in scheduling server resources, hindering the development of AI technology. Summary of the Invention

[0003] In view of the above problems, the present application provides a circuit board, a computing system, a computing cabinet and a computing method.

[0004] According to a first aspect of the present application, a circuit board is provided, comprising: a switching unit; an accelerator card unit connected to the switching unit, a first memory of the accelerator card unit being used to store first model parameters of a first expert model; an extended storage unit connected to the switching unit, the extended storage unit being used to store second model parameters of a second expert model; a calling frequency of the first expert model and a calling frequency of the second expert model are different; a monitoring unit connected to the accelerator card unit and connected to a first control device via multiple buses; the data transmission speeds of the multiple buses are different, so that the monitoring unit can transmit the calling frequency information of the accelerator card unit calling different expert models to the first control device at different transmission speeds.

[0005] According to the second aspect of the present application, a computing system is provided, comprising: a plurality of the above-mentioned circuit boards; and a first control device for collecting call frequency information of the expert models configured in the plurality of circuit boards, and determining the scheduling strategies of the plurality of circuit boards based on the received call frequency information.

[0006] According to the third aspect of the present application, a computer cabinet is provided, comprising: a plurality of the above-mentioned computing systems; and a second control device for receiving call frequency information from the first control device of the plurality of computing systems, and determining a scheduling strategy for the plurality of computing systems based on the received call frequency information.

[0007] According to a fourth aspect of the present application, a calculation method is provided, which is applied to a circuit board; the method includes: an accelerator card unit of the circuit board obtains a first model parameter from the memory of the accelerator card unit of the circuit board, and / or obtains a second model parameter from the extended storage unit of the circuit board via the switching unit of the circuit board; the accelerator card unit of the circuit board calls at least one of a first expert model and a second expert model to perform a model processing task based on the first model parameter and / or the second model parameter, and the calling frequency of the first expert model is higher than the calling frequency of the second expert model; the monitoring unit of the circuit board transmits the calling frequency information of the accelerator card unit of the circuit board calling different expert models to the first control device at different transmission speeds.

[0008] According to an embodiment of the present application, during operation, the accelerator card unit can promptly obtain the first model parameters from the first memory, and can also obtain the second model parameters from the additional extended storage unit via the exchange unit. In this way, by expanding the memory accessible to the accelerator card unit, the present application avoids the accelerator card unit consuming a long time to call model parameters from other storage spaces, thereby reducing the time it takes for the accelerator card unit to call model parameters, thereby fully utilizing the computing power of the accelerator card unit, and further increasing the computing power of the server without adding circuit boards, thereby saving resource consumption.

[0009] On this basis, the monitoring unit is connected to the accelerator card unit and can collect call frequency information for expert models with different call frequencies. Simultaneously, the monitoring unit can also be connected to the first control device via multiple buses, and transmit different call frequency information to the first control device via the multiple buses at different speeds. This allows information that requires rapid processing (such as call frequency information for frequently called expert models) to be transmitted to the first control device at a relatively high speed, enabling the first control device to promptly schedule resources for the accelerator card unit, improving the timeliness and efficiency of resource scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0011] Figure 1 A schematic diagram of a circuit board according to a first embodiment of the present application is shown.

[0012] Figure 2 A schematic diagram of a circuit board according to a second embodiment of the present application is shown.

[0013] Figure 3 A schematic diagram of a circuit board according to a third embodiment of the present application is shown.

[0014] Figure 4A schematic diagram of a computing system according to an embodiment of the present application is shown.

[0015] Figure 5 A schematic diagram of a first circuit board and a second circuit board according to a first embodiment of the present application is shown.

[0016] Figure 6 A schematic diagram of a first circuit board and a second circuit board according to a second embodiment of the present application is shown.

[0017] Figure 7 A schematic diagram of the connection of units in different circuit boards according to an embodiment of the present application is shown.

[0018] Figure 8 A schematic diagram of the connection between the first control device and the circuit board according to an embodiment of the present application is shown.

[0019] Figure 9 A schematic diagram of a computer cabinet according to a first embodiment of the present application is shown.

[0020] Figure 10 A schematic diagram of a computer cabinet according to a second embodiment of the present application is shown.

[0021] Figure 11 A connection diagram of a computer cabinet according to an embodiment of the present application is shown.

[0022] Figure 12 A schematic diagram illustrating connections between a circuit board of a first computing system and a circuit board of a second computing system according to an embodiment of the present application is shown.

[0023] Figure 13 A schematic diagram of a calculation method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0024] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0025] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0027] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0028] Figure 1 A schematic diagram of a circuit board according to a first embodiment of the present application is shown.

[0029] like Figure 1 As shown, the circuit board of this embodiment includes a switching unit SW, an acceleration card unit AI, a first memory AIS of the acceleration card unit, an extended storage unit ES and a monitoring unit MON.

[0030] The AI accelerator card unit can be a GPU. The AI accelerator card unit can perform model processing tasks based on mixed expert models (MoE). The AI accelerator card unit can be deployed with expert models and can use the deployed expert models to perform model processing tasks. Model processing tasks can refer to tasks that use expert models to process data.

[0031] The accelerator card unit AI is electrically connected to the first memory AIS of the accelerator card unit AI. The first memory AIS can store first model parameters of the first expert model. The accelerator card unit AI can obtain the first model parameters from the first memory AIS. For example, the first memory AIS can be a compression attached memory module (CAMM). The CAMM can support multiple memory types. Specifically, the first memory AIS can adopt the particle form of the sixth generation graphics double data rate synchronous dynamic random access memory (Graphics Double Data Rate 6, GDDR6), which has the advantages of high speed, high capacity, adjustable bandwidth, and low cost.

[0032] The accelerator card unit AI can also be electrically connected to the switching unit SW. For example, the switching unit SW can be a high-speed serial computer expansion bus standard signal Peripheral Component Interconnect Express (PCIE) switch. The switching unit SW can be electrically connected to the extended storage unit ES. The extended storage unit ES can store second model parameters of the second expert model. In this way, the accelerator card unit AI can obtain the second model parameters from the extended storage unit ES via the switching unit SW. It should be understood that the accelerator card unit AI can read the first model parameters from the first memory AIS and the second model parameters from the extended storage unit ES simultaneously, or can read the first model parameters and the second model parameters separately at different time periods, and this application is not limited to this. Furthermore, in an embodiment of the present application, in addition to the first memory AIS, the delay for the accelerator card unit AI to access other memories (such as the memory of the extended storage unit) can be increased to prevent the accelerator card unit AI from obtaining data from other memories and affecting the timeliness of data obtained from the first memory AIS.

[0033] The accelerator card unit AI can also be electrically connected to a monitoring unit MON. For example, the monitoring unit MON can be implemented based on an embedded controller (Baseboard Management Controller, BMC), a field programmable gate array (FPGA), or the like. The monitoring unit MON can collect information related to the accelerator card unit AI's calls to the expert model, such as operating status information and call frequency information. The call frequency information can refer to the frequency of expert model calls by the accelerator card unit AI per unit time. This call frequency can be determined based on the frequency of access to model parameters stored on the circuit board. Furthermore, the monitoring unit MON can also be electrically connected to a first control device CTD1 independent of the circuit board via multiple buses L1, ..., Ln, and transmit the monitored information via the multiple buses L1, ..., Ln, where n is an integer greater than 1. Furthermore, since the data transmission speeds of the multiple buses L1, ..., Ln differ, buses with different transmission speeds can be selected to transmit the aforementioned information. For example, the multiple buses L1, ..., Ln may include a Universal Serial Bus (USB) bus, an LVDS (Low-Voltage Differential Signaling) Tunneling Protocol and Interface (LTPI) bus, and the like. In one embodiment of the present application, the first expert model and the second expert model are invoked at different frequencies. Among the multiple buses L1, ..., Ln, the bus with a relatively high transmission speed may be used to transmit the expert model with a relatively high invocation frequency among the first and second expert models.

[0034] Based on this, during operation, the accelerator card unit AI can promptly obtain the first model parameters from the first memory AIS, and can also obtain the second model parameters from the additional extended storage unit ES via the switching unit SW. In this way, the present application avoids the accelerator card unit AI consuming a lot of time to call model parameters from other storage spaces by expanding the memory accessible to the accelerator card unit AI, thereby reducing the time it takes for the accelerator card unit AI to call model parameters, thereby fully utilizing the computing power of the accelerator card unit AI, and further increasing the computing power of the server without adding circuit boards, thus saving resource consumption.

[0035] Furthermore, the monitoring unit MON is connected to the accelerator card unit AI and can collect call frequency information for expert models with different call frequencies invoked by the accelerator card unit AI. Simultaneously, the monitoring unit MON can also be connected to the first control device CTD1 via multiple buses L1, ..., Ln, and transmit different call frequency information to the first control device CTD1 via the multiple buses L1, ..., Ln at different speeds. This allows information that requires rapid processing (such as call frequency information for frequently invoked expert models) to be transmitted to the first control device CTD1 at a relatively high speed, enabling the first control device CTD1 to promptly schedule resources for the accelerator card unit AI, thereby improving the timeliness and efficiency of resource scheduling.

[0036] In this embodiment of the present application, the ports of the switch unit SW can be configured to support different protocols based on firmware information. For example, some ports of the switch unit SW can be configured as PCIE ports, while others can be configured as Compute Express Link (CXL) ports. In this way, the switch unit SW can communicate with the accelerator card unit AI via PCIE signals and with the extended storage unit ES via CXL signals.

[0037] In an embodiment of the present application, the circuit board may further include a control unit. The control unit may be a central processing unit (CPU). The control unit is electrically connected to the switching unit SW and sends instructions to the switching unit SW. For example, the instructions sent by the control unit to the switching unit SW may be based on PCIE signals, specifically PCIEx16 signals.

[0038] The control unit may be pre-configured with an initial scheduling strategy. For example, the initial scheduling strategy may be sent to the control unit by the first control device. This initial scheduling strategy may be implemented based on a predetermined program. Based on the scheduling strategy, the control unit may send scheduling instructions and other data to the switching unit SW. Under the control of the scheduling instructions, the switching unit SW may electrically connect other units to the control unit so that data from the control unit is sent to the other units, thereby scheduling the other units. For example, the switching unit SW may send control instructions from the control unit to the accelerator card unit AI to schedule the accelerator card unit AI to perform model processing tasks.

[0039] Under the control of a scheduling instruction, the switching unit SW can also send the second model parameters from the extended storage unit ES to the accelerator card unit AI. Based on the second model parameters, the accelerator card unit AI can also invoke the second expert model to perform model processing tasks and obtain the model processing results of the second expert model. The CXL link must have sufficient bandwidth to match the GPU requirements to prevent bandwidth from becoming a resource bottleneck for invoking the extended storage unit ES.

[0040] Based on this, under the control of the control instructions, the accelerator card unit AI can obtain the first model parameters from the first memory AIS, or obtain the second model parameters from the extended storage unit ES via the switching unit SW. Thereafter, the accelerator card unit AI can call the first expert model based on the obtained first model parameters to perform the model processing task and obtain the model processing results of the first expert model; it can also call the second expert model based on the obtained second model parameters to perform the model processing task and obtain the model processing results of the second expert model; it can also call the first expert model and the second expert model simultaneously to perform the model processing task, thereby obtaining the model processing results of the first expert model and the model processing results of the second expert model, respectively. The frequency of calling the first expert model is higher than the frequency of calling the second expert model. In this way, during operation, the accelerator card unit AI can prioritize reading the first model parameters from the first memory AIS, which can be accessed at a relatively high speed, to perform the model processing task.

[0041] Based on this, the time it takes for the accelerator card unit AI to read the first model parameters from the first memory AIS is shorter than the time it takes to read the second model parameters from the extended storage unit ES via the switching unit SW. Thus, storing the first model parameters, which require frequent access, in the first memory AIS, which the accelerator card unit AI can access more quickly, facilitates frequent access of the first expert parameters by the accelerator card unit AI. Second model parameters, which require less frequent access, are stored in the extended storage unit, which has a relatively slower access speed. This allows the accelerator card unit AI's computing power to be fully utilized for frequent processing of the first expert parameters, reducing the waste of computing power in the accelerator card unit AI. This reduces resources consumed to increase server computing power, improves the efficiency of the accelerator card unit AI in acquiring model parameters, and reduces the power consumption generated by the accelerator card unit AI in acquiring model parameters. Furthermore, pre-storing the less frequently accessed second model parameters in the extended storage unit ES avoids the need to temporarily schedule the storage location of the second model parameters when needed, which consumes additional scheduling time. This allows the model parameters of the less frequently accessed expert model to be read from the extended storage unit ES in a timely manner. This improves the efficiency of the accelerator card unit AI in executing model processing tasks.

[0042] Furthermore, the monitoring unit MON can collect information on different call frequencies of different expert models by the accelerator card unit AI and transmit the information to the first control device via multiple buses L1, ..., Ln. For example, the monitoring unit MON can collect call frequency information for the first expert model when the accelerator card unit AI calls the first expert model, and can also collect call frequency information for the second expert model when the accelerator card unit AI calls the second expert model.

[0043] For example, the multiple buses L1, ..., Ln may include a first bus and a second bus. The monitoring unit MON and the first control device CTD1 may be electrically connected via the first bus and the second bus, and may send the call frequency information of the first expert model to the first control device via the first bus at a first transmission speed, and send the call frequency information of the second expert model to the second control device via the second bus at a second transmission speed. The first transmission speed is higher than the second transmission speed. In this way, for the first expert model with a relatively high call frequency, a bus with a relatively faster transmission speed is used for transmission, so that the first control device can timely obtain the call status of the high-frequency call expert model by the acceleration card unit AI, so that the resource scheduling for the acceleration card unit AI can be carried out in a timely manner, thereby improving the timeliness of the scheduling resources and improving the efficiency of the scheduling resources.

[0044] The embodiments of the present application are not limited thereto. The monitoring unit MON may further select multiple buses L1, ..., Ln based on the storage locations of the model parameters of different expert models, thereby transmitting different call frequency information to the first control device CTD1 via the multiple buses L1, ..., Ln. For example, the monitoring unit MON may select the first bus to transmit the first call frequency information of the model parameters stored in the first memory AIS when the accelerator card unit AI reads the model parameters from the first memory, or may select the second bus to transmit the second call frequency information of the model parameters stored in the extended storage unit ES when the accelerator card unit AI reads the model parameters from the extended storage unit ES.

[0045] The first control device CTD1 may pre-store a mapping relationship between the first call frequency information, the second call frequency information, and the scheduling strategy. For example, the mapping relationship may be in the form of a mapping table, etc. When the first call frequency information and the second call frequency information are received, the mapping relationship stored in the first control device CTD1 may be used to determine the scheduling strategy associated with the first call frequency information and the second call frequency information, and the queried scheduling strategy may be sent to the control unit. In this way, based on the association between the storage location of the model parameters and the call frequency of the expert model described above, the call frequency information of the expert model to which the model parameters belong may be transmitted using the bus corresponding to the storage location for the specific storage location of the model parameters, thereby avoiding the problem that the speed of the bus used to transmit the call frequency information does not correspond to the call frequency information due to a change in the call frequency of the expert model in some cases, thereby affecting the timeliness of transmitting the call frequency information of the expert model with high frequency calls. The embodiments of the present application are not limited to this. The above-mentioned mapping relationship can also be stored in the terminal device, and the first control device CTD1 can also send the first call frequency information and the second call frequency information to the terminal device. The terminal device performs the above-mentioned operations to obtain the scheduling strategy, and sends the scheduling strategy to the first control device CTD1, and then the first control device CTD1 sends the scheduling strategy to the control unit.

[0046] In other embodiments of the present application, each model processing task may include multiple model processing flows. In different model processing flows, the frequency with which the accelerator card unit AI calls the same expert model may be different. For example, in the first model processing flow of the accelerator card unit AI, because the call frequency of the first expert model is higher than the call frequency of the second expert model, the first memory AIS may store the first model parameters, and the extended storage unit ES may store the second model parameters, so that the accelerator card unit AI can promptly obtain the model parameters of the frequently called expert model. In this case, the monitoring unit MON may send different call frequency information to the first control device CTD1 via multiple buses L1, ..., Ln, so that the first control device CTD1 determines the second model processing flow after the first model processing flow based on the received call frequency information. For example, the first control device CTD1 may pre-store a first mapping relationship between call frequency information and model processing flow, and a second mapping relationship between model processing flow and scheduling policy. For example, the first mapping relationship and the second mapping relationship may be in the form of a mapping table. In this way, the first control device CTD1 can use the first mapping relationship based on the received call frequency information to determine that the current model processing process is the first model processing process, and determine the second model processing process that is after the first model processing process based on the order of multiple model processing processes in the first mapping relationship. Then, according to the second model processing process, the second mapping relationship can be used to determine the scheduling strategy corresponding to the second model processing process. Thereafter, when the first model processing process ends, the determined scheduling strategy can be sent to the control unit. Among them, the end of the first model processing process can be determined when the accelerator card unit AI stops calling the expert model. In another embodiment of the present application, the call frequency information can also be sent to the terminal device via the first control device CTD1, so that the terminal device can determine the scheduling strategy based on the call frequency information.

[0047] The storage location of at least one of the first model parameter and the second model parameter in the second model processing flow is different from the storage location in the first model processing flow. For example, in the second model processing flow, the call frequency of the first expert model and the call frequency of the second expert model may change relative to the first model processing flow. Consequently, the model parameters stored in the first memory AIS may change, for example, from storing the first model parameters to storing the second model parameters. Correspondingly, the model parameters stored in the extended storage unit ES may change from the second model parameters to the first model parameters. In this case, the storage locations of the first model parameters and the second model parameters change. For example, after receiving a scheduling policy from the first control device CTD1, the control unit may update its previously deployed initial scheduling policy to a new scheduling policy. Under the control of this scheduling policy, the control unit may send the first model parameters to the extended storage unit ES via the switching unit SW to write the first model parameters to the extended storage unit, thereby changing the storage location of the first model parameters, and / or send the second model parameters to the accelerator card unit AI via the switching unit SW to write the second model parameters to the first memory, thereby changing the storage location of the second model parameters. For example, after writing the first model parameters to the extended storage unit ES, the first model parameters can overwrite the second model parameters previously stored in the extended storage unit ES. After writing the second model parameters to the first memory, the second model parameters can overwrite the first model parameters previously stored in the extended storage unit ES. By changing the storage location of the model parameters, the transmission speed of the call frequency information of the first and second expert models can be adjusted simultaneously, thereby ensuring that the transmission speed of the call frequency information of the first and second expert models matches the second model processing flow.

[0048] Specifically, the storage locations of the first and second model parameters are changed. The first model parameters are stored in the extended storage unit ES as the model parameters of the first expert model, which is invoked less frequently, while the second model parameters are stored in the first internal memory AIS as the model parameters of the second expert model, which is invoked more frequently. Consequently, in a second model processing flow that differs from the first model processing flow, the first expert model becomes the expert model for infrequent invocation, while the second expert model becomes the expert model for frequent invocation. Consequently, the invocation frequency information of the second expert model can be transmitted at the first transmission speed, while the invocation frequency information of the first expert model can be transmitted at the second transmission speed.

[0049] The embodiments of the present application are not limited to this. In the embodiments of the present application, the first model processing flow and the second model processing flow may each belong to different model processing tasks. For example, the calculation results obtained in the first model processing flow may not be used as input information for the second model processing flow. Specifically, the second model processing flow may be a model processing flow for a model processing task executed by the AI accelerator card unit of another circuit board, and so on.

[0050] On this basis, in order to address the problem in some solutions that models deployed on a single GPU (such as a deep learning model architecture based on a self-attention mechanism (Transformer)) can only execute multiple model processing processes of the model processing task in series according to a fixed pipeline during the execution of the same model processing task, making it difficult to schedule GPU-related resources in a timely manner, the above-mentioned solution of the present application can decouple the model processing tasks that are executed in series, and use multiple circuit boards to execute the model processing processes that were originally executed in series in the form of a pipeline in parallel. In the model processing processes that are executed in parallel, the monitoring units of the respective circuit boards that execute the model processing processes in parallel can send the call frequency information corresponding to the model parameters at the corresponding transmission speed according to the storage location of the model parameters. For example, the monitoring unit transmits the call frequency information of the expert model with a relatively high call frequency at a relatively high speed. In this way, even if the acceleration card unit AI of any circuit board receives the scheduling strategy of the model processing flow belonging to other circuit boards and thus executes the same model processing task in parallel with other circuit boards, the first control device CTD1 can also promptly obtain the acceleration card unit AI of each of the multiple circuit boards that process the same model processing task in parallel based on the positions of the model parameters stored in each of the multiple circuit boards, and send the calling frequency information of the high-frequency called expert model deployed on each circuit board to the first control device CTD1 in a timely manner, so that the first control device CTD1 can promptly perform resource scheduling for multiple circuit boards based on the received calling frequency information of the high-frequency called expert model, thereby improving the timeliness of scheduling resources and improving the efficiency of scheduling resources.

[0051] Figure 2 A schematic diagram of a circuit board according to a second embodiment of the present application is shown.

[0052] like Figure 2As shown, the expansion control unit CTRL may include an expansion control subunit ES_1 and multiple expansion storage subunits. For example, the multiple expansion storage subunits may include a first expansion storage subunit ES_21 and a second expansion storage subunit ES_22. For example, the expansion control subunit ES_1 may be a memory extension controller (MXC). The expansion storage subunit may utilize Double Data Rate Synchronous Dynamic Random Access Memory (DDR) granules. This memory represents the physical form of a memory resource pool extended based on the CXL protocol, maximizing its capacity based on bandwidth. In addition to data exchange between the accelerator card unit AI and the storage subunit, the control unit CTRL can also exchange data with the storage subunit via the switch unit SW and the CXL 3.1 Global Interconnect Manager (GIM). Memory resource functions are accessible to the control unit CTRL. This allows for dual acceleration based on the CPU as a general computing node and the GPU as an intelligent computing node, addressing memory capacity shortages and other memory-related issues in both general computing and intelligent computing. The extended control subunit ES_1 can control at least one of the multiple extended storage subunits to store data transmitted via the switching unit SW, and can also send the data stored in at least one of the multiple extended storage subunits to the switching unit SW for transmission to other units. By setting the extended control subunit ES_1 and the extended storage subunit on the circuit board, the extended control subunit ES_1 and the extended storage subunit can be used to store model parameters with low call frequency, so that when the computing power of the circuit board is insufficient, the model parameters stored in the extended storage subunit can be sent to other circuit boards for calculation. In this way, the strategy of exchanging memory for computing power adopted in this application can solve the problem of insufficient computing power, greatly improve the quality of activated experts on a single circuit board, and greatly increase the number of inactivated experts. In addition, in this application, the new architecture of the extended control subunit ES_1 and the extended storage subunit is adopted and combined with the CXL cache consistency protocol to expand the memory of the accelerator card unit AI, and then the latency problem can be improved by adopting multi-level cache and increasing the cache capacity through the extended control subunit ES_1 and the extended storage subunit. It should be understood that, Figure 2 The number of the extended control subunit ES_1 and the extended storage subunit is only an example, and this application does not limit the number of the extended control subunit ES_1 and the extended storage subunit.

[0053] In this application, in addition to the first and second expert models, the accelerator card unit AI can also call a third expert model to perform model processing tasks. The first expert model is called more frequently than the third expert model, and the third expert model is called more frequently than the second expert model. The second memory CTRLS of the control unit CTRL can be used to store third model parameters of the third expert model. The second memory CTRLS can be in the form of ordinary DDR particles, which is adapted to the platform on which the CPU is located. Furthermore, when using DDR, using an LGA (Land Grid Array, LGA) base reduces signal transmission loss and increases signal rate.

[0054] The accelerator card unit AI can also obtain the third model parameters from the second memory CTRLS under the control of the control instruction, and based on the third model parameters, call the third expert model to perform the model processing task to obtain the model processing result of the third expert model. For example, the first expert model, the second expert model and the third expert model can be divided based on a predetermined call frequency range. In an embodiment of the present application, the first expert model, the second expert model and the third expert model can be used to perform the same type of model processing tasks. For example, the expert models set on the same circuit board can be used to process the same type of minimum text unit (token). For example, the first expert model, the second expert model and the third expert model are all used to process punctuation marks, or are all used to process words, etc.

[0055] The second memory CTRLS of the control unit CTRL can store the first model parameter, the second model parameter and the third model parameter. In an embodiment of the present application, the control unit CTRL can pre-schedule the resources of the first memory AIS of the accelerator card unit AI and the resources of the extended storage unit ES to store the model parameters before the accelerator card unit AI executes the model processing task. For example, the control unit CTRL can send a scheduling instruction, the first model parameter and the second model parameter to the switching unit SW according to the initial scheduling strategy. Under the control of the scheduling instruction, the switching unit SW can send the first model parameter to the accelerator card unit AI to write the first model parameter into the first memory AIS of the accelerator card unit AI, and can also send the second model parameter to the extended storage unit ES to write the second model parameter into the extended storage unit ES.

[0056] Figure 3 A schematic diagram of a circuit board according to a third embodiment of the present application is shown.

[0057] like Figure 3 As shown, the circuit board of this embodiment may further include a communication unit INT in addition to the above-mentioned control unit CTRL, switching unit SW, acceleration card unit AI, extended storage unit ES and monitoring unit MON.

[0058] The communication unit INT may be a network interface card (NIC), may be electrically connected to the switching unit SW, receive data from the control unit CTRL via the switching unit SW, and transmit the received data to the other circuit boards based on the communication connection with the other circuit boards.

[0059] The monitoring unit MON can be electrically connected to a first interface CON1, which can be a high-speed connector. The monitoring unit MON can output bus signals of various speeds via the first interface CON1. For example, the first interface CON1 can include a USB interface, LTPI, etc. The USB interface can be used to transmit USB signals. The LTPI interface can be used to transmit two-wire serial bus (Inter-Integrated Circuit, I2C) signals, Universal Asynchronous Receiver / Transmitter (UART) signals, General Purpose Input / Output (GPIO) signals, and other signals.

[0060] The switching unit SW may be electrically connected to the second interface CON2, which may be a high-speed connector. The switching unit SW may be connected to a switching unit SW of another circuit board via the second interface CON2.

[0061] The acceleration card unit AI can be electrically connected to the third interface CON3, which can be a high-speed connector. The acceleration card unit AI can be connected to acceleration card units AI of other circuit boards via the third interface CON3.

[0062] The communication unit INT may be electrically connected to the fourth interface CON4, which may be a high-speed connector. The communication unit INT may be connected to a communication unit INT of another circuit board via the fourth interface CON4.

[0063] Specifically, after the circuit board is powered on and begins operation, it can function not only as a computing core but also as a GPU and memory pool. When the accelerator card AI unit is operating, in addition to obtaining model parameters for the first memory AIS, it can also obtain model parameters for the second memory CTRLS and the extended storage unit ES. For example, the accelerator card AI unit can call the Compute Unified Device Architecture (CUDA) to allocate an area on the first memory AIS. It can also call CUDA to allocate the same amount of space on the second memory CTRLS, initialize data, and copy the initialized data from the first memory AIS to the second memory CTRLS. Furthermore, the Uniform Memory Access (UMA) technology can automatically take over the memory placement task. This allows the UMA to automatically handle the process by simply calling the fixed application programming interface (API) in CUDA.

[0064] However, when there are multiple nodes (such as the circuit board mentioned above) in the entire chassis and cabinet, since the processors of each node access memory through a shared bus, as the number of processors increases, the bus will become a performance bottleneck, resulting in slower memory access speeds and affecting overall system performance. In addition, the memory that the control unit CTRL can carry is limited. Therefore, calling the memory of the extended storage unit ES becomes particularly important and convenient. Through CXL bus technology, the CXL resources of the control unit CTRL can be converted into a memory expansion resource pool, allowing the accelerator card unit AI to easily call the extended memory resources. The prerequisites for calling are:

[0065] CXL interface integration: The model processing subunit AI_H supports the CXL.memory protocol or can access the extended storage subunit through the extended control subunit ES_1. The CXL.memory protocol allows the control unit CTRL and the accelerator card unit AI to directly access or pool the memory of the extended storage subunit.

[0066] ② Physical connection: CXL extended memory is connected to the accelerator card unit AI and the control unit CTRL through high-speed interconnects (such as CXL links), ensuring low latency and high bandwidth.

[0067] ③ Virtualization support: Map the CXL extended memory to a unified virtual address space, and the accelerator card unit AI accesses the memory of the extended storage subunit through its built-in memory management unit (MMU).

[0068] ④ Page table management: The accelerator card unit AI driver collaborates with the control unit CTRL to manage the page table of the memory of the extended storage subunit and processes the address translation request of the accelerator card unit AI through the input-output memory management unit (IOMMU) or MMU.

[0069] ⑤CXL.cache protocol: The accelerator card unit AI supports cache consistency and uses CXL.cache to maintain the cache status between the accelerator card unit AI and the control unit CTRL to avoid data inconsistency.

[0070] ⑥ Hardware consistency domain: A consistency domain is built within the circuit board, allowing the accelerator card unit AI and the control unit CTRL to transparently access the extended storage subunit without explicit refresh.

[0071] ⑦ At the driver and operating system layer: The control unit CTRL recognizes access to the extended storage subunit as an allocatable resource, provides an API for management by the CXL.mem module driver of the Linux system, and can modify the preset driver of the accelerator card unit AI so that the accelerator card unit AI can support memory allocation of the extended storage subunit, and enable the accelerator card unit AI to process DMA (Direct Memory Access) requests and address mapping.

[0072] After the memory of the extended storage subunit is called, the system still needs further optimization. In this way, based on the detected call frequencies of the first, second, and third expert models, the locations of the model parameters stored in the first memory AIS, the extended storage unit ES, and the second memory CTRLS can be adjusted, placing frequently accessed data in the first memory AIS and storing less frequently accessed data in the extended storage subunit. Data records such as call frequencies will be acquired by the monitoring unit MON, which can send the acquired call frequency information to the terminal device to update the scheduling policy of the control unit CTRL.

[0073] On a single circuit board, a small-scale hybrid expert model can be deployed based on the capacity of the first memory AIS, the extended storage subunit, and the second memory control unit CTRLS. The number of expert models that can be deployed on a single circuit board can be calculated using the following formula:

[0074] Maximum number of expert models = [(capacity of the first memory AIS - overhead of the 2GB circuit board) + (total memory capacity that can be called by the accelerator card unit AI × predetermined model parameter call factor)] / (number of expert model parameters × number of bytes processed by the accelerator card unit AI per unit time + activation value occupied value) × (50%~70%) (1)

[0075] It can be seen from the above formula that 50 expert models can be deployed on a single circuit board. Compared with the 20 expert models in the related technology, the number of model deployments is increased, which facilitates improving the computing power of the circuit board of this application.

[0076] Furthermore, multiple groups of expert models can be deployed on a single circuit board, and each group of expert models includes expert models of the same type. Specifically, the architecture and workflow of MOE can be expressed by the following two formulas:

[0077] MOE (x)=∑K(Gi(x)Ei(x)) (2)

[0078] G(x)=TopK(Softmax(Wg(x)+€)): (3)

[0079] Where Ei(x) is the model processing result of the input token after passing through the i-th expert model. Gi(x) is the output weight of the i-th expert model (between 0 and 1). G(x) represents the total model processing result. TopK selects the top K largest values from the input data. Softmax represents the activation function. Wg represents the predetermined weight. x represents the input data. € represents the predetermined variable.

[0080] From the architecture and workflow of MOE, it can be seen that in addition to having fewer activation parameters and fast inference speed, it is also highly scalable. By increasing or decreasing the number of expert models, the parameter scale of the expert model can be adjusted at will to adapt to more different tasks.

[0081] Based on the working principle of the MOE architecture, this application further deploys 1 to 2 expert models for executing the same type of model processing tasks for a single circuit board CD, which can further reduce the parameters activated by a single circuit board CD, thereby better allocating memory resources to the expert models deployed on a single circuit board CD, reducing the communication overhead caused by data interaction between multiple circuit boards CDs when the computing system CS is working, and avoiding memory-related problems such as video memory fragmentation.

[0082] Compared to the related art of excessively stacking GPUs and increasing GPU power consumption, this application further refines the hardware architecture based on the MOE model. This not only increases the number of expert models in the MOE model to improve computing power, but also makes each computing hardware node more refined. Instead of deploying the expert models in the MOE model on all GPUs in the entire machine cluster, the expert models will be deployed on a confirmed single computing chip based on the characteristics of the business processing. This means that the large feedforward neural network (FFN) network in the Transformer model is replaced with multiple small expert models (also FFN structures). A sparse activation mechanism is then used to activate only one or two expert models of the same type during each inference, greatly reducing the number of computational parameters required for inference. This greatly improves the inference efficiency of the entire MOE model. While increasing the number of stored parameters, it also improves the interpretability of model processing and the ability to obtain local information. It should be noted that each expert model is also an FFN structure. The expert models of this application can be compressed, for example, in the direction of compression, such as sparsity, distillation, and quantization.

[0083] Based on this, the number of experts in this application can be adjusted according to system needs. Since only the same type of expert model is activated at a time in this application, the parallel capability of the expert model needs to be improved to handle dedicated transactions, which places higher requirements on the memory capacity and bandwidth of the current accelerator card unit AI.

[0084] Figure 4 A schematic diagram of a computing system according to an embodiment of the present application is shown.

[0085] like Figure 4 As shown, the computing system of this embodiment includes multiple circuit boards CD and a first control device CTD1. The circuit boards CD here can be any of the circuit boards CD mentioned above. A single computing system can be deployed in a single chassis.

[0086] The first control device CTD1 can collect call frequency information from expert models configured on multiple circuit boards CD and transmit this call frequency information to a terminal device. For example, the first control device CTD1 can receive the call frequency information from the monitoring unit MON of the circuit board CD and determine a scheduling policy for the circuit board based on the call frequency information. For example, a mapping relationship between the call frequency information and the scheduling policy can be pre-set, and the mapping relationship can be used to determine the scheduling policy corresponding to the call frequency information.

[0087] Optionally, the plurality of circuit boards CD include a first circuit board CD1 and a second circuit board CD2. The model parameters stored in the first circuit board CD1 and the model parameters stored in the second circuit board CD2 are respectively used to perform different types of model processing tasks. For example, the model parameters stored in the first circuit board CD1 and the model parameters stored in the second circuit board CD2 can be used to process different types of data. Specifically, the model parameters stored in the first circuit board CD1 and the model parameters stored in the second circuit board CD2 are used to process different types of minimum units (tokens). For example, the model parameters of the first circuit board CD1 can be used to process punctuation marks, while the model parameters of the second circuit board CD2 can be used to process words, and so on.

[0088] In an embodiment of the present application, the model parameters stored in the first circuit board CD1 may be target model parameters for executing a target model processing task. The first circuit board CD1 is electrically connected to the second circuit board CD2 and may send the target model parameters to the second circuit board CD2. After receiving the target model parameters, the second circuit board CD2 may call the target expert model to execute the target type model processing task based on the target model parameters and obtain the target model processing result. Based on this, by scheduling the computing power resources of the second circuit board CD2 to assist in processing the target model parameters of the first circuit board CD1, it is possible to avoid problems such as the first circuit board CD1 failing due to the first circuit board CD1 calling the target expert model too frequently, thereby improving the reliability of the circuit board CD and improving the redundancy of the computing system.

[0089] Figure 5 A schematic diagram of a first circuit board and a second circuit board according to a first embodiment of the present application is shown.

[0090] exist Figure 5In the illustrated embodiment, the first circuit board CD1 includes a first control unit CTRL1 and an accelerator card unit AI1 electrically connected via a switch unit SW1 of the first circuit board. The first control unit CTRL1 can send a first scheduling instruction, a first control instruction, and target model parameters to the switch unit SW1 of the first circuit board according to a first target scheduling policy. Under the control of the first scheduling instruction, the switch unit SW1 of the first circuit board can send the first control instruction and target model parameters to the accelerator card unit AI1 of the first circuit board based on the first target scheduling policy. It should be understood that in this application, the first circuit board CD1 and the second circuit board CD2 are both the circuit boards described above, and therefore, both have structures corresponding to the circuit boards described above. For example, the accelerator card unit AI1 of the first circuit board and the accelerator card unit AI2 of the second circuit board are both the accelerator card units described above, and the same applies to the other units in the first circuit board CD1 and the second circuit board CD2, which are not described in detail here. Similarly, the circuit boards described in the following sections of this application are similar to the circuit boards described above, and the units of the circuit boards described in the following sections are also similar to those described above, and are not described in detail here.

[0091] The second circuit board CD2 includes an accelerator card unit AI2. The accelerator card unit AI1 of the first circuit board is electrically connected to the accelerator card unit AI2 of the second circuit board. Under the control of a first control instruction, the accelerator card unit AI1 can send target model parameters to the accelerator card unit AI2 of the second circuit board, thereby writing the target model parameters into the memory of the accelerator card unit AI2 of the second circuit board. The memory here refers to the first memory of the accelerator card unit.

[0092] Figure 6 A schematic diagram of a first circuit board CD1 and a second circuit board CD2 according to a second embodiment of the present application is shown.

[0093] exist Figure 6 In the illustrated embodiment, the first circuit board CD1 includes a first control unit CTRL1 and an accelerator card unit AI1 electrically connected via a switch unit SW1 of the first circuit board. The first control unit CTRL1 can send a second scheduling instruction, a second control instruction, and target model parameters to the switch unit SW1 of the first circuit board according to the second target scheduling policy.

[0094] The second circuit board CD2 may include a switching unit SW2 of the second circuit board and an extended storage unit ES2 of the second circuit board. The switching unit SW1 of the first circuit board is electrically connected to the switching unit SW2 of the second circuit board and can, under the control of the second scheduling instruction and based on the second target scheduling policy, send target model parameters to the switching unit SW of the second circuit board. The switching unit SW of the second circuit board is electrically connected to the extended storage unit ES of the second circuit board. In this way, the switching unit SW1 of the first circuit board can send the target model parameters to the extended storage unit ES of the second circuit board via the switching unit SW2 of the second circuit board, so that the target model parameters are written to the extended storage unit ES of the second circuit board.

[0095] The embodiments of the present application are not limited thereto. The first circuit board CD1 may further include a first communication unit, and the second circuit board CD2 may further include a second communication unit. The switching unit SW1 of the first circuit board may further be electrically connected to the first communication unit. In this way, the first communication unit may receive the target model parameters via the switching unit SW1 of the first circuit board. The first communication unit is communicatively connected to the second communication unit and may send the target model parameters to the second communication unit. The second communication unit is electrically connected to the switching unit SW2 of the second circuit board and may send the target model parameters to the extended storage unit ES2 of the second circuit board via the switching unit SW2 of the second circuit board.

[0096] Figure 7 A schematic diagram of the connection of units in different circuit boards according to an embodiment of the present application is shown.

[0097] like Figure 7 As shown, accelerator card units on different circuit boards in a computing system can be connected to each other. For example, the first accelerator card unit can be connected to the second accelerator card unit, the first accelerator card unit and the second accelerator card unit can be connected to the Nth accelerator card unit, and so on. In this way, each accelerator card unit can be connected to other accelerator card units, thereby facilitating data transmission between accelerator card units. The first accelerator card unit may refer to the accelerator card unit on the first circuit board, and the other accelerator card units are similarly described and will not be further described here.

[0098] Similarly, switching units on different circuit boards in a computing system can be connected to each other. For example, the first switching unit can be connected to the second switching unit, the first switching unit and the second switching unit can be connected to the Nth switching unit, and so on. In this way, each switching unit can be connected to other switching units, thereby facilitating data transmission between switching units. The first switching unit may refer to the switching unit on the first circuit board, and the other switching units are similarly connected, and are not described in detail here.

[0099] Similarly, the communication units of different circuit boards in the computing system can be connected to each other. For example, the first communication unit can be connected to the second communication unit, the first communication unit and the second communication unit can be connected to the Nth communication unit, and so on. In this way, each communication unit can be connected to other communication units, thereby facilitating data transmission between communication units. Among them, the first communication unit may refer to the communication unit of the first circuit board, and the other communication units are similar, which will not be described here. It should be noted that for the sake of illustration, Figure 7 In the figure, the first communication unit and the Nth communication unit are connected by a black line, but it should be understood that this is only for illustration, and in fact the communication units may be communicatively connected.

[0100] Furthermore, in the present application, the switching units of multiple circuit boards are connected to form a FABRIC (structure) link. Within a single computing system, this network architecture has good nanosecond-level low-latency performance. On this basis, the accelerator card units of multiple circuit boards are connected to realize peer-to-peer communication (P2P) between the accelerator card units, thereby forming a dual-ring network with the above-mentioned FABRIC network architecture. In this way, model parameters with high call frequencies can be transmitted between the connected accelerator card units, and model parameters with high call frequencies can also be transmitted between the connected switching units. The network formed by the connected accelerator card units and the network formed by the connected switching units are redundant with each other, and the amount of data related to the reduce function of the link can be reduced to half of the original amount, which can meet both the general low-latency transmission requirements and the special needs of the communication library. In addition, the above-mentioned FABRIC link can also be used to transmit data stored in the extended storage unit and the data in the memory of the control unit. Since the network formed between the accelerator card units does not require a third-party unit for data transfer, this can avoid the delay caused by the data transfer of the third-party unit. Therefore, the dedicated link used to transmit high-frequency data can also be regarded as an exclusive link for the accelerator card units to call each other's memory.

[0101] The data throughput of the network composed of connected communication units is larger than that of the above-mentioned dual-ring network, but data needs to be demodulated during data transmission, so it can be used to transmit data with a call frequency other than high-frequency call data.

[0102] On this basis, three balanced topologies can be formed through the connection methods between the switching units, accelerator card units and communication units of the above-mentioned different circuit boards, thereby reducing the amount of data related to the reduce function of the link, which is conducive to optimizing the resource allocation of the system.

[0103] Optionally, the first control device may send target call frequency information of the target expert model from the first circuit board to the terminal device. For example, the target call frequency information may be sent to the terminal device via an RJ45 (Registered Jack) network interface. After receiving the target call frequency information, the terminal device may display the target call frequency information through a visual interface so that relevant personnel can send a target scheduling policy to the first control device through the terminal device based on the displayed target call frequency information. In response to receiving the target scheduling policy corresponding to the target call frequency information from the terminal device, the first control device may send the target scheduling policy to the first circuit board to control the first circuit board to send target model parameters to the second circuit board, thereby implementing resource scheduling of the second circuit board. In an embodiment of the present application, the target scheduling policy may also be determined based on a combination of other information such as GPU load information, first memory occupancy information, second memory control unit occupancy information, circuit board power consumption information, and the like. The method is similar to that described above and will not be elaborated upon here.

[0104] For some exemplary server hardware architectures, due to their high complexity and difficulty in management, they will result in high operation and maintenance costs, high management difficulty and high costs. In addition, there are a large number of nodes in this server hardware architecture, and the computing power of some nodes may not be fully utilized, resulting in a waste of resources and affecting the computing efficiency of the architecture. In addition, this server hardware architecture is subject to hardware limitations. Since the hardware resources of a single server have an upper limit, it is difficult to further expand the hardware resources of the server after the hardware resources reach the upper limit, resulting in limited computing power of the server. In addition, since a large number of resources are deployed on the same server, the failure of a server node may cause the entire system to crash. In addition, the above-mentioned server hardware architecture is inefficient when processing a large number of small memory read and write tasks. The present application adopts hardware-aware routing technology and dynamic monitoring technology to determine the status information such as the call frequency of the circuit board, and then when the computing power of the accelerator card unit of any circuit board is insufficient, the model parameters stored in the circuit board can be sent to other circuit boards for processing. In this way, when any circuit board calls the first target expert model for executing the target model processing task of the target type, the circuit board can send the model parameters of the second target expert model for the target model processing task to other circuit boards. In this way, the above-mentioned circuit board can execute part of the data processing process of the target model processing task, while other circuit boards can call the second target expert model to execute the data processing process of other parts of the target model processing task, realizing pipeline parallel processing (Pipeline Parallelism) context tasks, and solving the problem of low parallel computing capability of large models in related technologies. In addition, the distribution of expert models between the above-mentioned circuit board and other circuit boards can be dynamically adjusted by means of the above-mentioned update adjustment strategy, thereby realizing dynamic batching (Dynamic Batching), solving the problem of insufficient computing power caused by insufficient memory of large models in related technologies, and balancing the differences in computing time of different expert models, thereby improving computing efficiency.

[0105] In this way, by achieving cross-device, cross-system and even cross-architecture scheduling, we can break through the limitations of the hardware architecture on the computing power scheduling of the accelerator card unit, and make full use of the resources of the accelerator card unit, reduce resource waste, and lower costs. It can solve the problems of insufficient large model parameters, few expert models, and low capabilities in related technologies, and can solve the problem of fixed resource waste caused by the current large model deployment method.

[0106] Optionally, multiple circuit boards are stacked in the vertical direction in sequence according to the number of calls to the configured expert model. The first circuit board is adjacent to the second circuit board, and the number of calls to the expert model of the first circuit board is higher than the number of calls to the expert model of the second circuit board. In this way, when the number of calls to the expert model of the first circuit board is higher, by scheduling the resources of the second circuit board adjacent to the first circuit board to perform model processing tasks based on the model parameters of the first circuit board, it is possible to avoid problems such as the first circuit board calling the target expert model too frequently, which may cause the first circuit board to malfunction, thereby improving the reliability of the circuit board. It should be understood that in the embodiment of the present application, the resources of other circuit boards called by the first circuit board are not limited to the circuit boards adjacent to the first circuit board, but may also include circuit boards with lower number of calls to other expert models. Based on this, for a single computing system, the circuit board of the present application uses a single accelerator card unit to connect to a single control unit, and performs special deployment of the expert model on a single circuit board to realize a single circuit board as a small-scale basic model. On this basis, dynamic routing of model parameters between multiple circuit boards within a single computing system can constitute a medium-scale model. By deploying a medium-sized model in this way, the number of expert models within the computing system, the resources that the experts can dispatch, and the data flow and gating network between expert models can be optimized. Furthermore, in embodiments of the present application, a hierarchical communication strategy can be adopted, deploying frequently interacting expert groups on the same circuit board or computing system, utilizing NVLink communication technology or PCIE Fabric interconnect technology to increase communication speed.

[0107] Optionally, the first control device is electrically connected to any one of the plurality of circuit boards via a plurality of buses, wherein the plurality of buses have different data transmission speeds.

[0108] The first control device can collect the call frequency information of different types of expert models from any circuit board via multiple buses. The type of the expert model represents the historical call frequency of the expert model. For example, the type of expert model can be divided according to the historical call frequency and call frequency range of the expert. For example, the first expert model, the second expert model, and the third expert model described above are different types of expert models. In an embodiment of the present application, the call frequency information of an expert model with a high call frequency (such as the call frequency information of the first expert model) and the call frequency information of an expert model with a low call frequency (such as the call frequency information of the second expert model) can be collected from any circuit board via multiple buses. The collected call frequency of the expert model can be sent to the terminal device so that the circuit board can be scheduled to call the expert model to perform the model processing task in a timely manner according to the number of calls of different types of expert models.

[0109] Figure 8A schematic diagram of the connection between the first control device and the circuit board according to an embodiment of the present application is shown.

[0110] like Figure 8 As shown, the multiple buses include a first bus L1 and a second bus L2. The data transmission speed of the first bus L1 is higher than that of the second bus L2. The monitoring unit MON of the circuit board CD can collect the call frequency information of the expert model of the circuit board CD. The first control device CTD1 may include an acquisition unit ACQ, which can be electrically connected to the monitoring unit MON via the first bus L1 and the second bus L2, and collect first call frequency information of the first type of expert model from the monitoring unit MON via the first bus L1, and collect second call frequency information of the second type of expert model from the monitoring unit MON via the second bus L2. The first control device CTD1 may also include a device control unit ECT. The device control unit ECT can send the first call frequency information and the second call frequency information to the terminal device TE.

[0111] The acquisition unit ACQ may include a first acquisition subunit and a second acquisition subunit. The first bus L1 may be the aforementioned USB bus. The monitoring unit MON may be electrically connected to the first acquisition subunit via the first bus L1, and the first acquisition subunit may be connected to a USB hub (USB hub). The USB bus allows for rapid data transmission to the USB hub for real-time processing, after which the USB hub transmits the collected information to the device control unit ECT. The second bus L2 may be electrically connected to the second acquisition subunit. The second bus L2 may be the aforementioned LTPI bus. The FPGA may collect model parameters with medium call frequencies from the monitoring unit MON via the second bus L2. For model parameter call frequency information with low call frequencies, the FPGA may collect this information via I2C signals or CPIO signals.

[0112] Optionally, the first control device CTD1 further includes a system switching unit SSW. For example, the system switching unit can be implemented based on a PCIE switch unit. The system switching unit SSW is electrically connected to the switching units SW of each of the multiple circuit boards CD. The device control unit ECT is electrically connected to the system switching unit SSW and is configured to send a target scheduling policy to the system switching unit SSW, so that the switching units of the circuit boards CD receive the target scheduling policy via the system switching unit, and thus write the target scheduling policy to the control unit via the switching unit. For example, the switching unit of the first circuit board of the first circuit board receives the target scheduling policy via the system switching unit, and thus writes the target scheduling policy to the first control unit via the switching unit of the first circuit board.

[0113] Figure 9A schematic diagram of a computer cabinet according to a first embodiment of the present application is shown.

[0114] like Figure 9 As shown, the computer cabinet of this embodiment includes multiple computing systems CS and a second control device CTD2. The computing system CS here can be any of the computing systems mentioned above.

[0115] The second control device CTD2 may receive call frequency information from the first control devices of the plurality of computing systems CS and may determine a scheduling strategy based on the received call frequency information. The specific method is similar to that described above and will not be repeated here. For example, the computer cabinet may correspond to a server cabinet.

[0116] Figure 10 FIG2 shows a schematic diagram of a computer cabinet according to a second embodiment of the present application. Figure 11 A connection diagram of a computer cabinet according to an embodiment of the present application is shown.

[0117] like Figure 10 and Figure 11 As shown, the computer cabinet also includes a collection device CLD. The collection device CLD can be a remote power distribution unit (RPDU). The collection device CLD can be electrically connected to the circuit boards of multiple computing systems CS and collect power consumption information of the circuit boards. In one embodiment of the present application, the collection device CLD can send the power consumption information directly to the terminal device. In another embodiment, the collection device CLD can be electrically connected to the second control device CTD2 and send the power consumption information to the terminal device via the second control device CTD2. After receiving the power consumption information, the terminal device can determine a scheduling strategy based on the power consumption information and send the scheduling strategy to the second control device CTD2. The second control device CTD2 can send the scheduling strategy to the circuit board in the computing system CS, thereby scheduling the resources of the circuit board.

[0118] Optionally, the second control device CTD2 can be an industrial-grade CPU based on the edge series. This CPU features a wide operating temperature range and fast response. The second control device CTD2 can include a system control unit SCT, a first fabric switching unit SYC1, and a second fabric switching unit SYC2. The first fabric switching unit SYC1 can be implemented as a PCIE switch. The second fabric switching unit SYC2 can be implemented as an Ethernet switch (ETH SW). The switching units SW of the circuit boards of multiple computing systems CS can be connected to the first fabric switching unit SYC1, and data can be transmitted via the first fabric switching unit SYC1. For example, the switching units SW can also be connected to the first fabric switching unit SYC1 via the system switching unit. The communication units of the circuit boards of multiple computing systems CS can be connected to the second fabric switching unit SYC2, and data can be transmitted via the second fabric switching unit SYC2. In one embodiment of the present application, the system control unit SCT can receive the scheduling policy corresponding to the power consumption information from a terminal device and send it to the switching unit SW of the circuit board via the first fabric switching unit SYC1, so that the scheduling policy can be written to the control unit of the circuit board via the switching unit SW. In another embodiment of the present application, the system control unit SCT can receive the above-mentioned scheduling strategy corresponding to the power consumption information from the terminal device, and send it to the communication unit INT of the circuit board via the second architecture switching unit SYC2, so as to write the scheduling strategy into the control unit of the circuit board via the communication unit INT and the switching unit SW connected to the communication unit INT.

[0119] Optionally, the multiple computing systems include a first computing system and a second computing system. The first computing system includes a circuit board of the first computing system. The second computing system includes a circuit board of the second computing system. The acquisition device CLD is electrically connected to the circuit board of the first computing system and the circuit board of the second computing system, and can collect power consumption information of the circuit board of the first computing system and the circuit board of the second computing system. Then, the acquisition device CLD can send the power consumption information to the terminal device, and send the scheduling policy corresponding to the power consumption information received from the terminal device to the circuit board of the first computing system. The circuit board of the first computing system can send the model parameters of the circuit board of the first computing system to the circuit board of the second computing system according to the scheduling policy corresponding to the power consumption information. The circuit board of the second computing system can call the expert model of the circuit board of the first computing system based on the model parameters of the circuit board of the first computing system to perform the model processing task, and obtain the model processing results of the circuit board of the first computing system.

[0120] Multiple architectures can also be connected via the first architecture switching unit SYC1 and the second architecture switching unit SYC2 described above. This provides physical conditions for network clustering between multiple chassis within the same architecture, as well as between multiple architectures. Furthermore, RDMA technology can be used to enable mutual access between multiple first control devices, ensuring efficient all-to-all communication and reducing latency in gradient synchronization between experts. Furthermore, this allows for PCIE networking between multiple chassis within a cabinet, making it more applicable in low-latency scenarios. The second control device CTD2 can also include a rack management controller (RMC). The RMC can oversee and manage the entire architecture. For example, the RMC can identify computing systems (CSs) within the architecture with high power consumption. These computing systems can be systems with high memory usage or high expert model call frequency in accelerator card units. The RMC can send power consumption information to terminal devices, receive feedback on scheduling policies, and add these scheduling policies to the image file loaded by the RMC. The RMC can transmit power consumption information via the RJ45 network interface. Resources of high-power computing systems can be prioritized, and redundancy can be provided using other computing systems to ensure stable operation. Multiple server chassis can be placed within a single cabinet, and two types of switches, the second fabric switching unit SYC1 and the second fabric switching unit SYC2, can be used to create multiple ports to exchange resources between chassis and nodes. For the MOE model, localized communication can be used for optimization, deploying frequently interacting expert groups within the same rack. Technologies such as NVLINK can be used for optimization, and asynchronous updates can be implemented for global parameters of non-critical paths (such as the gating network) to reduce communication congestion.

[0121] Figure 12 A schematic diagram illustrating connections between a circuit board of a first computing system and a circuit board of a second computing system according to an embodiment of the present application is shown.

[0122] like Figure 12 As shown, the circuit board TCD1 of the first computing system includes a control unit TCTRL1 of the circuit board of the first computing system and a communication unit TINT1 of the circuit board of the first computing system electrically connected via the switching unit TSW1 of the circuit board of the first computing system, and the circuit board TCD2 of the second computing system includes a communication unit TINT2 of the circuit board of the second computing system and an extended storage unit TES2 of the circuit board of the second computing system electrically connected via the switching unit TSW2 of the circuit board of the second computing system.

[0123] The control unit TCTRL1 of the circuit board of the first computing system can send the first target scheduling instruction, the first target control instruction and the model parameters of the circuit board TCD1 of the first computing system to the switching unit TSW1 of the circuit board of the first computing system according to the scheduling strategy corresponding to the power consumption information.

[0124] In one embodiment of the present application, the switching unit TSW1 of the circuit board of the first computing system can, under the control of the second target scheduling instruction, send the model parameters of the circuit board TCD1 of the first computing system to the communication unit TINT1 of the circuit board of the first computing system based on the scheduling strategy corresponding to the power consumption information, so as to send the model parameters of the circuit board TCD1 of the first computing system to the communication unit TINT2 of the circuit board of the second computing system via the communication unit TINT1 of the circuit board of the first computing system, so as to write the model parameters of the circuit board TCD1 of the first computing system received via the communication unit TINT2 of the circuit board of the second computing system into the extended storage unit TES2 of the circuit board of the second computing system.

[0125] The circuit board TCD2 of the second computing system may further include a switching unit TSW2 of the circuit board of the second computing system. The communication unit TINT2 of the circuit board of the second computing system may write the model parameters of the circuit board TCD1 of the first computing system into the extended storage unit TES2 of the circuit board of the second computing system via the switching unit TSW2 of the circuit board of the second computing system.

[0126] The embodiments of the present application are not limited to this. The circuit board TCD2 of the second computing system may also include an accelerator card unit of the circuit board of the second computing system. The accelerator card unit of the circuit board of the second computing system may be electrically connected to the switching unit TSW2 of the circuit board of the second computing system. Thus, in another embodiment of the present application, the communication unit TINT2 of the circuit board of the second computing system may also write the model parameters of the circuit board TCD1 of the first computing system into the memory of the accelerator card unit of the circuit board of the second computing system.

[0127] In the embodiments of the present application, multiple computer cabinets can be deployed, and the second control devices of multiple computer cabinets can be interconnected to enable data exchange between the multiple computer cabinets. The high-speed network card signal of the communication unit can be transmitted electrically within the cabinet, and can also be transmitted optically in long-distance transmission scenarios across cabinets. Using a network group cluster method, it is possible to achieve interconnection of hundreds, thousands, or even tens of thousands of cards across chassis and cabinets, and in data nodes.

[0128] Figure 13 A schematic diagram of a calculation method according to an embodiment of the present application is shown.

[0129] like Figure 13As shown, the calculation method of this embodiment can be applied to a circuit board, and the calculation method can include: operations S1310 to S1330.

[0130] In operation S1310 , the acceleration card unit of the circuit board obtains first model parameters from a memory of the acceleration card unit of the circuit board, and / or obtains second model parameters from an extended storage unit of the circuit board via a switching unit of the circuit board.

[0131] In operation S1320, the accelerator card unit of the circuit board calls at least one of the first expert model and the second expert model to perform a model processing task based on the first model parameter and / or the second model parameter to obtain a model processing result;

[0132] In operation S1330 , the monitoring unit of the circuit board transmits information on the frequency of calling different expert models by the acceleration card unit of the circuit board to the control device at different transmission speeds.

[0133] In the embodiment of the present application, operations S1310 to S1330 are similar to the operations performed by the circuit board described above.

[0134] For example, the monitoring unit of the circuit board transmits the call frequency information of the acceleration card unit of the circuit board to different expert models to the control device at different transmission speeds, which may include: the monitoring unit selects multiple buses based on the storage locations of the model parameters of different expert models, and thus transmits different call frequency information to the control device via multiple buses, so that the control device determines the scheduling strategy of the circuit board based on the received call frequency information.

[0135] For example, in a first model processing flow of the accelerator card unit, the memory is used to store first model parameters, and the extended storage unit is used to store second model parameters. The control device determining a scheduling strategy for the circuit board based on the received call frequency information may include: the control device determining a second model processing flow following the first model processing flow based on the received call frequency information, and determining the scheduling strategy based on the second model processing flow, wherein at least one of the first model parameter and the second model parameter is stored in a different location in the second model processing flow than in the first model processing flow.

[0136] For example, the calculation method may also include: the control unit of the circuit board receives a scheduling strategy from the control device; under the control of the scheduling strategy, the control unit sends the first model parameter to the extended storage unit via the switching unit to write the first model parameter into the extended storage unit to change the storage location of the first model parameter; and / or, sends the second model parameter to the accelerator card unit via the switching unit to write the second model parameter into the memory to change the storage location of the second model parameter; wherein, after the storage location is changed, the transmission speed of the call frequency information of the first expert model and the second expert model respectively matches the second model processing flow.

[0137] For example, the multiple buses include a first bus and a second bus; transmitting different call frequency information to the control device via the multiple buses may include: the monitoring unit sends the call frequency information of the model parameters stored in the memory to the control device via the first bus at a first transmission speed; and / or sends the call frequency information of the model parameters stored in the extended storage unit to the control device via the second bus at a second transmission speed; wherein the first transmission speed is higher than the second transmission speed.

[0138] For example, the above-mentioned calculation method may also include the following steps: when the accelerator card unit of the circuit board performs a model processing task based on one of the first model parameter and the second model parameter: when the accelerator card unit of the circuit board performs a model processing task based on one of the first model parameter and the second model parameter: the accelerator card unit of the circuit board sends the other model parameter of the first model parameter and the second model parameter to other circuit boards to write the other model parameter to the other circuit boards, so that the other circuit boards perform the model processing task based on the other model parameter. It should be understood that the calculation method of the present application is not limited to this and will not be described in detail here.

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0140] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

[0141] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A circuit board, characterized in that: include: exchange unit; an accelerator card unit connected to the switching unit, wherein a first memory of the accelerator card unit is used to store first model parameters of a first expert model; an extended storage unit connected to the exchange unit, wherein the extended storage unit is used to store a second model parameter of the second expert model; the calling frequency of the first expert model is higher than the calling frequency of the second expert model; A monitoring unit is connected to the acceleration card unit and is connected to the first control device via multiple buses; the data transmission speeds of the multiple buses are different, so that the monitoring unit can transmit the call frequency information of the acceleration card unit calling different expert models to the first control device at different transmission speeds.

2. The circuit board according to claim 1, wherein: The acceleration card unit is used to: obtain the first model parameters from the first memory to call the first expert model to perform the model processing task; and / or obtain the second model parameters from the extended storage unit via the exchange unit to call the second expert model to perform the model processing task.

3. The circuit board according to claim 2, wherein: The monitoring unit is further configured to collect information on different calling frequencies of the accelerator card unit for different expert models, and transmit the information on different calling frequencies to the first control device via the multiple buses.

4. The circuit board according to claim 3, wherein: The monitoring unit is also used to select the multiple buses respectively based on the storage locations of the model parameters of the different expert models, so as to transmit the different call frequency information to the first control device respectively via the multiple buses, so that the first control device determines the scheduling strategy of the circuit board based on the received call frequency information.

5. The circuit board according to claim 4, characterized in that In the first model processing flow of the acceleration card unit, the first memory is used to store the first model parameters, and the extended storage unit is used to store the second model parameters; The monitoring unit is further configured to send the different call frequency information to the first control device via the multiple buses, so that the first control device determines a second model processing flow following the first model processing flow based on the received call frequency information, and determines the scheduling strategy based on the second model processing flow; The storage location of at least one of the first model parameter and the second model parameter in the second model processing flow is different from the storage location in the first model processing flow.

6. The circuit board according to claim 5, characterized in that Also includes: A control unit, connected to the switching unit, configured to: receiving the scheduling policy from the first control device; Under the control of the scheduling policy, sending the first model parameter to the extended storage unit via the switching unit to write the first model parameter into the extended storage unit to change the storage location of the first model parameter; and / or, sending the second model parameters to the accelerator card unit via the switching unit, so as to write the second model parameters into the first memory, thereby changing the storage location of the second model parameters; After the storage location is changed, the transmission speed of the call frequency information of the first expert model and the second expert model respectively matches the processing flow of the second model.

7. The circuit board according to claim 6, wherein: The control unit is further configured to: After receiving the scheduling policy, the initial scheduling policy deployed by itself is updated to the scheduling policy.

8. The circuit board according to claim 7, wherein: The control unit is further configured to: sending an initial scheduling instruction, the first model parameter, and the second model parameter to the switching unit according to the initial scheduling strategy; The switching unit is configured to: sending the first model parameters to the accelerator card unit to write the first model parameters into a first memory of the accelerator card unit; The second model parameters are sent to the extended storage unit to write the second model parameters into the extended storage unit.

9. The circuit board according to any one of claims 6 to 8, characterized in that: The second memory of the control unit is used to store third model parameters of the third expert model; The accelerator card unit is also used for: Under the control of a control instruction received from the control unit via the switching unit, acquiring the third model parameter from the second memory; and Based on the third model parameters, the third expert model is called to perform a model processing task.

10. The circuit board according to claim 9, wherein: The calling frequency of the first expert model is higher than the calling frequency of the third expert model, and the calling frequency of the third expert model is higher than the calling frequency of the second expert model.

11. The circuit board according to any one of claims 3 to 8, wherein: The plurality of buses include a first bus and a second bus; the monitoring unit is connected to the first control device via the first bus and the second bus; The monitoring unit is further configured to: sending, via the first bus, the call frequency information of the model parameters stored in the first memory to the first control device at a first transmission speed; and / or sending, via the second bus, the call frequency information of the model parameters stored in the extended storage unit to the first control device at a second transmission speed; The first transmission speed is higher than the second transmission speed.

12. A computing system comprising: A plurality of circuit boards according to any one of claims 1 to 11; as well as The first control device is used to collect call frequency information of the expert models configured in the multiple circuit boards, and determine the scheduling strategies of the multiple circuit boards based on the received call frequency information.

13. The computing system according to claim 12, wherein: The plurality of circuit boards include a first circuit board and a second circuit board, wherein the first circuit board is used to store target model parameters; the model parameters stored in the first circuit board and the model parameters stored in the second circuit board are respectively used to perform different types of model processing tasks; The first circuit board is electrically connected to the second circuit board, and is configured to send the target model parameters to the second circuit board; The second circuit board is used to call the target expert model to execute the target type model processing task based on the target model parameters.

14. The computing system according to claim 13, wherein: The first control device is further configured to: Sending target call frequency information of the target expert model from the first circuit board to a terminal device; In response to receiving a target scheduling policy corresponding to the target call frequency information from the terminal device, the target scheduling policy is sent to the first circuit board to control the first circuit board to send the target model parameters to the second circuit board.

15. The computing system according to claim 14, wherein: The first control device further includes a device control unit and a system switching unit, wherein the system switching unit is electrically connected to the respective switching units of the plurality of circuit boards; In which, the device control unit is electrically connected to the system switching unit, and is used to send the target scheduling policy to the system switching unit, so that the switching unit of the first circuit board receives the target scheduling policy via the system switching unit, and thus writes the target scheduling policy into the first circuit board via the switching unit of the first circuit board.

16. The computing system according to any one of claims 14 to 15, characterized in that: The accelerator card unit of the first circuit board is used to send the target model parameters to the accelerator card unit of the second circuit board, so as to write the target model parameters into the first memory of the accelerator card unit of the second circuit board.

17. The computing system according to claim 16, wherein: The target scheduling strategy includes a first target scheduling strategy; The switching unit of the first circuit board is configured to send the target model parameters to the accelerator card unit of the first circuit board based on the first target scheduling strategy.

18. The computing system according to any one of claims 14 to 15, characterized in that: The target scheduling strategy includes a second target scheduling strategy; The switching unit of the first circuit board is used to send the target model parameters to the switching unit of the second circuit board based on the second target scheduling strategy, so as to send the target model parameters to the extended storage unit of the second circuit board via the switching unit of the second circuit board, so as to write the target model parameters into the extended storage unit of the second circuit board.

19. The computing system according to any one of claims 13 to 15, wherein: The first circuit board is adjacent to the second circuit board, and the number of times the expert model of the first circuit board is called is higher than the number of times the expert model of the second circuit board is called.

20. The computing system according to any one of claims 13 to 15, wherein: The first control device includes: a collection unit electrically connected to the monitoring unit of any circuit board via a first bus and a second bus, configured to collect first call frequency information from the monitoring unit via the first bus, and collect second call frequency information from the monitoring unit via the second bus; A device control unit is configured to determine a scheduling strategy for any circuit board based on the first call frequency information and the second call frequency information.

21. A computer cabinet, characterized in that: include: A plurality of computing systems according to any one of claims 12 to 20; as well as The second control device is configured to receive call frequency information from the first control devices of the plurality of computing systems, and determine scheduling strategies for the plurality of computing systems based on the received call frequency information.

22. The computer cabinet according to claim 21, wherein: The plurality of computing systems includes a first computing system and a second computing system; The computer cabinet further includes an acquisition device electrically connected to the circuit board of the first computing system and the circuit board of the second computing system, for: collecting power consumption information of a circuit board of the first computing system and a circuit board of the second computing system; Sending the power consumption information to the terminal device; sending the scheduling policy corresponding to the power consumption information received from the terminal device to a circuit board of the first computing system; the circuit board of the first computing system being configured to send the model parameters of the circuit board of the first computing system to the circuit board of the second computing system according to a scheduling policy corresponding to the power consumption information; The circuit board of the second computing system is used to call the expert model of the circuit board of the first computing system to perform a model processing task based on the model parameters of the circuit board of the first computing system.

23. The computer cabinet according to claim 22, wherein: The circuit board of the first computing system includes a switching unit of the circuit board of the first computing system and a communication unit of the circuit board of the first computing system electrically connected, and the circuit board of the second computing system includes a communication unit of the circuit board of the second computing system and an expansion storage unit of the circuit board of the second computing system electrically connected via the switching unit of the circuit board of the second computing system; The switching unit of the circuit board of the first computing system is configured to send the model parameters of the circuit board of the first computing system to the communication unit of the circuit board of the first computing system based on a scheduling strategy corresponding to the power consumption information, so as to send the model parameters of the circuit board of the first computing system to the communication unit of the circuit board of the second computing system via the communication unit of the circuit board of the first computing system, so as to write the model parameters of the circuit board of the first computing system into the extended storage unit of the circuit board of the second computing system via the communication unit of the circuit board of the second computing system.

24. A calculation method, characterized in that Applicable to a circuit board as claimed in any one of claims 1 to 11, The method comprises: The accelerator card unit of the circuit board obtains the first model parameter from the memory of the accelerator card unit of the circuit board, and / or obtains the second model parameter from the extended storage unit of the circuit board via the switching unit of the circuit board; The accelerator card unit of the circuit board calls at least one of a first expert model and a second expert model to perform a model processing task based on the first model parameter and / or the second model parameter, wherein the calling frequency of the first expert model is higher than the calling frequency of the second expert model; The monitoring unit of the circuit board transmits the calling frequency information of calling different expert models by the acceleration card unit of the circuit board to the control device at different transmission speeds.

25. The calculation method according to claim 24, characterized in that: The monitoring unit of the circuit board transmits the call frequency information of the acceleration card unit of the circuit board calling different expert models to the control device at different transmission speeds, including: The monitoring unit of the circuit board selects multiple buses based on the storage locations of the model parameters of the different expert models, and transmits different call frequency information to the control device via the multiple buses, so that the control device determines the scheduling strategy of the circuit board based on the received call frequency information.

26. The calculation method according to claim 25, characterized in that In the first model processing flow of the acceleration card unit, the memory is used to store the first model parameters, and the extended storage unit is used to store the second model parameters; The control device determines a scheduling strategy for the circuit board based on the received call frequency information, including: The control device determines, based on the received call frequency information, a second model processing flow following the first model processing flow, and determines the scheduling strategy based on the second model processing flow; The storage location of at least one of the first model parameter and the second model parameter in the second model processing flow is different from the storage location in the first model processing flow.

27. The calculation method according to claim 26, characterized in that: The method further comprises: The control unit of the circuit board receives the scheduling strategy from the control device; Under the control of the scheduling policy, the control unit of the circuit board sends the first model parameter to the extended storage unit via the switching unit to write the first model parameter into the extended storage unit, thereby changing the storage location of the first model parameter; and / or sends the second model parameter to the accelerator card unit via the switching unit to write the second model parameter into the memory, thereby changing the storage location of the second model parameter; After the storage location is changed, the transmission speed of the call frequency information of the first expert model and the second expert model respectively matches the processing flow of the second model.

28. The calculation method according to any one of claims 25 to 27, characterized in that: The plurality of buses include a first bus and a second bus; The transmitting the different call frequency information to the control device via the multiple buses respectively includes: The monitoring unit of the circuit board sends the call frequency information of the model parameters stored in the memory to the control device via the first bus at a first transmission speed; and / or sends the call frequency information of the model parameters stored in the extended storage unit to the control device via the second bus at a second transmission speed; The first transmission speed is higher than the second transmission speed.

29. The calculation method according to any one of claims 25 to 27, characterized in that: The method also includes, when the accelerator card unit of the circuit board performs a model processing task based on one of the first model parameter and the second model parameter: the accelerator card unit of the circuit board sends the other model parameter of the first model parameter and the second model parameter to other circuit boards to write the other model parameter to the other circuit boards, so that the other circuit boards perform the model processing task based on the other model parameter; wherein the number of calls to the expert model of the circuit board is higher than the number of calls to the expert models of the other circuit boards.

Citation Information

Patent Citations

  • Single GPU (Graphics Processing Unit) chip architecture system based on multiple types of extended memories

    CN118247120A

  • Method and apparatus for DMA between accelerator cards, and accelerator card, acceleration platform and medium

    WO2025073257A1