Processor, processor communication method and computer equipment
By integrating the IB network card component into the GPU and adopting the NVLink protocol, the problems of long communication paths and high power consumption between AI server nodes are solved, and efficient and low-latency communication is achieved.
Patent Information
- Application Number
- CN202511179979.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-30
AI Technical Summary
In existing technologies, GPU communication between AI server nodes relies on motherboard routing and PCIe protocols, resulting in long communication paths, high power consumption, large delays, and a shortage of PCIe bandwidth resources, affecting communication efficiency.
The IB controller, physical interface, link controller, etc. of the IB network card are integrated into the I/O module of the GPU. The communication module and the GPU module are directly connected through the CoWoS-S chip packaging technology, and the NVLink protocol is used for communication to avoid dependence on the PCIe bus and CPU.
It significantly improves communication efficiency, reduces communication latency, alleviates PCIe bandwidth competition pressure, reduces power consumption, and improves system energy efficiency.
Smart Images

Figure CN120723682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chip design, and in particular to a processor, a processor communication method, and a computer device. Background Art
[0002] Currently, the number of parameters in large models is increasing, the scale of distributed systems used to train large models is growing, and the number of computing nodes is increasing. This makes the communication between GPUs (Graphic Process Units) in artificial intelligence server nodes increasingly complex, and the requirements for communication throughput are becoming increasingly higher. At the same time, the communication process requires low communication latency.
[0003] Currently, communication between GPUs in AI server nodes primarily relies on independent InfiniBand (IB) network cards on the motherboard based on the PCIe (peripheral component interconnect express) protocol. Data flows through the GPU's internal interface, the PCIe controller, the motherboard wiring, and the IB network card. The use of motherboard wiring results in a longer communication path, increasing power consumption. Furthermore, the PCIe protocol itself requires serial and parallel conversion of data transmission formats, resulting in significant communication delays. Furthermore, the communication process competes for limited PCIe bandwidth resources, impacting other communication processes. Summary of the Invention
[0004] In view of this, the present invention provides a processor, a processor communication method and a computer device to solve the problem that the communication process transmits data through the motherboard wiring, resulting in a long communication path, and the PCIe protocol itself requires serial and parallel conversion of the data transmission format, which increases communication power consumption and communication delay.
[0005] In a first aspect, the present application provides a processor, the processor comprising: a computing module and a communication module; The computing module is connected to the communication module via a first communication link. The computing module is configured to, when the processor has a data transmission demand, obtain a first data packet to be transmitted based on a first communication protocol, and transmit the first data packet to be transmitted to the communication module via the first communication link. The first communication protocol is a communication protocol corresponding to the first communication link. The communication module is connected to other processors through a second communication link. The communication module is used to perform protocol conversion on the first data packet to be transmitted, obtain a second data packet to be transmitted based on the second communication protocol, and transmit the second data packet to be transmitted to other processors through the second communication link, wherein the second communication protocol is the communication protocol corresponding to the second communication link.
[0006] In a second aspect, the present application provides a processor communication method, the method comprising: When the processor has a data transmission demand, obtaining a first data packet to be transmitted based on a first communication protocol, wherein the first communication protocol is a communication protocol corresponding to the first communication link; Performing protocol conversion on the first data packet to be transmitted to obtain a second data packet to be transmitted based on a second communication protocol, wherein the second communication protocol is a communication protocol corresponding to the second communication link; The second data packet to be transmitted is transmitted to other processors through the second communication link.
[0007] In a third aspect, the present application provides a processor communication device, the device comprising: a data packet acquisition module, configured to acquire, when the processor has a data transmission demand, a first data packet to be transmitted based on a first communication protocol, wherein the first communication protocol is a communication protocol corresponding to the first communication link; a protocol conversion module, configured to perform protocol conversion on the first data packet to be transmitted to obtain a second data packet to be transmitted based on a second communication protocol, wherein the second communication protocol is a communication protocol corresponding to the second communication link; The data packet transmission module is used to transmit the second data packet to be transmitted to other processors through the second communication link.
[0008] In a fourth aspect, the present application provides a computer device comprising: a storage component and a processing component, the storage component and the processing component being communicatively connected to each other, the storage component storing computer instructions, and the processing component executing the computer instructions to thereby execute the processor communication method of the above-mentioned second aspect or any corresponding embodiment thereof.
[0009] In a fifth aspect, the present application provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the processor communication method of the above-mentioned second aspect or any corresponding embodiment thereof.
[0010] In a sixth aspect, the present application provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the processor communication method of the above-mentioned second aspect or any corresponding embodiment thereof.
[0011] Through this application, the communication module is encapsulated into the processor, and the data of the computing module is transmitted to other processors through the communication module. In addition, the communication module can convert the first data packet to be transmitted from the first communication protocol to the second communication protocol, and transmit the second data packet to be transmitted after the protocol conversion to other processors. This can solve the problem that the communication process transmits data through the motherboard wiring, resulting in a long communication path, and the PCIe protocol itself requires serial and parallel conversion of the data transmission format, which increases communication power consumption and communication delay. The processor is encapsulated with a communication module, and GPU communication between nodes is carried out through the communication module, which no longer relies on the PCIe bus and the central processing unit, alleviates the pressure of PCIe bandwidth competition and improves system energy efficiency, significantly improves communication efficiency and reduces communication delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the specific implementation methods of this application or the technical solutions in related technologies, the following is a brief introduction to the drawings required for use in the specific implementation methods or related technical descriptions. Obviously, the drawings described below are some implementation methods of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 is a schematic diagram of a traditional inter-node GPU communication method according to an embodiment of the present application; Figure 2 is a structural diagram of a processor according to an embodiment of the present application; Figure 3 is a schematic diagram of a GPU chip package according to an embodiment of the present application; Figure 4 is a schematic diagram of data packet protocol conversion according to an embodiment of the present application; Figure 5 is a schematic diagram of an inter-node GPU communication method according to an embodiment of the present application; Figure 6 is a flowchart of a processor communication method according to an embodiment of the present application; Figure 7 is a structural block diagram of a processor communication device according to an embodiment of the present application; Figure 8 It is a schematic diagram of the hardware structure of the computer device of an embodiment of the present application. DETAILED DESCRIPTION
[0014] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0015] A large language model (LLM) is an artificial intelligence (AI) algorithm model with over 100 million parameters. It can handle a wide range of tasks, such as answering questions, generating text, understanding language meaning, and performing translation. The increasing number of parameters in large models makes model training increasingly difficult, and the distributed systems used to train them are becoming increasingly larger. The number of compute nodes in distributed systems is also increasing. A compute node refers to each computing unit in the distributed system used to train large models, such as an AI server (an AI server equipped with a GPU card or GPU module capable of AI training). This increasing number of compute nodes complicates communication between GPUs in different compute nodes, placing increasing demands on communication throughput and decreasing latency tolerance.
[0016] Currently, GPU communication within AI server nodes is mainly achieved through the NVLink protocol, which is a communication protocol between GPUs. GPU communication between AI server nodes is mainly achieved through IB network cards based on the PCIe protocol, such as Figure 1As shown in the figure, nodes 1 and 2 are AI server nodes. If nodes 1 and 2 communicate using traditional inter-node GPU communication, the data transmission path is: compute module A in node 1, GPU cache (L1 cache, L2 cache, L3 cache), memory controller, video memory, GPU PCIe bus, central processing unit (CPU), IB network card PCIe bus, IB network card controller, physical (PHY) interface, optical module, optical module in node 2, physical interface, IB network card controller, IB network card PCIe bus, video memory, memory controller, GPU cache, and compute module B. An example of an optical module is QSFP (Quad Small Form-factor Pluggable), a small form-factor pluggable optical module interface based on a four-lane architecture. This communication method relies on PCIe bandwidth allocation and CPU process scheduling, making software design complex. Furthermore, the communication process requires long motherboard traces and conversion between PCIe and IB protocols, resulting in long data latency. Furthermore, the PCIe bus bandwidth resources of current AI servers are limited. For example, some AI servers have 256 PCIe 5.0 channels, with a total bidirectional bandwidth of 256 × 8 GB / s = 2048 GB / s. Table 1 shows the PCIe bandwidth usage of different components in an AI server. The AI server includes a PCIe switch (PCIe switching chip), eight GPUs, 16 solid-state drives, and an IB network card. The IB network card is used for GPU communication between AI server nodes. The total PCIe bandwidth occupied by these components is 1920 GB / s, resulting in a PCIe bus utilization rate of up to 94%, severely impacting system stability and recoverability and making the AI server's circuit architecture virtually unscalable.
[0017] Table 1 AI server PCIe bandwidth allocation
[0018] Based on the above, GPU communication between AI server nodes is achieved through the PCIe bus, connected to an independent IB network card (or high-speed Ethernet card) on the motherboard. This method relies on the PCIe protocol for data communication. This communication method has the following drawbacks: data is transmitted through multiple stages, resulting in a long transmission path and requiring data conversion between the PCIe and IB protocols. The PCIe protocol itself requires serial and parallel data transmission formats, which significantly increases data transmission latency and affects data transmission efficiency. Furthermore, the PCIe controller requires data scheduling by the central processor, further increasing data transmission time. Furthermore, various PCIe devices share the PCIe bus, such as GPUs, IB network cards, and solid-state drives. Simultaneous communication among these devices competes for limited PCIe bandwidth resources, which are relatively limited. Using the PCIe bus for GPU communication further exacerbates this resource shortage. Finally, the PCBs (Printed Circuit Boards) of current AI servers typically have at least 20 layers. Independent IB network cards require a dedicated PCIe slot during design, increasing design complexity, increasing the number of connection lines, and increasing power consumption.
[0019] Based on the foregoing, embodiments of the present application provide a processor that integrates the IB controller, physical interface, and link controller of an IB network card (NIC) into the GPU's I / O (input / output) module as a communication module, eliminating the independent IB NIC's RDMA (remote direct memory access) controller, IB power module, and PCIe controller. This processor utilizes advanced CoWoS-S (Chip-on-Wafer-on-Substrate Silicon) packaging technology, an advanced chip packaging technology that uses silicon as an intermediary layer, to achieve direct connectivity between the communication module and other GPU modules within the same package. This processor utilizes a communication module for inter-node GPU communication, eliminating reliance on the PCIe bus and CPU, alleviating PCIe bandwidth competition and improving system energy efficiency. Furthermore, it utilizes the NVLink bus, which is faster and has lower latency than PCIe, for communication, significantly improving communication efficiency and reducing latency.
[0020] According to an embodiment of the present application, a processor is provided, such as Figure 2 As shown, the processor includes: a computing module and a communication module; The computing module is connected to the communication module via a first communication link. The computing module is configured to, when the processor has a data transmission demand, obtain a first data packet to be transmitted based on a first communication protocol, and transmit the first data packet to be transmitted to the communication module via the first communication link. The first communication protocol is a communication protocol corresponding to the first communication link. The communication module is connected to other processors through a second communication link. The communication module is used to perform protocol conversion on the first data packet to be transmitted, obtain a second data packet to be transmitted based on the second communication protocol, and transmit the second data packet to be transmitted to other processors through the second communication link, wherein the second communication protocol is the communication protocol corresponding to the second communication link.
[0021] Specifically, this embodiment integrates the IB controller, physical interface, link controller, etc. of an independent IB network card into the processor as a communication module, such as a graphics processor. The communication module in the processor is also an IB network card encapsulated in the processor. In addition, this embodiment deletes the RDMA controller, IB power module, and PCIe controller of the independent IB network card, and realizes a direct connection between the communication module and the computing module of the processor in the same package through the advanced CoWoS-S chip packaging technology. The packaging adopts a silicon interposer. The chip packaged by this packaging technology can provide an interconnection bandwidth between the GPU and the I / O chip greater than 200GB / s, and a communication delay of less than 100 picoseconds, which is suitable for various application scenarios such as artificial intelligence training and reasoning. Figure 2 As shown, the computing module in the processor is connected to the communication module via a first communication link. The first communication link is, for example: Figure 3 As shown, the first communication link includes a cache, a video memory controller, a video memory, a communication protocol bus between the graphics processor, and a silicon intermediate layer, and the cache includes an L1 cache, an L2 cache, and an L3 cache.
[0022] The computing module in the processor can communicate with the computing modules of processors in other nodes through the communication module. The communication process includes: when the processor has a data transmission demand, the computing module obtains a first data packet to be transmitted based on a first communication protocol, and transmits the first data packet to be transmitted to the communication module through a first communication link. The first communication protocol is a communication protocol corresponding to the first communication link, for example, the NVLink protocol.
[0023] The communication module is connected to other processors via a second communication link, such as Figure 2As shown, the second communication link is, for example, an optical module and an external network, and the optical module is, for example, an OSFP (Quad Small Form-factor Pluggable) optical module. The other processors have the same structure as the processor, including a computing module and a communication module, and the computing module and the communication module are both connected via a first communication link. The communication module of the processor is connected to the communication modules of other processors via a second communication link. The process of the communication module transmitting the first data packet to be transmitted to other processors includes: the communication module performs protocol conversion on the first data packet to be transmitted, obtains a second data packet to be transmitted based on a second communication protocol, and transmits the second data packet to be transmitted to the other processors via the second communication link. The second communication protocol is the communication protocol corresponding to the second communication link, for example, the IB protocol. After receiving the second data packet to be transmitted, the other processors will use their own communication modules to perform protocol conversion. The computing modules of the other processors can obtain data from the converted data packet and perform corresponding calculations. In addition, this embodiment can also update the GPU driver to cooperate with the above-mentioned communication method, such as removing the RDMA and PCIe protocol drivers.
[0024] The above method allows the processor to no longer rely on the PCIe bus when using the communication module for data transmission, saving the PCIe bandwidth occupied by the IB network card in Table 1 and reducing the PCIe bandwidth utilization rate to 87%, so that PCIe resources have a certain degree of redundancy and backup.
[0025] This embodiment provides a processor that encapsulates a communication module into the processor, transmits data from the computing module to other processors through the communication module, and the communication module can convert a first data packet to be transmitted from a first communication protocol to a second communication protocol, and transmit the second data packet to be transmitted after the protocol conversion to other processors. The processor is encapsulated with a communication module, and GPU communication between nodes is performed through the communication module, which no longer relies on the PCIe bus and the central processing unit, alleviates the pressure of PCIe bandwidth competition and improves system energy efficiency, significantly improves communication efficiency and reduces communication delay. It solves the problem that the communication process transmits data through the motherboard wiring, resulting in a long communication path, and the PCIe protocol itself requires serial and parallel conversion of data transmission formats, which increases communication power consumption and communication delay.
[0026] As an optional embodiment, the communication module is further configured to, when the processor has a data reception demand, receive a first data packet to be received based on the second communication protocol through a second communication link, perform protocol conversion on the first data packet to be received to obtain a second data packet to be received based on the first communication protocol, and send the second data packet to be received to the computing module through the first communication link; The calculation module is further used to obtain the data to be received according to the second data packet to be received, and to calculate the data to be received.
[0027] Specifically, other processors can send data packets to the processor, and the processor receives the data packets through the communication module, performs protocol conversion, and then transmits them to the computing module.
[0028] When a processor has a data receiving demand, it indicates that another processor has sent a first data packet to be received to the processor. The second communication protocol is the communication protocol corresponding to the second communication link, for example, the IB protocol. The communication module receives the first data packet to be received based on the second communication protocol through the second communication link, performs protocol conversion on the first data packet to be received, obtains the second data packet to be received based on the first communication protocol, and sends the second data packet to be received to the computing module through the first communication link. The first communication protocol is the communication protocol corresponding to the first communication link, for example, the NVLink protocol.
[0029] The calculation module parses the received second data packet to be received to obtain the data to be received, and calculates the data to be received. For example, the calculation module: Figure 3 As shown in the figure, the computing module includes a streaming multiprocessor (SM), a basic computing unit (SP), a tensor core (Tensor Core), a ray tracing core (RayTracing Core), etc. Among them, the basic computing unit is the smallest unit of the GPU that performs floating-point operations and integer operations; the streaming multiprocessor mainly performs process task scheduling; the tensor core is a unit designed for matrix operations, mainly used for matrix operations during model training and inference; the main function of the ray tracing core is to display three-dimensional rendering functions.
[0030] In this embodiment, the computing module obtains the data to be received from other processors through the communication module and performs data calculations, no longer relying on the PCIe bus and the central processing unit, alleviating the pressure of PCIe bandwidth competition and improving system energy efficiency. In addition, the communication module uses the NVLink protocol to communicate with the video memory through a silicon interposer with almost negligible impedance, significantly improving communication efficiency and reducing communication latency.
[0031] As an optional embodiment, the calculation module includes: a link training synchronization unit; A link training synchronization unit, configured to train the first communication link to obtain adjustment parameters; The link training synchronization unit is further configured to adjust the first communication link according to the adjustment parameter.
[0032] Specifically, if Figure 3As shown, the computing module includes a link training synchronization unit. This unit is responsible for training and establishing GPU-related communication links. Once the GPU communication link is established, it can communicate with the IB network card. In addition, the link training synchronization unit is used to maintain the stability and reliability of the communication data link.
[0033] The link training synchronization unit trains the first communication link to obtain adjustment parameters. For example, the link training synchronization unit automatically negotiates and determines optimal communication parameters, i.e., adjustment parameters, on the first communication link, including but not limited to signal amplitude, equalization settings, timing offset, and pre-emphasis / de-emphasis levels. The adjustment parameters enable the first communication link to adapt to different channel characteristics and environmental conditions and to maximize compensation for signal attenuation and jitter. The link training synchronization unit adjusts the first communication link based on the adjustment parameters.
[0034] The Link Training Synchronization Unit also guides the devices at both ends of the first communication link to complete the necessary handshake protocols and state machine transitions, ensuring that the sender and receiver reach agreement at the physical layer and successfully establish a reliable, bidirectional communication channel that meets specific rate requirements. After the link is established, the Link Training Synchronization Unit continuously monitors the signal quality (such as bit error rate and eye opening) and timing deviation of the first communication link. When performance degradation or environmental changes (such as temperature drift and voltage fluctuations) are detected, it dynamically triggers retraining or fine-tuning of synchronization parameters to maintain a low bit error rate, high reliability, and signal integrity for the link, ensuring stable and error-free high-speed data transmission between GPUs or between GPUs and video memory.
[0035] In addition, when the link training synchronization unit detects a serious error or interruption in the first communication link, it can coordinate and initiate a link reset and reconstruction process to try to restore communication.
[0036] In this embodiment, the link training synchronization unit is responsible for training and establishing the first communication link, and is also used to maintain the first communication link to ensure the stability and reliability of the data transmission process.
[0037] As an optional embodiment, the processor further includes: a power module; a power supply module, configured to provide a first power supply to the communication module upon receiving a first preset instruction; The power supply module is further configured to provide a second power supply to the communication module upon receiving a second preset instruction, wherein the power of the second power supply is greater than that of the first power supply.
[0038] Specifically, if Figure 3 As shown, the processor also includes a power module. Using the processor's power module to power the communication module can reduce circuit complexity, and the processor can also centrally schedule power distribution, making power distribution more efficient.
[0039] During the computational process, when the communication module is not required to transmit data, the processor sends a first predetermined instruction to the power module. For example, the first predetermined instruction instructs the power module to reduce the power provided to the communication module. Upon receiving the first predetermined instruction, the power module stops supplying the main power supply Vcc (second power supply) to the communication module and only provides the backup power supply Vsb (first power supply) to the communication module, thereby reducing the power consumption of the communication module by at least 70%. The second power supply has a greater power than the first power supply.
[0040] When the computing module completes data calculations and needs to transmit a data packet, the processor sends a second predetermined instruction to the power module. For example, the second predetermined instruction instructs the power module to increase the power provided to the communication module. The power module then provides the primary power source (Vcc), or secondary power source, to the communication module, allowing the communication module to operate normally and transmit data packets.
[0041] In addition, if Figure 3 As shown, the processor also includes a high-speed serial computer expansion bus standard bus controller, an NVME (Non Volatile Memory Host Controller Interface Specification) interface, and an SMBus (System Management Bus) module. The NVME interface is used by the GPU to manage hard drive access; the SMBus module is primarily used to check GPU temperature and fan status. The power module provides power to all modules; the high-speed serial computer expansion bus standard bus controller is used for communication between the GPU and the central processing unit. Optical modules such as OSFP are used to communicate with the processor via external networks, such as the network between AI server nodes.
[0042] In this embodiment, the power supply module of the communication module is deleted. Using the power supply module of the processor to power the communication module can reduce the complexity of the circuit, and the processor can also uniformly schedule power distribution, making power distribution more efficient.
[0043] As an optional embodiment, the communication module includes: a first protocol conversion unit; a first protocol conversion unit, configured to obtain data to be received from the first data packet to be received, obtain preset field information according to the first communication protocol, and obtain a first intermediate data packet based on the data to be received and the preset field information, wherein the preset field information is used to describe parameters that need to be referenced during transmission of the first data packet to be received; The first protocol conversion unit is further configured to determine first verification information of the first intermediate data packet, and obtain a second data packet to be received based on the first intermediate data packet and the first verification information.
[0044] Specifically, in this embodiment, it is necessary to convert the first data packet to be received based on the second communication protocol transmitted from other processors into the second data packet to be received based on the first communication protocol, so a first protocol conversion unit is added. The first protocol conversion unit is, for example: Figure 3 As shown, the first protocol conversion unit is a field programmable gate array (FPGA). The first communication protocol is the communication protocol corresponding to the first communication link, such as the NVLink protocol. The second communication protocol is the communication protocol corresponding to the second communication link, such as the IB protocol. Therefore, the first protocol conversion unit can be an FPGA programming module that converts between the NVLink protocol and the IB protocol.
[0045] Combine Figure 4 The protocol conversion process of the first protocol conversion unit is described. The format of the first data packet to be received is as follows: Figure 4 The first protocol conversion unit obtains the data to be received, i.e., the data in the data packet format of IB, from the first data packet to be received, and removes all other parts of the first data packet to be received. The first protocol conversion unit obtains preset field information according to the NVLink protocol, such as: the header field (header), the downstream header field (Downstream Header), the address extension field (Address Extension), and the byte enable field (ByteEnable) of the data packet, wherein the header field is used to define core information such as the data packet type, address, length, and transaction identifier; the downstream header field is used to guide the downstream forwarding path of the data packet in the multi-hop network; the address extension field is used to extend the address bit width to support a larger addressing space; the byte enable field is used to mark valid data bytes to support partial write operations. The above preset field information is used to describe the parameters that the first data packet to be received needs to refer to during the transmission process.
[0046] The first protocol conversion unit splices the above-mentioned data to be received and the preset field information to obtain a first intermediate data packet, and determines the first verification information of the first intermediate data packet. For example, the first verification information is a CPC (Cyclic Redundancy Check) checksum, and the CPC checksum of all data in the first intermediate data packet is calculated as the first verification information.
[0047] The first protocol conversion unit writes the first verification information into the first intermediate data packet to obtain Figure 4The second to-be-received data packet in the NVLink data packet format.
[0048] In this embodiment, the first protocol conversion unit obtains a first intermediate data packet based on the to-be-received data and preset field information, determines first verification information for the first intermediate data packet, and obtains a second to-be-received data packet based on the first intermediate data packet and the first verification information. The first protocol conversion unit performs protocol conversion on the to-be-received data packet to facilitate the calculation module to obtain the to-be-received data from the to-be-received data packet and perform calculations. Furthermore, the first verification information is added to the data packet to ensure the accuracy of the data in the data packet.
[0049] As an optional embodiment, the communication module includes: a second protocol conversion unit; a second protocol conversion unit, configured to obtain data to be transmitted from the first data packet to be transmitted, obtain a routing header and a transmission header according to the second communication protocol, and obtain a second intermediate data packet according to the data to be transmitted, the routing header, and the transmission header; The second protocol conversion unit is further configured to determine second verification information of the second intermediate data packet, and obtain a second data packet to be transmitted according to the second intermediate data packet and the second verification information.
[0050] Specifically, in this embodiment, it is necessary to convert the first data packet to be transmitted based on the first communication protocol transmitted from the computing module into a second data packet to be transmitted based on the second communication protocol, so a second protocol conversion unit is added. The second protocol conversion unit is, for example: Figure 3 As shown, the second protocol conversion unit is a field programmable gate array (FPGA). The first communication protocol is the communication protocol corresponding to the first communication link, such as the NVLink protocol. The second communication protocol is the communication protocol corresponding to the second communication link, such as the IB protocol. Therefore, the second protocol conversion unit can be an FPGA programming module for converting between the NVLink protocol and the IB protocol. The second protocol conversion unit can be the same protocol conversion unit as the first protocol conversion unit, such as Figure 3 shown.
[0051] Combine Figure 4 The protocol conversion process of the second protocol conversion unit is described. The format of the first data packet to be transmitted is as follows: Figure 4The second protocol conversion unit obtains the data to be transmitted, i.e., the data in the NVLink packet format, from the first data packet to be transmitted, and removes all other portions of the first data packet to be transmitted. The second protocol conversion unit obtains the routing header and transport header according to the IB protocol. Routing headers include, for example, the Local Routing Header (LRH) and the Global Routing Header (GRH). Transport headers include, for example, the Base Transport Header (BTH) and the Extended Transport Header (ETH). The Local Routing Header is used to determine intra-subnet routing, the Global Routing Header is used to determine inter-subnet global routing, the Base Transport Header is used for transmission control, reliability, and queue addressing, and the Extended Transport Header is used to extend transmission functionality.
[0052] The second protocol conversion unit splices the above-mentioned data to be transmitted, routing header and transmission header to obtain a second intermediate data packet, and determines the second verification information of the second intermediate data packet. For example, the second verification information is a CPC checksum, and the CPC checksum of all data in the second intermediate data packet is calculated as the second verification information.
[0053] The second protocol conversion unit writes the second verification information into the second intermediate data packet to obtain Figure 4 The second data packet to be transmitted in the IB data packet format.
[0054] In this embodiment, the second protocol conversion unit obtains a second intermediate data packet based on the data to be transmitted, the routing header, and the transmission header, determines second verification information for the second intermediate data packet, and obtains a second data packet to be transmitted based on the second intermediate data packet and the second verification information. The second protocol conversion unit performs protocol conversion on the data packet to be transmitted, facilitating other processors to obtain the data to be transmitted from the data packet to be transmitted and perform calculations. Furthermore, the second protocol conversion unit adds the first verification information to the data packet to ensure the accuracy of the data in the data packet.
[0055] As an optional embodiment, the communication module further includes: a communication control unit, a link control unit and a physical layer device; The communication control unit is connected to the second protocol conversion unit and is used to transmit the first data packet to be transmitted to the second protocol conversion unit; The communication control unit is connected to the link control unit and is used to obtain the transmission strategy of the second data packet to be transmitted from the link control unit when the second data packet to be transmitted is obtained; The communication control unit is connected to the physical layer device and is used to transmit the second data packet to be transmitted to the preset network through the physical layer device according to the transmission strategy, so that other processors can obtain the second data packet to be transmitted in the preset network through the second communication link.
[0056] Specifically, the communication module also includes: a communication control unit, a link control unit and a physical layer device. Figure 3 As shown, the communication control unit is, for example, a network card controller, the link control unit is, for example, a link controller, and the physical layer device is, for example, a physical interface.
[0057] The communication control unit, such as an IB controller or other network card controller, is the central processor of the IB network card and is responsible for managing and coordinating the various functions of the IB network card. The communication control unit is connected to the second protocol conversion unit. After receiving the first data packet to be transmitted from the computing module, the communication control unit transmits the first data packet to be transmitted to the second protocol conversion unit. For example, the second protocol conversion unit is an FPGA programming module for converting the NVLink protocol and the IB protocol. The second protocol conversion unit converts the first data packet to be transmitted based on the first communication protocol into a second data packet to be transmitted based on the second communication protocol, and transmits the second data packet to be transmitted to the communication control unit.
[0058] The link control unit is primarily responsible for managing and maintaining the physical link connections between nodes in the IB network, ensuring the efficiency, reliability, and stability of data transmission. The link control unit is also responsible for controlling the data transmission format. Furthermore, the link control unit also controls data flow, handles exceptions or errors such as congestion, and intelligently schedules network resources. Therefore, the communication control unit can generate a transmission strategy for transmitting data packets, such as data transmission format, exception handling method, and flow control method. The communication control unit is connected to the link control unit, and upon receiving the second data packet to be transmitted from the second protocol conversion unit, obtains the transmission strategy for the second data packet to be transmitted from the link control unit. This transmission strategy is generated by the communication control unit.
[0059] Physical layer devices, such as physical interfaces, primarily transmit digital data packets through the physical interface, such as 100Gps (Gigabits per second) and 400Gps. Furthermore, the physical interface converts digital signals into analog signals for transmission. The communication control unit is connected to the physical layer device and, based on the transmission strategy, transmits the second data packet to be transmitted through the physical layer device to the preset network, allowing other processors to access the second data packet on the preset network via the second communication link. The preset network, for example, is the external network, which is the network between AI server nodes.
[0060] Furthermore, the RDMA engine in the IB NIC reduces the CPU's need to directly access host memory, enabling efficient data transfer between nodes and significantly reducing latency and CPU overhead. However, this still requires the PCIe bus. The PCIe controller primarily facilitates communication between the IB NIC and the CPU, and the IB power supply primarily provides power to the IB NIC. The communication module in this embodiment does not require an RDMA controller, so the corresponding RDMA functionality of the link control unit is also removed. This removal also reduces circuit complexity and reduces circuit board wiring. Similarly, the PCIe controller and IB power supply are also removed.
[0061] Based on the above, the data transmission path between node 1 and node 2 is adjusted as follows: Figure 5 As shown in the figure, the computing module A in node 1, the GPU cache (L1 cache, L2 cache, L3 cache), the video memory controller, the video memory, the communication protocol bus between the graphics processors, the silicon interposer, the network card controller, the physical interface, the optical module, the optical module in node 2, the physical interface, the network card controller, the silicon interposer, the communication protocol bus between the graphics processors, the video memory, the video memory controller, the GPU cache, and the computing module B. This allows inter-node processor communication to bypass the PCIe bus and CPU, achieving near-zero PCIe bandwidth usage (for IB network cards). NVLink communication offers significantly higher bandwidth than PCIe. For example, NVLink 4.0 boasts a total bandwidth of 900GB / s, while PCIe 5.0 x16 boasts a mere 128GB / s. NVLink's transmission efficiency is significantly higher than PCIe. Furthermore, the link doesn't require a PCIe controller or CPU. Given the typical PCIe bus access latency of approximately 780 nanoseconds, this approach can reduce transmission latency by over 1.6 microseconds, significantly improving signal transmission efficiency. This path omits the PCIe controller and motherboard wiring found in traditional architectures, instead enabling data exchange directly through in-package interconnects.
[0062] In this embodiment, the communication control unit is used to obtain the transmission strategy of the second data packet to be transmitted from the link control unit when the second data packet to be transmitted is obtained, and transmit the second data packet to be transmitted to the preset network through the physical layer device according to the transmission strategy. The transmission strategy is used to ensure the efficiency, reliability and stability of data transmission. In addition, the above transmission process does not use the PCIe bus and does not occupy PCIe bandwidth.
[0063] According to an embodiment of the present application, an embodiment of a processor communication method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in the above-mentioned processor, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0064] In this embodiment, a processor communication method is provided. Figure 6 is a flowchart of a processor communication method according to an embodiment of the present application, such as Figure 6 As shown, the process includes the following steps: Step S601: When the processor has a data transmission demand, a first data packet to be transmitted based on a first communication protocol is obtained, wherein the first communication protocol is a communication protocol corresponding to a first communication link.
[0065] Specifically, the computing module in the processor can communicate with the computing modules of the processors in other nodes through the communication module. The communication process includes: when the processor has a data transmission demand, the computing module obtains a first data packet to be transmitted based on a first communication protocol, and transmits the first data packet to be transmitted to the communication module through a first communication link. The first communication protocol is a communication protocol corresponding to the first communication link, for example, the NVLink protocol.
[0066] Step S602: performing protocol conversion on the first data packet to be transmitted to obtain a second data packet to be transmitted based on a second communication protocol, wherein the second communication protocol is a communication protocol corresponding to the second communication link.
[0067] Specifically, the process of the communication module transmitting the first data packet to be transmitted to the other processor includes: the communication module performing protocol conversion on the first data packet to be transmitted to obtain a second data packet to be transmitted based on a second communication protocol, and transmitting the second data packet to be transmitted to the other processor via a second communication link. The second communication protocol is a communication protocol corresponding to the second communication link, for example, the IB protocol.
[0068] Step S603: Transmit the second data packet to be transmitted to other processors via the second communication link.
[0069] Specifically, the communication module is connected to other processors via a second communication link, such as Figure 2As shown, the second communication link is, for example, an optical module and an external network, and the optical module is, for example, an OSFP (Quad Small Form-factor Pluggable, four-channel small hot-pluggable optical module) optical module. The structure of other processors is the same as that of the processor, and they all include a computing module and a communication module. The computing module and the communication module are connected through the first communication link, and the communication module of the processor is connected to the communication module of other processors through the second communication link. After receiving the second data packet to be transmitted, the other processors will use their own communication modules to perform protocol conversion, and the computing modules of other processors can obtain data from the converted data packet and perform corresponding calculations. In addition, this embodiment can also update the GPU driver to cooperate with the above-mentioned communication method, such as removing the RDMA, PCIe protocol driver, etc.
[0070] The above-mentioned processor communication method integrates the communication control unit, protocol conversion unit, link control unit and physical layer equipment of the IB network card into the processor, avoiding the additional cost and space occupation of an independent IB network card. The current processor does not require structural design and production line assembly, which saves production costs. Advanced packaging technology is used to achieve short-distance and low-latency interconnection between the computing module and the communication module. The interconnection within the package is a dedicated channel, which does not occupy PCIe bandwidth and can support independent high-bandwidth communication between the GPU and the IB network. Compared with PCB routing, the packaging technology of the silicon interposer makes the interconnection energy consumption between the IB network card and the GPU close to zero. The existing solution consumes about 7W of power for the PCIe controller and about 25W to 40W for the RDMA controller. Therefore, this embodiment can save at least 33W of power consumption as a whole, improving the energy efficiency of the entire system.
[0071] This embodiment provides a processor communication method that encapsulates a communication module within a processor and transmits data from a computing module to other processors via the communication module. Furthermore, the communication module can convert a first data packet to be transmitted from a first communication protocol to a second communication protocol and transmit the converted second data packet to the other processors. This method addresses the issues of long communication paths caused by transmitting data via motherboard wiring, as well as the PCIe protocol's inherent requirement for serial-to-parallel data transmission formats, which increases communication power consumption and latency.
[0072] As an optional embodiment, the method further includes: When the processor has a data receiving demand, receiving a first data packet to be received based on the second communication protocol through the second communication link, performing protocol conversion on the first data packet to be received, and obtaining a second data packet to be received based on the first communication protocol; The data to be received is obtained according to the second data packet to be received, and the data to be received is calculated.
[0073] Specifically, other processors can send data packets to the processor, and the processor receives the data packets through the communication module, performs protocol conversion, and then transmits them to the computing module.
[0074] When a processor has a data receiving demand, it indicates that another processor has sent a first data packet to be received to the processor. The second communication protocol is the communication protocol corresponding to the second communication link, for example, the IB protocol. The communication module receives the first data packet to be received based on the second communication protocol through the second communication link, performs protocol conversion on the first data packet to be received, obtains the second data packet to be received based on the first communication protocol, and sends the second data packet to be received to the computing module through the first communication link. The first communication protocol is the communication protocol corresponding to the first communication link, for example, the NVLink protocol.
[0075] The calculation module parses the received second data packet to be received to obtain the data to be received, and calculates the data to be received. For example, the calculation module: Figure 3 As shown, the computing module includes a stream multiprocessor, a basic computing unit, a tensor core, a ray tracing core, etc. Among them, the basic computing unit is the smallest unit of the GPU that performs floating-point operations and integer operations; the stream multiprocessor mainly performs process task scheduling; the tensor core is a unit designed for matrix operations, mainly used for matrix operations during model training and inference; the main function of the ray tracing core is to display three-dimensional rendering functions.
[0076] In this embodiment, the computing module obtains the data to be received from other processors through the communication module and performs data calculations, no longer relying on the PCIe bus and the central processing unit, alleviating the pressure of PCIe bandwidth competition and improving system energy efficiency. In addition, the communication module uses the NVLink protocol to communicate with the video memory through a silicon interposer with almost negligible impedance, significantly improving communication efficiency and reducing communication latency.
[0077] As an optional embodiment, the processor communication method may further include steps A1 to A4. Steps A1 to A4 are performed by the communication module, which integrates high-speed interconnection protocols of various processor manufacturers.
[0078] In step A1, the communication module obtains the processor set included in all AI server nodes.
[0079] Specifically, processors such as GPUs. Currently, there are many GPU manufacturers and products. GPU cards from different manufacturers are incompatible with each other. They can only achieve high-speed interconnection of multiple GPUs of their own products through their own high-speed channels.
[0080] The communication module in this embodiment functions as a "high-speed processor switch," integrating high-speed interconnect protocols from various GPU vendors, such as NVLink, XGMI, and MLU-Link. A processor cluster is a collection of processors from different vendors within each AI server node, requiring support for high-speed interconnection among these processors.
[0081] Step A2: Determine the model of each processor in the processor set.
[0082] Specifically, different processor models may be produced by different manufacturers. The communication module automatically determines the processor model to match the corresponding manufacturer.
[0083] Step A3: Match the model of each processor with the high-speed interconnection protocol of each processor manufacturer through the communication module to obtain a matched high-speed interconnection protocol.
[0084] Step A4: Implementing high-speed interconnection communication between the processor sets according to the matched high-speed interconnection protocol.
[0085] Specifically, based on the models of the connected processors, the communication module automatically matches each processor with a specific high-speed interconnect protocol, obtaining a matching high-speed interconnect protocol for each processor. Based on this matching high-speed interconnect protocol, high-speed interconnect communication is achieved between the processors. This is achieved by unpacking and encapsulating the high-speed interconnect protocols of processors from different vendors.
[0086] In addition, the method provided in this application can further upgrade the communication module based on customer needs, such as customer networking needs, data transmission needs, etc., so that it is not only compatible with the processor, but also supports other function card sets, such as HBA (Host Bus Adapter) cards.
[0087] In this embodiment, a communication module is used to match each processor with the high-speed interconnection protocol of each GPU manufacturer, determine the matched high-speed interconnection protocol for each processor, and implement a technical solution for high-speed interconnection between processors from different manufacturers based on the matched high-speed interconnection protocol. This achieves interconnection protocol compatibility and high-speed interconnection between processors from different manufacturers.
[0088] This embodiment also provides a processor communication device for implementing the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0089] This embodiment provides a processor communication device, such as Figure 7 Shown, including: The data packet acquisition module 701 is configured to acquire a first data packet to be transmitted based on a first communication protocol when the processor has a data transmission demand, wherein the first communication protocol is a communication protocol corresponding to the first communication link; The protocol conversion module 702 is configured to perform protocol conversion on the first data packet to be transmitted to obtain a second data packet to be transmitted based on a second communication protocol, wherein the second communication protocol is a communication protocol corresponding to the second communication link; The data packet transmission module 703 is configured to transmit the second data packet to be transmitted to other processors via the second communication link.
[0090] As an optional embodiment, the device further includes: a data packet receiving module, configured to receive, when the processor has a data receiving demand, a first data packet to be received based on a second communication protocol via a second communication link, and perform protocol conversion on the first data packet to be received to obtain a second data packet to be received based on the first communication protocol; The data acquisition module is used to obtain the data to be received according to the second data packet to be received, and to calculate the data to be received.
[0091] The further functional description of each of the above modules is the same as that of the above corresponding embodiments and will not be repeated here.
[0092] The processor communication device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processing component and a storage component that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0093] The present application also provides a computer device having the above Figure 7 The processor communication device shown.
[0094] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present application. Figure 8As shown, the computer device includes: one or more processing components 10, a storage component 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processing component can process the instructions executed in the computer device, including instructions stored in or on the storage component to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processing components and / or multiple buses can be used together with multiple storage components and multiple storage components. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processing component system). Figure 8 A processing component 10 is taken as an example.
[0095] The processing component 10 may be a central processing unit, a network processing unit, or a combination thereof. The processing component 10 may further include an integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0096] The storage component 20 stores instructions that can be executed by at least one processing component 10, so that the at least one processing component 10 executes the method shown in the above embodiment.
[0097] The storage component 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the storage component 20 may include a high-speed random access storage component, and may also include a non-transient storage component, such as at least one disk storage component, a flash memory device, or other non-transient solid-state storage component. In some optional embodiments, the storage component 20 may optionally include a storage component remotely located relative to the processing component 10, and these remote storage components may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0098] The storage component 20 may include a volatile storage component, such as a random access storage component; the storage component may also include a non-volatile storage component, such as a flash storage component, a hard disk or a solid-state drive; the storage component 20 may also include a combination of the above types of storage components.
[0099] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0100] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processing component, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash storage component, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of storage components. It can be understood that the computer, processing component, microprocessor component controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processing component or hardware, the method shown in the above embodiment is implemented.
[0101] Part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0102] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the present application.
Claims
1. A processor, characterized in that: The processor includes: a computing module and a communication module; The computing module is connected to the communication module via a first communication link, and the computing module is configured to, when the processor has a data transmission demand, obtain a first data packet to be transmitted based on a first communication protocol, and transmit the first data packet to be transmitted to the communication module via the first communication link, wherein the first communication protocol is a communication protocol corresponding to the first communication link; The communication module is connected to other processors through a second communication link. The communication module is used to perform protocol conversion on the first data packet to be transmitted, obtain a second data packet to be transmitted based on a second communication protocol, and transmit the second data packet to be transmitted to the other processors through the second communication link, wherein the second communication protocol is the communication protocol corresponding to the second communication link.
2. The processor according to claim 1, wherein: The communication module is further configured to, when the processor has a data reception requirement, receive, through the second communication link, a first data packet to be received based on the second communication protocol, perform protocol conversion on the first data packet to be received to obtain a second data packet to be received based on the first communication protocol, and send the second data packet to be received to the computing module through the first communication link; The calculation module is further configured to obtain data to be received according to the second data packet to be received, and to calculate the data to be received.
3. The processor according to claim 1, wherein: The calculation module includes: a link training synchronization unit; The link training synchronization unit is used to train the first communication link to obtain adjustment parameters; The link training synchronization unit is further configured to adjust the first communication link according to the adjustment parameter.
4. The processor according to claim 1, wherein: The processor further includes: a power supply module; The power supply module is configured to provide a first power supply to the communication module upon receiving a first preset instruction; The power supply module is further configured to provide a second power supply to the communication module upon receiving a second preset instruction, wherein the power of the second power supply is greater than that of the first power supply.
5. The processor according to claim 2, wherein: The communication module includes: a first protocol conversion unit; The first protocol conversion unit is configured to obtain data to be received from the first data packet to be received, obtain preset field information according to the first communication protocol, and obtain a first intermediate data packet based on the data to be received and the preset field information, wherein the preset field information is used to describe parameters that need to be referenced during transmission of the first data packet to be received; The first protocol conversion unit is further configured to determine first verification information of the first intermediate data packet, and obtain the second data packet to be received based on the first intermediate data packet and the first verification information. The processor according to claim 1 , wherein: The communication module includes: a second protocol conversion unit; The second protocol conversion unit is configured to obtain data to be transmitted from the first data packet to be transmitted, obtain a routing header and a transmission header according to the second communication protocol, and obtain a second intermediate data packet according to the data to be transmitted, the routing header, and the transmission header; The second protocol conversion unit is further configured to determine second verification information of the second intermediate data packet, and obtain the second data packet to be transmitted based on the second intermediate data packet and the second verification information.
7. The processor according to claim 6, wherein: The communication module further includes: a communication control unit, a link control unit and a physical layer device; The communication control unit is connected to the second protocol conversion unit, and is used to transmit the first data packet to be transmitted to the second protocol conversion unit; The communication control unit is connected to the link control unit, and is configured to obtain a transmission strategy of the second data packet to be transmitted from the link control unit when the second data packet to be transmitted is obtained; The communication control unit is connected to the physical layer device and is used to transmit the second data packet to be transmitted to the preset network through the physical layer device according to the transmission strategy, so that the other processors obtain the second data packet to be transmitted on the preset network through the second communication link.
8. A processor communication method, characterized in that: The method comprises: When the processor has a data transmission demand, obtaining a first data packet to be transmitted based on a first communication protocol, wherein the first communication protocol is a communication protocol corresponding to the first communication link; Performing protocol conversion on the first data packet to be transmitted to obtain a second data packet to be transmitted based on a second communication protocol, wherein the second communication protocol is a communication protocol corresponding to the second communication link; The second data packet to be transmitted is transmitted to other processors through the second communication link.
9. The method according to claim 8, characterized in that The method further comprises: When the processor has a data receiving demand, receiving a first data packet to be received based on the second communication protocol through the second communication link, performing protocol conversion on the first data packet to be received, and obtaining a second data packet to be received based on the first communication protocol; Obtain data to be received according to the second data packet to be received, and perform calculations on the data to be received.
10. A computer device, characterized in that: include: A storage component and a processing component, wherein the storage component and the processing component are communicatively connected to each other, the storage component stores computer instructions, and the processing component executes the processor communication method according to any one of claims 8 to 9 by executing the computer instructions.
Citation Information
Patent Citations
Accelerating device and server
CN113946537A
Method and system for realizing data exchange between GPUs (Graphics Processing Unit) and protocol conversion chip
CN117827735A
Processor communication method, device, equipment, system and storage medium
CN118626431A
Event-driven architecture for interactive systems and applications
CN120066328A
Computing device, server, data processing method, and storage medium
WO2025138849A1
Cited By
Processor interconnection circuit and processor interconnection configuration method
CN121255711A
Network card control method and system and electronic equipment
CN122247930A