General Graphics Processing System, Computing Device, and Distributed System

By integrating the switching module in a general graphics processing system, the problem of interconnected communication overhead in distributed systems is solved, efficient data exchange and transmission is achieved, and computing performance is improved.

CN114066707BActive Publication Date: 2025-05-27T-HEAD (SHANGHAI) SEMICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010787539.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-07
Publication Date
2025-05-27
Estimated Expiration
2040-08-07

AI Technical Summary

Technical Problem

In distributed systems, the interconnection communication overhead between computing nodes and the interconnection communication overhead between internal devices have become bottlenecks in computing power growth. Traditional PCIe and TCP/IP interconnection technologies have limitations in bandwidth, network latency and scale expansion.

Method used

A general graphics processing system is designed, including a computing unit, a cache, a storage controller and a switching module. The switching module receives the identification and data source address of the target to be accessed through multiple interfaces, determines the interface based on the interconnection information, and performs data reading and sending, so as to realize efficient data exchange and transmission.

Benefits of technology

Through a general graphics processing system integrating switching modules, the data transmission capability between multiple graphics processing units is improved, network delay is reduced, computing performance is improved, and the computing burden of computing units is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114066707B_ABST
    Figure CN114066707B_ABST
Patent Text Reader

Abstract

Disclosed are a general-purpose graphics processing system, a computing device, and a distributed system. The general-purpose graphics processing system includes: a computing unit; a cache; a storage controller coupled to the cache; a switching module including a plurality of interfaces for receiving an identifier of a target to be accessed and a source address of first data to be written, determining a first interface among the plurality of interfaces according to the identifier of the target to be accessed and pre-stored interconnection information, reading the first data to be written from the cache according to the source address, and sending the first data to be written via the first interface; and a connection unit for coupling the computing unit, the storage controller, the cache, and the switching module. According to an embodiment of the present disclosure, the general-purpose graphics processing system integrated with the switching module is no longer a simple end device, but has networking and switching capabilities, so that networking no longer solely relies on external switches and routers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning, and in particular, to a general graphics processing system, a computing device, and a distributed system. Background Art

[0002] Deep learning is one of the most remarkable technologies that have emerged again in the past decade and has achieved many breakthroughs and practical applications in the fields of speech, images, big data, biomedical technology, etc. To meet the needs of more complex application scenarios, the scale of deep learning models has become increasingly large. The growth of model parameters makes it difficult for the computing resources required for model training to be satisfied by the computing resources of a single node. Therefore, for complex models, a distributed system including multiple computing nodes is usually used for model training.

[0003] However, in a distributed system, the interconnection communication overhead between computing nodes and the interconnection communication overhead between components within a computing node will both become bottlenecks for the growth of computing power. Therefore, the computing power does not increase linearly with the increase in the number of nodes. Traditional PCIe and TCP / IP interconnection technologies have been proven to have limitations in terms of bandwidth, network latency, scalability, etc., and are not suitable for such distributed systems. Thus, the interconnection communication problem has become one of the core difficult problems in the design of distributed systems. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to provide a general graphics processing system, a computing device, and a distributed system to solve the problems existing in the prior art.

[0005] According to the first aspect of the embodiments of the present disclosure, an embodiment of the present disclosure provides a general graphics processing system, including:

[0006] A computing unit;

[0007] A cache;

[0008] A storage controller coupled to the cache;

[0009] A switching module including a plurality of interfaces, configured to receive an identifier of a target to be accessed and a source address of first data to be written, determine a first interface among the plurality of interfaces according to the identifier of the target to be accessed and pre-stored interconnection information, read the first data to be written from the cache according to the source address, and send the first data to be written via the first interface;

[0010] A connection unit configured to couple the computing unit, the storage controller, the cache, and the switching module.

[0011] In some embodiments, the identifier of the target to be accessed and the source address are from a data operation request submitted by the computing unit or the storage controller.

[0012] In some embodiments, the switching module further includes:

[0013] A transmission engine, configured to encode the first data to be written according to a specified transport layer / network layer communication protocol;

[0014] A switching unit, configured to determine the first interface according to the identifier of the target to be accessed and the interconnection information, continue to encode the first data to be written according to the Ethernet communication protocol and the physical layer protocol, and send the encoded data via the first interface.

[0015] In some embodiments, the transmission engine supports multiple transport layer / network layer communication protocols, and the transmission engine selects the specified transport layer / network layer communication protocol from the multiple transport layer / network layer communication protocols to encode the first data to be written.

[0016] In some embodiments, the transmission engine includes:

[0017] A RoCE protocol processing module, configured to encode the first data to be written based on the RoCEv2 communication protocol;

[0018] A proprietary protocol processing module, configured to encode the first data to be written based on a proprietary protocol.

[0019] In some embodiments, the RoCE protocol processing module includes a TOE unit, and the TOE unit is configured to encode the first data to be written according to the IP / TCP / UDP protocol.

[0020] In some embodiments, the RoCE protocol processing module includes an IB unit, and the IB unit is configured to establish end-to-end data transmission based on the IB protocol.

[0021] In some embodiments, the RoCE protocol processing module includes a predicate interface.

[0022] In some embodiments, the proprietary protocol processing module includes a proprietary protocol unit and a proprietary protocol driver interface. The proprietary protocol unit is configured to encode and decode data according to a proprietary protocol, and the proprietary protocol driver interface is configured to provide a driver and a hardware interface.

[0023] In some embodiments, the identifier of the target to be accessed includes the identifier of the target graphics processing unit connected via the first interface and the identifier of the target application program.

[0024] In some embodiments, the switching unit is further configured to receive second data to be written via a second interface, determine a target address according to the identifier of the target to be accessed, and write the second data to be written to the target address.

[0025] In some embodiments, the source address and the target address are respectively specific storage addresses of a source application and a target application in their respective application memory spaces.

[0026] In some embodiments, the switching module is integrated into a network card processor.

[0027] In some embodiments, the target graphics processing unit and the general graphics processing system are located in different computing nodes.

[0028] In some embodiments, the plurality of interfaces are Ethernet interfaces.

[0029] In a second aspect, an embodiment of the present disclosure provides a computing device, including a plurality of computing nodes, where each computing node includes a coupled memory, a general-purpose processor, and the general graphics processing system according to any one of the above, and wherein each computing node is coupled to the general graphics processing systems of at least one other computing node through its own general graphics processing system.

[0030] In some embodiments, the computing nodes are encapsulated in the same silicon wafer, and the plurality of computing nodes are integrated together through a printed circuit board.

[0031] In some embodiments, the computing device is configured to perform a training task of a deep learning model.

[0032] In a third aspect, an embodiment of the present disclosure provides a distributed system, including a plurality of the computing devices according to any one of the above, and data transmission is performed among the plurality of computing devices through an external bridging device.

[0033] In a fourth aspect, an embodiment of the present disclosure provides a distributed system, including a plurality of computing nodes, where each computing node includes a coupled memory, a general-purpose processor, and the general graphics processing system according to any one of the above, and wherein each computing node is coupled to the general graphics processing systems of at least one adjacent other computing node through its own general graphics processing system, and a 3D-Torus interconnection network is formed.

[0034] In a fifth aspect, an embodiment of the present disclosure provides a cloud server, including the computing device according to any one of the above.

[0035] In a sixth aspect, an embodiment of the present disclosure provides a method implemented in a general graphics processing system, including:

[0036] Receiving an identifier of a target to be accessed and a source address of data to be written;

[0037] Determine a first interface among multiple interfaces according to the identifier of the target to be accessed and pre-stored interconnection information;

[0038] Read the data to be written from the cache according to the source address, and send the data to be written via the first interface.

[0039] According to the embodiments of the present disclosure, the general graphics processing system of the integrated switching module is no longer a simple end device, but has networking and data exchange capabilities, enabling networking to no longer rely solely on external switches and routers. Since the general graphics processing system can read data from other devices or write data to other devices without relying on external switches and routers, it can improve the data exchange ability between the general graphics processing system and other devices, and also helps to improve the computing performance of the general graphics processing system. Brief Description of the Drawings

[0040] Through the description of the embodiments of the present disclosure with reference to the following drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0041] Figure 1 is a hierarchical structure diagram of a data center;

[0042] Figure 2 is a three-dimensional structure diagram of a data center;

[0043] Figure 3 is a schematic structural diagram of a cloud server with a general structure of a data center.

[0044] Figure 4a is a structural diagram of a cloud server for performing a model training task in the prior art;

[0045] Figure 4b shows in Figure 4a an improved solution for a cloud server;

[0046] Figure 5 is a schematic structural diagram of a general graphics processing system provided according to an embodiment of the present disclosure;

[0047] Figure 6 shows Figure 5 a more specific functional structural diagram of the switching module in

[0048] Figure 7 is a structural diagram of a message of a proprietary protocol;

[0049] Figure 8 is Figure 5 a schematic diagram of the interconnection structure of the general graphics processing system shown in

[0050] Figure 9 and 10 are respectively schematic application diagrams of a computing node including an embodiment of the present disclosure;

[0051] Figure 11 is a flowchart of a method for a general graphics processing system according to an embodiment of the present disclosure. Detailed implementation manners

[0052] The present disclosure will be described below based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these details. In order to avoid obscuring the essence of the present disclosure, well-known methods, processes, and procedures are not described in detail. Additionally, the drawings are not necessarily drawn to scale.

[0053] The following terms are used herein:

[0054] OSI model: The International Organization for Standardization (ISO) has developed the OSI (Open System Interconnection) model. The OSI model is an abstract model system that divides the work of network communication into seven layers, namely the physical layer, data link layer, network layer, transport layer, session layer, presentation layer, and application layer. Among them, the physical layer performs the actual final signal transmission, transmitting the electronic signal of the bit stream through the physical medium. The data link layer is responsible for network addressing, error detection, and correction. In this layer, the header and trailer are added to the data packet to form a frame. The network layer determines the path selection and forwarding of data. In this layer, the network header (NH) is added to the data packet to form a packet. The transport layer determines to provide a reliable connection from end to end. In this layer, the transport header (TH) is added to the data to form a data packet. The session layer provides mechanisms for establishing and maintaining communication between applications, including access authentication and session management. The presentation layer mainly solves the problem of the syntax representation of user information. It provides formatted representation and data conversion services. The application layer provides interface services between the network and user application software.

[0055] Ethernet technology: The essence of this technology is a media access control technology at the data link layer. It can be combined with various physical layer technologies to form various Ethernet technology access systems. For example, combined with VDSL on telephone copper cables, it forms the EoVDSL technology; combined with passive optical networks, it generates the EPON technology; in a wireless environment, it develops into the WLAN technology.

[0056] RoCE Protocol: It is an extended protocol of Ethernet technology in the field of private networks. Although Ethernet technology has always dominated the global interconnected Internet, it reveals many drawbacks in high-bandwidth and low-latency private networks. With the rise of the concept of network convergence, in the DCB (Data Center Bridging) standard released by the IETF (Internet Engineering Task Force), lossless links based on RDMA / Infiniband have been solved, and Ethernet finally has its own standard in the field of private networks. At the same time, the concept of RoCE (RDMA over Converged Ethernet) has also been proposed. After version upgrades (from RoCEv1 to RoCEv2), new NICs (Network Interface Controllers) and switches above 10 Gb basically integrate RoCE support. RoCE v1 (Layer 2) operates at the data link layer (Ehternet LinkLayer). RoCE v2 operates at the network layer / transport layer (UDP / IPv4 or UDP / IPv6).

[0057] Data center

[0058] Figure 1 Shows a hierarchical structure diagram of a data center as a scenario applied in an embodiment of the present disclosure.

[0059] A data center is a specific network of devices for global collaboration, used to transfer, accelerate, display, compute, and store data information on the Internet network infrastructure. In future development, data centers will also become assets for enterprise competition. With the widespread application of data centers, artificial intelligence and the like are increasingly applied to data centers. As an important technology of artificial intelligence, neural networks have been widely applied to big data analysis operations in data centers.

[0060] In traditional large data centers, the network structure is usually Figure 1 The three-layer structure shown, that is, the hierarchical inter-networking model. This model includes the following three layers:

[0061] Access Layer 103: Sometimes also referred to as the Edge Layer, it includes access switches 130 and various servers 140 connected to the access switches. Each server 140 is a processing and storage entity in the data center, and the processing and storage of a large amount of data in the data center are completed by these servers 140. The access switch 130 is a switch used to connect these servers to the data center. One access switch 130 connects multiple servers 140. The access switches 130 are usually located at the top of the rack, so they are also called Top of Rack switches, and they are physically connected to the servers.

[0062] Aggregation Layer 102: Sometimes also called the Distribution Layer, it includes aggregation switches 120. Each aggregation switch 120 connects multiple access switches and also provides other services, such as firewalls, intrusion detection, network analysis, etc.

[0063] Core Layer 101: It includes core switches 110. The core switches 110 provide high-speed forwarding for packets entering and leaving the data center and provide connectivity for multiple aggregation layers. The network of the entire data center is divided into an L3 layer routing network and an L2 layer routing network. The core switches 110 usually provide a flexible L3 layer routing network for the network of the entire data center.

[0064] Normally, the aggregation switch 120 is the demarcation point between the L2 and L3 layer routing networks. Below the aggregation switch 120 is the L2 network, and above is the L3 network. Each group of aggregation switches manages a Point Of Delivery (POD), and each POD is an independent VLAN network. When a server migrates within a POD, it does not need to modify its IP address and default gateway because one POD corresponds to one L2 broadcast domain.

[0065] The Spanning Tree Protocol (STP) is usually used between the switch 120 and the access switch 130. STP makes only one aggregation layer switch 120 available for a VLAN network, and other aggregation layer switches 120 are used only when a failure occurs (the dotted lines in the above figure). That is to say, at the aggregation layer, horizontal expansion cannot be achieved because even if multiple aggregation switches 120 are added, only one is working.

[0066] Figure 2 shows Figure 1 the physical connections of the components in the hierarchical data center. As Figure 2As shown in the figure, a core switch 110 is connected to multiple aggregation switches 120, an aggregation switch 120 is connected to multiple access switches 130, and an access switch 130 is connected to multiple servers 140.

[0067] Cloud server

[0068] The cloud server 140 is the real device in the data center. Since the cloud server 140 operates at a high speed to execute various tasks such as matrix calculation, image processing, machine learning, compression, search and sorting, etc., in order to efficiently complete the above various tasks, the cloud server 140 usually includes a central processing unit (CPU) and various acceleration units, such as Figure 3 shown. The acceleration unit is, for example, an acceleration unit dedicated to neural networks, a data transfer unit (DTU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). The following Figure 3 takes the example shown below to introduce each acceleration unit separately.

[0069] Acceleration unit 230 dedicated to neural networks: It is an architecture that adopts data-driven parallel computing and is a processing unit for processing a large number of operations (such as convolution, pooling, etc.) of each neural network node. Since the data and intermediate results in a large number of operations (such as convolution, pooling, etc.) of each neural network node are closely related during the entire calculation process and are often used, with the existing CPU architecture, because the memory capacity inside the CPU core is very small, a large amount of frequent access to off-core memory is required, resulting in low processing efficiency. By using an acceleration unit, each core has on-chip memory with a storage capacity suitable for neural network computing, avoiding frequent access to off-core memory, which can greatly improve processing efficiency and computing performance.

[0070] Data transfer unit (DTU) 260: It is a wireless terminal device specifically used to convert serial port data into IP data or convert IP data into serial port data for transmission through a wireless communication network. The main function of the DTU is to transmit the data of remote devices back to the background center wirelessly. At the front end, the DTU is connected to the customer's device through an interface. After the DTU is powered on and runs, it first registers to the mobile GPRS network, and then establishes a socket connection with the background center set in the DTU. The background center is the server of the socket connection, and the DTU is the client of the socket connection. Therefore, the DTU and the background software are used together. After the connection is established, the front-end device and the background center can perform wireless data transmission through the DTU.

[0071] Graphics processing unit (GPU) 240: is a processor that specializes in image and graphics related computing. The use of GPU overcomes the shortcoming of too little space for computing units in the CPU, and adopts a large number of computing units dedicated to graphics computing, so that the graphics card reduces the dependence on the CPU and undertakes some of the computationally intensive image processing work that the CPU originally undertook.

[0072] Application-specific integrated circuit (ASIC): refers to an integrated circuit designed and manufactured to meet the needs of specific users and specific electronic systems. Since this integrated circuit is customized according to user requirements, its structure is often adapted to specific user requirements.

[0073] Field Programmable Gate Array (FPGA): It is a product further developed on the basis of programmable devices such as PAL and GAL. It appears as a semi-custom circuit in the field of Application Specific Integrated Circuit (ASIC), which not only solves the shortcomings of custom circuits, but also overcomes the shortcomings of the limited number of gate circuits of the original programmable devices.

[0074] Based on the above-mentioned general structure of the cloud server, system operation and maintenance personnel will select different acceleration units in the cloud server according to different tasks in practice. Figure 4a The structure diagram of a cloud server performing a model training task is shown. As shown in the figure, the cloud server 140 is composed of multiple nodes 141 and multiple switches 144 located outside the nodes 141. Each node 141 is coupled to at least one external switch 144, and multiple switches 144 are coupled to each other, thereby forming a network structure, in which data exchange is realized between nodes 141 via switches 144.

[0075] The structure of an exemplary node 141 is shown in the figure. The node 141 includes task units 143 and 142. Task units 142 and 143 have the same structure. Task unit 143 is taken as an example for introduction here. Task unit 143 includes a memory 1411, a CPU 1412, an interconnection switching unit 1413, a GPU 1414, a network interface controller (NIC) 1415, a GPU 1416, and a network interface controller 1417. The functions of CPU 1412, GPU 1414, and GPU 1416 can be referred to as described above. It should be understood that, according to the model training task to be performed, the cloud server 140 will include a suitable number of task units, and the CPU and GPU in each task unit will adopt a suitable ratio.

[0076] The following takes task unit 143 as an example for introduction. The memory 1411 is used to store instructions and data, which can be a random access memory or a flash memory, etc. The interconnection and switching unit 1413 has two functions: providing a physical connection between the GPU 1414, GPU 1416, network interface controllers 1415 and 1417, and the CPU 1412, and providing data forwarding capabilities between them. For example, through the interconnection and switching unit 1413 and via the network interface controller 1415, the CPU 1412 can obtain instructions and data from the outside, and then provide the instructions and data for the GPU 1414 to use. The network interface controllers 1415 and 1417 are connected to the external switch 144.

[0077] The working process of the task unit 141 can be as follows: The central processing unit 1412 obtains instructions and data from the outside via the network interface controller 1415 and the network interface controller 1417 and stores them in the memory 1411 via the interconnection and switching unit 1413, and then reads the instructions and data from the memory 1411 and distributes them to the GPUs 1414 and 1416 for parallel computing of image data by the GPUs 1414 and 1416.

[0078] It should be noted that the node 141 can have multiple product manifestations. For example, the node 141 is implemented on the same silicon chip (i.e., a system-on-chip), and the network structure composed of multiple nodes 141 and multiple switches 144 is implemented as an integrated device. Another example is that the node 141 is implemented as a computer, and the network structure composed of multiple nodes 141 and multiple switches 144 constitutes a computer cluster structure.

[0079] In the above structure, the interconnection communication overhead will become a bottleneck for the growth of computing power. For example, due to the interconnection communication overhead and network latency caused by the interconnection and switching unit 1413 and the external switch 144, the computing power of the task unit 143 is usually weaker than the sum of the computing powers of the GPUs 1414 and 1416. Therefore, in order to improve the overall computing power of the system, it is necessary to further reduce the interconnection communication overhead. Figure 4b Shown in Figure 4a Based on the improvement scheme. In Figure 4bIn this case, the interconnect switching unit 1413 and the CPU 1412 are integrated together, so that the data exchange requirements between the CPU 1412 and the network interface controllers 1415 and 1417 can be eliminated. Similarly, since the interconnect switching unit 1423 and the CPU 1422 are integrated together, the data exchange requirements between the CPU 1422 and the network interface controllers 1425 and 1427 are eliminated. However, this solution still has the following limitations: Since the network interface controllers 1415 and 1417 can adopt the latest RDMA technology to achieve a data bandwidth of 800G, but the switch 144 usually uses the existing Ethernet switches on the market, and such Ethernet switches mainly support a data bandwidth of 25G - 100G. The mismatch between the two results in the inability to maximize the computing power.

[0080] General graphics processing system provided by the embodiments of the present disclosure

[0081] Embodiments of the present disclosure provide a Figure 5 General-purpose computing system on graphics processing 500 as shown in Figure 3 to replace the GPU in

[0082] As shown in the figure above, the general-purpose graphics processing system 500 adds a switching module 501. The switching module 501 includes hardware and software implementations, and its function is similar to the Figure 4a and 4b interconnect switching units in

[0083] to implement the data exchange function. Outside the switching module 501, the general-purpose graphics processing system 500 further includes a plurality of computing units 502, a connection unit 503, a storage controller 504, and a cache 505.

[0084] The connection unit 503 is used to couple the switching module 501, the plurality of computing units 502, and the storage controller 504. The main functions of the connection unit 503 are data transmission and interface conversion. Data transmission transfers data from one component to another component, and interface conversion means converting the received data and then outputting it.

[0084] The computing unit 502 is used to complete computing tasks related to image data. Multiple computing units 502 can execute computing tasks in parallel, thus improving the computing performance of the general graphics processing system 500. The computing unit 502 can read instructions and data from the cache 505 via the memory controller 504 to execute computing tasks. The computing unit 502 further includes an instruction fetch unit, a decoding unit, an arithmetic computing unit, and registers and caches required for execution that are not shown in the figure. The instruction unit is used to read instructions and data from the memory. The decoding unit is used to parse the instructions. The logical computing unit is used to perform actual arithmetic operations. These components cooperate to complete the computing tasks of the computing unit 502.

[0085] As shown in the figure, the memory controller 504 and the cache 505 are coupled. The memory controller 504 and the cache 505 can be integrated into one memory device or be separate devices as shown in the figure. The memory controller 504 performs necessary control over the access to the cache 505. For example, when the computing unit 502 accesses the cache 505 via the memory controller 504, the memory controller 504 converts the read and write commands issued by the computing unit 502 into signals that the cache 505 can recognize, and also completes address decoding and data format conversion between the computing unit 502 and the memory controller 504. Similarly, when the switching module 501 accesses the cache 505 via the memory controller 504, the memory controller 504 also performs corresponding access control.

[0086] The switching module 501 includes multiple interfaces 5011, and is coupled to other devices, such as other general graphics processing systems, via the multiple interfaces 5011, and exchanges data with other devices via the multiple interfaces 5011. Specifically, if the current general graphics processing system wants to write data to another general graphics processing system coupled to it, the switching module 50 will receive a write operation request, which usually includes the identifier of the target to be accessed and the source address of the data to be written. The first interface among the multiple interfaces 5011 is determined according to the identifier of the target to be accessed and the pre-stored interconnection information. The data to be written is read from the cache 505 according to the source address, and the data to be written is sent via the first interface. If the current general graphics processing system wants to read data from another general graphics processing system coupled to it, the switching module 50 will receive a read operation request, and the read operation request usually includes the identifier of the target to be accessed and the source address, and the source address represents the storage address of the data to be read. The first interface among the multiple interfaces 5011 is determined according to the identifier of the target to be accessed and the pre-stored interconnection information, and the data to be written to the source address is received via the first interface. Among them, the interconnection information records the connection information between the current general graphics processing system and other general graphics processing systems.

[0087] Write operation requests and read operation requests can come from application programs executed by the computing unit 502, or from the storage controller 504. The target to be accessed can be a target application program executed in other general-purpose graphics processing systems coupled to the current general-purpose graphics processing system. Therefore, the identifier of the target to be accessed usually includes the identifier of the general-purpose image processing unit to which the target application program belongs and the identifier of the target application program itself.

[0088] Other general-purpose graphics processing systems coupled to the current general-purpose graphics processing system also have a switching module 501. For write operation requests and read operation requests, the switching module 501 needs to determine the target address in the memory space of the target application program. The target address corresponds to the source address, which is a specific storage address in the application memory space of the source application program in the cache. The target address is a specific storage address in the application memory space applied for by the target application program in the cache. If it is a write operation request, the switching module 501 will directly read the data from the source address and write it to the target address. If it is a read operation request, the switching module 501 will read the data from the target address and write it to the source address.

[0089] In this embodiment, the general-purpose graphics processing system has a built-in data exchange function. Therefore, multiple graphics processing units can form a network structure and complete data forwarding without external and internal switches, thereby improving the data transmission ability between multiple graphics processing units and reducing network latency, thus improving the computing performance of the graphics processing unit.

[0090] Furthermore, both reading and writing data are independently completed by the switching module without using the resources of the computing unit, so it helps to reduce the computing burden of the computing unit, thus improving the computing performance of the graphics processing unit.

[0091] Continuing to refer to the figure shown above, as an alternative embodiment, the switching module 501 includes a plurality of internally coupled interfaces 5011, a switching unit 5012, and a transmission engine 5013.

[0092] The interface 5011 is, for example, an Ethernet interface compliant with the 802.3 standard. The interface 5011 can be implemented using an Ethernet interface chip, and its main function is to transmit bitstream signals at the physical layer. Transmitting bitstream signals includes sending and receiving bitstream signals. When sending bitstream signals, frame data is received from the data link layer and then converted into bitstream signals for output. When receiving bitstream signals, bitstream signals are received from the transmission medium connected to the interface 5011 and converted into frame data to be provided to the data link layer. The interface 5011 usually also includes a serializer / deserializer (SerDes) based serial-parallel converter to achieve the transmission of high-speed serial signals. When sending bitstream signals, the serial-parallel converter converts parallel bitstream signals into serial bitstream signals and sends them to the interface 5011 at the receiving end via the transmission medium. When receiving bitstream signals, the serial-parallel converter converts serial bitstream signals into parallel bitstream signals. Currently, integrating a serial-parallel converter in the interface to achieve the transmission of high-speed serial signals has become the mainstream choice for many commercial products.

[0093] The transmission engine 5013 is used to encode data according to the specified transport layer / network layer communication protocol, send the encoded data to the switching unit 5012, and at the same time receive the data to be decoded from the switching unit 5012 and decode the data according to the specified transport layer / network layer communication protocol.

[0094] The switching unit 5012 is used to determine the first interface according to the identifier of the target to be accessed and the interconnection information, encode the data according to the Ethernet communication protocol and the physical layer protocol, and send the encoded data via the first interface. At the same time, the switching unit 5012 receives data from the first interface, decodes the data according to the Ethernet communication protocol and the physical layer protocol, and sends the decoded data to the transmission engine 5013.

[0095] Of course, the operation of determining the first interface according to the identifier of the target to be accessed and the interconnection information can also be performed by the transmission engine 5013, and then the transmission engine 5013 passes the identifier of the first interface to the switching unit 5012.

[0096] It should be noted that the communication protocols referred to in this article are usually protocol families. For example, as mentioned above, the OSI model includes multiple layers, and multiple communication protocols are involved in each layer. For example, at the data link layer, there are specifications such as ARP, RARP, IEEE802.3, PPP, CSDA / CD, RoCE, etc.; at the network layer, there are specifications such as IP, ICMP, RIP, and IGMP, etc.; at the transport layer, there are specifications such as TCP and UDP, etc. The data transmission implemented by the switching module 501 needs to be encoded and decoded according to the OSI model specifications or other model specifications.

[0097] In some embodiments, at least some components of the switching module 501 may be integrated into a network card processor, and sometimes the interface 5011 may also be integrated into the network card processor.

[0098] Figure 6 shows Figure 5 a more specific functional structure diagram of the switching module in. As shown in the figure, the switching module 501 includes a switching unit 5012, a transmission engine 5013, and multiple interfaces 5011. The transmission engine 5013 includes a RoCE protocol processing module 601, a proprietary protocol processing module 602, and a driver module 603. The switching unit 5012 includes an Ethernet processing unit 6031 and multiple MAC controllers 6032.

[0099] Among them, the RoCE protocol processing module 601 implements the functions of the transport layer using the RoCEv2 communication protocol, and the proprietary protocol processing module 602 uses a custom proprietary protocol. The RoCE protocol processing module 601 further includes a TOE unit 6011, an IB unit 6012, and a verbs interface (verbs inf) 6013. The proprietary protocol processing module 602 further includes a dedicated protocol driver interface 6021 and a proprietary protocol unit 6022. The driver module 603 refers to the hardware driver. If the operation of determining the first interface according to the identifier of the target to be accessed and the interconnection information is performed by the transmission engine 5013, the driver module 603 may determine the first interface according to the identifier of the target to be accessed and the interconnection information, and transmit the identifier information of the first interface to other functional modules. Of course, the present disclosure is not limited to this.

[0100] The TOE unit 6011 implements data encoding and decoding of traditional network layer and transport layer communication protocols, and these communication protocols include one or several of IP, UDP, DHCP, ICMP, and ARP. When the TOE unit 6011 receives a data packet from the switching unit 5012, it strips the packet header information from the received data, and the packet header information is, for example, an IP packet header, a UDP packet header, an ARP packet header, and so on. It also checks for packet header type errors and packet errors according to the communication protocol type. Only error-free data packets will be sent to the IB unit 6012. When it receives a data packet from the IB unit 6012, it adds a packet header according to the communication protocol type and sends the data packet with the added packet header to the Ethernet processing unit 6031.

[0101] The IB unit 6012 implements the data transmission function based on the IB protocol. The IB unit 6012 parses the work queue element (WQE), and then performs send message, receive message, read / write atomic operation, etc. of the RoCEv2 protocol according to the work queue element to achieve end-to-end data transmission. The work queue element is written by other processing units, for example, by the computing unit 502 or the memory controller 504. The IB unit 6012 also checks data integrity according to the packet sequence number (PSN) of each data packet.

[0102] The predicate interface 6013 is various interfaces between the driver and the hardware. It includes memory read / write and atomic operations based on the RoCEv2 protocol, and these operations are not executed via the processor.

[0103] The proprietary protocol processing module 602 further includes a dedicated protocol driver interface 6021 and a proprietary protocol unit 6022. The dedicated protocol driver interface 6021 is a custom dedicated protocol processing module above the data link layer, which does not include the processing of IP / TCP / UDP protocols, so as to reduce the resource overhead caused by processing heavy protocol headers. The proprietary protocol directly encodes the target address, security key and packet sequence number (PSN) of its operation after the MAC frame header.

[0104] As shown in the figure, the switching unit 5012 includes an Ethernet processing unit 6031 and multiple MAC controllers 6032. The MAC controller 6032 refers to a device that implements the data link layer function. The MAC controller 6032 receives data packets from the Ethernet processing unit 6031, and sends the processed data packets to the interface 5011 according to the protocol specifications of the data link layer. At the same time, the MAC controller 6032 receives data packets from the interface 5011 and sends the processed data packets to the Ethernet processing unit 6031. The MAC controller 6032 also checks errors in the frame data and discards damaged data packets at the same time.

[0105] The Ethernet processing unit 6031 is a functional layer of the data link layer, which is used to encode and decode data according to the Ethernet protocol. When sending data, the Ethernet processing unit 6031 needs to obtain the MAC address of the first interface, then encode the data according to the MAC address, and then send the encoded data to one of the MAC controllers 6032 according to the MAC address. When receiving data, the Ethernet processing unit 6031 parses the received data packet and performs verification on it. After the verification is completed, the Ethernet processing unit 6031 will send the data packet to the RoCE protocol processing module 601 or the proprietary protocol processing module 602 according to the indication information contained in the MAC header. The indication information in the MAC indicates the communication protocol used by the data at the transport layer.

[0106] In this embodiment, the RoCEv2 protocol is adopted to achieve direct transfer of buffers between the application programs at both ends, without the intervention of the operating system and the protocol stack. Thus, ultra-low latency and ultra-high throughput data transmission can be achieved, and the processing resources of the computing unit are basically not occupied.

[0107] It should be noted that the RoCEv2 adopted in this embodiment is a direct memory access technology. However, in addition to RoCEv2, other direct memory access technologies can also be adopted in this embodiment. The direct memory access technology is usually applied in computer systems, but in this embodiment, it is applied to the general graphics processing system, and this processing method can improve the computing performance of the general graphics processing system.

[0108] Figure 7 It is a message structure diagram of an exemplary proprietary protocol. As shown in the figure, the message of the proprietary protocol is divided into a MAC header and the data itself. The MAC header includes the recipient MAC address, the sender MAC address, and the Ethernet type. Taking the above embodiment as an example. When the computing unit 502 or the memory controller 504 writes data to another general graphics processing system, the sender MAC address should be the MAC address of the first interface 5011 determined according to the identifier of the target to be accessed, and the recipient MAC address should be the MAC address of the second interface of another general graphics processing system connected to the first interface 5011. The Ethernet type refers to the protocol type used by the previous layer. Since the switching unit 5012 processes the communication protocols on each layer in the order from top to bottom of the OSI model, when writing data, the protocol type used by the previous layer, for example, the network layer is the IP protocol, the TCP protocol, or a custom proprietary protocol. The data itself includes the target address, the security key, and the data to be written to the target address divided into groups. The security key provides security protection in the data integrity function.

[0109] The encoding and decoding of the MAC header are implemented in the Ethernet processing unit 6031. For example, after the Ethernet processing unit 6031 receives a data packet from the proprietary protocol engine 6015, it encodes the MAC header before the data packet, and the encoding and decoding of the data packet are performed in the proprietary protocol engine 6015.

[0110] Figure 8 is Figure 5Schematic diagram of the interconnection structure of the general graphics processing system shown. As described above, the interconnection structure includes computing nodes 800 and 900. Computing node 800 includes two task units 801 and 802. Task unit 801 includes a memory 8011, a processor 8012, and general graphics processing systems 8013 and 8014 implemented based on the embodiments of the present disclosure, which are coupled via a bus. Task unit 802 includes a memory 8022, a processor 8021, general graphics processing systems 8023 and 8024, which are coupled via a bus. Processors 8012 and 8021 are coupled. Computing node 900 includes task units 901 and 902. Task unit 901 includes a memory 9011, a processor 9012, general graphics processing systems 9013 and 9014, which are coupled via a bus. Task unit 902 includes a memory 9021, a processor 9022, general graphics processors 9023 and 9024, which are coupled via a bus. Processors 9012 and 9022 are coupled via a bus.

[0111] In the figure, computing nodes 800 and 900 are coupled depending on the interfaces inside their respective internal general graphics processing systems. As shown in the figure, computing node 800 uses the interfaces in general graphics processing system 8013 to be coupled to general graphics processing systems 9013, 9014, 9023 and 9024 in computing node 900 respectively.

[0112] Based on this interconnection structure, the computing nodes communicate with each other. For example, processor 8012 can access the cache in general graphics processing system 9013 via general graphics processing system 8013 or 8014. For another example, if memory controllers adapted to RDMA are deployed in both general graphics processing systems 8013 and 9013, the memory controller of general graphics processing system 8013 can directly write data to or read data from the storage controller in general graphics processing system 9013.

[0113] Figure 9 and 10 are respectively application schematic diagrams of computing nodes including an embodiment of the present disclosure.

[0114] As Figure 9 shown, computing nodes 11 to 14 are integrated into an integrated device through a printed circuit board (PCB) 1, and computing nodes N1 to N4 are integrated into an integrated device through a printed circuit board N. Inside the integrated device, coupled communication is achieved through a switching module. In terms of the ability to scale up, between the integrated devices, coupled communication is achieved through the printed circuit board. At this time, the board-level transmission rate can reach 800G. In terms of the ability to scale out, between the integrated device and the bridging device of the data center, the switching module in the computing node can provide an output transmission capacity of 100G to the bridging device of the data center.

[0115] Optionally, each computing node can be implemented as a system on a chip (SoC), and the interconnection structure composed of computing nodes can be encapsulated as an integrated device.

[0116] Figure 10 A 3D-Torus interconnection network is shown. In each computing node, a NOC and a NIU are used to couple the cache, the general-purpose graphics processing system, and the switching module together. The NOC can be defined as a multiprocessing system based on network communication implemented on a single chip. The NOC has the following advantages over using a bus: 1) It has good address space scalability, and theoretically, the number of resource nodes that can be integrated is not limited; 2) It provides good parallel communication capabilities, thus improving data throughput and overall performance; 3) The NOC uses packet switching as the basic communication technology and uses global asynchronous - local synchronous. The NIU (NOC Interface Units) is the NOC interface unit that provides transaction layer interconnection services between IP cores.

[0117] In terms of architecture, each switching module is interconnected with neighboring computing nodes through its multiple interfaces, and a 3D-Torus interconnection network is completely constructed by the switching modules within the computing nodes, without relying on external switching devices.

[0118] In summary, according to the embodiments of the present disclosure, the general-purpose graphics processing system integrated with a switching module is no longer a simple end device, but has networking and data exchange capabilities, enabling networking not to solely rely on external switches and routers, thereby improving the computing performance of the general-purpose graphics processing system.

[0119] In a further embodiment, the general-purpose graphics processing system integrates a lightweight RDMA protocol and a custom proprietary protocol, which improves efficiency, provides strong expansion capabilities, reduces the overall resource overhead of the system, and improves the system integration level.

[0120] It should be emphasized that the embodiments of the present disclosure do not impose any restrictions on the manufacturing process of the general-purpose graphics processing system. For example, the switching module and other devices can be designed and manufactured separately to form independent components, and then encapsulated together through an integration process. The switching module and other devices can also be integrally formed into an independent component. Furthermore, in the manufacturing process, the computing unit, cache, memory controller, switching module, and connection unit can be implemented on the same wafer, and further can be implemented on one or more dies. For example, the computing unit, cache, memory controller, and connection unit are implemented on one die, while the switching module is implemented on another die.

[0121] Data structure in the embodiments of the present disclosure

[0122] As shown in Reference Table 1, the interconnection information defines the interconnection relationships between multiple general graphics processing systems. The local application identifier is the identifier of the application program that initiates a data operation request, and the target application identifier is the identifier of the application program that receives the data operation request. Both the local application identifier and the target application identifier include the identifier of the graphics processing unit to which the application program belongs. The interface is the interface of the graphics processing unit to which the application program belongs.

[0123] Table 1 Interconnection Information Table

[0124]

[0125] Method implemented in the above general graphics processing system

[0126] Figure 11 The flowchart shows the implementation of communication for a general graphics processing system according to an embodiment of the present disclosure. The flowchart includes steps S100 and S300. Steps S100 and S300 can be implemented in a switching module as shown in Figure 5 shown.

[0127] In step S100, the identifier of the target to be accessed and the source address of the data to be written are received.

[0128] In step S200, a first interface among multiple interfaces is determined according to the identifier of the target to be accessed and the pre-stored interconnection information.

[0129] In step S300, the data to be written is read from the cache according to the source address, and the data to be written is sent via the first interface.

[0130] In some implementation manners, the data to be written and the target address are received via a second interface among multiple interfaces of the general graphics processing system, and the data to be written is written to the target address.

[0131] In some implementation manners, before sending the data to be written via the first interface, a specific communication protocol is selected from multiple different communication protocols, and then the data to be written is encoded according to the specific communication protocol, so as to send the encoded data via the first interface.

[0132] In some implementation manners, the multiple communication protocols include the RDMA RoCEv2 communication protocol and a custom proprietary protocol. The RDMA RoCEv2 communication protocol realizes end-to-end data transmission at the transport layer, while the custom proprietary protocol directly encodes the first data after the MAC header, that is, the overhead of the TCP / UDP / IP headers is omitted.

[0133] In some implementations, the target address is one of the following devices: The source address and the target address are specific storage addresses of the source application and the target application in their respective application memory spaces.

[0134] Commercial value of the embodiments of the present disclosure

[0135] Traditional general-purpose graphics processing systems do not have networking and switching capabilities. The embodiments of the present disclosure provide a general-purpose graphics processing system with networking and switching capabilities, so that a distributed system for model training can be constructed by using the general-purpose graphics processing system without external switches and routers. Therefore, the embodiments of the present disclosure have application prospects and commercial value.

[0136] Those skilled in the art can understand that the present disclosure can be implemented as a system, a method, and a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode), and can also be implemented in the form of a combination of software and hardware. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code.

[0137] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium is, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of the computer-readable storage medium include: an electrical connection of one or more specific wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory, a magnetic memory, or any suitable combination of the above. In this article, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by, or in combination with, a processing unit, device, or component.

[0138] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any other suitable combination. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program used by, or in combination with, an instruction system, device, or component.

[0139] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., and any suitable combination of the foregoing.

[0140] The computer program code for performing the embodiments of the present disclosure can be written in one or more programming languages or combinations. The programming languages include object-oriented programming languages such as JAVA, C++, and can also include conventional procedural programming languages such as C. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0141] The foregoing are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A general graphics processing system, comprising: a computing unit; a cache; a memory controller coupled to the cache; a switching module including a plurality of internally coupled interfaces, a switching unit, and a transmission engine, the switching unit and the transmission engine being configured to receive an identifier of a target to be accessed and a source address of a first data to be written sent by the computing unit or the memory controller, determine a first interface among the plurality of interfaces according to the identifier of the target to be accessed and pre-stored interconnection information, read the first data to be written from the cache according to the source address, and after encoding the first data to be written, send the first data to be written to a device coupled to the general graphics processing system via the first interface, the device including the general graphics processing system; a connection unit configured to couple the computing unit, the memory controller, the cache, and the switching module.

2. The general graphics processing system according to claim 1, wherein, the identifier of the target to be accessed and the source address are from a data operation request submitted by the computing unit or the memory controller.

3. The general graphics processing system according to claim 1, wherein, the switching module further includes: a transmission engine configured to encode the first data to be written according to a specified transport layer / network layer communication protocol; a switching unit configured to determine the first interface according to the identifier of the target to be accessed and the interconnection information, continue to encode the first data to be written according to an Ethernet communication protocol and a physical layer protocol, and send the encoded data via the first interface.

4. The general graphics processing system according to claim 3, wherein, the transmission engine supports a plurality of transport layer / network layer communication protocols, and the transmission engine selects the specified transport layer / network layer communication protocol from the plurality of transport layer / network layer communication protocols to encode the first data to be written.

5. The general graphics processing system according to claim 4, wherein, the transmission engine includes: a RoCE protocol processing module configured to encode the first data to be written based on the RoCEv2 communication protocol; a proprietary protocol processing module configured to encode the first data to be written based on a proprietary protocol.

6. The general graphics processing system according to claim 5, wherein, the RoCE protocol processing module includes a TOE unit configured to encode the first data to be written according to the IP / TCP / UDP protocol.

7. The general graphics processing system according to claim 5, wherein, the RoCE protocol processing module includes an IB unit configured to establish end-to-end data transmission based on the IB protocol.

8. The general graphics processing system according to claim 5, wherein, the RoCE protocol processing module includes a predicate interface.

9. The general graphics processing system according to claim 5, wherein, the proprietary protocol processing module includes a proprietary protocol unit and a dedicated protocol driver interface, the proprietary protocol unit is configured to encode and decode data according to a proprietary protocol, and the dedicated protocol driver interface is configured to provide a driver and a hardware interface.

10. The general graphics processing system according to claim 1, wherein, the identifier of the target to be accessed includes the identifier of the target graphics processing unit connected via the first interface and the identifier of the target application program.

11. The general graphics processing system according to claim 1, wherein, the exchange module is further configured to receive second data to be written via a second interface, determine a target address according to the identifier of the target to be accessed, and write the second data to be written to the target address.

12. The general graphics processing system according to claim 11, wherein, the source address and the target address are respectively specific storage addresses of the source application program and the target application program in their respective application memory spaces.

13. The general graphics processing system according to claim 1, wherein, the exchange module is integrated into a network card processor.

14. The general graphics processing system according to claim 10, wherein, the target graphics processing unit and the general graphics processing system are located in different computing nodes.

15. The general graphics processing system according to claim 1, wherein, the multiple interfaces are Ethernet interfaces.

16. A computing device, comprising a plurality of computing nodes, the computing nodes including a coupled memory, a general-purpose processor, and the general graphics processing system according to any one of claims 1 to 15, wherein, the computing node is coupled to the general graphics processing system of at least one other computing node through its own general graphics processing system.

17. The computing device according to claim 16, wherein, the computing nodes are encapsulated in the same silicon chip, and a plurality of the computing nodes are integrated together through a printed circuit board.

18. The computing device according to claim 16, wherein, the computing device is used to execute the training task of the deep learning model.

19. A distributed system, comprising a plurality of computing devices according to any one of claims 16 to 18, and data transmission is performed among the plurality of computing devices through an external bridging device.

20. A distributed system, comprising a plurality of computing nodes, the computing nodes including a coupled memory, a general-purpose processor, and the general graphics processing system according to any one of claims 1 to 15, wherein, the computing node is coupled to the general graphics processing system of at least one adjacent other computing node through its own general graphics processing system, and a 3D-Torus interconnection network is formed.

21. A cloud server, comprising the computing device according to any one of claims 16 to 18.

22. A method implemented in a general graphics processing system, comprising: receiving the identifier of the target to be accessed and the source address of the data to be written sent by a computing unit or a storage controller; determining a first interface among a plurality of interfaces according to the identifier of the target to be accessed and pre-stored interconnection information; reading the data to be written from the cache according to the source address, encoding the data to be written, and sending the data to be written to a device coupled to the general graphics processing system via the first interface, the device including a general graphics processing system.

Citation Information

Patent Citations

  • Data storage method and network interface card

    CN104063344A

  • Data transmission method and data transmission device

    CN111159075A