Interconnection and exchange architecture and method between u2u chips for torus and fat tree fusion calculation
By integrating torus and fat tree topology in the inter-chip interconnection switching architecture, flexible switching of data transmission paths between chips is achieved, and the problem of difficult to meet the needs of multiple computing tasks in the existing technology is solved, and the performance and efficiency of data transmission are improved.
Patent Information
- Application Number
- CN202510244124.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, topological connections between chips are difficult to meet a variety of computing task requirements, especially in the case of latency-sensitive real-time computing tasks and high bandwidth requirements.
A u2u inter-chip interconnection switching architecture for torus and fat tree fusion computing is proposed. By switching between the first topology structure (torus ring grid topology) and the second topology structure (fat tree fat tree topology) to flexibly adjust the data transmission path between chips.
It realizes flexible switching of interconnection modes according to different task requirements, which not only meets the needs of delay-sensitive real-time computing tasks, but also provides high bandwidth support when the amount of data is huge, effectively improving the overall performance of data transmission.
Smart Images

Figure CN120151207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer communication technologies, and particularly to a u2u inter-chip interconnection and switching architecture, method, device, and medium for the integrated computing of torus and fat tree. Background Art
[0002] With the continuous growth of high-performance computing requirements, for chips undertaking computing task processing, their performance, flexibility, etc. are facing huge challenges. In the prior art, chips process computing tasks according to a certain topological connection method. When facing different computing task requirements, this topological connection method has certain limitations and is difficult to meet diverse task requirements. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide a u2u inter-chip interconnection and switching architecture, method, device, and medium for the integrated computing of torus and fat tree that overcome the above problems or at least partially solve the above problems.
[0004] To solve the above problems, embodiments of the present invention disclose a u2u inter-chip interconnection and switching architecture for the integrated computing of torus and fat tree. The architecture includes multiple clusters; each cluster includes multiple chips networked and interconnected using a first topological structure; the multiple clusters are networked and connected using a second topological structure; the first topological structure includes a torus ring grid topological structure, and the second topological structure includes a fat tree topological structure.
[0005] The chip is configured to transmit data to be transmitted to a target chip through the communication protocol corresponding to the first topological structure, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topological structure.
[0006] Optionally, the chip is configured to determine an evaluation parameter of the data to be transmitted, and transmit the data to be transmitted to a target chip through the communication protocol corresponding to the first topological structure according to the evaluation parameter, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topological structure.
[0007] Optionally, the chip is configured to obtain a transmission parameter of the data to be transmitted and determine the evaluation parameter according to the transmission parameter.
[0008] Optionally, the cluster includes a hardware switch, and the chip includes a mode switching control unit. The hardware switch is respectively connected to the chip and the mode control unit;
[0009] The mode switching control unit is configured to control the hardware switch to switch the signal path between the first topology and the second topology.
[0010] Optionally, the chip further includes a cache. The mode switching control unit is further configured to store the data being transmitted by the chip into the cache when it is necessary to switch the signal path between the first topology and the second topology, and transmit the data stored in the cache to the target chip through the switched communication protocol when the signal path switching is completed.
[0011] Optionally, the chip further includes a communication protocol conversion unit, which is configured to convert the communication protocol of the data to be transmitted from the initial communication protocol to the communication protocol corresponding to the switched topology.
[0012] Optionally, the transmission parameters include at least one of a source address, a destination address, a data traffic value, and a transmission frequency.
[0013] Optionally, the evaluation parameters include at least one of a latency requirement, a bandwidth requirement, and a load value.
[0014] The present invention also discloses a method for data transmission between chips, which is applied to a u2u chip - to - chip interconnection and switching architecture for torus - and - fat - tree fusion computing. The architecture includes multiple clusters; each cluster includes multiple chips interconnected by using the first topology; the multiple clusters are connected by using the second topology. The method is used for the above - mentioned u2u chip - to - chip interconnection and switching architecture for torus - and - fat - tree fusion computing. The method includes:
[0015] Transmitting the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the first topology, or transmitting the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the second topology; the first topology includes a torus ring grid topology, and the second topology includes a fat tree topology.
[0016] Optionally, the step of transmitting the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the first topology, or transmitting the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the second topology includes:
[0017] Determining the evaluation parameters of the data to be transmitted;
[0018] Transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology according to the evaluation parameter, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology.
[0019] Optionally, determining the evaluation parameter of the data to be transmitted includes:
[0020] Obtain the transmission parameter of the data to be transmitted;
[0021] Determine the evaluation parameter according to the transmission parameter.
[0022] Optionally, the cluster includes a hardware switch, the chip includes a mode switching control unit, and the hardware switch is respectively connected to the chip and the mode control unit; the method further includes:
[0023] Control the hardware switch to switch the signal path between the first topology and the second topology through the mode switching control unit.
[0024] Optionally, the chip further includes a cache, and the method further includes:
[0025] When it is necessary to switch the signal path between the first topology and the second topology through the mode switching control unit, store the data being transmitted by the chip into the cache;
[0026] When the signal path switching is completed, transmit the data stored in the cache to the target chip through the switched communication protocol.
[0027] Optionally, the chip further includes a communication protocol conversion unit, and further includes:
[0028] Convert the communication protocol of the data to be transmitted from the initial communication protocol to the communication protocol corresponding to the switched topology through the communication protocol unit.
[0029] Optionally, the transmission parameter includes at least one of a source address, a destination address, a data traffic value, and a transmission frequency.
[0030] Optionally, the evaluation parameter includes at least one of a delay requirement, a bandwidth requirement, and a load value.
[0031] The present invention also discloses an electronic device, including: a processor, a memory, and a computer program stored on the memory and capable of running on the processor, and when the computer program is executed by the processor, the steps of the above-mentioned u2u chip - to - chip interconnection and exchange method for torus - fat tree fusion calculation are implemented.
[0032] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned u2u chip-to-chip interconnection and switching method for torus and fat tree integrated computing are implemented.
[0033] The embodiments of the present invention have the following advantages:
[0034] The present invention discloses a u2u chip-to-chip interconnection and switching architecture for torus and fat tree integrated computing. When the chip in the present invention receives data to be transmitted, it can transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology structure, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure. The present invention can flexibly switch the interconnection mode according to different task requirements, which not only meets the needs of real-time computing tasks sensitive to latency, but also provides high-bandwidth support in the case of a large amount of data, effectively improving the overall performance of processing the data to be transmitted and improving the data transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a structural block diagram of a u2u chip-to-chip interconnection and switching architecture for torus and fat tree integrated computing provided by an embodiment of the present invention;
[0036] Figure 2 is a structural block diagram of a cluster provided by an embodiment of the present invention;
[0037] Figure 3 is a structural block diagram of a chip provided by an embodiment of the present invention;
[0038] Figure 4 is a structural block diagram of a second topology structure provided by an embodiment of the present invention;
[0039] Figure 5 is a structural block diagram of a u2u chip-to-chip interconnection and switching method for torus and fat tree integrated computing provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] With the continuous growth of high-performance computing requirements, such as application scenarios like large-scale scientific computing and deep learning training, extremely high requirements are imposed on the performance, scalability, and reliability of computing systems. Traditional computing interconnection network topologies cannot meet different task requirements.
[0042] One of the core concepts of the embodiments of the present invention is that the chip in the present invention can transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology structure, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure, and can flexibly switch the interconnection mode according to different task requirements, which not only meets the requirements of real-time computing tasks sensitive to latency, but also provides high-bandwidth support in the case of a large amount of data, effectively improving the overall performance of processing the data to be transmitted.
[0043] Referring to Figure 1 , a structural block diagram of a u2u chip-to-chip interconnection and switching architecture for torus and fat tree fusion computing provided by an embodiment of the present invention is shown. The architecture may include multiple clusters 101; each cluster 101 includes multiple chips networked and interconnected using a first topology structure 102; multiple clusters are networked and connected using a second topology structure 103; the first topology structure includes a torus ring grid topology structure, and the second topology structure includes a fat tree topology structure;
[0044] A chip is configured to transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology structure, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure.
[0045] In the embodiments of the present invention, each node in the torus ring grid topology structure is usually connected to multiple adjacent nodes, forming multiple data transmission paths. When a certain link fails, the data can be transmitted through other paths, thereby improving the reliability and fault tolerance of the network. For example, in a two-dimensional torus ring grid, a node can communicate with adjacent nodes through horizontal and vertical links. Even if a certain link is disconnected, it can still reach the target node through other links.
[0046] Since the distance between nodes is relatively short and there are multiple parallel paths, the transmission delay of data in the network is low. For some application scenarios with high real-time requirements, such as data interaction in high-performance computing, the torus ring grid topology structure can provide better performance.
[0047] The torus ring grid topology structure can be expanded by adding nodes and links. When expanding, only new nodes need to be connected to the existing grid, which has little impact on the original network. Therefore, it has good scalability.
[0048] The core feature of the fat tree topology structure is that its link bandwidth increases as the network scale increases. In a large-scale network, it can provide a high aggregation bandwidth to meet the needs of a large amount of data parallel transmission. For example, in a data center network, the fat tree topology structure can easily handle high-speed data exchange between multiple servers.
[0049] The hierarchical structure of the fat-tree topology makes the routing algorithm relatively simple. Data can be forwarded according to fixed rules based on the destination address, without the need for complex path selection algorithms, reducing routing overhead and computational complexity. The fat-tree topology adopts a modular design, and each module can be independently constructed and managed, making the construction and maintenance of the network more convenient. At the same time, it also improves the scalability and flexibility of the system.
[0050] In this interconnection and switching architecture, multiple chips are first networked and interconnected through the first topology to form a cluster. For example, cluster 1011 can include 1011 1 ...1011 n and other multiple chips. 1011 1 ...1011 n and other multiple chips are connected through the first topology 102. Cluster 101n can include 101n 1 ...101n n and other multiple chips. 101n 1 ...101n n and other multiple chips are connected through the first topology 102. The first topology provides a specific way and rules for the connection between chips. Each chip conducts data transmission and interaction within its respective cluster according to the communication protocol corresponding to the first topology.
[0051] After forming multiple clusters networked by the first topology, these clusters are networked and connected through the second topology 103. The communication protocol corresponding to the second topology defines the rules and ways of data transmission between clusters, realizing the interconnection and interoperability between different clusters. Each cluster, as a node in the Fat Tree structure, is connected to each other according to its specific connection method. The communication protocol between clusters ensures the correct transmission of data between the chips of different clusters, including data routing selection, flow control, etc.
[0052] When a chip has data to transmit, it will carry key data information in the data to be transmitted. These information can include the identifier of the destination chip, the type of data, transmission priority, data source address, destination address, demand information, etc. The control logic inside the chip will parse these data information to determine the transmission direction of the data and the appropriate topology.
[0053] Once the topology is determined, the chip will process and transmit the data according to the corresponding communication protocol. During the transmission process, ensure that the data format, transmission rate, etc. meet the requirements of the communication protocol of the selected topology. If it is cross-topology transmission, protocol conversion is also required to ensure the smooth transmission of data.
[0054] When data is transmitted from one cluster to another through a second topology, before entering the second topology network, the chip needs to convert the data from the protocol format of the first topology to the protocol format of the second topology to adapt to different transmission rules. The present invention can quickly convert the data format of the original protocol to the data format of the target protocol when data enters from one interconnection mode to another, ensuring the smooth transmission of data and avoiding data errors.
[0055] The present invention discloses a u2u chip - to - chip interconnection and switching architecture for the fusion of torus and fat tree computing. The chips in the present invention can transmit data to be transmitted to a target chip through the communication protocol corresponding to the first topology, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology, and can flexibly switch the interconnection mode according to different task requirements, which not only meets the needs of real - time computing tasks sensitive to latency but also provides high - bandwidth support in the case of a large amount of data, effectively improving the overall performance of processing the data to be transmitted.
[0056] In an embodiment of the present invention, a chip is configured to determine evaluation parameters of data to be transmitted and transmit the data to be transmitted to a target chip through the communication protocol corresponding to the first topology according to the evaluation parameters, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology.
[0057] In an embodiment of the present invention, the evaluation parameters are quantitative indicators for measuring the characteristics of the data to be transmitted and the current network environment. The chip decides whether to use the communication protocol corresponding to the first topology or the communication protocol corresponding to the second topology for data transmission based on these evaluation parameters.
[0058] The chip determines whether to transmit data through the communication protocol corresponding to the first topology or the communication protocol corresponding to the second topology according to the determined evaluation parameters. Generally speaking, if the evaluation parameters indicate that it is suitable for fast and low - latency transmission within the cluster, the first topology will be selected; if higher bandwidth and better scalability are required to handle a large amount of data or cross - cluster transmission, the second topology will be selected.
[0059] If the data to be transmitted is a real - time sensor data with extremely high latency requirements, and the current load within the cluster is low and the bandwidth is sufficient, the chip judges according to the evaluation parameters and selects the communication protocol corresponding to the first topology for transmission, taking advantage of the low - latency characteristic of the first topology to ensure that the data can quickly reach the target chip within the same cluster.
[0060] If the data to be transmitted is a large dataset with very high bandwidth requirements and the target chip is located in another cluster, the chip will consider that the first topology may not be able to provide sufficient bandwidth, while the second topology has higher bandwidth and better scalability, and thus will choose the communication protocol corresponding to the second topology for transmission.
[0061] In an embodiment of the present invention, a chip is used to obtain the transmission parameters of the data to be transmitted and determine the evaluation parameters according to the transmission parameters.
[0062] In the embodiments of the present invention, the transmission parameters refer to the basic parameters directly related to the data to be transmitted, which describe the characteristics of the data itself and the transmission requirements. These parameters are the direct manifestation of the data's own attributes and the requirements of the application scenario, and are usually determined when the data is generated or ready to be transmitted. Common transmission parameters include the size of the data volume, data type, transmission time requirement, data source, etc.
[0063] Different types of data and transmission time requirements have different sensitivities to latency. For example, data with high real-time requirements (such as real-time video and audio streams) is usually very sensitive to latency and requires a low-latency transmission environment. The chip will determine the latency requirement according to the data type and transmission time requirement. If it is real-time video call data and requires no obvious lag in the picture, the chip will determine its latency requirement to be low latency, perhaps requiring the latency to be controlled within dozens of milliseconds; while for non-real-time data transmission tasks, such as file downloads, the latency requirement is relatively low.
[0064] The size of the data volume and the transmission time requirement directly affect the bandwidth requirement. The chip will calculate the required bandwidth according to the size of the data volume and the allowed transmission time. For example, if a file of size 100MB needs to be transmitted within 10 seconds, the chip will calculate that the required bandwidth is at least 80Mbps (100MB = 100×8×1024×1024 bits, and the bandwidth for transmission in 10 seconds is 100×8×1024×1024÷10 bits per second, approximately 80Mbps). At the same time, the data type will also affect the bandwidth requirement. For example, high-definition video data usually requires higher bandwidth than text data.
[0065] If the data volume to be transmitted is large while the remaining available bandwidth of the current network link is small, then the chip will consider that this transmission will significantly increase the network load and the load value will increase accordingly. For example, the total bandwidth of a certain link in the network is 1Gbps, and 800Mbps has been used currently. At this time, there is a 500Mbps data to be transmitted. The chip will judge that this transmission will make the load of this link close to saturation, and the load value will be evaluated as high.
[0066] In an embodiment of the present invention, the cluster includes a hardware switch, and the chip includes a mode switching control unit. The hardware switch is respectively connected to the chip and the mode control unit; the mode switching control unit is used to control the hardware switch to switch the signal path between the first topology and the second topology.
[0067] In the embodiment of the present invention, as Figure 2 , a structural block diagram of a cluster 1011 provided by an embodiment of the present invention is shown. The cluster 1011 may include chips 10111, 10112, 10113, 10114, 10115, 10116, 10117, 10118, 10119. The chips are connected through u2u interfaces, that is, the chips are connected using the first topology. The cluster 1011 may include a hardware switch A1, and the chip 10111 may include a mode control unit 101111. In one example, when the chip 10111 needs to send data to be transmitted to the chip 10119, if the current second topology is adopted and it is determined that the current data to be transmitted is more suitable for transmission using the first topology, the mode control unit may control the hardware switch A1 to implement switching of the signal path between the first topology and the second topology, that is, to open the torus topology signal path inside the cluster of the chip 10111 and the chip 10119, and close the fat tree topology signal path outside the cluster.
[0068] In an embodiment of the present invention, the chip further includes a cache. The mode switching control unit is further used to store the data being transmitted by the chip in the cache when it is necessary to switch the signal path between the first topology and the second topology, and when the signal path switching is completed, transmit the data stored in the cache to the target chip through the switched communication protocol.
[0069] In the embodiment of the present invention, as Figure 3 , an architecture block diagram of a chip 10111 provided by an embodiment of the present invention is shown. The chip may further include a cache 101112. The mode switching control unit 101111 in the chip continuously monitors information such as the operating state of the system and related instructions, and determines whether it is necessary to switch the signal path between the first topology and the second topology. Once it is determined that a topology switch is required, the mode switching control unit 101111 issues an instruction to store the data being transmitted by the chip in the cache 101112, which is to prevent data loss or errors during the signal path switching process. The cache has the characteristics of high-speed reading and writing and can quickly receive and store data. For example, when a chip that is performing video decoding needs to switch the topology, it will temporarily store the video data segment being processed in the cache.
[0070] After the signal path is switched, the chip will check whether the switch is completed through specific mechanisms, which may include self-checking the new topology, checking whether the relevant status registers are set correctly, or verifying whether the signal path is unobstructed by sending some test signals. Only after confirming that the signal path switch is completely successful will the next data transmission operation be carried out.
[0071] After the signal path switch is completed and confirmed to be error-free, the chip will read the previously stored data from the cache. Then, according to the switched communication protocol, the data will be processed such as being encapsulated and converted accordingly to make it conform to the new communication format and requirements. Finally, the processed data will be transmitted to the target chip. For example, if the switched communication protocol adopts a new coding method and data frame format, the chip will repackage the data in the cache according to these rules and then send it to the target chip through the corresponding physical interface. The chip can ensure the integrity and accuracy of the data during the topology switch and achieve normal data communication and system function operation under different topologies.
[0072] In an embodiment of the present invention, the chip further includes a communication protocol conversion unit, which is used to convert the communication protocol of the data to be transmitted from the initial communication protocol to the communication protocol corresponding to the switched topology.
[0073] In the embodiment of the present invention, such as Figure 3 , the chip may further include a protocol conversion unit 101113. After the signal path switch is completed and confirmed to be error-free, the communication protocol conversion unit 101113 can read the previously stored data to be transmitted from the cache and analyze the initial communication protocol followed by these data.
[0074] According to the target communication protocol corresponding to the switched second topology, the communication protocol conversion unit 101113 performs a series of processing on the data, such as modifying the data frame format, adjusting the coding method, adding or deleting specific control information, etc., to convert the data into a format conforming to the target communication protocol.
[0075] The chip sends the converted data to the target chip through the switched signal path according to the communication protocol corresponding to the second topology.
[0076] During the transmission process, some error detection and correction mechanisms can be adopted to ensure that the data reaches the target chip accurately and without error. After the data is successfully transmitted to the target chip, the chip system will update the relevant status information, indicating that the topology switch and the data transmission process are completed. Then, according to the new topology and communication protocol, subsequent data interaction and task processing can continue. Through this step, the chip can effectively handle the storage, protocol conversion, and transmission of data during the topology switch, ensuring the stable operation of the system and the correct communication of data.
[0077] In an embodiment of the present invention, the transmission parameters include at least one of a source address, a destination address, a data traffic value, and a transmission frequency.
[0078] In an embodiment of the present invention, the source address refers to the starting position of the data to be transmitted or the address information of the source device. For example, in network communication, the source address may be the IP address of a certain computer or server. The chip obtains this address information through relevant network interfaces and protocols to know where the data is sent from.
[0079] The destination address refers to the target position or the address of the device to which the data is to be transmitted. Taking network communication as an example again, the destination address can be the IP address of another computer, server, router, etc. The chip uses network configuration and relevant mechanisms to obtain this destination address to clarify the final destination of the data.
[0080] In an embodiment of the present invention, the source address and the destination address refer to the address of the sending chip and the address of the receiving chip.
[0081] The data traffic value refers to the quantity or scale of the data to be transmitted, which can be achieved by counting the size of the data. For example, the number of bytes or data packets of the data is calculated within a certain time interval. The data traffic value reflects the scale of data transmission and is of great significance for subsequent evaluation and processing.
[0082] The chip will also obtain the frequency of data transmission, that is, the transmission frequency, which can be determined by recording the number of data transmissions within a certain time. Different application scenarios may have different requirements for the transmission frequency. The chip's acquisition of this parameter helps to more accurately understand the characteristics of data transmission.
[0083] In an embodiment of the present invention, the evaluation parameters include at least one of a delay requirement, a bandwidth requirement, and a load value.
[0084] In the embodiments of the present invention, the latency requirement refers to the maximum time interval allowed from when data is generated in the source chip and starts to be transmitted until it arrives at the target chip completely and correctly. It reflects the requirement for the timeliness of data transmission in a specific application scenario. Different applications have significantly different sensitivities to latency. For example, applications such as real-time audio and video communication and high-frequency trading are extremely sensitive to latency and require extremely low latency to ensure service quality; while applications such as file downloading and bulk data storage have relatively loose requirements for latency.
[0085] The latency requirement can be determined based on factors such as the data traffic value and transmission frequency in the transmission parameters. If the data traffic is large and the transmission frequency is high, in order to ensure the real-time and effectiveness of the data, the requirement for latency will be relatively strict, that is, the allowed data transmission latency time is short. For example, in real-time video stream transmission, in order to ensure the smoothness of the video, the requirement for latency is very low, and the chip will determine the appropriate latency requirement according to parameters such as the current data traffic and transmission frequency.
[0086] The bandwidth requirement refers to the network transmission capacity required to ensure that data can be transmitted at the expected rate and efficiency during data transmission, usually measured by the amount of data that can be transmitted per unit time (such as bits per second, bps). Different types of data transmission tasks have very different requirements for bandwidth. For example, high-definition video streams, migration of large data sets, etc. require high bandwidth to ensure smooth data transmission; while ordinary text information transmission, simple instruction interaction, etc. have relatively low requirements for bandwidth.
[0087] The data traffic value in the transmission parameters is a key factor in determining the bandwidth requirement. The larger the data traffic, the higher the bandwidth required to transmit data quickly and efficiently. The chip will comprehensively evaluate the bandwidth size required to meet data transmission based on the obtained data traffic value and transmission frequency information. For example, when transmitting large data files, a larger bandwidth is required to ensure the transmission speed, and the chip will determine the appropriate bandwidth requirement according to the specific transmission parameters.
[0088] The load value is used to reflect the working load level of the chip itself and the entire network environment. It can be measured from multiple dimensions: Chip load: mainly refers to the occupancy of internal computing resources of the chip (such as CPU cores, GPU units, etc.) and the usage ratio of storage resources such as caches and memories. When the chip load is too high, its ability to process new data will be limited, which may lead to an increase in data processing and transmission latency.
[0089] Network link load: refers to the bandwidth usage of each communication link in the network. If the load of a certain link is too high, it means that the bandwidth of this link has approached or reached its maximum carrying capacity, and problems such as congestion and packet loss may occur when data is transmitted on this link, thus affecting the performance of the entire data transmission.
[0090] Transmission parameters such as the source address, destination address, and data traffic value will affect the load conditions of devices (such as routers, switches, etc.) on the transmission path. The chip will evaluate the load values of each node during the transmission process based on these parameters. For example, if the data traffic between the source address and the destination address is large, the intermediate devices will bear a large load. The chip will determine the appropriate load values based on this information to reasonably allocate resources and perform data transmission scheduling.
[0091] In one example, if the evaluation parameter includes the network link load, the network link loads of each communication link under the first topology structure and the second topology structure can be determined, and the communication link with the minimum network link load is taken as the link for the chip to transmit data.
[0092] In one example, if the evaluation parameter includes the delay requirement, the source chip node will send communication requests to the target chip node simultaneously in two topology architectures. The time when the requests are returned is used to determine the routing path, and the path with the lowest delay is selected for data forwarding. If the delays of all paths under the torus topology are greater than a certain fixed threshold, it indicates that the network traffic under the torus topology is very large at this time, and the data forwarding is immediately switched to the Fat tree mode.
[0093] Such as Figure 4 , which shows the architecture block diagram of another u2u chip - to - chip interconnection and switching architecture 10 for the integration of torus and fat tree computing provided by the embodiment of the present invention. The second topology structure 103 may include a core layer, an aggregation layer, and an access layer. The clusters are respectively connected to the leaf switches deployed in the corresponding core layer. For example, cluster 1011 is connected to leaf switch leaf1,... cluster 1018 is connected to leaf switch leaf8. The switches deployed in the core layer are respectively connected to multiple backbone switches deployed in the aggregation layer. For example, leaf1 is respectively connected to spine1, spine2, spine3, spine4,... leaf3 is respectively connected to spine1, spine2, spine3, spine4; the backbone switches deployed in the aggregation layer are respectively connected to the core servers in the core layer. For example, spine is respectively connected to computing core1, computing core2,... spine4 is respectively connected to computing core1, computing core2.
[0094] In some non-real-time scenarios, such as simple graphics display, 2D image processing, low-resolution video playback, offline data analysis, and batch data processing, the data traffic is relatively small in these scenarios. Using the torus topology combined with the U2U interface between chips can meet the communication requirements. However, when performing artificial intelligence deep learning training and updating model parameters during gradient synchronization, the data traffic between nodes increases instantaneously. At this time, it is necessary to determine whether to perform topology switching based on whether the actual torus topology can meet the minimum latency of data transmission between nodes.
[0095] The present invention discloses a U2U inter-chip interconnection and switching architecture for the fusion of torus and fat tree computing. The chips in the present invention can transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology structure, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure, and can flexibly switch the interconnection mode according to different task requirements, which not only meets the requirements of real-time computing tasks sensitive to latency, but also provides high-bandwidth support in the case of huge data volume, effectively improving the overall performance of processing the data to be transmitted.
[0096] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
[0097] As Figure 5 , it shows a U2U inter-chip interconnection and switching method provided by an embodiment of the present invention. This method is applied to the above-mentioned U2U inter-chip interconnection and switching architecture for the fusion of torus and fat tree computing. The architecture includes multiple clusters; the cluster includes multiple chips interconnected by using the first topology structure; the multiple clusters are connected by using the second topology structure. The first topology structure includes a torus ring grid topology structure, and the second topology structure includes a fat tree topology structure. This method may include:
[0098] Step 201, transmit the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the first topology structure, or transmit the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the second topology structure.
[0099] The present invention discloses a method for data transmission between chips. In the present invention, a chip can transmit data to be transmitted to a target chip through a communication protocol corresponding to a first topology structure, or transmit the data to be transmitted to the target chip through a communication protocol corresponding to a second topology structure, and can flexibly switch the interconnection mode according to different task requirements, which not only meets the requirements of real-time computing tasks sensitive to latency, but also provides high-bandwidth support in the case of a large amount of data, effectively improving the overall performance of processing the data to be transmitted.
[0100] In an embodiment of the present invention, transmitting the data to be transmitted in the chip to the target chip through a communication protocol corresponding to a first topology structure, or transmitting the data to be transmitted in the chip to the target chip through a communication protocol corresponding to a second topology structure includes: determining an evaluation parameter of the data to be transmitted; and transmitting the data to be transmitted to the target chip through a communication protocol corresponding to a first topology structure or transmitting the data to be transmitted to the target chip through a communication protocol corresponding to a second topology structure according to the evaluation parameter.
[0101] In an embodiment of the present invention, determining an evaluation parameter of the data to be transmitted includes: obtaining a transmission parameter of the data to be transmitted; and determining the evaluation parameter according to the transmission parameter.
[0102] In an embodiment of the present invention, the cluster includes a hardware switch, and the chip includes a mode switching control unit, and the hardware switch is respectively connected to the chip and the mode control unit; the method further includes: controlling, by the mode switching control unit, the hardware switch to switch a signal path between a first topology structure and a second topology structure.
[0103] In an embodiment of the present invention, the chip further includes a cache, and the method further includes: when it is necessary to switch a signal path between a first topology structure and a second topology structure, storing, by the mode switching control unit, the data being transmitted by the chip into the cache; and when the signal path switching is completed, transmitting the data stored in the cache to the target chip through the switched communication protocol.
[0104] In an embodiment of the present invention, the chip further includes a communication protocol conversion unit, and further includes: converting, by the communication protocol conversion unit, the communication protocol of the data to be transmitted from an initial communication protocol to a communication protocol corresponding to a switched topology structure.
[0105] In an embodiment of the present invention, the transmission parameter includes at least one of a source address, a destination address, a data traffic value, and a transmission frequency.
[0106] In an embodiment of the present invention, the evaluation parameter includes at least one of a latency requirement, a bandwidth requirement, and a load value.
[0107] The present invention discloses a u2u inter-chip interconnection and switching method for the integrated computing of torus and fat tree. The chips in the present invention can transmit data to be transmitted to a target chip through the communication protocol corresponding to the first topology structure, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure, and can flexibly switch the interconnection mode according to different task requirements, which not only meets the requirements of real-time computing tasks sensitive to latency, but also provides high-bandwidth support in the case of a huge amount of data, effectively improving the overall performance of processing the data to be transmitted.
[0108] It should be noted that, for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0109] The embodiments of the present invention also provide an electronic device, including:
[0110] It includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it realizes each process of the above-mentioned u2u inter-chip interconnection and switching method embodiment for the integrated computing of torus and fat tree between chips, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0111] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it realizes each process of the above-mentioned u2u inter-chip interconnection and switching method embodiment for the integrated computing of torus and fat tree, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0112] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0113] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, apparatuses, or computer program products. Therefore, the embodiments of the present invention can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0114] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0117] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0118] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0119] The above has introduced in detail a u2u chip - to - chip interconnection and switching architecture for torus - and - fat - tree - integrated computing, a u2u chip - to - chip interconnection and switching method, device and medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A u2u chip interconnection switching architecture for torus and fat tree fusion computing, characterized in that: The architecture includes multiple clusters; the clusters include multiple chips that are networked and interconnected using a first topology structure; the multiple clusters are networked and connected using a second topology structure; the first topology structure includes a torus ring mesh topology structure, and the second topology structure includes a fat tree topology structure; The chip is used to transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology structure, or to transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure.
2. The architecture according to claim 1, characterized in that The chip is used to determine evaluation parameters of the data to be transmitted, and transmit the data to be transmitted to the target chip through the communication protocol corresponding to the first topology structure according to the evaluation parameters, or transmit the data to be transmitted to the target chip through the communication protocol corresponding to the second topology structure.
3. The architecture according to claim 2, characterized in that: The chip is used to obtain transmission parameters of the data to be transmitted, and determine the evaluation parameters according to the transmission parameters.
4. The architecture according to claim 1, characterized in that: The cluster includes a hardware switch, the chip includes a mode switching control unit, and the hardware switch is connected to the chip and the mode control unit respectively; The mode switching control unit is used to control the hardware switch to switch the signal path between the first topology structure and the second topology structure.
5. The architecture according to claim 4, characterized in that The chip also includes a cache, and the mode switching control unit is also used to store the data being transmitted by the chip in the cache when it is necessary to switch the signal path between the first topology and the second topology, and when the signal path switching is completed, transmit the data stored in the cache to the target chip via the switched communication protocol.
6. The architecture according to claim 4, characterized in that: The chip further includes a communication protocol conversion unit, which is used to convert the communication protocol of the data to be transmitted from an initial communication protocol to a communication protocol corresponding to the switched topology structure.
7. The architecture according to claim 3, characterized in that: The transmission parameters include at least one of a source address, a destination address, a data flow value, and a transmission frequency.
8. The architecture according to claim 2, characterized in that: The evaluation parameter includes at least one of a delay requirement, a bandwidth requirement, and a load value.
9. A u2u chip interconnection and switching method for torus and fat tree fusion computing, characterized in that: The method is applied to a U2U chip interconnection exchange architecture for torus and fat tree fusion computing, the architecture comprising multiple clusters; the clusters comprising multiple chips interconnected by a first topology structure; the multiple clusters are networked and connected by a second topology structure, the first topology structure comprising a torus ring mesh topology structure, and the second topology structure comprising a fat tree topology structure; The method comprises: The data to be transmitted in the chip is transmitted to the target chip through the communication protocol corresponding to the first topology structure, or the data to be transmitted in the chip is transmitted to the target chip through the communication protocol corresponding to the second topology structure.
10. The method according to claim 9, characterized in that The method of transmitting the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the first topology structure, or transmitting the data to be transmitted in the chip to the target chip through the communication protocol corresponding to the second topology structure, comprises: Determining evaluation parameters of the data to be transmitted; The data to be transmitted is transmitted to the target chip through the communication protocol corresponding to the first topology structure according to the evaluation parameter, or the data to be transmitted is transmitted to the target chip through the communication protocol corresponding to the second topology structure.
11. The method according to claim 10, characterized in that The determining of the evaluation parameters of the data to be transmitted includes: Acquire transmission parameters of the data to be transmitted; The evaluation parameter is determined based on the transmission parameter.
12. The method according to claim 9, characterized in that The cluster includes a hardware switch, the chip includes a mode switching control unit, and the hardware switch is connected to the chip and the mode control unit respectively; the method further includes: The mode switching control unit controls the hardware switch to switch a signal path between the first topology structure and the second topology structure.
13. The method according to claim 12, characterized in that The chip also includes a cache, and the method further includes: When it is necessary to switch the signal path between the first topology structure and the second topology structure, the mode switching control unit stores the data being transmitted by the chip into the cache; When the signal path switching is completed, the data stored in the cache is transmitted to the target chip via the switched communication protocol.
14. The method according to claim 13, characterized in that The chip also includes a communication protocol conversion unit, and further includes: The communication protocol unit converts the communication protocol of the data to be transmitted from the initial communication protocol to the communication protocol corresponding to the switched topology structure.
15. The method according to claim 10, characterized in that The transmission parameters include at least one of a source address, a destination address, a data flow value, and a transmission frequency.
16. The method according to claim 11, characterized in that The evaluation parameter includes at least one of a delay requirement, a bandwidth requirement, and a load value.
17. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the u2u chip interconnection and switching method for torus and fat tree fusion computing as described in any one of claims 9 to 16 are implemented.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the U2U chip interconnection and switching method for torus and fattree fusion computing as described in any one of claims 9 to 16 are implemented.