Wafer level communication structure, communication method and AI chip
By introducing cross-regional communication components and communication system into the wafer-level communication structure, the problem of low data transmission efficiency of cross-regional chips in the prior art is solved, and high bandwidth and efficient data transmission are achieved.
Patent Information
- Application Number
- CN202510104439.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, data transmission between cross-region chips adopts serial transmission method, which is easy to lose data, has a low transmission rate and a small bandwidth, making it difficult to meet the needs of artificial intelligence computing.
A wafer-level communication structure is provided, including at least two dies, a message switching network group and a cross-region communication component. Communication between two message switching network groups is realized through cross-region communication components, and a communication handshake is established between the communication components to control data transmission and reception to avoid data loss and improve transmission efficiency.
It realizes efficient data transmission, avoids data loss, improves transmission rate and bandwidth, and meets the needs of artificial intelligence computing.
Smart Images

Figure CN120090994A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit technology, and particularly to a wafer-level communication structure, a communication method and an AI chip. Background Art
[0002] With the rapid development of artificial intelligence, the computing power of a single chip is difficult to meet the requirements, so a super-large chip design is formed by combining multiple chips. However, at present, the serial transmission method is adopted for data transmission between chips in different regions. The serial transmission method is prone to data loss during data transmission, has a low transmission rate and a small bandwidth, and is difficult to meet the usage requirements. Summary of the Invention
[0003] Aiming at the defects in the prior art, the present invention provides a wafer-level communication structure that can effectively avoid data loss during transmission and achieve high-speed and high-bandwidth communication.
[0004] A wafer-level communication structure provided by the present application, the wafer-level communication structure includes:
[0005] A wafer, on which at least two bare dies are provided, and a packet switching network group is correspondingly provided for each bare die. The packet switching network group includes a plurality of on-chip packet networks arranged in a matrix;
[0006] A cross-region communication component, the cross-region communication component includes a first communication group and a second communication group. The first communication group is arranged around one of the packet switching network groups, and the second communication group is arranged around another packet switching network group. The first communication group and the second communication group are correspondingly arranged. The cross-region communication component further includes a source synchronous bus, and the source synchronous bus connects the first communication group and the second communication group to enable the first communication group and the second communication group to establish a communication handshake;
[0007] When transmitting data, one of the first communication group and the second communication group is a data sender, and the other is a data receiver;
[0008] When the number of packet network data cached by the data receiver is greater than or equal to a preset threshold, the data receiver feeds back a first handshake message to the data sender, and the data sender pauses sending data according to the first handshake message;
[0009] When the number of packet network data cached by the data receiver is lower than the preset threshold, the data receiver feeds back a second handshake message to the data sender, and the data sender continues to send data according to the second handshake message.
[0010] In one aspect, the first communication group includes a first sending module and a first receiving module, the second communication group includes a second sending module and a second receiving module, the first sending module and the second receiving module are connected to the source synchronous bus and establish a communication handshake, the first receiving module and the second sending module are connected to the source synchronous bus and establish a communication handshake, the source synchronous bus is used to align the message network data of the on-chip message network with the edge of the clock signal;
[0011] The first sending module sends message network data to the second receiving module. When the number of message network data cached by the second receiving module is greater than or equal to the preset threshold, the second receiving module feeds back the first handshake information to the first sending module, and the first sending module suspends sending data according to the first handshake information. When the number of message network data cached by the second receiving module is lower than the preset threshold, the second receiving module feeds back the second handshake information to the first sending module, and the first sending module continues to send data according to the second handshake information.
[0012] The second sending module sends message network data to the first receiving module. When the number of cached data in the first receiving module is greater than or equal to the preset threshold, the first receiving module feeds back the first handshake information to the second sending module, and the second sending module suspends sending data according to the first handshake information. When the number of message network data cached in the first receiving module is lower than the preset threshold, the first receiving module feeds back the second handshake information to the second sending module, and the second sending module continues to send data according to the second handshake information.
[0013] In one aspect, the cross-region communication component also includes an independent cache module, which is respectively arranged corresponding to the first sending module, the first receiving module, the second sending module and the second receiving module, and the independent cache module includes several independent cache blocks, and at least some of the cache blocks have different storage data capacities.
[0014] In one aspect, the cross-region communication component also includes a priority arbitrator, one priority arbitrator corresponds to one independent cache module, and the priority arbitrator adjusts the data sending priority in the independent cache module so that the data in each cache block cyclically obtains the sending permission of the priority arbitrator.
[0015] In one aspect, the source synchronous bus includes: a data sending source, a buffer module, a signal alignment module, and a data target module. The data sending source is used to send packet network data and a clock signal. A plurality of buffer modules are provided, and the plurality of buffer modules are arranged between the data sending source and the data target module. The target module outputs the packet network data. The signal alignment module is arranged between the buffer modules and is used to synchronize the rising edges of the packet network data and the clock signal when the deviation between the packet network data and the clock signal is greater than half a clock cycle.
[0016] In one aspect, multiple groups of cross-region communication components are provided, and the multiple groups of cross-region communication components are arranged between the two packet switching network groups.
[0017] In one aspect, the first communication group includes a first upstream router port, a first downstream router port, a second upstream router port, and a second downstream router port. The first upstream router port and the second downstream router port are arranged on the periphery of one packet switching network group. The first downstream router port is correspondingly connected to the first upstream router port, and the second downstream router port is correspondingly connected to the second upstream router port;
[0018] The second communication group includes a third upstream router port, a third downstream router port, a fourth upstream router port, and a fourth downstream router port. The third upstream router port and the fourth downstream router port are arranged on the periphery of the other packet switching network group. The third downstream router port is correspondingly connected to the third upstream router port, and the fourth downstream router port is correspondingly connected to the fourth upstream router port;
[0019] Wherein, the first upstream router port and the first downstream router port establish a credit connection. The first upstream router port pre-stores a first data upper limit value. When the number of transmitted data is equal to the first data upper limit value, the first upstream router port suspends transmitting data to the first downstream router port. The second upstream router port and the second downstream router port establish a credit connection. The second upstream router port pre-stores a second data upper limit value. When the number of transmitted data is equal to the second data upper limit value, the second upstream router port suspends transmitting data to the second downstream router port;
[0020] The third upstream router port and the third downstream router port establish a credit connection. The third upstream router port pre-stores a third data upper limit value. When the number of transmitted data is equal to the third data upper limit value, the third upstream router port suspends transmitting data to the third downstream router port. The fourth upstream router port and the fourth downstream router port establish a credit connection. The fourth upstream router port pre-stores a fourth data upper limit value. When the number of transmitted data is equal to the fourth data upper limit value, the fourth upstream router port suspends transmitting data to the fourth downstream router port.
[0021] In one aspect, credit counters are provided for the first upstream router port, the second upstream router port, the third upstream router port, and the fourth upstream router port;
[0022] When transmitting data, the credit counter of the first upstream router port counts to a first initial value. For each data sent by the first upstream router port, the credit counter of the first upstream router port is decremented by 1. For each data sent by the first downstream router port, the first downstream router port returns a credit to the first upstream router port, and the credit counter of the first upstream router port is incremented by 1;
[0023] The credit counter of the second upstream router port counts to a second initial value. For each data sent by the second upstream router port, the credit counter of the second upstream router port is decremented by 1. For each data sent by the second downstream router port, the second downstream router port returns a credit to the second upstream router port, and the credit counter of the second upstream router port is incremented by 1;
[0024] The credit counter of the third upstream router port counts to a third initial value. For each data sent by the third upstream router port, the credit counter of the third upstream router port is decremented by 1. For each data sent by the third downstream router port, the third downstream router port returns a credit to the third upstream router port, and the credit counter of the third upstream router port is incremented by 1;
[0025] The credit counter of the fourth upstream router port counts to a fourth initial value. For each data sent by the fourth upstream router port, the credit counter of the fourth upstream router port is decremented by 1. For each data sent by the fourth downstream router port, the fourth downstream router port returns a credit to the fourth upstream router port, and the credit counter of the fourth upstream router port is incremented by 1.
[0026] In addition, to solve the above problems, the present application also provides a wafer-level communication method, which is applied to the wafer-level communication structure as described above. The wafer-level communication method includes:
[0027] Start communication, and the data sender provides packet network data to the data receiver;
[0028] After receiving the packet network data, the data receiver caches the packet network data;
[0029] When the number of cached packet network data is greater than or equal to a preset threshold, the data receiver feeds back a first handshake message to the data sender, and the data sender pauses sending data according to the first handshake message;
[0030] When the number of cached packet network data is lower than the preset threshold, the data receiver feeds back a second handshake message to the data sender, and the data sender continues to send data according to the second handshake message.
[0031] In addition, to solve the above problems, the present application also provides an AI chip, which includes the wafer-level communication structure as described above.
[0032] The beneficial effects of the present invention are reflected in: realizing communication between two packet exchange network groups through a cross-region communication component, and at least combining two bare dies for use. By establishing a communication handshake between the first communication group and the second communication group, when the number of packet network data cached by the data receiver is greater than or equal to the preset threshold during data transmission, the data sender pauses sending data according to the first handshake message, avoiding exceeding the upper limit of the cache capacity of the data receiver and thus avoiding data loss. And when the number of packet network data cached by the data receiver is lower than the preset threshold, the data sender continues to send data according to the second handshake message. And through the connection setting of the source synchronous bus, the bandwidth between the first communication group and the second communication group is increased, realizing high-bandwidth communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale.
[0034] Figure 1 It is a schematic diagram of the overall structure of the wafer-level communication structure in the present application;
[0035] Figure 2 It is a schematic diagram of the internal structure of the communication component in the wafer-level communication structure of the present application;
[0036] Figure 3 This is a schematic diagram of the structure of the router port corresponding to the first communication group in the wafer-level communication structure of the present application;
[0037] Figure 4 This is a schematic diagram of the structure of the router port corresponding to the second communication group in the wafer-level communication structure of the present application;
[0038] Figure 5 This is a schematic diagram of the internal structure of the source synchronous bus in the wafer-level communication structure of the present application;
[0039] Figure 6 This is a schematic diagram of the process steps of the wafer-level communication method in the present application.
[0040] In the drawings, 100, wafer; 200, cross-region communication component; 300, message switching network group;
[0041] 210, first communication group; 220, second communication group; 230, independent cache module; 240, priority arbiter; 250, source synchronous bus; 310, on-chip message network; Data, message network data; clk, clock signal;
[0042] 211, first transmission module; 212, first reception module; 222, second transmission module; 221, second reception module; 2501, first bus segment; 2502, second bus segment; 251, data transmission source; 252, buffer module; 253, signal alignment module; 254, data target module; 2511, first D flip-flop; 2512, first delay circuit unit; 2521, buffer; 2531, second D flip-flop; 2532, second delay circuit unit; 2533, first inversion circuit; 2541, second inversion circuit; 2542, independent cache unit; 213, first upstream router port; 214, first downstream router port; 215, second upstream router port; 216, second downstream router port; 223, third upstream router port; 224, third downstream router port; 225, fourth upstream router port; 226, fourth downstream router port. Detailed implementation manners
[0043] Hereinafter, embodiments of the technical solution of the present invention will be described in detail with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and thus are only examples and cannot be used to limit the protection scope of the present invention.
[0044] It should be noted that unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meanings understood by those skilled in the art to which the present invention belongs.
[0045] As shown in Figure 1 the figure, the present application provides a wafer-level communication structure, which includes a wafer 100 and a cross-region communication component 200. The wafer 100 refers to a silicon wafer used to fabricate silicon semiconductor circuits, and its raw material is silicon. High-purity polysilicon is dissolved and doped with silicon crystal seeds, and then slowly pulled out to form a cylindrical single-crystal silicon. After the silicon ingot is ground, polished, and sliced, silicon wafers 100 are formed, also known as the wafer 100.
[0046] At least two bare dies are provided on the wafer 100. A message switching network group 300 is correspondingly provided for each bare die. The message switching network group 300 includes a plurality of on-chip message networks 310 arranged in a matrix; a bare die can be understood as a die on the wafer 100 that has not been encapsulated yet. A die has a complete circuit bare die, and there is a certain distance between the bare dies. The on-chip message network 310 is a Network on Chip (NOC) based on message switching.
[0047] The cross-region communication component 200 includes a first communication group 210 and a second communication group 220. The first communication group 210 is arranged around a message switching network group 300, and the second communication group 220 is arranged around another message switching network group 300. The first communication group 210 and the second communication group 220 are correspondingly arranged. The cross-region communication component 200 further includes a source synchronous bus 250. The source synchronous bus 250 connects the first communication group 210 and the second communication group 220 and enables the first communication group 210 and the second communication group 220 to establish a communication handshake; generally speaking, the first communication group 210 and the second communication group 220 are directly connected, and the communication handshake between the first communication group 210 and the second communication group 220 is realized through the source synchronous bus 250.
[0048] When transmitting data, one of the first communication group 210 and the second communication group 220 is the data sender, and the other is the data receiver; the process of data transmission is mutual. It can be from the first communication group 210 to the second communication group 220, or from the second communication group 220 to the first communication group 210.
[0049] When the number of message network data cached by the data receiver is greater than or equal to a preset threshold, the data receiver feeds back a first handshake message to the data sender, and the data sender pauses sending data according to the first handshake message; this can avoid excessive data transmission and exceed the cache capacity of the data receiver.
[0050] When the number of packet network data cached by the data receiver is lower than a preset threshold, the data receiver feeds back second handshake information to the data sender, and the data sender continues to send data according to the second handshake information. The first handshake information and the second handshake information can be represented by high and low levels. For example, the first handshake information is a high level and the second handshake information is a low level.
[0051] For example, the data sender and the data receiver communicate and handshake through the slow_en signal. After the number of data cached by the data receiver reaches the set preset threshold, a high-level slow_en is generated and the high-level slow_en is output to the data sender. After receiving the high-level slow_en input signal, the data sender stops sending data. After the handshake slow_en signal becomes low, the data sender resumes data sending, which can effectively avoid the problems of internal data congestion and data packet loss in the data receiver.
[0052] In this embodiment, the communication between two packet switching network groups 300 is implemented through the cross-region communication component 200, and at least two bare dies are used together. By establishing a communication handshake between the first communication group 210 and the second communication group 220, when the number of packet network data cached by the data receiver is greater than or equal to the preset threshold during data transmission, the data sender pauses sending data according to the first handshake information, avoiding exceeding the upper limit of the cache capacity of the data receiver, thereby avoiding data loss. And when the number of packet network data cached by the data receiver is lower than the preset threshold, the data sender continues to send data according to the second handshake information, effectively ensuring that the data can be continuously sent, avoiding congestion in the communication network and also avoiding packet loss. And through the connection setting of the source synchronous bus, the bandwidth between the first communication group 210 and the second communication group 220 is increased to achieve high-bandwidth communication.
[0053] It should be noted that this embodiment mainly realizes the credit handshake communication between the first communication group 210 and the second communication group 220 through the source synchronous bus. In this embodiment, the source synchronous bus can achieve communication with a frequency above 500M. The bandwidth can be understood as the product of the communication frequency and the data bit width. More channel ports are used for communication, thereby increasing the communication bandwidth.
[0054] In an embodiment of the present application, multiple sets of cross-region communication components 200 are provided, and multiple sets of cross-region communication components 200 are arranged between two packet switching network groups 300. As Figure 1 shown, for the 3*3 packet switching network group 300, three sets of cross-region communication components 200 are provided. It can be understood that more channel settings are beneficial to achieving high speed and high bandwidth. Moreover, the flexibility of signal transmission can be improved, and more communication paths can be selected.
[0055] As Figure 2 shown, in an embodiment of the present application, the first communication group 210 includes a first sending module 211 and a first receiving module 212, and the second communication group 220 includes a second sending module 222 and a second receiving module 221. The first sending module 211 and the second receiving module 221 are connected to the source synchronous bus 250 to establish a communication handshake through the source synchronous bus 250. The first receiving module 212 and the second sending module 222 are connected to the source synchronous bus 250 to establish a communication handshake through the source synchronous bus 250. The receiving module is mainly used to receive data, and the sending module is mainly used to send data. It can be seen from this that in this embodiment, both the first communication group 210 and the second communication group 220 are provided with a sending module and a receiving module.
[0056] The first sending module 211 sends packet network data to the second receiving module 221. When the number of packet network data cached in the second receiving module 221 is greater than or equal to a preset threshold, the second receiving module 221 feeds back the first handshake information to the first sending module 211, and the first sending module 211 suspends sending data according to the first handshake information. When the number of packet network data cached in the second receiving module 221 is lower than the preset threshold, the second receiving module 221 feeds back the second handshake information to the first sending module 211, and the first sending module 211 continues to send data according to the second handshake information.
[0057] The second sending module 222 sends packet network data to the first receiving module 212. When the number of cached data in the first receiving module 212 is greater than or equal to a preset threshold, the first receiving module 212 feeds back the first handshake information to the second sending module 222, and the second sending module 222 suspends sending data according to the first handshake information. When the number of packet network data cached in the first receiving module 212 is lower than the preset threshold, the first receiving module 212 feeds back the second handshake information to the second sending module 222, and the second sending module 222 continues to send data according to the second handshake information. It should be noted that the preset threshold can be set according to the cache capacity of the receiving module.
[0058] The first communication group 210 can send data to the second communication group 220 or receive data from the second communication group 220. Similarly, the second communication group 220 can also send data to the first communication group 210 or receive data from the first communication group 210. The data sending channel and the data receiving channel are independent of each other to avoid interference during the sending and receiving processes.
[0059] In an embodiment of the present application, the cross-region communication component 200 further includes an independent cache module 230. The independent cache module 230 is respectively arranged corresponding to the first sending module 211, the first receiving module 212, the second sending module 222, and the second receiving module 221. The independent cache module 230 includes several independent cache blocks, and the storage data capacities of at least some of the cache blocks are different. It is possible that an independent cache module 230 is respectively arranged inside the first sending module 211, the first receiving module 212, the second sending module 222, and the second receiving module 221, that is, the independent cache module 230 is arranged inside them; or, the first sending module 211, the first receiving module 212, the second receiving module 221, and the second sending module 222 are respectively connected to an independent cache module 230.
[0060] The caching principle of the cache block can be the FIFO (First Input First Output) queue mechanism. This can avoid the situation where some data in the independent cache module 230 cannot be transmitted for a long time. Moreover, by having different storage data capacities in the cache blocks, different data can be stored in different cache blocks, making full use of the storage space of the cache blocks. For data with a larger bit width, it can be stored in a cache block with a larger capacity, and for data with a smaller bit width, it can be stored in a cache block with a smaller capacity. In this way, different message information can be stored using the independent cache module 230 to achieve cross-time-domain processing of data. The cache blocks vary according to the actual usage scenarios, thereby saving hardware resources and reducing the area occupied by the cross-region communication component 200 in a targeted manner.
[0061] In an embodiment of the present application, the cross-region communication component 200 further includes a priority arbiter 240. One priority arbiter 240 corresponds to one independent cache module 230. The priority arbiter 240 adjusts the data sending priority in the independent cache module 230, enabling the data in each cache block to cyclically obtain the sending permission of the priority arbiter 240. The English abbreviation of the priority arbiter 240 is ARBIT (Round Robin arbiter). The priority arbiter 240 is a technology for resource access control among multiple requesters, especially suitable for shared resources, parallel computing, or multi-processor systems. It mainly ensures that each requester can obtain an equal opportunity to access resources within a certain period of time, thereby avoiding a certain requester monopolizing resources for a long time.
[0062] For example, when transmitting data, the priority arbiter 240 first determines the priority and whether there are other requests after judging the transmission priority; if it is found that there are no other requests at this time, the priority arbiter 240 will respond to the transmission request of the message and allow it to be transmitted. If there are other high-priority transmission requests at this time, the priority arbiter 240 will respond to the transmission request with a higher priority. After responding to the message transmission request, the priority arbiter 240 will lower its priority and allow other message transmission requests to occupy the interface. This effectively solves the problem of a certain message transmission occupying the interface for a long time.
[0063] In one embodiment of the present application, the cross-region communication component 200 also includes a source synchronous bus 250, which corresponds to the first sending module 211, the first receiving module 212, the second sending module 222, and the second receiving module 221, respectively. The source synchronous bus 250 is used to align the message network data of the on-chip message network 310 with the edge of the clock signal clk. The source synchronous bus 250 is referred to as SSB (Source Synchronous Bus) in English, and the source synchronous bus 250 can be set in the first sending module 211, the first receiving module 212, the second sending module 222, and the second receiving module 221. Due to the long line, the rising edge of the message network data and the rising edge of the clock signal clk cannot be aligned, and it is easy to cause the message network data to be unable to be effectively restored. Through the setting of the source synchronous bus 250, the phase relationship between the message network data and the clock signal clk can be calibrated. The source synchronous bus 250 includes multiple synchronous bus units, Uint0 to UintN. Through the setting of multiple synchronous bus units, the data bit width of the source synchronous bus 250 is increased on the basis of the bandwidth being the product of the frequency and the data bit width, thereby further achieving high bandwidth.
[0064] like Figure 5 As shown, in one embodiment of the present application, the source synchronous bus 250 includes: a data sending source 251, a buffer module 252, a signal alignment module 253 and a data target module 254. The data sending source 251 is used to send message network data and a clock signal clk. The buffer module 252 is provided with multiple buffer modules 252, and the multiple buffer modules 252 are arranged between the data sending source 251 and the data target module 254. The data target module 254 outputs the message network data. The signal alignment module 253 is arranged between the buffer modules 252. The signal alignment module 253 is used to synchronize the message network data Data and the rising edge of the clock signal clk when the deviation between the message network data and the clock signal clk is greater than half a clock cycle, so as to ensure that the message network data Data can be accurately restored.
[0065] The data sending source 251 is used to align the rising edge of the message network data Data and the clock signal clk. The buffer module 252 (Buffers) is used to allow the driving ability of the signal to be improved within half a clock cycle to cross a relatively long distance. The buffer module 252 can be provided with a plurality of buffers 2521. When the signal needs to cross a distance that spans more than half a clock cycle, the message network data Data and the rising edge of the clock signal clk are resynchronized through the signal alignment module 253 to solve the deviation problem between the clock signal clk and the data signal during the transmission process. Through the data destination module 254, the input message network data is captured using the falling edge of the clock signal clk, and the message network data is restored. The data destination module 254 further includes an independent cache unit 2542. The restored message network data is stored in the independent cache unit 2542, and a group of synchronous bus units corresponds to one independent cache unit 2542. After it is confirmed that all the independent cache units 2542 store data, the data of all the independent cache units 2542 is allowed to be accessed simultaneously, which can avoid the clock signal skew between the data and reduce the problem of inconsistent data arrival times caused by different paths passed.
[0066] The data sending source 251 includes a first D flip-flop 2511 and a first delay circuit unit 2512. The signal alignment module 253 includes a second D flip-flop 2531, a second delay circuit unit 2532, and a first inversion circuit 2533. The buffer module 252 is provided with two buffer lines, one data line and one clock line. The data destination module 254 further includes a second inversion circuit 2541. The first delay circuit unit 2512 is used to reduce the delay difference between the data signal and the clock signal after passing through the first D flip-flop 2511. The second delay circuit unit 2532 is used to reduce the delay difference between the data signal and the clock signal after passing through the second D flip-flop 2531. The first D flip-flop 2511 receives the clock signal clk, the D port of the first D flip-flop 2511 receives the message network data, one line is connected to the Q end of the first D flip-flop 2511, and the other is connected to the clock signal clk through the first delay circuit unit 2512. Thus, the accuracy of data transfer can be improved between the first communication group 210 and the second communication group 220, and the message network data Data can be accurately restored.
[0067] It should be further noted that the source synchronous bus 250 includes two parts, namely the first bus segment 2501 and the second bus segment 2502. The first bus segment 2501 includes a data transmission source 251, a buffer module 252, and a signal alignment module 253. The second bus segment 2502 includes a data target module 254. The first bus segment 2501 is arranged at the end for sending data, that is, the first bus segment 2501 is arranged in the first transmission module 211 and the second transmission module 222. The second bus segment 2502 is arranged at the end for receiving data, that is, the second bus segment 2502 is arranged in the first receiving module 212 and the second receiving module 221.
[0068] As Figure 3 and Figure 4 shown, in an embodiment of the present application, the first communication group 210 includes a first upstream router port 213, a first downstream router port 214, a second upstream router port 215, and a second downstream router port 216. The first upstream router port 213 and the second downstream router port 216 are arranged on the periphery of a packet switching network group 300. The first downstream router port 214 is correspondingly connected to the first upstream router port 213, and the second downstream router port 216 is correspondingly connected to the second upstream router port 215.
[0069] The second communication group 220 includes a third upstream router port 223, a third downstream router port 224, a fourth upstream router port 225, and a fourth downstream router port 226. The third upstream router port 223 and the fourth downstream router port 226 are arranged on the periphery of the opposite packet switching network group 300. The third downstream router port 224 is correspondingly connected to the third upstream router port 223, and the fourth downstream router port 226 is correspondingly connected to the fourth upstream router port 225.
[0070] Among them, the first upstream router port 213 and the first downstream router port 214 establish a credit connection. The first upstream router port 213 pre-stores a first data upper limit value. When the number of transmitted data is equal to the first data upper limit value, the first upstream router port 213 suspends transmitting data to the first downstream router port 214. The second upstream router port 215 and the second downstream router port 216 establish a credit connection. The second upstream router port 215 pre-stores a second data upper limit value. When the number of transmitted data is equal to the second data upper limit value, the second upstream router port 215 suspends transmitting data to the second downstream router port 216.
[0071] The third upstream router port 223 and the third downstream router port 224 establish a credit connection. The third upstream router port 223 pre-stores a third data upper limit value. When the number of transmitted data is equal to the third data upper limit value, the third upstream router port 223 pauses transmitting data to the third downstream router port 224. The fourth upstream router port 225 and the fourth downstream router port 226 establish a credit connection. The fourth upstream router port 225 pre-stores a fourth data upper limit value. When the number of transmitted data is equal to the fourth data upper limit value, the fourth upstream router port 225 pauses transmitting data to the fourth downstream router port 226, thereby avoiding data congestion. The first data upper limit and the second data upper limit, as well as the third data upper limit and the fourth data upper limit, may be the same or different.
[0072] In an embodiment of the present application, the first upstream router port 213, the second upstream router port 215, the third upstream router port 223, and the fourth upstream router port 225 are all provided with credit counters.
[0073] When transmitting data, the credit counter of the first upstream router port 213 counts to a first initial value. For each data sent by the first upstream router port 213, the credit counter of the first upstream router port 213 is decremented by 1. For each data sent by the first downstream router port 214, the first downstream router port 214 returns a credit to the first upstream router port 213, and the credit counter of the first upstream router port 213 is incremented by 1. Thereby avoiding data overflow in the downstream router and reducing data congestion.
[0074] The credit counter of the second upstream router port 215 counts to a second initial value. For each data sent by the second upstream router port 215, the credit counter of the second upstream router port 215 is decremented by 1. For each data sent by the second downstream router port 216, the second downstream router port 216 returns a credit to the second upstream router port 215, and the credit counter of the second upstream router port 215 is incremented by 1.
[0075] The credit counter of the third upstream router port 223 counts to a third initial value. For each data sent by the third upstream router port 223, the credit counter of the third upstream router port 223 is decremented by 1. For each data sent by the third downstream router port 224, the third downstream router port 224 returns a credit to the third upstream router port 223, and the credit counter of the third upstream router port 223 is incremented by 1.
[0076] The credit counter of the fourth upstream router port 225 counts to the fourth initial value. For each data sent by the fourth upstream router port 225, the credit counter of the fourth upstream router port 225 is decremented by 1. For each data sent by the fourth downstream router port 226, the fourth downstream router port 226 returns a credit to the fourth upstream router port 225, and the credit counter of the fourth upstream router port 225 is incremented by 1. This can avoid data overflow in the downstream router. The first initial value, the second initial value, the third initial value, and the fourth initial value can be equal or unequal. The first initial value, the second initial value, the third initial value, and the third initial value are the values when the downstream router's data storage is full.
[0077] Taking the first communication group 210 as an example for illustration, the first communication group 210 sometimes needs to obtain data from the on-chip packet network 310 and then transmit the data to the on-chip packet network 310 of the peer; sometimes it also needs to transmit the data of the on-chip packet network 310 of the peer to the local on-chip packet network 310. For this reason, the first communication group 210 is provided with a first upstream router port 213, a first downstream router port 214, a second upstream router port 215, and a second downstream router port 216. The first upstream router port 213 and the first downstream router port 214 are set corresponding to the first sending module 211, and the second upstream router port 215 and the second downstream router port 216 are set corresponding to the first receiving module 212. For the first sending module 211, the first upstream router port 213 is set in the on-chip packet network 310, and the first downstream router port 214 is set in the first communication group 210. For the first receiving module 212, the second upstream router port 215 is set in the first communication group 210, and the second downstream router port 216 is set in the on-chip packet network 310 。
[0078] The router port setting scheme in the second communication group 220 can refer to the setting scheme of the first communication group 210. Generally, the router is set in the on-chip packet network, and the corresponding router ports are set in cooperation in the cross-region communication component.
[0079] As Figure 6 shown, the present application also provides a wafer-level communication method. The wafer-level communication method is applied to the wafer-level communication structure as described above. The wafer-level communication method includes:
[0080] Step S10, start communication, and the data sender provides packet network data to the data receiver; the first communication group 210 can be the data sender or the data receiver; similarly, the second communication group 220 can be the data sender or the data receiver.
[0081] Step S20: After receiving the packet network data, the data receiver caches the packet network data. The independent cache module 230 can be used to cache the packet network data.
[0082] Step S30: When the number of cached packet network data is greater than or equal to a preset threshold, the data receiver feeds back the first handshake information to the data sender, and the data sender pauses sending data according to the first handshake information. The preset threshold can be adjusted as needed.
[0083] Step S40: When the number of cached packet network data is lower than the preset threshold, the data receiver feeds back the second handshake information to the data sender, and the data sender continues to send data according to the second handshake information.
[0084] For example, the data sender and the data receiver communicate and handshake through the slow_en signal. After the number of data cached by the data receiver reaches the set preset threshold, a high-level slow_en is generated and the high-level slow_en is output to the data sender. After receiving the high-level slow_en input signal, the data sender stops sending data. After the handshake slow_en signal becomes low level, the data sender resumes sending data.
[0085] In this embodiment, the communication between two packet exchange network groups 300 is implemented through the cross-region communication component 200, and at least two bare dies are used together. By establishing a communication handshake between the first communication group 210 and the second communication group 220, when the number of packet network data cached by the data receiver is greater than or equal to the preset threshold during data transmission, the data sender pauses sending data according to the first handshake information, avoiding exceeding the upper limit of the cache capacity of the data receiver, thus avoiding data loss. And when the number of packet network data cached by the data receiver is lower than the preset threshold, the data sender continues to send data according to the second handshake information, effectively ensuring that data can be continuously sent, avoiding congestion in the communication network and packet loss. And through the connection setting of the source synchronous bus, the bandwidth between the first communication group and the second communication group is increased, realizing high-speed and high-bandwidth communication.
[0086] This application also provides an AI chip, which includes the wafer-level communication structure as described above.
[0087] For the specific embodiments and beneficial effects of the AI chip in this application, please refer to the above wafer-level communication structure and will not be elaborated here.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. A wafer-level communication structure, characterized in that: The wafer-level communication structure comprises: A wafer, wherein at least two bare cores are arranged on the wafer, and one of the bare cores is correspondingly provided with a message switching network group, wherein the message switching network group includes a plurality of on-chip message networks arranged in a matrix; A cross-region communication component, the cross-region communication component includes a first communication group and a second communication group, the first communication group is arranged at the periphery of one of the message switching network groups, the second communication group is arranged at the periphery of another of the message switching network groups, the first communication group and the second communication group are arranged correspondingly, the cross-region communication component also includes a source synchronous bus, the source synchronous bus connects the first communication group and the second communication group, so that the first communication group and the second communication group establish a communication handshake; When transmitting data, one of the first communication group and the second communication group is a data sender, and the other is a data receiver; When the number of message network data cached by the data receiving party is greater than or equal to a preset threshold, the data receiving party feeds back first handshake information to the data sending party, and the data sending party suspends sending data according to the first handshake information; When the number of message network data cached by the data receiving party is lower than the preset threshold, the data receiving party feeds back second handshake information to the data sending party, and the data sending party continues to send data according to the second handshake information.
2. The wafer-level communication structure according to claim 1, characterized in that: The first communication group includes a first sending module and a first receiving module, the second communication group includes a second sending module and a second receiving module, the first sending module and the second receiving module are connected to the source synchronous bus and establish a communication handshake, the first receiving module and the second sending module are connected to the source synchronous bus and establish a communication handshake, the source synchronous bus is used to align the message network data of the on-chip message network with the edge of the clock signal; The first sending module sends message network data to the second receiving module. When the number of message network data cached by the second receiving module is greater than or equal to the preset threshold, the second receiving module feeds back the first handshake information to the first sending module, and the first sending module suspends sending data according to the first handshake information. When the number of message network data cached by the second receiving module is lower than the preset threshold, the second receiving module feeds back the second handshake information to the first sending module, and the first sending module continues to send data according to the second handshake information. The second sending module sends message network data to the first receiving module. When the number of cached data in the first receiving module is greater than or equal to the preset threshold, the first receiving module feeds back the first handshake information to the second sending module, and the second sending module suspends sending data according to the first handshake information. When the number of message network data cached in the first receiving module is lower than the preset threshold, the first receiving module feeds back the second handshake information to the second sending module, and the second sending module continues to send data according to the second handshake information.
3. The wafer-level communication structure according to claim 2, characterized in that: The cross-region communication component also includes an independent cache module, which is respectively set corresponding to the first sending module, the first receiving module, the second sending module and the second receiving module. The independent cache module includes several independent cache blocks, and at least some of the cache blocks have different storage data capacities.
4. The wafer-level communication structure according to claim 3, characterized in that: The cross-region communication component also includes a priority arbitrator, one priority arbitrator corresponds to one independent cache module, and the priority arbitrator adjusts the data sending priority in the independent cache module so that the data in each cache block cyclically obtains the sending permission of the priority arbitrator.
5. The wafer-level communication structure according to claim 1, characterized in that: The source synchronous bus includes: a data sending source, a buffer module, a signal alignment module and a data target module. The data sending source is used to send message network data and a clock signal. There are multiple buffer modules, and the multiple buffer modules are arranged between the data sending source and the data target module. The target module outputs the message network data. The signal alignment module is arranged between the buffer modules. The signal alignment module is used to synchronize the message network data and the rising edge of the clock signal when the deviation between the message network data and the clock signal is greater than half a clock cycle.
6. The wafer-level communication structure according to claim 1, characterized in that: The cross-region communication components are provided in multiple groups, and the multiple groups of cross-region communication components are provided between the two message exchange network groups.
7. The wafer-level communication structure according to claim 1, characterized in that: The first communication group includes a first upstream router port, a first downstream router port, a second upstream router port, and a second downstream router port, wherein the first upstream router port and the second downstream router port are arranged at the periphery of the message switching network group, the first downstream router port is connected and arranged correspondingly to the first upstream router port, and the second downstream router port is connected and arranged correspondingly to the second upstream router port; The second communication group includes a third upstream router port, a third downstream router port, a fourth upstream router port and a fourth downstream router port, wherein the third upstream router port and the fourth downstream router port are arranged around another message switching network group, the third downstream router port is connected and arranged correspondingly to the third upstream router port, and the fourth downstream router port is connected and arranged correspondingly to the fourth upstream router port; Wherein, the first upstream router port establishes a credit connection with the first downstream router port, the first upstream router port pre-stores a first data upper limit value, and when the number of transmitted data is equal to the first data upper limit value, the first upstream router port suspends data transmission to the first downstream router port, and the second upstream router port establishes a credit connection with the second downstream router port, the second upstream router port pre-stores a second data upper limit value, and when the number of transmitted data is equal to the second data upper limit value, the second upstream router port suspends data transmission to the second downstream router port; The third upstream router port establishes a credit connection with the third downstream router port, and the third upstream router port pre-stores a third data upper limit value. When the number of transmitted data is equal to the third data upper limit value, the third upstream router port suspends transmitting data to the third downstream router port. The fourth upstream router port establishes a credit connection with the fourth downstream router port, and the fourth upstream router port pre-stores a fourth data upper limit value. When the number of transmitted data is equal to the fourth data upper limit value, the fourth upstream router port suspends transmitting data to the fourth downstream router port.
8. The wafer-level communication structure according to claim 7, characterized in that: The first upstream router port, the second upstream router port, the third upstream router port and the fourth upstream router port are each provided with a credit counter; When transmitting data, the credit counter of the first upstream router port counts to a first initial value, each time the first upstream router port sends a piece of data, the credit counter of the first upstream router port is reduced by 1, each time the first downstream router port sends a piece of data, the first downstream router port returns a credit to the first upstream router port, and the credit counter of the first upstream router port is increased by 1; The credit counter of the second upstream router port counts to a second initial value, each time the second upstream router port sends a piece of data, the credit counter of the second upstream router port is reduced by 1, each time the second downstream router port sends a piece of data, the second downstream router port returns a credit to the second upstream router port, and the credit counter of the second upstream router port is increased by 1; The credit counter of the third upstream router port counts to a third initial value, each time the third upstream router port sends a piece of data, the credit counter of the third upstream router port is reduced by 1, each time the third downstream router port sends a piece of data, the third downstream router port returns a credit to the third upstream router port, and the credit counter of the third upstream router port is increased by 1; The credit counter of the fourth upstream router port counts to a fourth initial value, each time the fourth upstream router port sends a piece of data, the credit counter of the fourth upstream router port is reduced by 1, each time the fourth downstream router port sends a piece of data, the fourth downstream router port returns a credit to the fourth upstream router port, and the credit counter of the fourth upstream router port is increased by 1.
9. A wafer-level communication method, characterized in that: The wafer-level communication method is applied to the wafer-level communication structure according to any one of claims 1 to 8, and the wafer-level communication method comprises: Initiating communication, the data sender provides the data receiver with message network data; After receiving the message network data, the data receiver caches the message network data; When the number of cached message network data is greater than or equal to a preset threshold, the data receiving party feeds back first handshake information to the data sending party, and the data sending party suspends sending data according to the first handshake information; When the number of cached message network data is lower than the preset threshold, the data receiving party feeds back second handshake information to the data sending party, and the data sending party continues to send data according to the second handshake information.
10. An AI chip, characterized in that: The AI chip includes a wafer-level communication structure as described in any one of claims 1 to 8.
Citation Information
Cited By
Data transmission method, device and equipment of network-on-chip, and storage medium
CN120804026A