Supporting multiple functions of a network-on-chip using a buffer
A unified buffer in the network-on-chip addresses communication challenges by supporting credit management, clock-domain crossing, and data upscaling, reducing footprint, power consumption, and latency, and lowering costs in system-on-chip designs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-02
AI Technical Summary
The increased complexity and inefficiency in communication between subsystems within a system-on-chip due to differences in clock domains, voltage domains, power domains, and data widths, leading to higher footprint, power consumption, and latency, as well as increased development and verification costs.
Implementing a network-on-chip with a unified buffer that supports credit management, asynchronous clock-domain crossing, and data upscaling, allowing for a single storage element to manage these functions, reducing the need for multiple storage elements and enabling the insertion of a level shifter before the buffer to compensate for propagation delays.
The unified buffer reduces the network-on-chip's footprint, power consumption, and latency while lowering manufacturing, development, and integration costs, while maintaining efficient communication across different domains and data widths.
Smart Images

Figure US2024049319_02042026_PF_FP_ABST
Abstract
Description
SUPPORTING MULTIPLE FUNCTIONS OF A NETWORK-ON-CHIP USING A BUFFERBACKGROUND
[0001] An electronic device can be implemented with a system-on-chip (SoC), which can provide many features of the electronic device. An example system-on-chip can include multiple subsystems, such as a central processing unit (CPU), a graphics processing unit (GPU), and / or an image processing unit (1PU). Technological advancements enable a system-on-chip to be designed with a larger quantity of subsystems to further expand feature capabilities of the electronic device. The larger quantity of subsystems, however, can increase a complexity for establishing communications between these subsystems.SUMMARY
[0002] Techniques and apparatuses are described for supporting multiple functions of a networkon-chip using a buffer. In example aspects, a network-on-chip uses a credit-based protocol and facilitates communication between different subsystems associated with different clock domains, different voltage domains, different power domains, different data widths, or some combination thereof. The network-on-chip includes a buffer, which provides storage for credit management, storage for an asynchronous clock-domain crossing, and storage for data upscaling. While other network-on-chips or other system-on-chips can have individual storage elements to support each function, the buffer acts as a single storage element that can support each of these functions. With the buffer, the network-on-chip can have a smaller footprint, can consume less power, can have less latency, can be simpler to integrate, which saves development and verification time, and can be cheaper to manufacture compared to other network-on-chips that utilize multiple storage elements. This single storage-element architecture also enables an insertion point of a level shifter to be placed before the buffer. As such, any propagation delay associated with the level shifter can be compensated for using the buffer.
[0003] Aspects described below include a first method performed by a network-on-chip. The method includes receiving, by a buffer of the network-on-chip, data from a first subsystem that is coupled to the network-on-chip. The receiving of the data occurs based on a first clock domain associated with the first subsystem. The data has a first data width associated with the first subsystem. The method also includes upscaling, using the buffer, the data from the first data width to a second data width associated with a second subsystem to generate upscaled data. The second subsystem is coupled to the network-on-chip. The method also includes sending the upscaled data from the buffer to the second subsystem. The sending of the data occurs based on a second clockdomain associated with the second subsystem. The second clock domain is different than the first clock domain.
[0004] Aspects described below include a second method perfonned by a write manager of a network-on-chip. The method includes receiving, from a first subsystem that is coupled to the network-on-chip, different sets of beats associated with different packets. The different packets are associated with different virtual channels. The receiving occurs based on a first clock domain associated wi th the first subsystem. The method also includes writing the different sets of beats to different address locations of a buffer of the network-on-chip. The method additionally includes sending, based on an address location of the different address locations becoming full or storing a last beat of one of the different packets, first control information to enable corresponding beats stored within the address location to be read from the buffer and passed to a second subsystem that is coupled to the network-on-chip.
[0005] Aspects described below also include an apparatus including a network-on-chip with a buffer and a write manager. The network-on-chip is configured to perform, using the buffer and / or the write manager, any of the described methods.
[0006] Aspects described below also include a system with means for supporting multiple functions of a network-on-chip.BRIEF DESCRIPTION OF DRAWINGS
[0007] Apparatuses for and techniques for supporting multiple functions of a network-on-chip using a buffer are described with reference to the following drawings. The same numbers are used throughout the drawings to reference like features and components:FIG. 1 illustrates an example environment in which supporting multiple functions of a network-on-chip using a buffer can be implemented;FIG. 2 illustrates example functions of a network-on-chip that are supported using a buffer;FIG. 3 illustrates an example implementation of a computing device that can implement aspects of supporting multiple functions of a network-on-chip using a buffer;FIG. 4 illustrates example components of a network-on-chip;FIG. 5 illustrates an example implementation of a buffer capable of supporting multiple functions of a network-on-chip;FIG. 6 illustrates an example scheme implemented by a write manager of a network-on- chip;FIG. 7 illustrates an example implementation of a network-on-chip including a buffer capable of supporting multiple functions of the network-on-chip;FIG. 8 illustrates an example writing scheme for upscaling data using a buffer of a network-on-chip;FIG. 9 illustrates a first example method for supporting multiple functions of a networkon-chip using a buffer;FIG. 10 illustrates a second example method for supporting multiple functions of a network-on-chip using a buffer; andFIG. 11 illustrates an example computing system embodying, or in which techniques may be implemented that enable use of, supporting multiple functions of a network-on-chip using a buffer.DETAILED DESCRIPTION
[0008] Technological advancements enable a system-on-chip (SoC) to be designed with a larger quantity of subsystems to further expand feature capabilities of an electronic device. Example subsystems can include a central processing unit (CPU), a graphics processing unit (GPU), and / or an image processing unit (IPU). The larger quantity of subsystems, however, can increase a complexity for establishing communications between these subsystems. This is particularly challenging as different subsystems can utilize different clock domains, different voltage domains, different power domains, and / or different data widths.
[0009] To address this issue, a system-on-chip can be implemented with one or more networkon-chips (NOCs), which provide an interface for the subsystems to communicate with each other. Network-on-chips can use a credit-based protocol to manage traffic efficiently for different traffic classes. The credit-based protocol can provide improved throughput and reduced latency compared to other types of protocols. In addition to supporting the credit-based protocol, a network-on-chip can communicate data across different voltage domains and across different clock domains associated with different subsystems by performing voltage-domain crossing and clock-domain crossing, respectively. Furthermore, in situations in which different subsystems utilize different data widths, the network-on-chip can appropriately scale a width of the data. This scaling can include downscaling data from a wider width to a narrower width or upscaling the data from a narrower width to a wider width.
[0010] To support the credit-based protocol, the clock-domain crossing, and the data upscaling, some network-on-chips use three different storage elements to manage these three different functions. For example, a network-on-chip can include a first storage element to handle the buffering for the clock-domain crossing; a second storage element that provides buffering for the data upscaling; and a third storage element that provides buffering for the credit-based protocol. Each of these three storage elements are implemented within a data path of the network-on-chip.This means that the three storage elements can each store data that the network-on-chip is to propagate from one subsystem to another subsystem.
[0011] Use of multiple storage elements increases a footprint of the network-on-chip and reduces the power efficiency due to an amount of power that can leak from these multiple storage elements. The multiple storage elements can also impact a latency of the network-on-chip as one or more additional clock cycles can be required to propagate data across the above listed multiple storage elements. Further costs are incurred in the development, integration, and verification of the multiple storage elements, which can impact production time and increase a cost of the system- on-chip.
[0012] In some implementations, the first storage element is inserted before a level shifter of the network-on-chip that handles the voltage-domain crossing. In this architecture, any delays associated with the level shifter can negatively impact the latency of the network-on-chip and can place additional constraints on the clock domain of a receiving subsystem. In general, it is desirable to design a network-on-chip that can support the credit-based protocol, the clock-domain crossing, the data upscaling, the voltage-domain crossing, and the power-domain crossing while reducing a footprint of the network-on-chip, increasing its power efficiency, and decreasing its latency.
[0013] To address this challenge, techniques are described for supporting multiple functions of a network-on-chip using a buffer. In example aspects, a network-on-chip uses a credit-based protocol and facilitates communication between different subsystems associated with different clock domains, different voltage domains, different power domains, different data widths, or some combination thereof. The network-on-chip includes a buffer, which provides storage for credit management, storage for an asynchronous clock-domain crossing, and storage for data upscaling. While other network-on-chips or other system-on-chips can have individual storage elements to support each function, the buffer acts as a single storage element that can support each of these functions. With the buffer, the network-on-chip can have a smaller footprint, can consume less power, can have less latency, and can be cheaper to manufacture, develop, integrate, and / or verity’ compared to other network-on-chips that utilize multiple storage elements. This single storageelement architecture also enables an insertion point of a level shifter to be placed before the buffer. As such, any propagation delay associated with the level shifter can be compensated for using the buffer.Operating Environment
[0014] FIG. 1 is an illustration of an example environment 100 in which supporting multiple functions of a network-on-chip using a buffer can be implemented. In the example environment 100, a computing device 102 provides features and / or services for a user 104. Although depicted as a smartphone, the computing device 102 can include other types of devices, including those described with respect to FIG. 3. The computing device 102 includes at least one system-on-chip (SOC) 106 (SOC 106). The system-on-chip 106 can be implemented with electronic circuitry, a microprocessor, memory, input-output (I / O) control logic, communication interfaces, firmware, and / or software useful to provide functionalities of the computing device 102.
[0015] The system-on-chip 106 includes multiple subsystems 108-1, 108-2... 108-S, where S represents a positive integer. The subsystems 108 can also be referred to as agents, modules, intellectual-property blocks (IP blocks), intellectual-property cores, or virtual components. Example subsystems 108 can include a central processing unit (CPU), a graphics processing unit (GPU), an image processing unit (IPU), a modem, a digital signal processor (DSP), a neural processing unit (NPU), a power processing unit (PPU), a display, a processor, a memory, a sensor, an analog circuit, a digital circuit, components that handle application-specific processing functions, and so forth. To facilitate independent operation, the subsystems 108 can have independent clock domains, independent voltage domains, independent power domains, independent operating data widths, or some combination thereof, as further described with respect to FIG. 2. To perform one or more functions of the computing device 102, at least one of the subsystems 108 can communicate with at least another one of the subsystems 108. In some situations, two or more subsystems 108 can communicate with two or more other subsystems 108 during a same time period.
[0016] The system-on-chip 106 also includes at least one network-on-chip (NOC) 110 (NOC 110) to provide a communication network between the subsystems 108. In some implementations, the system-on-chip 106 includes multiple network-on-chips 110 (e.g., multiple instances of a network-on-chip 110), which connect different sets of subsystems 108 together. The network-on- chip 110 can also be considered another subsystem 108 of the system-on-chip 106. The components of the system-on-chip 106 (e.g., the subsystems 108 and the network-on-chip 110) can alternatively be implemented within other types of integrated circuits or embedded systems, such as a microchip, an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a digital signal processor (DSP), a programmable system-on-chip (PSoC), system-in-package (SiP), controller, and so forth.
[0017] In example implementations, the network-on-chip 110 manages traffic between two subsystems 108 using a credit-based protocol 112. The credit-based protocol 112 involves a transmitting subsystem 108 keeping track of a quantity of credits it has available and sending data to the network-on-chip 110 when it has available credit. The network-on-chip 110 returns credit to the transmitting subsystem 108 after processing the data it receives from the transmitting subsystem 108. With the credit-based protocol 112, the network-on-chip 1 10 can realize ahigh er throughput compared to other types of protocols as the transmitting subsystem 108 does not have to wait for a response from the network-on-chip 110 prior to sending data.
[0018] To support the credit-based protocol 112, the network-on-chip 110 has to be capable of storing at least a credit’s worth of data (e.g., a credit’s worth of packets). To achieve peak network throughput, the network-on-chip 110 can provide additional storage to account for a round-trip latency of the credit-based protocol 112. The round-trip latency represents the quantity of clock cycles it takes for the network-on-chip 110 to process data received from the transmitting subsystem 108 and to return the credit to the transmitting subsystem 108. The round-trip latency also accounts for any delays associated with a data path and / or a credit return channel.
[0019] The network-on-chip 110 can improve a quality-of-service (QoS) 114 of the system-on- chip 106 by utilizing multiple virtual channels (VCs) 116-1 to 116-V, where V represents a positive integer. The virtual channels 116 represent independent logical channels that use a same physical channel (or link). Different virtual channels 116 can have different priorities. The network-on-chip 110 can include logic that selects an appropriate virtual channel for a given type of data. Different implementations of the network-on-chip 110 can support a different quantity of virtual channels 116. In example implementations, the quantity of virtual channels 116 (e.g., V) can be greater than or equal to 2, 3, 4, 8, 16, and so forth.
[0020] Consider an example in which the user 104 uses the computing device 102 to perform two different tasks at a same time. A first task can involve playing a video. A second task can involve downloading a file from the Internet and saving the file to a memory of the computing device 102. To provide a target quality-of-service 114 for the user 104, the network-on-chip 110 utilizes a first virtual channel 116 for performing the first task and a second virtual channel 1 16 for performing the second task. In this case, the first virtual channel 116 has a higher priority than the second virtual channel. This enables the computing device 102 to play the video with high resolution and / or play the video continuously without interruptions (or pauses) while the file is downloaded in the background. In this way, the system-on-chip 106 can improve the overall user experience and realize a target quality-of-service 114. The virtual channels 1 16 can also alleviate head-of-line blocking (HOLB) and enable the network-on-chip 110 to realize higher throughput while keeping the design resources and routing wires to a reduced number.
[0021] The network-on-chip 110 includes a buffer 118, which provides persistent and / or non- transitory data storage (e.g., in contrast to mere signal transmission). Any suitable type of component can be used to implement the buffer, including a flip-flop, a register, a memory cell, and so forth. The buffer 118 can be implemented using volatile memory, such as a cache memory, a random-access memory, a portion of a memory array, or flip-flops.
[0022] A size of the buffer 118 is designed to support the credit-based protocol 112 and to support the quantity of virtual channels 116. which is represented by the variable V. This means that the buffer 118 is adequately sized to account for the round-trip delay associated with the credit-based protocol 112. Furthermore, the buffer 118 is adequately sized to support the multiple virtual channels 116. The buffer 118 can alternatively be referred to as a credit buffer or a storage element. As the buffer 118 can implement aspects of supporting multiple functions of a network- on-chip, the buffer 118 can also be referred to as a “unified” buffer.
[0023] The network-on-chip 1 10 propagates data across a data path, which includes the buffer 118. In general, the buffer 118 represents a single storage element within the data path that is capable of storing data. The data path does not include another storage element capable of storing the data. The network-on-chip 110 can include other storage elements capable of storing auxiliary data that enables the network-on-chip 110 to manage and / or control the flow of data. These other storage elements, however, are not directly implemented within the data path and therefore do not store the data.
[0024] The network-on-chip 110 can provide many functions to facilitate communications between subsystems 108 in addition to those associated with the credit-based protocol 1 12 and the quality -of-service 114. The buffer 118 can provide the necessary storage to support these functions, as further described with respect to FIG. 2.
[0025] FIG. 2 illustrates example functions that are performed by the network-on-chip 110 and are supported using the buffer 118. In the example shown in FIG. 2. the network-on-chip 110 provides a communication network between a first subsystem 108-1 and a second subsystem 108-2. The first subsystem 108-1 represents a transmitter 202 (e.g., a source or an initiator), which sends data to the second subsystem 108-2 through the network-on-chip 110. The second subsystem 108-2 represents a receiver 204 (e.g.. a destination or a target), which receives the data that is sent by the first subsystem 108-1.
[0026] The subsystems 108-1 and 108-2 can operate with different clock domains 206, different voltage domains 208, and / or different pow er domains 210. Variations in the clock domains 206,the voltage domains 208, and the power domains 210 can be based on the different functionalities provided by the subsystems 108 and / or can be based on different implementations of the subsystems 108. In this example, the first subsystem 108-1 is associated with (e.g., uses or performs an operation based on) a first clock domain 206-1, a first voltage domain 208-1. and a first power domain 210-1. The second subsystem 108-2 is associated with (e.g., uses or performs an operation based on) a second clock domain 206-2, a second voltage domain 208-2, and a second power domain 210-2.
[0027] The clock domains 206-1 and 206-2 enable the subsystems 108-1 and 108-2 to perform operations based on clock signals that are generated by different sources (e.g., generated by different clock generators or different phase-locked loops). The clock signals associated with different clock domains 206 can have similar or different frequencies and / or phases. Example frequencies of the clock signals can include frequencies at approximately 400 megahertz (MHz), 500 MHz, 800 MHz, 1 GHz, 1.4 GHz, 2 GHz, and so forth. Other clock frequencies are also possible. The different clock domains 206-1 and 206-2 provide additional flexibility in designing the system-on-chip 106 and positioning the subsystems 108-1 and 108-2 within the system-on- chip 106. For instance, with the different clock domains 206-1 and 206-2, the subsystems 108-1 and 108-2 can be positioned relatively far apart compared to subsystems 108 that share a same clock domain. As the subsystems 108-1 and 108-2 can operate in different clock domains 206-1 and 206-2, the network-on-chip 110 is designed to support asynchronous communications between the subsystems 108-1 and 108-2.
[0028] The voltage domains 208-1 and 208-2 enable the subsystems 108-1 and 108-2 to use different power supply voltages. The power domains 210-1 and 210-2 enable the subsystems 108-1 and 108-2 to independently power on or off. With the different clock domains 206 and the different voltage domains 208, the subsystems 108-1 and 108-2 can utilize dynamic voltage and frequency scaling (DVFS) to improve a pow er efficiency of the system-on- chip 106. With the different power domains 210, the subsystems 108-1 and 108-2 can also reduce power consumption of the system-on-chip 106 by powering off when not in use.
[0029] The subsystems 108-1 and 108-2 can also perform communications using different data widths 212-1 and 212-2. The data widths 212 represent a quantity of bits that the subsystems 108-1 and 108-2 are designed to send or receive at one time across a virtual channel 116. Explained another way, the data width 212 represents a quantity of bits within a beat of data (e.g., a width of a beat). From the transmitter 202’s perspective, a beat represents a portion of a packet that can be w ritten to the buffer 118 of the network-on-chip 1 10 in a single writeoperation. From the receiver 204’s perspective, a beat represents a portion of a packet that can be read from the buffer 118 in a single read operation.
[0030] To support the various data widths 212-1 and 212-2, the network-on-chip 110 can perform downscaling and / or upscaling. With downscaling, the network-on-chip 110 can send, to the receiver 204, a beat of data having a data width 212-2 that is narrower than the data width 212-1 of a beat that is received from the transmitter 202. With upscaling, the network-on-chip 1 10 can send, to the receiver 204 a beat of data having a data width 212-2 that is wider than the data width 212-1 of a beat that is received from the transmitter 202.
[0031] The network-on-chip 110 can also provide voltage-domain crossing 220 using at least one level shifter 222 (e.g., a level converter, a logic level shifter, or a voltage level translator). The level shifter 222 is coupled betw een the first subsystem 108-1 and the buffer 118. For the voltagedomain crossing 220, the level shifter 222 translates data from a first logic level associated with the first voltage domain 208-1 to a second logic level associated with the second voltage domain 208-2. In some implementations, the network-on-chip 110 includes additional delay elements to manage the propagation delay associated with the level shifter 222.
[0032] The buffer 118 acts as a single storage element capable of providing storage for credit management 214, clock-domain crossing 216, and data upscaling 218. In this way, the buffer 118 can support the credit-based protocol 112 and can enable the network-on-chip 110 to support communications between subsystems 108 having different clock domains 206 and between subsystems 108 utilizing different data widths 212. The buffer 118 also supports the quality-of- service 114 by providing storage for the virtual channels 116.
[0033] The buffer 118 further enables an insertion point of the level shifter 222 to be placed before the buffer 118. As such, any propagation delay associated with the level shifter 222 can be compensated for using the buffer 118. This insertion point also makes it easier to perform the voltage-domain crossing 220 as it does not put additional constraints on the second clock domain 206-2 of the receiver 204.
[0034] With the buffer 118, the network-on-chip 110 can have a smaller footprint, can consume less power, and can have less latency compared to other network-on-chips that utilize multiple storage elements. The area savings with respect to the storage can be at least 30% (e.g., 40%, 50%, or more) in some cases. In general, reducing the quantity of storage elements reduces the power leakage, thereby improving a power efficiency of the network-on-chip 110. The latency is improved by removing an intermediate storage stage, which could otherwise require an additional clock cycle to perform a write and read operation. The buffer 118 can also be cheaper and easierto manufacture, design, integrate, and / or verify compared to other network-on-chips that use multiple storage elements. The computing device 102 is further described with respect to FIG. 3.
[0035] FIG. 3 illustrates an example computing device 102. The computing device 102 is illustrated with various non-limiting example devices including a desktop computer 102-1, a tablet 102-2, a laptop 102-3, a television 102-4, a computing watch 102-5, computing glasses 102-6, a gaming system 102-7, a router 102-8, and a vehicle 102-9. Other devices may also be used, such as a home sendee device, a smart speaker, a smart thermostat, a baby monitor, a Wi-Fi™ router, a drone, a trackpad, a drawing pad, a netbook, an e-reader, a home automation and control system, a wall display, and another home appliance. Note that the computing device 102 can be wearable, non-wearable but mobile, or relatively immobile (e.g., desktops and appliances).
[0036] The computing device 102 includes at least one system-on-chip 106. The system-on- chip 106 includes the subsystems 108-1 to 108-S and the network-on-chip 110. The network-on- chip 110 includes the buffer 118, which acts as a single storage element for supporting multiple functions of the network-on-chip 1 10. These multiple functions include the credit management 214, the clock-domain crossing 216, the data upscaling 218, the voltage-domain crossing 220. the power-domain crossing, and / or the quality-of-service 114.
[0037] The computing device 102 can also include a network interface 302 for communicating data over wired, wireless, or optical networks. For example, the network interface 302 may communicate data over a local-area-network (LAN), a wireless local-area-network (WLAN), a personal-area-network (PAN), a wire-area-network (WAN), an intranet, the Internet, a peer-to- peer network, point-to-point network, a mesh network, Bluetooth™, and the like. The computing device 102 may also include a display 304. Example components of the network-on-chip 110 are further described with respect to FIG. 4.Supporting Multiple Functions of a Network-on-Chip using a Buffer
[0038] FIG. 4 illustrates example components of a network-on-chip 110. In the depicted configuration, the network-on-chip 110 includes the level shifter 222 and the buffer 118. An example implementation of the buffer 118 is further described with respect to FIG. 5. The network-on-chip 110 also includes at least one write manager 402, at least one read manager 404, and at least one address management circuit 406.
[0039] The write manager 402 controls the writing of data from the transmitter 202 to the buffer 118. More specifically, the wri te manager 402 writes data to the buffer 118 in a structured way that can automatically upscale the data. The write manager 402 can be implemented usingmultiple address trackers 408-1 to 408-V and multiple segment trackers 406-1 to 406-V. Each address tracker 408 and each segment tracker 410 is associated with one of the virtual channels 116. The address trackers 408 keep track of address locations within the buffer 118 that are currently allocated (or assigned) to a virtual channel 116. The segment tracker 410 identifies a next segment within an address location that is available for storing a next beat that is received using the virtual channel 1 16.
[0040] The write manager 402 also provides, to the read manager 404, an indication if an address location within the buffer 118 is full and / or when a last beat of a packet has been received. An example operation of the write manager 402 is further described with respect to FIG. 6.
[0041] With the indication provided by the write manager 402, the read manager 404 can control the reading of the data from the buffer 118 and the propagation of the data to the receiver 204. The read manager 404 includes additional storage elements, which facilitates communicating auxiliary information across different clock domains 206. The address management circuit 406 performs address allocation (e.g., address assignment) and address deallocation (e.g., address release) for the buffer 118. An example implementation of the buffer 118 is further described with respect to FIG. 5.
[0042] FIG. 5 illustrates an example implementation of the buffer 118. The buffer 118 has multiple address locations 502-1, 502-2... 502 -M. Each address location 502 has multiple segments 504-1, 504-2. . . 504-N. A segment 504 can alternatively be referred to as a beat storage cell. To implement upscaling, a width 506 of each segment 504 (e.g., a quantity of bits that can be stored within each segment 504) is equal to the first data width 212-1. The first segment 504-1 can represent a least-significant segment of the address location 502 and the Nth segment 504-N can represent a most-significant segment of the address location 502. In an example implementation, data is written to an address location 502 of the buffer 118 in an order that aligns with an increasing index of the segments 504 (e.g., from the first segment 504-1, to the second segment 504-2. . . to the Nth segment 504-N). Other write orders are also possible.
[0043] A width 508 of each address location 502 (e.g., a quantity of bits that can be stored within each address location 502) is equal to the second data width 212-2. This architecture enables data to be written to each segment 504 and enables an entire address location 502 to be read for data upscaling 218. In an example implementation, each segment 504 can store 64 bits and the address location 502 can store 256 bits. This means that each address location 502 can store four beats' w orth of data as provided by the transmitter 202. In another example implementation, the address location 502 can store 512 bits. This means that each address location 502 can store eight beats’ worth of data.
[0044] During operation, the write manager 402 ensures data written to an address location 502 is associated with one of the virtual channels 1 16. This prevents data from different virtual channels 116 being merged together, which could result in data corruption and / or ordering problems. Each address location 502 stores one or more beats associated with a single virtual channel 116. The current virtual channel 116 associated with an address location 502 can change over time as an address location 502 is deallocated (e.g., released) and reallocated (e.g., reassigned). At any one time, some of the address locations 502 can store beats associated with a same virtual channel 116. Other address locations 502 can store beats associated with different virtual channels 116.
[0045] Consider an example in which the transmitter 202 sends a first packet of data having a first quantity of beats and associated with afirst virtual channel 116. If the first quantity ofbeats within the first packet is greater than the quantity of segments 504 within an address location 502 (e.g., / V). the beats are stored using multiple address locations 502. In this case, these multiple address locations 502 are associated with the same virtual channel 116.
[0046] During this time, the transmitter 202 can also send a second packet of data having a second quantity of beats and associated with a second virtual channel 116. In this case the one or more address locations 502 that store the beats of the second packet are associated with the second virtual channel 116. With this architecture, the buffer 118 can support data upscaling 218, as further described with respect to FIG. 8. The write manager 402 also supports data upscaling 218. An example operation of the write manager 402 is further described with respect to FIG. 6.
[0047] FIG. 6 illustrates an example scheme 600 implemented by the write manager 402. At 602, the network-on-chip 110 receives a beat associated with a virtual channel 116. For example, the write manager 402 receives a beat from the transmitter 202. The beat is associated with a virtual channel. The beat can be one of many beats that form the same packet.
[0048] At 604, the write manager 402 detennines if an address location 502 is associated with the virtual channel 116 (e.g.. determines if the address location 502 has been assigned or allocated to the virtual channel 116). For example, the write manager 402 determines if an address tracker 408 that corresponds to the virtual channel 116 is empty. If the address tracker 408 is empty, the process proceeds to 606. In this case, the segment tracker 410 that corresponds to the virtual channel 116 points to a first segment 504-1. Otherwise, the process continues at 608. In this case, the segment tracker 410 can point to any segment 504 of an address location 502 that is indicated by the address tracker 408.
[0049] At 606, the write manager 402 assigns an address location to the virtual channel 116. For example, the write manager 402 updates the address tracker 408 corresponding to the virtualchannel 116 to point to an address location 502. This assignment can be based on information provided by the address management circuit 406, as further described with respect to FIG. 7.
[0050] At 608, the write manager 402 writes the beat to the buffer 118. More specifically, the write manager 402 writes the beat to a segment 504 of an address location 502 that is identified by the address tracker 408. The segment 504 is identified by the segment tracker 410. At 610, the write manager 402 updates the segment tracker 410. In an example implementation, the write manager 402 updates the segment tracker 410 by incrementing the segment tracker 410 to point to a next segment 504 associated with the address location 502.
[0051] At 612. the write manager 402 detennines if a last beat in a packet has been received. If this beat is not the last beat in the packet, the process continues at 614. At 614, the write manager 402 determines if the address location 502 is full. If the address location 502 is not full, the write manager 402 does nothing further, as indicated at 616. If the address location 502 is full at 614, the process continues to 618, which is further described below.
[0052] If, at 612, the write manager 402 determines that the beat represents a last beat in the packet, the process continues to 618. At 618, the write manager 402 resets the address tracker 408 and the segment tracker 410. For example, the write manager 402 clears the address tracker 408 (e.g., causes the address tracker 408 to be empty) and sets the segment tracker 410 to point to a first segment 504-1. At 620, the write manager 402 sends control information to the read manager 404 to enable the read manager 404 to initiate a read operation and to trigger address deallocation via the address management circuit 406.
[0053] The process can repeat at 602 for a next beat that is received. The next beat may be associated with the same virtual channel 116 as the previous beat or may be associated with a different virtual channel 116. An example implementation of the network-on-chip 110 is further described with respect to FIG. 7.
[0054] FIG. 7 illustrates an example implementation of the netw ork-on-chip 110. In the depicted configuration, the network-on-chip 110 includes the buffer 118, the level shifter 222, the write manager 402, the read manager 404, and the address management circuit 406. The level shifter 222 is coupled between the transmitter 202 (not shown) and the write manager 402. The write manager 402 is coupled to the read manager 404, the address management circuit 406, and the buffer 118. The buffer 118 is coupled to the read manager 404 and the receiver 204 (not shown). The level shifter 222. the write manager 402, and the buffer 118 form a data path 700 of the network-on-chip 110. The data path 700 indicates a path along which data propagates through the network-on-chip 110 from the first subsystem 108-1 to the second subsystem 108-2.
[0055] The read manager 404 includes at least one control queue 702, at least one route detection logic 704, at least one virtual-channel controller 706 (VC controller 706), and at least one arbiter 708. The control queue 702 and the virtual-channel controller 706 can implemented using storage elements. However, these storage elements are not part of the data path 700 and therefore do not store data that propagates along the data path 700 from the transmitter 202 to the receiver 204. This enables the control queue 702 and the virtual-channel controller 706 to be implemented using storage elements with smaller sizes than the buffer 118.
[0056] The control queue 702 is coupled between the route detection logic 704 and the write manager 402. The virtual-channel controller 706 is coupled to the buffer 118 and the arbiter 708. The virtual-channel controller 706 is also coupled to the address management circuit 406, as further explained below.
[0057] The control queue 702 can be implemented using an asynchronous first-in first-out (FIFO) queue. A width of the control queue 702 can be equal to a summation of an address-width of the buffer 118 and a width of control infonnation that is provided by the write manager 402. The address-width differs from the width 508. More specifically, the address-width is a quantity of bits that are required to identify or reference the address location 502-1 while the width 508 represents a quantity of bits that the address location 502-1 can store within the buffer 118. A depth of the control queue 702 can be equal to a depth of the buffer 118.
[0058] The route detection logic 704 enables the netw ork-on-chip 110 to appropriately route data from the buffer 118 to the receiver 204. The virtual-channel controller 706 can include multiple storage elements. Each storage element can be associated with a different virtual channel 116. These storage elements can identify which address locations 502 of the buffer 118 store data that is ready to be read and provided to the receiver 204 using a corresponding virtual channel 116. The virtual-channel controller 706 can control the reading of the data within the buffer 118 based on arbitration decisions provided by the arbiter 708. The virtual-channel controller 706 also communicates with the address management circuit 406 to initiate an address deallocation process.
[0059] The address management circuit 406 includes at least one release queue 710 and at least one address manager 712. The release queue 710 is coupled between the virtual-channel controller 706 and the address manager 712. The release queue 710 can be implemented using an asynchronous first-in first-out queue, for instance. A width of the release queue 710 is at least equal to an address-width of the buffer 118. A depth of the release queue 710 is at least equal to a depth of the buffer 118.
[0060] The release queue 710 stores addresses (e.g., address pointers) of the address locations 502 within the buffer 1 18 that are ready to be released. The address manager 712 handles address allocation and deallocation of the buffer 118. In particular, the address manager 712 ensures addresses of address locations 502 that are already allocated (or are already- assigned) and are in use are not reallocated until the address-deallocation process completes.
[0061] The components of the network-on-chip 1 10 are shown relative to the different voltage domains 208-1 and 208-2 and are also shown relative to the different clock domains 206-1 and 206-2. In contrast to other types of network-on-chips, the voltage-domain crossing at 714 and the clock-domain crossing at 716 occur at separate points along the data path 700. An insertion point of the level shifter 222, which is indicated generally at 714, provides a point of transition between the voltage domains 208-1 and 208-2. The other components of the network-on-chip 110 (e.g., the write manager 402, the buffer 118, the address management circuit 406, and the read manager 404) are considered to be in the second voltage domain 208-2.
[0062] Positioning the level shifter 222 (e.g.. the voltage-domain crossing 220) at a beginning portion of the data path 700 (e g., between the transmitter 202 and the write manager 402) provides several benefits. In a first example, any delays associated with the level shifter 222 can be compensated for by the buffer 118 by increasing a size of the buffer 118. In a second example, the level shifter 222 is positioned within the write path and therefore does not impact the read operation. As such, the level shifter 222 does not place restrictions or limitations on the second clock domain 206-2. This means that the network-on-chip 110 has additional flexibility in supporting receivers 204 with various clock domains 206, including clock domains 206 associated with higher clock frequencies.
[0063] At 716, the control queue 702 and the release queue 710 support the clock-domain crossing 216 by enabling auxiliary information to flow between the clock domains 206-1 and 206-2. The control queue 702 enables information to flow from the first clock domain 206-1 to the second clock domain 206-2 for scheduling read operations. The release queue 710 enables information to flow from the second clock domain 206-2 to the first clock domain 206-1 for deallocating addresses within the buffer 118. In general, the level shifter 222, the write manager 402, and the address management circuit 406 are considered to operate in the first clock domain 206-1. The read manager 404 is considered to operate in the second clock domain 206-2.
[0064] Although the buffer 118 is shown to be within the clock domain 206- 1 in FIG. 7, it can be considered to be inserted at 716 such that it is associated with both the first clock domain 206-1 and the second clock domain 206-2. This is because the buffer 118 performs write operations based on the first clock domain 206-1 and performs read operations based on the second clockdomain 206-2. This is another way in which the buffer 118 supports the clock-domain crossing 216 performed by the network-on-chip 1 10.
[0065] Although the control queue 702 and the release queue 710 represent other storage elements of the network-on-chip 110, the control queue 702 and the release queue 710 are not implemented within the data path 700 of the network-on-chip 110. In other words, the control queue 702 and the release queue 710 do not directly store data that propagates from the transmitter 202 to the receiver 204. In general, the buffer 118 represents a single storage element that is integrated within the data path 700 and stores this data. An architecture of the buffer 118 enable the buffer 118 to support the credit management 214, the clock-domain crossing 216, and the data upscaling 218. The data upscaling 218 is also supported by the write scheme performed by the write manager 402. This write scheme is further described with respect to FIG. 8.
[0066] During operation, the level shifter 222 receives write data 714 from the transmitter 202. The write data 714 has one or more voltages (or logic levels) that are based on the first voltage domain 208-1. The level shifter 222 performs level shifting to generate the write data 714 having voltages (or logic levels) based on the second voltage domain 208-2. The write manager 402 receives the level-shifted write data 714 from the level shifter 222.
[0067] Consider a case in which the write data 714 includes a first beat associated with a first virtual channel 116 and the buffer 118 is empty (e.g., does not store any previous data). In this case, the write manager 402 sends a request 716 to the address manager 712. Based on the request 716, the address manager 712 allocates a first address location 502, which is empty, and provides an empty -address pointer 718 to the write manager 402. The empty-address pointer 718 points to the first address location 502. The write manager 402 sets a first address tracker 408 to the empty-address pointer 718. The first address tracker 408 is associated with a same virtual channel 116 as the first beat.
[0068] The write manager 402 writes the first beat to the address location 502 within the buffer 118 as specified by the first address tracker 408. More specifically, the write manager 402 writes the first beat to the segment 504 within the address location 502 as specified by the first segment tracker 410, which is associated with a same virtual channel 116 as the first beat. The write scheme performed by the write manager 402 enables the buffer 118 to support the data upscaling 218.
[0069] Once the write manager 402 detemrines that an address location 502 of the buffer 118 is full and / or a last beat of a packet has been received for a virtual channel 116, the write manager 402 generates control information 720, which is provided to the read manager 404. At a high level, the control information 720 enables the buffer 118 to perform a read operation, enablesthe data to be properly routed to the receiver 204, and enables the address management circuit 406 to release the address location 502 within the buffer 118 that previously stored the data. In example implementations, the control information 720 can include an address (e.g., an address pointer) of the address location 502 that is ready to be read. The control information 720 can also specify a virtual channel 1 16 associated with the data stored by the address location 502 as well as routing information relating to a final destination of the data (e.g., routing information relating to the second subsystem 108-2). The control queue 702 stores the control information 720 within an entry'.
[0070] The route detection logic 704 receives the control information 720 from the control queue 702. Based on the control information 720, the route detection logic 704 can determine a route for sending the data that is stored within the buffer 118 to the receiver 204. The route detection logic 704 forwards the virtual channel 116 and the address of the data associated with the control information 720 to the virtual-channel controller 706. The virtual-channel controller 706 stores the address of the data within one of its buffers.
[0071] The virtual -channel controller 706 sends a request 722 to the arbiter 708 and receives a grant 724 for any of the virtual channels 116. The virtual-channel controller 706 initiates a read operation based on the grant 724. This can include enabling data to be read from an address location 502 associated with the grant 724. The read operation causes the buffer 118 to read the data stored within the specified address location 502 and generate the read data 726. The networkon-chip 110 provides the read data 726 to the receiver 204. In example implementations, the read data 726 represents upscaled data (e.g., data that is upscaled relative to the write data 714). In other words, the read data 726 has a wider data width, which is represented by the second data width 212-2, compared to the data width of the write data 714, which is represented by the first data width 212-1.
[0072] The virtual-channel controller 706 also generates a release signal 728 (release 728), which indicates an address that is ready to be released. The release queue 710 stores the address specified by the release signal 728. The address manager 712 accesses the address 730 that is stored within the release queue 710. The address manager 712 updates its information to indicate that the address 730 is empty and available. This completes the address deallocation process. After this time, the address manager 712 can set a next empty-address pointer 718 to point to the address 730. The address location 502 corresponding to the address 730 can now be reused to store data associated with a same virtual channel 116 as the previous data or to store data associated with a different virtual channel 116.
[0073] Although the operations of the network-on-chip 1 10 are described with respect to the writing and reading of a single beat, the network-on-chip 110 can support the writing and reading of multiple beats associated with different virtual channels 116 at a same time. As such, the write manager 402 can generate the control information 720 for the different virtual channels 116 and can propagate multiple instances of the control information 720 to the control queue 702. Also, the address management circuit 406 can perform deallocation of multiple addresses 730. The writing scheme performed by the write manager 402 is further described with respect to FIG. 8.
[0074] FIG. 8 illustrates an example writing scheme for performing the data upscaling 218 using the buffer 118. For explanation purposes, the buffer 118 is shown to include four address locations 502-1, 502-2, 502-3, and 502-4. Each address location 502 includes four segments 504-1, 504-2, 504-3, and 504-4. Other implementations of the buffer 118 can have a different quantity of addresses locations 502 and / or a different quantity of segments 504. For simplicity’, the buffer 118 shown in FIG. 8 starts out as empty and does not store any previous data.
[0075] At 800, the network-on-chip 110 is shown to receive, over time, a variety of different beats, which are associated with different packets 802 and different virtual channels 116. Time progresses from left to right at the top of FIG. 8. The network-on-chip 110 receives two beats 804- 1 and 804-2. which form at least a portion of a first packet 802- 1. The first packet 802- 1 propagates through the network-on-chip 110 using a first virtual channel 116-1. As such, the first packet 802-1 (and the beats 804-1 and 804-2 associated with the first packet 802-1) are associated with the first virtual channel 116-1. The network-on-chip 110 also receives beat 806-1, which forms at least a portion of a second packet 802-2. The second packet 802-2 is associated with a second virtual channel 116-2. Additionally, the network-on-chip 110 receives beats 808-1, 808-2, 808-3, 808-4, 808-5, and 808-6, which form at least a portion of a third packet 802-3. The beat 808-6 represents a last beat 810 of the third packet 802-3. The third packet 802-3 is associated with a third virtual channel 116-3.
[0076] Using the scheme 600 described with respect to FIG. 6. the write manager 402 writes the beats to the segments 504 of the address locations 502 in an order indicated by the arrows. For example, the write manager 402 assigns the first address location 502-1 to the first virtual channel 116-1 by setting a pointer of a first address tracker 408, which corresponds to the first virtual channel 116-1, to the first address location 502-1. This assignment is based on the empty- address pointer 718 provided by the address manager 712, as explained in FIG. 7. At the time the first beat 804-1 is received, a pointer of a first segment tracker 410, w hich corresponds to the first virtual channel 116-1, is set to the first segment 504-1. Accordingly, the write manager 402 writesthe beat 804-1 to a first segment 504-1 of the first address location 502-1 based on the first address tracker 408 and based on the first segment tracker 410. The write manager 402 updates the pointer of the first segment tracker 410 to point to the second segment 504-2, as explained at 610 in FIG. 6.
[0077] When the beat 804-2 is received, the write manager 402 writes the beat 804-2 to the second segment 504-2 of the first address location 502-1 based on the first address tracker 408 and the segment tracker 410. The write manager 402 updates the pointer of the first segment tracker 41 to point to the third segment 504-3. In this case, the first address location 502-1 is reserved for beats associated with the first virtual channel 116-1 until the first address location 502-1 is released via the address management circuit 406.
[0078] When the beat 806-1 is received, the write manager 402 assigns the second address location 502-2 to the second virtual channel 116-2 by setting a pointer of a second address tracker 408, which corresponds to the second virtual channel 116-2, to the second address location 502-2. At the time the beat 806-1 is received, a pointer of a second segment tracker 410, which corresponds to the second virtual channel 116-2, is set to the first segment 504-1. Accordingly, the wri te manager 402 writes the beat 806-1 to a first segment 504-1 of the second address location 502-2 based on the second address tracker 408 and based on the second segment tracker 410. The write manager 402 updates the pointer of the second segment tracker 410 to point to the second segment 504-2.
[0079] As seen in FIG. 8, the beat 806-1 is written to the second address location 502-2 instead of the first address location 502-1 because the beat 806-1 is associated with the second virtual channel 11 -2 while the beats 804-1 and 804-2 are associated with the first virtual channel 116-1. In this case, the second address location 502-2 is reserved for beats associated with the second virtual channel 116-2 until the second address location 502-2 is released via the address management circuit 406. This write scheme therefore avoids merging beats from different virtual channels 116 within a same address location 502.
[0080] Next, the write manager 402 writes the beats 808-1, 808-2, 808-3, and 808-4 to the segments 504-1, 504-2, 504-3, and 504-4 of the third address location 502-3. After writing the beat 808-4, the write manager 402 determines that the third address location 502-3 is full and provides control information 720 to the control queue 702. The write manager 402 can determine that the third address location 502-3 is full based on the segment tracker 410 representing a last available segment 504 within the third address location 502-3 or based on the segment tracker 410 being updated (e.g., incremented) to an invalid value.
[0081] The control information 720 enables the data within the third address location 502-3 to be read. Once read, the third address location 502-3 can be released and made empty for storing new'data associated with one of the virtual channels 116. In addition to providing the control information 720, the write manager 402 resets the third address tracker 408 (e.g., causes the third address tracker 408 to be empty ) and updates the third segment tracker 410 to point to the first segment 504-1.
[0082] When the beat 808-5 is received, the write manager 402 assigns a new address location, the fourth address location 502-4, to the third virtual channel 1 16-3 by setting a pointer of the third address tracker 408 to the fourth address location 502-4. At the time the beat 808-5 is received, the pointer of the third segment tracker 410 is set to the first segment 504-1. Accordingly, the write manager 402 writes the beat 808-5 to the first segment 504-1 of the fourth address location 502-4 based on the third address tracker 408 and based on the third segment tracker 410. The write manager 402 updates the pointer of the third segment tracker 410 to point to the second segment 504-2. A similar process of referencing the third address tracker 408 and the third segment tracker 410 is used to write the beat 808-6 to the segments 504-2 of the fourth address location 502-4.
[0083] As the beat 808-6 represents a last beat 810 of the packet 802-3, the write manager 402 determines that the beat 808-6 represents an end of the packet 802-3. The write manager 402 provides control information 720 to the control queue 702. This control information 720 enables the data within the address location 502-4 to be read. Once read, the fourth address location 502-4 can be released and made empty for storing new data associated with one of the virtual channels 116. In addition to providing the control information 720, the write manager 402 resets the third address tracker 408 and updates the third segment tracker 410 to point to the first segment 504-1.
[0084] The operations described above utilize an architecture of the buffer 118 to support data upscaling 218. In this example, data with narrow widths (e.g., widths 506 of the segments 504) are received from the transmitter 202 and are written to the buffer 118. However, data with wider widths (e.g., widths 508 of the address locations 502) are read from the buffer 118 and provided to the receiver 204.Example Methods
[0085] FIGs. 9 and 10 depict example methods 900 and 1000 for implementing aspects of supporting multiple functions of a network-on-chip using a buffer. Methods 900 and 1000 are shown as a set of operations (or acts) performed but not necessarily limited to the order or combinations in which the operations are shown herein. Further, any of one or more of the operations may be repeated, combined, reorganized, or linked to provide a wide array of additionaland / or alternate methods. In portions of the following discussion, reference may be made to the environment 100 of FIG. 1, and entities detailed in FIGs. 2, 4, and 7, reference to which is made for example only. The techniques are not limited to performance by one entity or multiple entities operating on one device.
[0086] At 902 in FIG. 9, data from a first subsystem that is coupled to a network-on-chip is received by a buffer of the network-on-chip. The receiving of the data occurs based on a first clock domain associated with the first subsystem. The data has a first data width associated with the first subsystem. For example, the buffer 118 of the network-on-chip 110 receives data, such as the write data 714, from the first subsystem 108-1, as shown in FIG. 7. The first subsystem 108-1 is coupled to the network-on-chip 110, as shown in FIG. 2. In this example, the first subsystem 108-1 acts as a transmitter 202.
[0087] The receiving of the data occurs based on the first clock domain 206-1, which is associated with the first subsystem 108-1, as shown in FIGs. 2 and 7. This means that the receiving of the data occurs based on a frequency and phase of a clock signal that is associated with the first clock domain 206-1. The receiving of the data can also occur based on a second voltage domain 208-2 and / or a second power domain 210-2, which is associated with the second subsystem 108-2 (e.g., a different subsystem 108 than the subsystem 108 associated with the first clock domain 206-1), as shown in FIG. 7.
[0088] The data has a first data width 212-1 associated with the first subsystem 108-1 as shown in FIG. 2. The first data width 212-1 represents a width (e.g., a quantity of bits) of a beat that is transmitted (e.g., sent or provided) by the first subsystem 108-1. The buffer 118 is designed to have multiple segments 504 within each address location 502. Each segment 504 has a width 506 (e.g., a bit width) that is equal to the first data width 202-1.
[0089] At 904, the data is upscaled, using the buffer, from the first data width to a second data width associated with a second subsystem to generate upscaled data. The second subsystem is coupled to the network-on-chip. For example, the network-on-chip 110 upscales the write data 714 from the first data width 212-1 to a second data width 212-2 to generate upscaled data using the buffer 118. The second data width 212-2 is associated with the second subsystem 108-2. The upscaled data is represented by the read data 726 in FIG. 7. Due to the upscaling, the upscaled data (e.g., the read data 726) has a wider width than the data (e.g., the write data 714). The second subsystem 108-2 is coupled to the network-on-chip 110, as shown in FIG. 2.
[0090] At 906, the upscaled data is sent from the buffer to the second subsystem. The sending of the upscaled data occurs based on a second clock domain associated with the second subsystem. The second clock domain is different than the first clock domain. For example, the network-on-chip 110 sends the upscaled data (e.g., the read data 726) from the buffer 118 to the second subsystem 108-2. The sending of the upscaled data occurs based on the second clock domain 206-2, which is associated with the second subsystem 108-2. The second clock domain 206-2 is different than the first clock domain 206-1. The second clock domain 206-2, for instance, can be associated with a different source (e.g., a different clock generator or a different phase-locked loop). In some cases, the second clock domain 206-2 is associated with a clock frequency and / or phase that is different than a clock frequency and / or phase of the first clock domain 206-1.
[0091] At 1002 in FIG. 10, different sets of beats associated with different packets are received from a first subsystem that is coupled to a network-on-chip. The different packets are associated with different virtual channels. The receiving occurs based on a first clock domain associated with the first subsystem. For example, the write manager 402 receives different sets of beats associated with different packets 802 from the first subsystem 108-1, as shown in FIG. 8. In this example, the different sets of beats correspond to the different packets 802, respectively. Each set of beats can represent any quantity of beats, such as a single beat or multiple beats. The first subsystem 108-1 is coupled to the network-on-chip 110, as shown in FIG. 2. The different packets 802 are associated with different virtual channels 116, as shown in FIG. 8. As such, the network- on-chip 110 utilizes different virtual channels 116 to pass the different packets 802 from the first subsystem 108-1 to the second subsystem 108-2. In this example, the different packets 802 respectively correspond to the different virtual channels 116. Although not explicitly mentioned in FIG. 10, there can be other packets 802 that correspond to a same virtual channel 116. In other words, the network-on-chip 110 can receive multiple packets 802 from the first subsystem 108-1 and these multiple packets 802 can be associated with a same virtual channel 1 16. The receiving occurs based on a first clock domain 212-1, which is associated with the first subsystem 108-1, as shown in FIGs. 2 and 7.
[0092] At 1004, the different sets of beats are written to different address locations of a buffer of the network-on-chip. For example, the write manager 402 writes the different sets of beats to different address locations 502 of the buffer 118, as shown in FIG. 8. The write manager 402 uses address trackers 408 and segment trackers 410 to implement this w riting scheme, which prevents beats associated with different virtual channels 116 from being stored within different segments 504 of a same address location 502 at any given time. In this example, the different sets of beats are respectively written to different address locations 502. As explained with respect to FIG. 8, however, there can be a set of beats, such as those associated with the packet 802-3, w hich are written to multiple address locations 502, such as the address locations 502-3 and 502-4.
[0093] At 1006, first control information is sent based on an address location of the different address locations becoming full or storing a last beat of one of the different packets. The first control infonnation enables corresponding beats stored within the address location to be read from the buffer by a second subsystem that is coupled to the network-on-chip. For example, the write manager 402 sends the first control information 720 to a read manager 404 based on an address location 502 of the different address locations 502 becoming full or storing a last beat 810 of one of the different packets 802. The first control information 720 enables corresponding beats stored within the address location 502 to be read from the buffer and passed to a second subsystem 108-2 that is coupled to the network-on-chip 110.Example Computing System
[0094] FIG. 11 illustrates various components of an example computing system 1100 that can be implemented as any type of client, server, and / or computing device as described with reference to the previous FIGs. 2 and 3 to implement aspects of supporting multiple functions of a networkon-chip 106 using a buffer 118.
[0095] The computing system 1100 includes communication devices 1102 that enable wired and / or wireless communication of device data 1104 (e.g., received data, data that is being received, data scheduled for broadcast, or data packets of the data). The device data 1104 or other device content can include configuration settings of the device, media content stored on the device, and / or information associated with a user of the device. Media content stored on the computing system 1100 can include any type of audio, video, and / or image data. The computing system 1100 includes one or more data inputs 1106 via which any type of data, media content, and / or inputs can be received.
[0096] The computing system 1 100 also includes communication interfaces 1108, which can be implemented as any one or more of a serial and / or parallel interface, a wireless interface, any ty pe of network interface, a modem, and as any other type of communication interface. The communication interfaces 1108 provide a connection and / or communication links between the computing system 1100 and a communication network by which other electronic, computing, and communication devices communicate data with the computing system 1100.
[0097] The computing system 1100 includes one or more processors 1110 (e.g., any of microprocessors, controllers, and the like), which process various computer-executable instructions to control the operation of the computing system 1100. Alternatively or in addition, the computing system 1100 can be implemented with any' one or combination of hardware, firmware, or fixed logic circuitry7that is implemented in connection with processing and controlcircuits which are generally identified at 1 112. Although not shown, the computing system 1100 can include a system bus or data transfer system that couples the various components within the device. A system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus that utilizes any of a variety of bus architectures.
[0098] The computing system 1 100 also includes a computer-readable medium 11 14 (CRM 1114), such as one or more memory' devices that enable persistent and / or non-transitory data storage (i.e., in contrast to mere signal transmission), examples of which include random access memory (RAM), non-volatile memory (e.g., any one or more of a read-only memory’ (ROM), flash memory, EPROM, EEPROM, etc.), and a disk storage device. The disk storage device may be implemented as any type of magnetic or optical storage device, such as a hard disk drive, a recordable and / or rewriteable compact disc (CD), any type of a digital versatile disc (DVD), and the like. The computing system 1100 can also include a mass storage medium device (storage medium) 11 16.
[0099] The computer-readable medium 1114 provides data storage mechanisms to store the device data 1104, as w ell as various device applications and any other types of information and / or data related to operational aspects of the computing system 1100. For example, an operating system can be maintained as a computer application with the computer-readable medium 1114 and executed on the processors 1110. The device applications may include a device manager, such as any form of a control application, software application, signal-processing and control module, code that is native to a particular device, a hardware abstraction layer for a particular device, and so on.
[0100] The computing system 1100 also includes at least one network-on-chip 110. The networkon-chip 110 includes the buffer 118 and the write manager 402. In some implementations, the processor 1110, the processing and control 1112, the computer-readable medium 1114, and / or the storage medium 1116 can represent one or more subsystems 108 of a system-on-chip 106. The buffer 118 can represent a single storage-element that provides sufficient storage for credit manager 214, clock-domain crossing 216, and data upscaling 218. The write manager 402 can write data to the buffer 118 in a particular way to support the data upscaling 218.Conclusion
[0101] Although techniques using, and apparatuses including, supporting multiple functions of a network-on-chip using a buffer have been described in language specific to features and / or methods, it is to be understood that the subject of the appended claims is not necessarily limitedto the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations of supporting multiple functions of a network-on-chip using a buffer.
[0102] Some Examples are described below.
[0103] Example 1: A method performed by a network-on-chip, the method comprising: receiving, by a buffer of the network-on-chip, data from a first subsystem that is coupled to the network-on-chip, the receiving of the data occurring based on a first clock domain associated with the first subsystem, the data having a first data width associated with the first subsystem; upscaling, using the buffer, the data from the first data width to a second data width associated with a second subsystem to generate upscaled data, the second subsystem coupled to the network-on-chip; and sending the upscaled data from the buffer to the second subsystem, the sending of the upscaled data occurring based on a second clock domain associated with the second subsystem, the second clock domain being different than the first clock domain.
[0104] Example 2: The method of example 1, further comprising: propagating the data along a data path of the network-on-chip, wherein: the data path comprises a single storage element capable of storing the data; and the single storage element comprises the buffer.
[0105] Example 3: The method of any previous example, wherein: the buffer comprises multiple address locations; each address location is associated with a first quantity of bits that is equal to the second data width; each address location comprises multiple segments; and each segment is associated with a second quantity of bits that is equal to the first data width.
[0106] Example 4: The method of any previous example, further comprising: performing, by the network-on-chip, a credit-based protocol to communicate with the first subsystem, wherein the buffer has a size that accommodates a round-trip latency associated with the credit-based protocol.
[0107] Example 5: The method of any previous example, further comprising:prior to the receiving, shifting, by a level shifter of the network-on-chip, a voltage level of the data from a first voltage domain associated with the first subsystem to a second voltage domain associated with the second subsystem.
[0108] Example 6: The method of any previous example, further comprising: releasing an address location of the buffer that stored the upscaled data based on the sending of the data to the second subsystem.
[0109] Example 7: The method of any previous example, wherein: the data comprises different sets of beats associated with different packets; the different packets are associated with different virtual channels; and the method further comprises: receiving, by a write manager of the network-on-chip, the different sets of beats, the receiving occurring based on the first clock domain; writing the different sets of beats to different address locations of the buffer; and sending, based on a first address location of the different address locations becoming full or storing a last beat of one of the different packets, first control information to enable corresponding beats stored within the first address location to be read from the buffer by the second subsystem.
[0110] Example 8: The method of example 7. wherein: the different sets of beats are respectively associated with the different packets; the different packets are respectively associated with the different virtual channels; and the writing of the different sets of betas to different address locations comprises writing the different sets of beats to different address locations, respectively.
[0111] Example 9: The method of example 7 or 8, further comprising: storing, within a control queue of the network-on-chip, the first control information, wherein the first control information comprises the first address location, a virtual channel associated with the corresponding beats, and routing information.
[0112] Example 10: A method performed by a write manager of a network-on-chip, the method comprising: receiving, from a first subsystem that is coupled to the network-on-chip, different sets of beats associated with different packets, the different packets associated with different virtual channels, the receiving occurring based on a first clock domain associated with the first subsystem; writing the different sets of beats to different address locations of a buffer of the networkon-chip; andsending, based on an address location of the different address locations becoming full or storing a last beat of one of the different packets, first control information to enable corresponding beats stored within the address location to be read from the buffer and passed to a second subsystem that is coupled to the network-on-chip.
[0113] Example 11 : The method of example 10, wherein the writing comprises preventing beats associated with the different virtual channels from being written to different segments of a same address location within the buffer.
[0114] Example 12: The method of example 10 or 11, further comprising operating the write manager based on a voltage domain associated with the second subsystem.
[0115] Example 13: The method of any one of examples 10 to 12, wherein: the receiving of the different sets of beats comprises receiving a first beat associated with a first packet of the different packets; the first packet is associated with a first virtual channel of the different virtual channels; and the waiting of the different sets of beats to the different address locations comprises: determining, based on a first address tracker of the write manager, that a first address location of the different address locations is associated with the first virtual channel, the first address tracker associated with the first virtual channel; writing the first beat to a segment of the first address location that is identified by a first segment tracker of the write manager, the first segment tracker associated with the first virtual channel; and updating the first segment tracker to point to a next segment of the first address location.
[0116] Example 14: The method of example 13, wherein the determining that the first address location is associated with the first virtual channel comprises: sending a request to an address management circuit of the network-on-chip; and receiving, from the address management circuit and based on the request, a pointer to the first address location, the first address location having been previously released by the address management circuit.
[0117] Example 15: The method of example 13 or 14, wherein: the receiving of the different sets of beats comprises receiving a second beat associated with a second packet of the different packets; the second packet is associated with a second virtual channel of the different virtual channels; andthe writing of the different sets of beats to different address locations comprises: determining, based on a second address tracker of the write manager, that a second address location of the different address locations is associated with the second virtual channel, the second address tracker associated with the second virtual channel; writing the second beat to a segment of the second address location that is identified by a second segment tracker of the write manager, the second segment tracker associated with the second virtual channel; and updating the second segment tracker to point to a next segment of the second address location.
[0118] Example 16: The method of any one of examples 13 to 15, further comprising: resetting the first address tracker and the first segment tracker based on the first address location becoming full or storing the last beat of one of the different packets.
[0119] Example 17: An apparatus comprising: a network-on-chip comprising: a buffer; and a write manager that is coupled to the buffer, the network-on-chip configured to perform, using the buffer and / or the write manager, any one of the methods of examples 1 to 9 and / or any one of the methods of examples 10 to 16.
[0120] Example 18: The apparatus of example 17, wherein: the network-on-chip is configured to propagate data along a data path of the network-on- chip; the data path comprises a single storage element capable of storing the data; and the single storage element comprises the buffer.
[0121] Example 19: The apparatus of example 17 or 18, further comprising: a first subsystem configured to transmit, to the network-on-chip, data having a first data width; and a second subsystem configured to receive, from the network-on-chip, the data having a second data width that is wider than the first data width, wherein the buffer comprises multiple address locations, each address location configured to store a first quantity of bits that is equal to the second data width, each address location comprising multiple segments, each segment configured to store a second quantity of bits equal to the first data width.
[0122] Example 20: The apparatus of example 19, further comprising: a level shifter coupled betw een the first subsystem and the write manager.
[0123] Example 21 : The apparatus of any one of examples 17 to 20, wherein the write manager comprises: multiple address trackers corresponding to multiple virtual channels supported by the network-on-chip, each address tracker of the multiple address trackers configured to identify an address location that is associated with a corresponding virtual channel; and multiple segment trackers correspond to the multiple virtual channels, each segment tracker of the multiple segment trackers configured to identify a segment for writing a next beat associated with the corresponding virtual channel.
Claims
CLAIMSWhat is claimed is:
1. A method performed by a network-on-chip, the method comprising: receiving, by a buffer of the network-on-chip, data from a first subsystem that is coupled to the network-on-chip, the receiving of the data occurring based on a first clock domain associated with the first subsystem, the data having a first data width associated with the first subsystem; upscaling, using the buffer, the data from the first data width to a second data width associated with a second subsystem to generate upscaled data, the second subsystem coupled to the network-on-chip; and sending the upscaled data from the buffer to the second subsystem, the sending of the upscaled data occurring based on a second clock domain associated with the second subsystem, the second clock domain being different than the first clock domain.
2. The method of claim 1, further comprising: propagating the data along a data path of the network-on-chip, wherein: the data path comprises a single storage element capable of storing the data; and the single storage element comprises the buffer.
3. The method of any previous claim, wherein: the buffer comprises multiple address locations; each address location is associated with a first quantity of bits that is equal to the second data width; each address location comprises multiple segments; and each segment is associated with a second quantity of bits that is equal to the first data width.
4. The method of any previous claim, further comprising: performing, by the network-on-chip, a credit-based protocol to communicate with the first subsystem, wherein the buffer has a size that accommodates a round-trip latency associated with the credit-based protocol.
5. The method of any previous claim, further comprising: prior to the receiving, shifting, by a level shifter of the network-on-chip, a voltage level of the data from a first voltage domain associated with the first subsystem to a second voltage domain associated with the second subsystem.
6. The method of any previous claim, further comprising: releasing an address location of the buffer that stored the upscaled data based on the sending of the data to the second subsystem.
7. The method of any previous claim, wherein: the data comprises different sets of beats associated with different packets; the different packets are associated with different virtual channels; and the method further comprises: receiving, by a write manager of the network-on-chip, the different sets of beats, the receiving occurring based on the first clock domain; writing the different sets of beats to different address locations of the buffer; and sending, based on a first address location of the different address locations becoming full or storing a last beat of one of the different packets, first control information to enable corresponding beats stored within the first address location to be read from the buffer by the second subsystem.
8. The method of claim 7, further comprising: storing, within a control queue of the network-on-chip, the first control information, wherein the first control information comprises the first address location, a virtual channel associated with the corresponding beats, and routing information.
9. A method performed by a write manager of a network-on-chip, the method comprising: receiving, from a first subsystem that is coupled to the network-on-chip, different sets of beats associated with different packets, the different packets associated with different virtual channels, the receiving occurring based on a first clock domain associated with the first subsystem; writing the different sets of beats to different address locations of a buffer of the networkon-chip; and sending, based on an address location of the different address locations becoming full or storing a last beat of one of the different packets, first control information to enable corresponding beats stored within the address location to be read from the buffer and passed to a second subsystem that is coupled to the network-on-chip.
10. The method of claim 9, wherein the writing comprises preventing beats associated with the different virtual channels from being written to different segments of a same address location within the buffer.
11. The method of claim 9 or 10, further comprising operating the write manager based on a voltage domain associated with the second subsystem.
12. The method of any one of claims 9 to 11, wherein: the receiving of the different sets of beats comprises receiving a first beat associated with a first packet of the different packets; the first packet is associated with a first virtual channel of the different virtual channels; and the writing of the different sets of beats to the different address locations comprises: determining, based on a first address tracker of the write manager, that a first address location of the different address locations is associated with the first virtual channel, the first address tracker associated with the first virtual channel; writing the first beat to a segment of the first address location that is identified by a first segment tracker of the write manager, the first segment tracker associated with the first virtual channel; and updating the first segment tracker to point to a next segment of the first address location.
13. The method of claim 12, wherein the determining that the first address location is associated with the first virtual channel comprises: sending a request to an address management circuit of the network-on-chip; and receiving, from the address management circuit and based on the request, a pointer to the first address location, the first address location having been previously released by the address management circuit.
14. The method of claim 12 or 13, wherein: the receiving of the different sets of beats comprises receiving a second beat associated with a second packet of the different packets; the second packet is associated with a second virtual channel of the different virtual channels; and the writing of the different sets of beats to different address locations comprises: determining, based on a second address tracker of the write manager, that a second address location of the different address locations is associated with the second virtual channel, the second address tracker associated with the second virtual channel; writing the second beat to a segment of the second address location that is identified by a second segment tracker of the write manager, the second segment tracker associated with the second virtual channel; and updating the second segment tracker to point to a next segment of the second address location.
15. The method of any one of claims 12 to 14, further comprising: resetting the first address tracker and the first segment tracker based on the first address location becoming full or storing the last beat of one of the different packets.
16. An apparatus comprising: a network-on-chip comprising: a buffer; and a write manager that is coupled to the buffer, the network-on-chip configured to perform, using the buffer and / or the write manager, any one of the methods of claims 1 to 8 and / or any one of the methods of claims 9 to 15.
17. The apparatus of claim 16, wherein: the network-on-chip is configured to propagate data along a data path of the network-on- chip; the data path comprises a single storage element capable of storing the data; and the single storage element comprises the buffer.
18. The apparatus of claim 16 or 17, further comprising: a first subsystem configured to transmit, to the network-on-chip, data having a first data width; and a second subsystem configured to receive, from the network-on-chip, the data having a second data width that is wider than the first data width. wherein the buffer comprises multiple address locations, each address location configured to store a first quantity of bits that is equal to the second data width, each address location comprising multiple segments, each segment configured to store a second quantity of bits equal to the first data width.
19. The apparatus of claim 18, further comprising: a level shifter coupled between the first subsystem and the write manager.
20. The apparatus of any one of claims 16 to 19, wherein the write manager comprises: multiple address trackers corresponding to multiple virtual channels supported by the network-on-chip, each address tracker of the multiple address trackers configured to identity an address location that is allocated to a corresponding virtual channel; and multiple segment trackers correspond to the multiple virtual channels, each segment tracker of the multiple segment trackers configured to identify a segment for writing a next beat associated with the corresponding virtual channel.
Citation Information
Patent Citations
Buffering data that flows between buses operating at different frequencies
EP1032882B1
Dynamically configuring store-and-forward channels and cut-through channels in a network-on-chip
US20170063609A1
Network-On-Chip Topology Generation
US20210168038A1
Elastic buffers
US20240281388A1
System, method and apparatus for enabling partial data transfers with indicators
WO2020161465A1