Multi-channel shared cache module, cache configuration method, and electronic device

CN122654029APending Publication Date: 2026-08-28XIANGDIXIAN COMPUTING TECH (CHONGQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610458138.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-08
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]本公开的目的是提供一种多通道共享缓存模块、缓存配置方法以及电子设备,以解决缓存无法合理分配的问题

Benefits of technology

[0003]本公开的目的是提供一种多通道共享缓存模块、缓存配置方法以及电子设备,以解决缓存无法合理分配的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654029A_ABST
    Figure CN122654029A_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-channel shared cache module, a cache configuration method and an electronic device. The multi-channel shared cache module is deployed in a network-on-chip and is used for access control of multiple access channels, wherein any access channel is used to realize access between a master module and a slave module. The multi-channel shared cache module comprises a shared cache control module. The shared cache control module is configured to, according to a preset control period, respectively for each access channel, allocate a cache space for the access channel according to access demand data of the master module of the access channel within a preset time period and processing capacity data of the slave module of the access channel at present, to obtain multiple cache areas corresponding to the multiple access channels respectively. The space size of the cache area of each access channel is proportional to the access demand data of the master module within the preset time period and inversely proportional to the processing capacity data of the slave module at present.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of on-chip network technology, and in particular to a multi-channel shared cache module, a cache configuration method, and an electronic device. Background Technology

[0002] Modern SoCs increasingly favor multi-core architectures. Due to resource and bandwidth considerations, on-chip networks typically involve a topology where multiple master modules access multiple slave modules. When there is a bandwidth mismatch between master and slave modules, a caching module is usually added at the point of mismatch to mitigate the bandwidth drop caused by the rate mismatch from high to low bandwidth, as well as the backpressure on upstream components. However, current caching schemes do not allocate cache appropriately based on actual access conditions. Summary of the Invention

[0003] The purpose of this disclosure is to provide a multi-channel shared cache module, a cache configuration method, and an electronic device to solve the problem of unreasonable cache allocation.

[0004] According to a first aspect of this disclosure, a multi-channel shared cache module is provided, deployed on an on-chip network, for access control of multiple access channels, wherein any access channel is used to enable access between a master module and a slave module; the multi-channel shared cache module includes a shared cache control module. The shared cache control module is configured to allocate cache space for each access channel according to a preset control cycle, based on the access demand data of the master module of the access channel within a preset time period and the current processing capacity data of the slave module of the access channel, thereby obtaining multiple cache areas corresponding to the multiple access channels respectively; wherein, the size of the cache area of ​​each access channel is directly proportional to the access demand data of its master module within a preset time period and inversely proportional to the current processing capacity data of its slave module.

[0005] In one embodiment, the shared cache control module is further configured to: upon receiving an access request, determine the target cache area based on the master module and slave module of the access request, and store the access request in the target cache area.

[0006] In one embodiment, the multi-channel shared cache module further includes a channel allocation module; The channel allocation module is configured to obtain access requests from the multiple buffers and send the access requests to the corresponding slave module.

[0007] In one implementation, the shared cache control module includes a reader; The channel allocation module is specifically configured to receive the current credit value of each slave module and send the credit value to the reader; The shared cache control module is specifically configured to use a reader to determine the target slave module based on the credit value of each slave module, determine the target cache area of ​​the target slave module from multiple cache areas, and obtain the access request from the target cache area and send it to the channel allocation module; the credit value represents the number of access requests that the slave module can currently receive. The channel allocation module is specifically configured to determine the target slave module of the received access request and send the received access request to the target slave module.

[0008] In one implementation, the access demand data includes access bandwidth; the processing capacity data includes a credit value, which represents the number of access requests that the module can currently receive.

[0009] In one implementation, the shared cache control module is specifically configured to allocate cache space for each access channel according to the following formula: based on the access bandwidth of the master module within a preset time period and the current credit value of the slave module within the access channel:

[0010] Where p is the buffer for the access channels of the corresponding master module i and slave module j; is the size of the cache p; M is the number of master modules; N is the number of slave modules; Bi is the access bandwidth of master module i (1≤i≤M) within a preset time period; Q represents the current credit value of module j (1≤j≤N); Q represents the total size of the cache.

[0011] In one implementation, the shared cache control module is further configured to apply back pressure to the main module corresponding to the cache area after the cache area is full; and to re-determine the multiple cache areas corresponding to the multiple access channels after a preset control period is reached.

[0012] According to a second aspect of this disclosure, a cache configuration method is provided, applied to a multi-channel shared cache module of an on-chip network. The multi-channel shared cache module is used to control access to multiple access channels, wherein any access channel is used to enable access between a master module and a slave module. The method includes: According to a preset control cycle, for each access channel, cache space is allocated to the access channel based on the access demand data of the main module within a preset time period and the current processing capacity data of the slave module within the access channel, resulting in multiple cache areas corresponding to the multiple access channels respectively; wherein, the size of the cache area of ​​each access channel is directly proportional to the access demand data of its main module within a preset time period and inversely proportional to the current processing capacity data of its slave module.

[0013] According to a third aspect of this disclosure, a graphics processing system is provided, including the multi-channel shared cache module of the first aspect.

[0014] According to a fourth aspect of this disclosure, an electronic device is provided, including the graphics processing system of the third aspect.

[0015] According to a fifth aspect of this disclosure, an electronic device is provided, including the electronic device of the fourth aspect. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the structure of an on-chip network provided in one embodiment of the present disclosure; Figure 2 This is a schematic diagram of another on-chip network structure provided in one embodiment of the present disclosure; Figure 3 This is a schematic diagram of the structure of a multi-channel shared cache module provided in one embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of a reader provided in one embodiment of the present disclosure; Figure 5 This is a schematic diagram of a graphics processing system structure provided in one embodiment of the present disclosure. Detailed Implementation

[0017] Before introducing the embodiments of this disclosure, it should be noted that: Some embodiments of this disclosure are described as processing flows. Although the various operational steps of the flow may be numbered sequentially, the operational steps may be performed in parallel, concurrently, or simultaneously.

[0018] The embodiments disclosed herein may use terms such as "first," "second," etc., to describe various features, but these features should not be limited by these terms. These terms are used merely to distinguish one feature from another.

[0019] The term “and / or” may be used in embodiments of this disclosure, and “and / or” includes any and all combinations of one or more of the associated features listed.

[0020] It should be understood that when describing the connection or communication relationship between two components, unless it is explicitly stated that the two components are directly connected or directly communicating, the connection or communication between the two components can be understood as a direct connection or communication, or it can be understood as an indirect connection or communication through an intermediate component.

[0021] To make the technical solutions and advantages of the embodiments of this disclosure clearer, the exemplary embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.

[0022] Modern SoCs increasingly favor multi-core architectures. Due to resource and bandwidth considerations, on-chip networks typically involve a topology where multiple master modules access multiple slave modules. When there is a bandwidth mismatch between master and slave modules, a caching module is usually added at the point of mismatch to mitigate the bandwidth drop caused by the rate mismatch from large to small bandwidth, as well as the backpressure on upstream components.

[0023] like Figure 1 As shown, due to the bandwidth mismatch between router A and router B, a caching module is needed. However, since main module 1 and main module 2 do not utilize the on-chip network exactly the same in terms of space and time, one module may continuously send multiple accesses, exhausting the cache module's resources and thus blocking the other module from sending requests, reducing the module's average access bandwidth and increasing access latency. A common solution to this problem is to allocate a separate cache space for main module 1 and main module 2 within the cache module. In this way, even if main module 1 exhausts its cache space, it will not prevent main module 2 from sending requests. However, in different scenarios, the two modules have different cache space requirements. Allocating a fixed cache space for each module will result in wasted or insufficient cache space. Even if virtual channel technology is used between router A and router B, allocating a virtual channel to each main module, a fixed cache space is still needed for each module. In this case, cache resources increase linearly with the number of main modules. That is, the current caching scheme does not allocate cache reasonably according to the actual access situation of each main module.

[0024] To address the aforementioned issues, this disclosure proposes a multi-channel shared cache module and a cache configuration method. This method proposes to allocate a cache space reasonably to each access channel based on the actual access situation of the master module and slave module of each access channel, thereby enabling multiple master modules to fully reuse the cache, reducing the cache area on the chip, and improving the utilization efficiency of the cache area.

[0025] Specifically, such as Figure 2As shown, this disclosure proposes a multi-channel shared cache module deployed on a network-on-a-chip (NIC) for access control of multiple access channels. Each access channel is used to enable access between a master module and a slave module. The multi-channel shared cache module can be deployed in a router or as an independent module on the NIC. This disclosure does not limit this. As long as there are multiple access channels, i.e., multiple master modules accessing multiple slave modules, and the bandwidth of the multiple master modules is greater than the bandwidth of the multiple slave modules, i.e., the bandwidth of the upstream device is greater than that of the downstream device, and a scenario where caching is needed to cache data from the upstream device, the multi-channel shared cache module proposed in this disclosure can be deployed between the upstream and downstream devices.

[0026] The following describes the functions of the multi-channel shared cache module proposed in this disclosure. The multi-channel shared cache module includes a shared cache control module.

[0027] The shared cache control module is configured to allocate cache space for each access channel according to a preset control cycle, based on the access demand data of the master module of the access channel within a preset time period and the current processing capacity data of the slave module of the access channel, thereby obtaining multiple cache areas corresponding to multiple access channels respectively; wherein, the size of the cache area of ​​each access channel is directly proportional to the access demand data of its master module within a preset time period and inversely proportional to the current processing capacity data of its slave module.

[0028] by Figure 2 Taking the content shown as an example, it includes two main modules and two slave modules. The two main modules are main module 1 and main module 2, and the two slave modules are slave module 1 and slave module 2. The input of the shared cache control module comes from the two main modules respectively, and the downstream output of the multi-channel shared cache module is the two slave modules. That is, the multi-channel shared cache module is used to control access to four access channels. The four access channels are the access channel of main module 1 to slave module 1, the access channel of main module 1 to slave module 2, the access channel of main module 2 to slave module 1, and the access channel of main module 2 to slave module 2. The four access channels transmit data through a shared physical link, that is, the physical link where the channel shared cache module is located.

[0029] The shared cache control module can allocate cache space to each access channel according to a preset control cycle, thereby obtaining the cache area for each access channel.

[0030] The following explains how the shared cache control module allocates a cache area for each access channel.

[0031] In one implementation, the access demand data includes access bandwidth; the processing capacity data includes a credit value, wherein the credit value represents the number of access requests that the module can currently receive.

[0032] The shared cache control module can obtain the access bandwidth of the master module within a preset time period and the current credit value of the slave module for each access channel, and then allocate cache space for the access channel based on the obtained information. Specifically, when obtaining the access bandwidth, the shared cache control module can use a counter to count the total number of bytes contained in the data packets of access requests sent by each master module within the preset time period, where the duration of the preset time period is no longer than the duration of the preset control cycle. The shared cache control module can receive the credit values ​​sent by each slave module from downstream devices; the credit values ​​are the instantaneous values ​​of the slave modules.

[0033] The information obtained can be seen in Table 1.

[0034]

[0035] Table 1 Furthermore, the shared cache control module is specifically configured to allocate cache space for each access channel according to the following formula (1): based on the access bandwidth of the main module within a preset time period and the current credit value of the slave module within the access channel: (1) Where p is the buffer for the access channels of the corresponding master module i and slave module j; is the size of the cache p; M is the number of master modules; N is the number of slave modules; Bi is the access bandwidth of master module i (1≤i≤M) within a preset time period; Let p be the current credit value of module j (1≤j≤N); Q is the total size of the cache; where p = M(j-1)+i (1≤i≤M, 1≤j≤N).

[0036] Taking the content shown in Table 1 above as an example, Assuming the total cache size is Q, the size of the four cache blocks is calculated using the formula as follows: buffer_size1=(B1 / (B1+B2))·(C2 / (C1+C2))·Q; buffer_size2=(B1 / (B1+B2))·(C1 / (C1+C2))·Q; buffer_size3=(B2 / (B1+B2))·(C2 / (C1+C2))·Q; buffer_size4=(B2 / (B1+B2))·(C1 / (C1+C2))·Q; In another implementation, the access demand data is specifically access latency; the processing capacity data is specifically a credit score.

[0037] The shared cache control module can obtain the access latency of the master module and the current credit value of the slave module for each access channel within a preset time period, and then allocate cache space for the access channel based on the obtained information. Specifically, when obtaining the access latency, the shared cache control module can use a counter to record the sending time and receiving time of each master module's access request within the preset time period. The difference between the response time and the sending time is used as the access latency of an access request. The average access latency of all access requests within the preset time period is then calculated as the access latency of the master module within the preset time period.

[0038] Furthermore, the shared cache control module is specifically configured to allocate cache space for each access channel according to the following formula (2): based on the access latency of the main module within a preset time period and the current credit value of the slave module within the access channel. (2) Where p is the buffer for the access channels of the corresponding master module i and slave module j; is the size of the cache p; M is the number of master modules; N is the number of slave modules; Li is the access latency of master module i (1≤i≤M) within a preset time period; Let p be the current credit value of module j (1≤j≤N); Q is the total size of the cache; where p = M(j-1)+i (1≤i≤M, 1≤j≤N).

[0039] In another implementation, the access demand data specifically refers to the access bandwidth; the processing capacity data specifically refers to the processing speed of the module.

[0040] The shared cache control module can obtain the access bandwidth of the master module within a preset time period and the current processing speed of the slave module for each access channel, and then allocate cache space to the access channel based on the obtained information. The processing speed can be the average response time of the slave module in processing an access request, or the response time of the slave module in processing the most recent access request. When obtaining the current processing speed of the slave module, it can be obtained directly from the processing speed reported by the slave module, or it can be obtained based on active monitoring. This disclosure does not limit this.

[0041] Furthermore, the calculation method for allocating cache space for each access channel by the shared cache control module based on the access bandwidth of the main module within the preset time period and the current processing speed of the module can be referred to the above formula (1). In this case, the credit value is simply replaced with the processing speed, and this disclosure will not elaborate further.

[0042] In addition, the access demand data can specifically be access latency; the processing capacity data can specifically be the processing speed of the module. In this case, when the shared cache control module is the cache space for each access channel, it can refer to the above formula (2), where only the credit value needs to be replaced with the processing speed. This disclosure will not elaborate further on this.

[0043] Those skilled in the art can select the appropriate type of access requirement data and the type of processing capability data according to actual needs to implement the above-mentioned cache space allocation scheme.

[0044] After the shared cache control module allocates a cache area for each access channel, the shared cache control module is also configured to: upon receiving an access request, determine the target cache area based on the master module and slave module of the access request, and store the access request in the target cache area.

[0045] For example, refer to Figure 3 still with Figure 2 Taking the two master modules and two slave modules shown as an example, the shared cache control module allocates four cache areas for the four access channels. If the master module of the received access request is master module 1 and the slave module is slave module 1, then the access request is stored in cache area 1. And so on, the shared cache control module can store the received access requests in different cache areas respectively.

[0046] like Figure 3 As shown, the multi-channel shared cache module includes a channel allocation module in addition to the shared cache control module.

[0047] The channel allocation module is configured to retrieve access requests from multiple buffers and send the access requests to the corresponding slave module.

[0048] Specifically, in one implementation, such as Figure 3 As shown, the shared cache control module includes a reader; The channel allocation module is specifically configured to receive the current credit value of each slave module and send the credit value to the reader; The shared cache control module is specifically configured to use a reader to determine the target slave module based on the credit value of each slave module, determine the target cache area of ​​the target slave module from multiple cache areas, and obtain the access request from the target cache area and send it to the channel allocation module; the credit value indicates the number of access requests that the slave module can currently receive; The channel allocation module is specifically configured to determine the target slave module of the received access request and send the received access request to the target slave module.

[0049] Specifically, when the shared cache control module uses the reader to determine the target slave module based on the credit value of each slave module, it can determine the slave module with the highest credit value as the target slave module.

[0050] by Figure 2 ,and Figure 3 Taking the content shown as an example, after receiving the credit values ​​of slave module 1 and slave module 2, if the reader determines that the credit value of slave module 1 is greater than that of slave module 2, then slave module 1 is determined as the target slave module. Cache 1 and cache 3 are used to cache access requests sent to slave module 1, so cache 1 and cache 3 are determined as the target cache. Furthermore, a round-robin method can be used to determine which cache 1 and cache 3 is selected as the final target cache, and the first access request is obtained from it and sent to slave module 1. Alternatively, a cache 1 and cache 3 can be selected as the final target cache based on the principle of "prioritizing authorization to objects that have not been selected recently".

[0051] In another implementation, the shared cache control module is specifically configured to determine the target cache from multiple caches in a round-robin manner, and obtain access requests from the target cache and send them to the channel allocation module; The channel allocation module is specifically configured to determine the target slave module of the received access request and send the received access request to the target slave module.

[0052] In another implementation, the shared cache control module is specifically configured to determine the cache with the longest recent unselected time from multiple caches based on the recent unselected time of each cache, and obtain access requests from the target cache and send them to the channel allocation module. The channel allocation module is specifically configured to determine the target slave module of the received access request and send the received access request to the target slave module.

[0053] In one implementation, the shared cache control module is further configured to apply back pressure to the main module corresponding to the cache area after the cache area is full; that is, it will not receive access requests sent by the main module to the access channel. After the preset control period is reached, multiple buffers corresponding to multiple access channels are reassigned. That is, after the duration of the preset control period, the above method is re-executed to reallocate buffers for each access channel.

[0054] Since the access requirements of each main module vary at different times, and the processing capacity of each slave module varies at different times, the cache area can be reallocated more reasonably by redistributing it according to the actual access requests of the main and slave modules every preset control period.

[0055] To facilitate understanding of this solution, the following will be combined with... Figure 3 The contents shown further describe the structure and function of the multi-channel shared cache module proposed in this disclosure.

[0056] Multi-channel shared cache module microarchitecture such as Figure 3 As shown. The multi-channel shared buffer module mainly includes: a shared buffer control module, a channel allocation module, and a shared link between the two modules. The shared buffer control module mainly consists of a counter, a decoder, a buffer area, and a reader. Access issued by the master module is transmitted in the logical layer of the on-chip network according to flits (Flow Control Units). The header of each flit contains a source ID and a destination ID. The source ID indicates which master module the flit comes from, and the destination ID indicates the slave module to which the flit is destined. Based on the flit at the input port, the shared buffer control module, in turn, writes the "header + payload" into the corresponding buffer area in the decoder according to the difference between the source ID and the destination ID. The reader, based on the slave module credit value returned by the shared link, prioritizes allocating data to slave modules with higher credit values. The credit value indicates the number of flits that the downstream module can still receive. The channel allocation module sends the flit to the corresponding slave module according to the destination ID in the header of the input flit.

[0057] The counter counts the number of valid bytes of load passing through this port within a certain period. Typically, all submodules of the shared buffer control module use the same clock source, so the counter's result can be used to represent the bandwidth of the port's flit. The counter's enable signal can be configured via software or set to be triggered by a hardware timer.

[0058] The decoder decodes based on the source ID and destination ID in the flit packet header. Different combinations of source ID and destination ID are allocated a separate buffer. (As above) Figure 3As shown, the combination of {SrcID, DstID} is encoded into a 2-bit number, corresponding to buffers 1-4 respectively. Each buffer has a start address and an end address; the end address of the previous buffer is the end address of the next buffer, and the start and end addresses determine the size of the buffer. For each buffer, the decoder maintains a set of read / write control logic, ensuring that the use of buffers by different master modules and different slave modules does not interfere with each other.

[0059] The reader, based on the slave module's credit value carried in the shared link, uses a credit-based arbitration method to read the data needed by the slave module from the buffer. Each module reads only one request per arbitration round. The specific arbitration strategy is as follows: the reader first selects the slave module with the highest credit value; then it sends a blocking signal to the buffers corresponding to the remaining slave modules, even if these buffers also have requests to send downstream; finally, it selects a buffer based on the principle of "prioritizing authorization to the least recently unselected object," and sends the first request from that buffer to the downstream module via the shared link. In the above steps, after blocking the buffer, in addition to the already described authorization method, fixed-order arbitration, round-robin arbitration, and rotation arbitration can also be used to authorize the unblocked buffers.

[0060] Taking a scenario with 2 master modules and 2 slave modules as an example, the reader is as follows: Figure 5 As shown. The shared link carries credit1 and credit2, representing credit values ​​from slave module 1 and slave module 2 respectively. The comp module compares the two values ​​and generates a port_mask signal. If credit1 is larger, port_mask=2'b10 is generated, directly masking ports_01 and_11. Therefore, the reader will only read the required flits from buffer 1 and buffer 3. Conversely, if credit1 is smaller, the reader will only read the required flits from buffer 2 and buffer 4. If credit1 and credit2 are equal, port_mask=2'b00 is generated, and no ports are masked. Finally, for all unmasked ports, one port is authorized based on the principle of "prioritizing authorization to objects that have not been selected recently," and a request from the corresponding buffer is sent to the downstream module through that port.

[0061] The channel allocation module decodes the DstID carried in the header of the upstream flit packet into a unique hot code. Based on the unique hot code, it authorizes the corresponding slave module, allocating a flit to only one slave module at a time. Simultaneously, the slave module's credit value is returned to the shared link from the channel allocation module's input port.

[0062] Because the bandwidth requirements of each master module for the slave modules differ in different scenarios, the shared cache control module's cache area needs to support dynamic adjustment. The specific steps are as follows: After the SoC enters a specific working mode, the counter can be enabled to start working through software configuration, or the hardware timer can be configured to be triggered for counting. After a certain period of time, the shared cache control module automatically counts the number of bytes contained in the transaction packets sent by the main module 1 and the main module 2 within a certain period of time, and uses them as access bandwidths B1 and B2. The hardware automatically records the credit values ​​C1 and C2 received by the shared cache control module from downstream slave module 1 and slave module 2. These credit values ​​are instantaneous values ​​fed back by the downstream slave modules during the calculation. Assuming the total cache size is Q, the size of the four cache blocks can be calculated using the formula: buffer_size1=(B1 / (B1+B2))·(C2 / (C1+C2))·Q; buffer_size2=(B1 / (B1+B2))·(C1 / (C1+C2))·Q; buffer_size3=(B2 / (B1+B2))·(C2 / (C1+C2))·Q; buffer_size4=(B2 / (B1+B2))·(C1 / (C1+C2))·Q; Back pressure is applied to the upstream of the shared cache control module. After the cache is emptied, the start and end addresses of each cache are modified according to the calculated buffer size parameters. After the modifications are completed and the back pressure on the upstream is released, the on-chip network will start working normally.

[0063] Based on the same inventive concept, this disclosure also proposes a cache configuration method applied to a multi-channel shared cache module of an on-chip network. The multi-channel shared cache module is used to control access to multiple access channels, wherein any access channel is used to implement access between a master module and a slave module. The method includes: According to the preset control cycle, for each access channel, cache space is allocated to the access channel based on the access demand data of the main module of the access channel within the preset time period and the current processing capacity data of the slave module of the access channel, resulting in multiple cache areas corresponding to multiple access channels respectively; wherein, the size of the cache area of ​​each access channel is directly proportional to the access demand data of its main module within the preset time period and inversely proportional to the current processing capacity data of its slave module.

[0064] The specific implementation of the above cache configuration method can be found in the description of the function of the multi-channel shared cache module above, and will not be repeated here.

[0065] Based on the same inventive concept, such as Figure 5 As shown, this disclosure also proposes a graphics processing system, which is as follows: Figure 5 As shown, it includes at least: A GPU core can be understood as the graphics processor mentioned above, used to process commands, such as drawing commands, and execute the image rendering pipeline based on the drawing commands. The GPU core mainly contains computing units, which execute the compiled instructions of shaders; these are programmable modules composed of numerous ALUs; a cache (memory) used to cache data from the GPU core to reduce memory access; and a controller (not shown in the diagram). In addition, the GPU core also has various functional modules, such as rasterization (a fixed stage in the 3D rendering pipeline), tilling (slicing a frame in TBR and TBDR GPU architectures), clipping (a fixed stage in the 3D rendering pipeline that clips primitives outside the viewing area or those not displayed on the back), and post-processing (scaling, clipping, rotating, etc., of the drawn image).

[0066] General-purpose DMA is used to perform data transfer between host memory and GPU memory. For example, for vertex data used in 3D drawing, general-purpose DMA moves vertex data from host memory to GPU memory. The on-chip network is used for data exchange between various masters and slaves on the SOC. The multi-channel shared cache module mentioned above is deployed on this on-chip network to execute the routing configuration method described above.

[0067] The application processor is used to schedule tasks of various modules on the SOC. For example, after the GPU finishes rendering a frame, it notifies the application processor, which then starts the display controller to display the image drawn by the GPU on the screen. The PCIe controller is the interface used for communication with the host computer. It implements the PCIe protocol, allowing the GPU (graphics card) to connect to the host computer via the PCIe interface. The host computer runs graphics APIs and graphics card drivers, among other programs. The memory controller is used to connect memory devices and store data on the SOC. The display controller is used to control the output of the frame buffer in memory to the monitor via a display interface (HDMI, DP, etc.); A video decoder is used to decode encoded video on the host hard drive into a displayable image; A video encoder is used to encode the raw video stream on the host hard drive into a specified format and return it to the host.

[0068] Based on the same inventive concept, this disclosure also provides an electronic component that includes the graphics processing system described in any of the above embodiments. In some use cases, the electronic component is presented as a graphics card; in other use cases, the electronic component is presented as a CPU motherboard.

[0069] This disclosure also provides an electronic device that includes the aforementioned electronic components. In some usage scenarios, the electronic device is in the form of a portable electronic device, such as a smartphone, tablet computer, or VR device; in other usage scenarios, the electronic device is in the form of a personal computer or game console.

[0070] Although preferred embodiments of this disclosure have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this disclosure.

[0071] Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its spirit and scope. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.

Claims

1. A multi-channel shared cache module, deployed on an on-chip network, for access control of multiple access channels, wherein any access channel is used to enable access between a master module and a slave module; the multi-channel shared cache module includes a shared cache control module; The shared cache control module is configured to allocate cache space for each access channel according to a preset control period, based on the access demand data of the master module of the access channel within a preset time period and the current processing capacity data of the slave module of the access channel, thereby obtaining multiple cache areas corresponding to the multiple access channels respectively; wherein... The size of the buffer space for each access channel is directly proportional to the access demand data of its main module within a preset time period, and inversely proportional to the current processing capacity data of its slave module.

2. The multi-channel shared cache module according to claim 1, The shared cache control module is further configured to: upon receiving an access request, determine the target cache area based on the master module and slave module of the access request, and store the access request in the target cache area.

3. The multi-channel shared cache module according to claim 2, wherein the multi-channel shared cache module further includes a channel allocation module; The channel allocation module is configured to obtain access requests from the multiple buffers and send the access requests to the corresponding slave module.

4. The multi-channel shared cache module according to claim 3, wherein the shared cache control module includes a reader; The channel allocation module is specifically configured to receive the current credit value of each slave module and send the credit value to the reader; The shared cache control module is specifically configured to use a reader to determine the target slave module based on the credit value of each slave module, determine the target cache area of ​​the target slave module from multiple cache areas, and obtain the access request from the target cache area and send it to the channel allocation module; the credit value represents the number of access requests that the slave module can currently receive. The channel allocation module is specifically configured to determine the target slave module of the received access request and send the received access request to the target slave module.

5. The multi-channel shared cache module according to claim 1, wherein the access demand data includes access bandwidth; and the processing capacity data includes a credit value, wherein the credit value represents the number of access requests that the module can currently receive.

6. The multi-channel shared cache module according to claim 5, The shared cache control module is specifically configured to allocate cache space for each access channel according to the following formula: based on the access bandwidth of the master module within a preset time period and the current credit value of the slave module within the access channel: in, p is the buffer for the access channels of the corresponding main module i and slave module j; is the size of the cache p; M is the number of master modules; N is the number of slave modules; Bi is the access bandwidth of master module i (1≤i≤M) within a preset time period; Q represents the current credit value of module j (1≤j≤N); Q represents the total size of the cache.

7. The multi-channel shared cache module according to claim 1, The shared cache control module is further configured to apply back pressure to the main module corresponding to the cache area after the cache area is full; and to re-determine the multiple cache areas corresponding to the multiple access channels after the preset control period is reached.

8. A cache configuration method applied to a multi-channel shared cache module of an on-chip network, wherein the multi-channel shared cache module is used to control access to multiple access channels, wherein any access channel is used to enable access between a master module and a slave module; the method includes: According to a preset control cycle, for each access channel, cache space is allocated to the access channel based on the access demand data of the main module within a preset time period and the current processing capacity data of the slave module within the access channel, resulting in multiple cache areas corresponding to the multiple access channels respectively; wherein, the size of the cache area of ​​each access channel is directly proportional to the access demand data of its main module within a preset time period and inversely proportional to the current processing capacity data of its slave module.

9. A graphics processing system comprising the multi-channel allocation module as described in any one of claims 1-7.

10. An electronic device comprising the graphics processing system of claim 9.

11. An electronic device comprising the electronic device of claim 10.