An on-chip cache device and a reading and writing method
By setting up read and write processing modules, cache modules and memory modules on the chip, the problems of wasted storage space and excessive power consumption in packet packet processing are solved, and more efficient storage space utilization and power consumption reduction are achieved.
Patent Information
- Application Number
- CN202010306644.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-04-17
AI Technical Summary
In existing chip designs, the problem of waste of storage space and excessive power consumption due to small packet packet processing, especially in big data processing scenarios, chip area and resource utilization are not optimized enough.
The on-chip cache device is adopted, including a read and write processing module, a cache module and a memory module. Through the read and write processing module, the message is written to the cache module and transferred to the memory module. The cache module temporarily caches the packets, realizes the temporary storage of multiple messages, improves the utilization of storage space, and reduces chip power consumption.
It improves the utilization rate of storage space, reduces the memory area and power consumption of the chip, and optimizes the resource utilization of the chip.
Smart Images

Figure CN113535633B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication chips, and particularly to an on-chip cache device and method. Background Art
[0002] In a network switching engine, due to the requirements for improving performance metrics and the limitations of chip area, resources, and industry, it is impossible for a chip to meet the demand for performance improvement by infinitely increasing the main frequency and stacking resources. Therefore, there is an urgent need for a way to achieve low redundancy and low area consumption of chip space to improve chip performance.
[0003] In existing chip designs, the utilization methods of the resource space of chips in different application scenarios are different. In the big data processing scenario, due to the generation of small packet messages, and the storage in the chip can only be performed by row address. When the chip processes small packet messages, only one small packet message can be stored at each row address, wasting a large amount of storage space. To process small packet messages, the storage area of the chip design is large, resulting in a problem of excessive power consumption. Summary of the Invention
[0004] Embodiments of this application provide an on-chip cache device and a read / write method.
[0005] Embodiments of this application provide an on-chip cache device, which includes:
[0006] A read / write processing module, a cache module, and a memory module; wherein, the read / write processing module is respectively connected to the cache module and the memory module, and the read / write processing module is used to store messages into the cache module and the memory module, read the messages stored in the cache module and the memory module, and transfer the messages cached in the cache module to the memory module; the cache module is connected to the memory module through the read / write processing module, and the cache module includes at least one cache register for temporarily caching messages; the memory module is connected to the read / write processing module and is used to transfer the messages cached in the cache module.
[0007] Embodiments of this application provide an on-chip cache reading method, which includes:
[0008] Storing the obtained message into the cache register of the cache module according to the row address; determining that all the cache registers corresponding to the row address are occupied, and transferring the message corresponding to the row address in the cache module to the memory module; reading the stored messages in the cache module and / or the memory module according to the message reading address.
[0009] In the embodiments of the present application, by providing a read / write processing module, a cache module, and a memory module on a chip, the read / write processing module writes packets into the cache module, and then writes the packets cached in the cache module into the memory module. The read / write processing module reads the packets stored in the cache module and the memory module, realizing the caching of packets in the memory module. By temporarily storing multiple packets through the cache module, the memory module can write multiple packets simultaneously, improving the utilization rate of the storage space, reducing the storage area of the chip, and reducing the power consumption of the chip.
[0010] More descriptions regarding the above embodiments and other aspects of the present application and their implementation manners are provided in the accompanying drawings, the specific implementation manners, and the claims. Brief Description of the Drawings
[0011] Figure 1 FIG. is a schematic structural diagram of an on-chip caching device provided by an embodiment of the present application;
[0012] Figure 2 FIG. is a schematic structural diagram of an on-chip caching device provided by an embodiment of the present application;
[0013] Figure 3 FIG. is a schematic diagram of a depth relationship provided by an embodiment of the present application;
[0014] Figure 4 FIG. is a configuration example diagram of an on-chip caching device provided by an embodiment of the present application;
[0015] Figure 5 FIG. is a flowchart of an on-chip caching method provided by an embodiment of the present application;
[0016] Figure 6 FIG. is a flowchart of another on-chip caching method provided by an embodiment of the present application. Detailed Description of the Embodiments
[0017] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily.
[0018] Since the memory in the chip stores packets by row address and only one packet is stored in each row of the memory, when the chip is applied to large data volume processing, due to the generation of small packet-length packets during the processing, the length of this packet is less than that of ordinary packets. When the chip writes and reads small packet-length packets, it causes a large amount of storage space waste in the chip memory. Correspondingly, a larger memory needs to be set in the chip, resulting in a larger chip area and too high power consumption. In the embodiment of the present application, by using the row address and column address in the cache to cache multiple small packet-length packets or small fragments of packets, and then storing the packets in the cache into the memory at one time, the problem that each small packet-length packet occupies one address in the memory is solved, the storage space utilization rate is improved, the chip area is reduced, and the chip power consumption is reduced.
[0019] Figure 1 is a schematic structural diagram of an on-chip cache device provided by an embodiment of the present application. Refer to Figure 1 , the embodiment of the present application can be applied to the situation of packet processing in the chip. This device can be implemented by software and / or hardware and is generally integrated in the chip. The device in the embodiment of the present application specifically includes the following modules: a read / write processing module 110, a cache module 120, and a memory module 130.
[0020] In the embodiment of the present application, the on-chip cache device in the chip can mainly be composed of a read / write processing module 110, a cache module 120, and a memory module 130. The read / write processing module 110 can be a hardware processing circuit that can obtain the packet from the access source and read the packet to the access source. The cache module 120 can be a cache queue composed of one or more cache registers that can cache the packet, and the memory module 130 can be a random access memory (RAM) that can exchange packets with the access source in the chip.
[0021] Among them, the read / write processing module 110 is respectively connected to the cache module 120 and the memory module 130. The read / write processing module 110 is used to store the packet into the cache module 120 and the memory module 130, read the packets stored in the cache module 120 and the memory module 130, and transfer the packet cached in the cache module 120 to the memory module 130.
[0022] Specifically, the read-write processing module 110 can control the writing or reading of messages from the chip. The read-write processing module 110 can be connected to a cache module 120 and a memory module 130. The read-write processing module 110 can write messages to the cache module 120 and the memory module 130. Exemplarily, the messages can be first written to the cache module 120. When the number of messages stored in the cache module 120 or the number of occupied cache registers 1201 exceeds a threshold, the messages stored in the cache module 120 can be transferred and stored in the memory module 130, so as to achieve the simultaneous writing of multiple messages in the memory module 130 and improve the utilization rate of the storage space for messages. The read-write processing module 110 can also read the messages stored in the cache module 120 and the memory module 130. For example, the messages can be first read from the cache module 120 and then from the memory module 130. The read-write processing module 110 can also first read the messages from the memory module 130 and then from the cache module 120.
[0023] The cache module 120 is connected to the memory module 130 through the read-write processing module 110. The cache module 120 includes at least one cache register 1201 for temporarily caching messages.
[0024] In the embodiment of the present application, the cache module 120 can be a cache register group, which can include at least one cache register 1201. Messages can be cached in the cache register 1201. Further, after the messages are cached, description characters of the cache register 1201 can be generated. The description characters can include the row address and column address where the message is stored. Among them, the row address and column address can identify the position of the cache register 1201 in the cache module 120. In the on-chip cache device, the cache module 120 can play a role of temporary storage. Messages can be stored in one or more cache registers 1201. When the message is a normal message, the message length of the message can exceed the bit width of the cache register 1201, and multiple cache registers 1201 can be used to store one message. When the message is a small-packet-length message, the message length of the message can be less than or equal to the bit width of the cache register 1201, and the message can be cached in one cache register 1201.
[0025] The memory module 130 is connected to the read-write processing module 110 and is used to transfer and store the messages cached by the cache module 120.
[0026] Specifically, the memory module 130 can store the messages received by the read-write processing module 110. The memory module 130 can store one or more messages at the same time when storing messages. For example, the read-write processing module 110 can store a threshold number of messages cached in the cache module into the memory module 130. The threshold number can ensure that the length of one or more messages stored at the same time does not exceed the bit width corresponding to the row address of the memory module, thereby ensuring successful message storage.
[0027] In an embodiment of the present application, an on-chip cache device may include a read-write processing module, a cache module and a memory module. The read-write processing module writes messages to the cache module, and then writes the messages cached by the cache module into the memory module. The read-write processing module reads the messages stored in the cache module and the memory module, thereby realizing the caching of messages in the memory module. Multiple messages are temporarily stored in the cache module, and the memory module can write multiple messages at the same time, thereby improving the utilization of storage space, reducing the storage area of the chip, and reducing the power consumption of the chip.
[0028] Figure 2 This is a schematic diagram of the structure of an on-chip cache device provided by an embodiment of the present application. The embodiment of the present application is specific based on the above-mentioned embodiment. The read and write processing modules respectively use a read processing unit and a write processing unit to implement the writing and reading operations of the message, and the setting parameters of the cache module and the memory module are specific. Figure 2 In the embodiment of the present application, the on-chip cache device may include a read / write processing module 110 , a cache module 120 and a memory module 130 .
[0029] Furthermore, the on-chip cache device may also include an idle address module 140, which is connected to the read-write processing module 110 and is used to manage the idle row addresses in the memory module 130. Each FIFO in the idle address module 140 stores the idle row addresses. The FIFO depth of the idle address module 140 is the same as the RAM depth in the memory module 130, wherein the RAM depth is determined by the read and write bandwidth of the RAM main frequency packet transmission rate in the memory module 130 and the reserved acceleration ratio.
[0030] Specifically, to improve the message processing efficiency, an idle address module 140 can be set up to uniformly manage the idle row addresses in the memory module 130. Among them, the idle row address can be the row address of the unoccupied storage space in the memory module 130. The idle row addresses in the idle address module can be stored in the form of a FIFO queue. The idle address module can include at least one FIFO queue. To meet the requirement of the idle address module to manage the addresses of all memory modules, the FIFO depth in the idle address module can be the same as the RAM depth in the memory module. Among them, the FIFO depth can represent the number of FIFO queues storing idle addresses in the idle address module, and the RAM depth of the memory module can represent the ability to process messages simultaneously and can represent the number of RAMs.
[0031] Furthermore, the RAM depth can be related to the chip requirements. Specifically, it can be determined by the read-write bandwidth of the message transmission rate in the main frequency of the chip and the reserved acceleration ratio. For example, assuming that the number of messages processed per beat is 4 and the acceleration ratio is 50%, then the depth of the RAM in the device can be 4 + 4 * 50% = 6. Figure 3 It is a schematic diagram of the depth relationship provided by the embodiment of the present application. Refer to Figure 3 , the idle address module 140 can include q freefifo queues. Correspondingly, the number of RAMs in the memory module 130 can be q, and the FIFO depth in the idle address module 140 is the same as the RAM depth in the storage module.
[0032] In the embodiment of the present application, the read-write processing module 110 includes a write processing unit 1101. The write processing unit 1101 is respectively connected to the cache module 120, the memory module 130, and the idle address module 140. The write processing unit 1101 is used to receive messages, obtain idle row addresses from the idle address module 140, store the messages into the cache module 120 according to the idle row addresses, and when all the cache registers 1201 corresponding to the idle row addresses in the cache module 120 are occupied, transfer the messages in the cache registers 1201 corresponding to the idle row addresses to the memory module 130.
[0033] Specifically, the read-write processing module 110 can use the write processing unit 1101 to write packets into the memory module 130, and can fragment and cache packets or cache small packet-length packets. The write processing unit 1101 can apply for an idle row address in the idle address module 140 according to the packet, and cache the packet into the cache module according to the idle row address. One idle row address can correspond to multiple cache registers in the cache module, and each cache register in the cache module can store a packet fragment or a small packet-length packet. When all the cache registers corresponding to an idle row address in the cache module store packet fragments or small packet-length packets, the column address of the cache register 1201 can be marked as 1. When all the column addresses corresponding to the idle row address are marked, all the packet fragments or small packet-length packets corresponding to the idle row address can be written into the memory module 130 through the write processing unit 1101. Further, when the write processing unit 1101 determines that all the cache registers 1201 corresponding to the idle row address are occupied, the row address and column address of the cached packet can be stored simultaneously in 130.
[0034] Further, on the basis of the above application embodiments, the device further includes: a packet body linked list module 150, which is respectively connected to the read processing unit 1102 and the write processing unit 1101, and is used to store the storage address of the packet. The storage addresses in the packet body linked list module 150 are stored in the form of a linked list, and the number of nodes in the linked list is determined according to the read delay beats of the address memory.
[0035] In the embodiments of the present application, the storage address of the packet can be stored through the packet body linked list module 150. The storage address can include a row address and a column address. The corresponding packet can be read in the cache module 120 and / or the memory module 130 through the row address and the column address. The packet body linked list module 150 can be specifically composed of registers, and can store a linked list head, a linked list tail, and a linked list pointer. The content in each linked list can be specifically the row address and column address of each stored packet. It can be understood that, due to the impact of the chip's packet processing performance and chip area, in order to improve the chip area utilization rate, the number of nodes in the linked list in the packet body linked list module 150 can be set according to the read delay beats of the address memory.
[0036] The read-write processing module 110 includes a read processing unit 1102. The read processing unit 1102 is respectively connected to the cache module 120, the memory module 130, and the idle address module 140. The read processing unit 1102 is used to read packets. According to the packet row address, it reads the stored packets in the cache module 120 and / or the memory module 130, and after reading the packet, it stores the packet row address in the idle address module 140.
[0037] Specifically, the read / write processing module 110 can use the read processing unit 1102 to implement the reading of packets. The read processing unit 1102 can respectively read the corresponding packets from the cache module 120 and the memory module 130 according to the packet address. Whenever a packet is read, the column address in the packet address can be marked as 0. Further, after all the packets corresponding to the row address of the packet address are read, the read processing unit 1102 can store the row address as a free row address in the free address module 140. Since the storage space corresponding to each row address can store multiple small packet-length packets or packet fragments, only after all the small packet-length packets corresponding to this row address are recycled can the pointer of this row address be released. Each read operation first indexes the cache module 120 after obtaining the corresponding descriptor from the linked list. If there is no corresponding row address, then the packet is read from the memory module 130, which can reduce the read / write frequency of the cache, reduce power consumption, and reduce access conflicts.
[0038] Further, on the basis of the above application embodiment, the number of storage units 1301 in the memory module 130 is determined by the access source read / write bus bit width and the packet length, and the memory module includes a single-access bus memory.
[0039] In the application embodiment, the memory area can be reduced by using a single-access bus memory to replace a multi-read / write bus memory. The memory in the memory module 130 can specifically be a single-access bus memory. The space size of each row address of the RAM in the memory module 130 can be determined by the bit width of the access source bus, and the number of storage units for storing small packet-length packets or packet fragments that can be divided in each row address can be determined according to the packet length. For example, assuming that the length of a small packet-length packet is 64 bytes, the width of a row address is 600 bytes, and a maximum of 6 small packet-length packets can be written in each cycle, the row address in the memory module 130 needs to be divided into 6 storage units for storing small packet-length packets or packet fragments.
[0040] Further, on the basis of the above application embodiment, the cache module 120 includes at least one cache register group, and the cache register group includes at least one cache register. Among them, the number of cache register groups in the cache module is determined by the read delay beats of the cache register descriptor and the reserved burst processing number, and the number of cache registers in the cache register group is determined based on the number of storage units corresponding to the address space of the row address in the memory module.
[0041] Specifically, the cache module 120 may include multiple register groups. Each cache register group may include multiple cache registers. The storage space of the cache registers may correspond to the column address of the RAM of the memory module 130. The storage space corresponding to the row address of the memory module 130 may be segmented by caching small packet-length packets or packet fragments in the cache registers. The small packet-length packets and packet fragments in multiple cache registers may be simultaneously stored at a row address of the RAM of the memory module 130. The storage space of one cache register is the same as the storage space corresponding to the storage unit in the memory module 130, and the cache register may correspond to the column address of the storage unit. The number of cache registers in the cache module 120 may be related to the hardware parameters of the memory module 130 and may be determined by reading the delay beats and reserved burst processing numbers through the cache register descriptor.
[0042] Further, based on the above application embodiments, the device further includes:
[0043] A packet body linked list module 150, which is respectively connected to the read processing unit 1102 and the write processing unit 1101, and is used to store the storage address of the packet. The storage addresses in the packet body linked list module 150 are stored in the form of a linked list, and the number of nodes in the linked list is determined according to the read delay beats of the address memory.
[0044] Further, based on the above application embodiments, the device further includes:
[0045] A read-write conflict module, which is respectively connected to the read processing unit 1102 and the write processing unit 1101, and is used to handle the abnormal read-write conflicts of packets in the cache module and / or the memory module.
[0046] In the prior art, the cache RAM read-write conflicts may include three states: read-write conflict, write-write conflict, and read-read conflict. In the embodiments of the present application, by implementing caching of small packet-length packets or packet fragments through the cache module, the idle address module 140 can ensure the allocation of different idle row addresses to ensure that write-write conflicts do not occur. When a read-write conflict occurs, it only occurs when a read operation and a write operation are required for the same row address simultaneously. When a conflict occurs, the idle address module needs to first remove the row address to be read, and at the same time, the write operation requires caching the write operation after a register group is full, reducing the probability of conflicts; when a read-read conflict occurs, it can be scheduled in a priority manner to ensure that the access source with a higher priority reads the data first. When there is a cache register access conflict, all conflicts are processed in a priority manner. The conflict handling mechanism is implemented in the read-write processing module.
[0047] Exemplarily, Figure 4 is a configuration example diagram of an on-chip cache device provided by the embodiments of the present application. SeeFigure 4 , in the embodiments of the present application, relevant parameters in the read-write processing module, cache module, and memory module are configured according to the number of access read-write sources of the chip and the read-write bandwidth requirements, and the processing methods for handling type conflicts, the implementation methods of the free address module, the number of RAMs, the number of cache registers, and the number of packet body linked list nodes are determined respectively, so as to implement an efficient on-chip cache device, improve the utilization rate of the chip storage space, reduce the chip area, and reduce the chip power consumption.
[0048] Figure 5 It is a flowchart of an on-chip cache method provided by the embodiments of the present application. The embodiments of the present application can be applied to the situation of message processing in a chip. This method can be executed by the on-chip cache device of the embodiments of the present application. This device can be implemented in software and / or hardware and is generally integrated in the chip. The method provided by the embodiments of the present application includes the following steps:
[0049] Step 210: Store the obtained message into the cache register of the cache module according to the row address.
[0050] Among them, the message can be a message generated during big data processing, and can include messages with normal lengths and small packet length messages, etc. The message can perform read-write operations on the on-chip cache device.
[0051] Specifically, the obtained message can be first cached into the cache register of the cache module. When caching the message, it can be stored based on the row address. When the message is a normal message, the length of the message can exceed the storage space of one cache register. The message can be sliced, and each message slice can be stored into each cache register of the cache module corresponding to the row address. Each cache register can store one message slice. When the message is a small packet length message, the small packet length message can be stored into one cache register of the cache module.
[0052] Step 220: Determine that all the cache registers corresponding to the row address are occupied, and transfer the message corresponding to the row address in the cache module to the memory module.
[0053] In the embodiments of the present application, one row address in the cache module can correspond to a cache register group, and the cache register group can include multiple cache registers. It can be determined whether all the cache registers corresponding to the row address are occupied by determining whether the column address corresponding to the row address is marked as 1. The message stored in the cache register corresponding to the row address can be transferred to the memory module, and the messages in multiple cache registers can be written into one row in the memory module RAM at the same time. It can be understood that the row addresses in the cache module and the memory module can be the same or have a corresponding relationship.
[0054] Step 230: Read the stored message of the cache module and / or the memory module according to the read message address.
[0055] The read message address may be a read address of a stored message, specifically a description character, including at least a row address and a column address of the stored message.
[0056] Specifically, the stored message can be read according to the read message address. Since in the embodiment of the present application, the message can be stored in the cache module and the memory module, it can be read respectively in the cache module and the memory module according to the read message. For example, the stored message can be read in the cache module first, and when it is determined that there is no message corresponding to the message address, the stored message can be read in the memory module. It can be understood that it can also be read in the cache module and the memory module respectively according to the read message address.
[0057] In an embodiment of the present application, the acquired message is stored in the cache register of the cache module according to the row address. When it is determined that all the cache registers corresponding to the row address are occupied, the message corresponding to the row address in the cache module is transferred to the memory module, and the stored message of the cache module and / or memory module is read based on the read message address, thereby realizing the simultaneous writing of multiple messages into the memory, improving the utilization rate of the storage space, reducing the storage area of the chip, and reducing the power consumption of the chip.
[0058] Figure 6 This is a flowchart of another on-chip cache method provided by an embodiment of the present application. The embodiment of the present application is specific based on the above embodiment. In the embodiment of the present application, the writing and reading process of the message is further detailed. Figure 6 , the on-chip cache method provided by the embodiment of the present application includes:
[0059] Step 310: Apply for a row address from the free address module based on the message.
[0060] The idle address module may be a module for uniformly managing idle row addresses, and the idle address module may store row addresses of unoccupied storage spaces.
[0061] In an embodiment of the present application, when a message for accessing a read / write source is received, a free address module may be used to apply for a row address for storing the message, and the row address may be a storage space corresponding to a cache module and / or a memory module.
[0062] Step 320: For each message, determine whether the cache register corresponding to the row address in the cache module is not fully occupied, and then store the message in the unoccupied cache register.
[0063] Specifically, the message can be cached in the cache module first. Since all the cache registers corresponding to the row address in the cache module are occupied, the message needs to be transferred from the cache module to the memory module. Before storing the message, it can be determined whether all the cache registers corresponding to the row address have been stored. When there are unoccupied cache registers, the message can be stored in the unoccupied cache register corresponding to the row address.
[0064] Step 330: Determine that all the cache registers corresponding to the row address are occupied, generate a write descriptor for each cache register, store the write descriptor as the read message address in the packet body linked list module, and store the message cached in the cache register in each storage unit of the memory module, where the number of cache registers corresponds to the number of storage units.
[0065] Among them, the write descriptor can be information describing the storage location of the message, and can specifically include the row address and column address of the message.
[0066] Specifically, when all the cache registers corresponding to a row address in the cache module are occupied, for example, when all the column addresses corresponding to a row address are marked as 1, it can be considered that a cache register group in the cache module has stored all the messages, and the messages of this register group can be transferred from the cache module to the memory module. To facilitate the search for stored messages, the write descriptor corresponding to the message can be stored in the packet body linked list module. In the embodiment of the present application, when the message is transferred from the cache module to the memory module, the message in a cache register in the cache module can be stored in the storage unit of the memory module. The number of cache registers corresponding to a row address is the same as the number of storage units, and the column address storage of the memory module can be realized through the cache register, improving the space occupancy rate of the message.
[0067] Step 340: Obtain the read message address from the packet body linked list module, where the read message address includes at least the row address and column address.
[0068] Among them, the read message address can be the storage address of the message and can be used to read the message.
[0069] Specifically, the storage address of the message in the cache module and the memory module can be located in the packet body linked list module, and the required read message address can be obtained by searching the linked list information in the packet body linked list module. Further, the read message address can be composed of the row address and column address, and can specifically correspond to a cache register in the cache module or a storage unit in the memory module.
[0070] Step 350: Determine whether the cache register of the cache module caches the stored message according to the read message address. If so, read the stored message in the cache module. If not, determine whether the memory module stores the stored message according to the read message address.
[0071] In the embodiment of the present application, the stored message can be read in the cache module first, and then the message can be read in the memory module, which can reduce the reading frequency of the message, reduce power consumption, and reduce the occurrence probability of access conflicts. Specifically, the corresponding cache register can be found according to the read message address. If the cache register does not store data, it means that the stored message is not stored in the cache module, and the corresponding stored message can be found in the memory module. If the cache register stores data, the data in the cache register is read out as the stored message.
[0072] Step 360: Read the stored message in the memory module when it is determined that the memory module stores the stored message.
[0073] Specifically, the storage unit can be found in the memory module according to the read message address. When the storage unit stores data, the data is read out as the stored message. It can be understood that when the storage unit does not store data, it is determined that the stored message does not exist, and an error reporting process can be performed.
[0074] Step 370: When all the stored messages in a group of cache registers in the cache module are read, store the row address corresponding to the cache register in the free address module.
[0075] Specifically, after the stored message is read, the storage space needs to be released. In the embodiment of the present application, when all the cache registers in a group in the high-speed cache module are read, it can be that after the stored messages in the cache registers corresponding to a row address are read, the cache registers corresponding to the row address can be cleared, and the row address is stored in the free address module as a free address.
[0076] Step 380: When all the stored messages in the storage unit corresponding to the row address in the memory module are read, store the row address in the free address module.
[0077] Specifically, since each row address in the memory module can store multiple small packet-length packets or packet fragments, only after all the small packet-length packets or packet fragments corresponding to that row address are completely recycled can the memory module release the pointer for that row address. A tag based on the cache depth can be used. When each column address is recycled, the tag for that column address is set to 0. Only after all the tags are 0 is the row address stored in the free address module. Each read operation first indexes the cache module after obtaining the corresponding descriptor from the packet body linked list module. If there is no corresponding row address, the stored packet is then read from the memory module.
[0078] In the embodiment of the present application, by applying for the application row address corresponding to the packet from the free address module, caching according to the row address in the cache module prior to it. When all the cache registers corresponding to the row address are occupied, the packet is transferred from the cache register in the cache module to the storage unit in the memory module, and the read packet address of the packet is stored in the packet body linked list module. When performing a read operation, the read packet address in the packet body linked list module is obtained, and the packet is read from the cache module and the memory module in sequence according to the packet address. When the stored packet in a group of cache registers or the storage unit corresponding to the row address is read, the row address is stored in the free address module, achieving unified management of free addresses, reducing the read and write frequency of the memory, reducing chip power consumption, and reducing access conflicts of packets.
[0079] Furthermore, on the basis of the above application embodiment, it further includes:
[0080] When a read-write conflict occurs, the free row address where the conflict occurs in the free address module is cleared. The read processing unit in the read-write processing module cannot read the packet based on the free row address, while the write processing unit in the read-write processing module writes the packet based on the free row address; when a read-read conflict occurs, the access source reads the packet in the read-write processing module according to the priority order.
[0081] In the prior art, the cache RAM read-write conflict can include three states: read-write conflict, write-write conflict, and read-read conflict. In the embodiment of the present application, by using the cache module to cache small packet-length packets or packet fragments, it can be ensured through the free address module that the issued free row addresses are different, and the write-write conflict does not occur. When a read-write conflict occurs, it only occurs when a row address needs to perform a read operation and a write operation simultaneously. When a conflict occurs, the free address module needs to first remove the row address to be read. At the same time, the write operation requires caching and writing after a register group is full, reducing the conflict probability; when a read-read conflict occurs, the priority method can be used for scheduling to ensure that the access source with a higher priority reads the data first. When there is a cache register access conflict, all conflicts are processed according to the priority method.
[0082] As described above, the above are only exemplary embodiments of the present application and are not intended to limit the protection scope of the present application.
[0083] Those skilled in the art should understand that the term user terminal covers any suitable type of wireless user equipment, such as a mobile phone, a portable data processing device, a portable network browser, or a vehicle-mounted mobile station.
[0084] Generally speaking, various embodiments of the present application can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices, although the present application is not limited thereto.
[0085] Embodiments of the present application can be implemented by a data processor of a mobile device executing computer program instructions, such as in a processor entity, or by hardware, or by a combination of software and hardware. The computer program instructions can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status-setting data, or source code or object code written in any combination of one or more programming languages.
[0086] Any block diagram of a logical process in the drawings of the present application can represent program steps, or can represent interconnected logical circuits, modules, and functions, or can represent a combination of program steps and logical circuits, modules, and functions. The computer program can be stored in a memory. The memory can have any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as but not limited to read-only memory (ROM), random access memory (RAM), optical memory devices and systems (digital versatile disc DVD or CD disc), etc. The computer-readable medium can include non-transitory storage media. The data processor can be any type suitable for the local technical environment, such as but not limited to a general-purpose computer, a dedicated computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a processor based on a multi-core processor architecture.
[0087] Through exemplary and non-limiting examples, a detailed description of the exemplary embodiments of the present application has been provided above. However, considering the accompanying drawings and the claims, various modifications and adjustments to the above embodiments will be obvious to those skilled in the art without departing from the scope of the present invention. Therefore, the proper scope of the present invention will be determined according to the claims.
Claims
1. An on-chip cache device, characterized in that, The device includes: a read-write processing module, a cache module, and a memory module; wherein, the read-write processing module is respectively connected to the cache module and the memory module, and the read-write processing module is used to store packets into the cache module and the memory module, read the packets stored in the cache module and the memory module, and transfer the packets cached in the cache module to the memory module; the cache module is connected to the memory module through the read-write processing module, and the cache module includes at least one cache register group, and each cache register group includes a plurality of cache registers for temporarily caching packets, wherein the number of the cache register groups is determined by the read delay beats of the cache register descriptor and the reserved burst processing number; the memory module is connected to the read-write processing module and is used to transfer the packets cached by the cache module. The number of storage units in the memory module is determined by the access source read-write bus width and the packet length, and the register in the memory module is a single-access bus memory.
2. The device according to claim 1, characterized in that, It further includes: an idle address module connected to the read-write processing module for managing the idle row addresses in the memory module. Each FIFO in the idle address module stores the idle row addresses, and the FIFO depth of the idle address module is the same as the RAM depth in the memory module, wherein the RAM depth is determined by the read-write bandwidth of the RAM main frequency packet transmission rate in the memory module and the reserved acceleration ratio.
3. The device according to claim 2, wherein The read-write processing module includes a write processing unit, and the write processing unit is respectively connected to the cache module, the memory module, and the idle address module. The write processing unit is used to receive packets, obtain idle row addresses from the idle address module, store the packets into the cache module according to the idle row addresses, and when all the cache register groups corresponding to the idle row addresses in the cache module are occupied, transfer the packets in the cache register groups corresponding to the idle row addresses to the memory module; The read-write processing module further includes a read processing unit, and the read processing unit is respectively connected to the cache module, the memory module, and the idle address module. The read processing unit is used to read packets, read the stored packets in the cache module and / or the memory module according to the packet addresses, and store the packet addresses into the idle address module after reading the packets.
4. The device according to claim 3, characterized in that The number of the cache registers in the cache register group is determined based on the number of storage units corresponding to the address space of the row addresses in the memory module.
5. The device according to claim 4, characterized in that The device further includes: a packet body linked list module respectively connected to the read processing unit and the write processing unit for storing the storage addresses of the packets. The storage addresses in the packet body linked list module are stored in the form of a linked list, and the number of nodes in the linked list is determined by the read delay beats of the address memory.
6. The device according to claim 4, wherein The device further includes: A read-write conflict module, which is respectively connected to the read processing unit and the write processing unit, is used to handle the abnormal read-write conflicts of packets in the cache module and / or the memory module.
7. A method for reading and writing on-chip cache, characterized in that, The method includes: Storing the obtained packet into the cache register group of the cache module according to the row address, wherein the number of the cache register groups is determined according to the read delay beats and the reserved burst processing number of the cache register descriptor; Determining that all the cache register groups corresponding to the row address are occupied, and transferring the packet corresponding to the row address in the cache module to the memory module; According to the row address and column address in the read packet address, preferentially reading the stored packet in the cache register group of the cache module; if there is no stored packet corresponding to the packet address in the cache register group, reading the stored packet from the memory module.
8. The method according to claim 7, wherein It further includes: Applying for a row address from the free address module based on the packet.
9. The method according to claim 8, wherein The storing the obtained packet into the cache register group of the cache module according to the row address includes: For each packet, determining that the cache register group corresponding to the row address in the cache module is not all occupied, and storing the packet into the unoccupied cache register group.
10. The method according to claim 9, wherein The determining that all the cache register groups corresponding to the row address are occupied and transferring the packet corresponding to the row address in the cache module to the memory module includes: Determining that all the cache register groups corresponding to the row address are occupied, generating a write descriptor for each cache register group, storing the write descriptor as the read packet address into the packet body linked list module, and storing the packets cached in the cache register group into each storage unit in the memory module, wherein the number of the cache register groups corresponds to the number of the storage units.
11. The method according to claim 10, characterized in that, The reading the stored packet in the cache module and / or the memory module according to the read packet address includes: Determining whether the cache register group of the cache module caches the stored packet according to the read packet address, if so, reading the stored packet in the cache module, if not, determining whether the memory module stores the stored packet according to the read packet address; Reading the stored packet in the memory module when it is determined that the memory module stores the stored packet.
12. The method according to claim 11, wherein It further includes: Obtaining the read packet address from the packet body linked list module, wherein the read packet address at least includes a row address and a column address.
13. The method according to claim 11, wherein It further includes: When all the stored packets in a group of cache register groups in the cache module are read, storing the row address corresponding to the cache register group into the free address module; and / or When all the stored packets in the storage unit corresponding to the row address in the memory module are read, storing the row address into the free address module.
14. The method according to any one of claims 7-11, characterized in that, It further includes: When a read-write conflict occurs, clearing the free row address where the conflict occurs in the free address module, the read processing unit in the read-write processing module cannot read packets based on the free row address, and at the same time, the write processing unit in the read-write processing module writes packets based on the free row address; When a read-read conflict occurs, the access source reads the message in the read-write processing module according to the priority order.
Citation Information
Patent Citations
Multi-channel FIFO (First In First Out) buffer and control method thereof
CN104407809A