Processing device, system and method for realizing high-performance cache based on single-port RAM (Random Access Memory)

By splitting large-capacity RAM into multiple single-port RAMs, and using technologies such as bank balanced scheduling and virtual channel division, the problem that single-port RAM in the existing technology cannot meet the high-throughput data processing, and a high-performance cache processing device and system is realized, which improves data bandwidth and adapts to the needs of different business scenarios.

CN120144486AActive Publication Date: 2025-06-13WUXI STARS MICRO SYSTEM TECHNOLOGIES CO LTD

Patent Information

Application Number
CN202510241737.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

In the prior art, single-port RAM cannot meet the needs of high throughput data processing, and pseudo-dual-port RAM and true dual-port RAM limit their application in resource-constrained scenarios due to their high area and power consumption.

Method used

By splitting large-capacity RAM into multiple single-port RAMs, and using technologies such as bank balanced scheduling, virtual channel division and linked list management, high-performance cache processing devices and systems are realized.

Benefits of technology

It reduces the difficulty of area and timing convergence of a single block of RAM, avoids read and write conflicts, realizes read and write operations within the same beat, improves data bandwidth, and adapts to the needs of different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144486A_ABST
    Figure CN120144486A_ABST
Patent Text Reader

Abstract

The invention provides a processing device and system for realizing high-performance cache based on a single-port RAM (Random Access Memory) and a data reading and writing method. The processing device comprises an array module for dividing a high-capacity RAM into n single-port RAM arrays; the bank balanced scheduling module is used for realizing read-write operation in the same beat through bank balanced scheduling and a conflict avoidance mechanism; and the freeaddrctrl module and the nextpoint module are used for managing a data address through a double linked list so as to realize dynamic address management. According to the technical scheme, the processing device and system and the data reading and writing method have the advantages that high performance is guaranteed, the area and power consumption are remarkably reduced, time sequence design is simplified, and the method is suitable for high-bandwidth data processing scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of memories, and particularly relates to a processing device, a system, and a data reading and writing method for implementing high-performance caching based on a single-port RAM (Random Access Memory). Background Art

[0002] With the continuous increase in network data bandwidth, the demand for caching in data processing systems is also increasing day by day. In a high-bandwidth environment, data caching needs to support higher throughput, usually requiring the processing ability of one or two packets per clock cycle. This poses higher requirements for the design of memories, especially in scenarios that support read and write operations simultaneously. In related memory technologies, single-port RAM (1pl1), pseudo-dual-port RAM (2p11), and true-dual-port RAM (2p22) are common choices. Single-port RAM is suitable for scenarios where read and write operations are not required to be performed simultaneously, but its limitation is that it cannot meet the requirements of high-throughput data processing. Pseudo-dual-port RAM and true-dual-port RAM are suitable for scenarios where read and write operations need to be performed simultaneously, especially in high-bandwidth data processing, but their relatively high area and power consumption limit their application in resource-constrained scenarios. Summary of the Invention

[0003] The purpose of this application is to provide a processing device, a system, and a data reading and writing method for implementing high-performance caching based on a single-port RAM. It aims to solve the problems of large area, high power consumption, difficult timing convergence, and inability to achieve continuous reading and writing in related technologies.

[0004] According to the first aspect of this application, a processing device for implementing high-performance caching based on a single-port RAM is provided, specifically including:

[0005] A Mem_array module, configured to divide a large-capacity RAM into n single-port RAMs on average to form a mem_array, where n is a natural number greater than 2;

[0006] A free addr ctrl module, configured to store the free addresses of each single-port RAM in an independent address management FIFO (a common data structure for managing address queues) for data write operations, and prefetch a number of addresses to accelerate address allocation;

[0007] A bank equalization scheduling module, configured to balance the loads between banks to ensure that the cached data volume of each bank is basically the same, so as to reduce bank conflicts during reading;

[0008] The VL_PKT_Q module is configured to divide different virtual channels VL (Virtual Channel) according to service requirements, create a doubly linked list within the corresponding virtual channel based on address information and related data, and concatenate the corresponding write data addresses in the form of sub-chains;

[0009] The next_point module is configured to update the address information of the doubly linked list.

[0010] In this application, by splitting a large-capacity RAM into multiple single-port RAMs, the area and timing convergence difficulty of a single RAM are reduced; through the bank equalization scheduling module, read and write operations are dynamically allocated to different banks, avoiding read-write conflicts and achieving read and write operations in the same cycle; in addition, this application increases the flexibility of timing modification through the VL_PKT_Q module to meet the requirements of different service scenarios. Through the free_addr_ctrl module and the next_point module, the address allocation and data read-write efficiency are optimized.

[0011] In an alternative embodiment, the device further includes:

[0012] The MUX module (Multiplexer) is configured to determine the destination virtual channel VL of the packet according to the packet information.

[0013] This embodiment improves the parallelism and flexibility of data processing through virtual channel division.

[0014] In an alternative embodiment, the bank equalization scheduling module is configured to implement read and write operations in the same cycle through a bank equalization scheduling and conflict avoidance mechanism. If a conflict occurs between the read bank and the write bank, the scheduler selects another idle bank for the write operation. This embodiment realizes read and write operations in the same cycle through an intelligent scheduling mechanism, effectively improving the data bandwidth.

[0015] In an alternative embodiment, the number of sub-chains of the VL_PKT_Q module is configured to be greater than or equal to the mem (Memory) read latency.

[0016] This embodiment sets the number of sub-chains to be greater than or equal to the mem read latency, increasing the flexibility of timing modification to meet the requirements of different service scenarios.

[0017] According to the second aspect of the present application, there is provided a system for implementing a high-performance cache based on a single-port RAM, including: a processing device for implementing a high-performance cache based on a single-port RAM according to the first aspect.

[0018] According to the third aspect of the present application, a method for data reading and writing of a high-performance cache based on a single-port RAM is provided, including:

[0019] Data initialization step: Write all the free addresses of each bank into the corresponding bankn_addr_fifo. After the initialization is completed, the bank balanced scheduling module schedules out the addresses from the corresponding fifo for subsequent data processing;

[0020] Data writing step: When external data is written, according to the header information, the data is allocated to the corresponding virtual channel VL through the MUX module, and a doubly linked list is created within VL according to the scheduled address information and the data;

[0021] Data storage step: While creating the linked list, write the address provided by the corresponding data into the mem_array;

[0022] Linked list information caching step: Cache the next_point information of the linked list for subsequent data reading.

[0023] In an alternative embodiment, the creating a doubly linked list within VL according to the scheduled address information and the data specifically includes:

[0024] Write the data of the first beat into the head pointer of the first sub-chain;

[0025] Write the data of the second beat into the head pointer of the second sub-chain;

[0026] And so on. There are a total of four sub-chains used to cover the read latency of the mem to ensure data continuity;

[0027] When the data of the fifth beat arrives, the data is written into the mem_array, and the next_point and tail_point of the first sub-chain are updated;

[0028] When the data of the sixth beat arrives, update the next_point and tail_point of the second sub-chain,

[0029] And so on, to complete the creation and update of the linked list.

[0030] In an alternative embodiment, the method further includes:

[0031] Data reading step: If the sending condition is satisfied and an external read request is generated, data needs to be read from the cache,

[0032] First, obtain the address of the data from the head information of the first sub-linked list and read the data,

[0033] Read the address information cached at next_point according to the next_point indication, and update it to head.

[0034] In the next cycle, read the head information of the second sub-chain. Similarly, read the address information corresponding to next_point and update it to the head of the second sub-chain.

[0035] And so on to complete the data reading and linked list update.

[0036] Data output step: Read the data cached in mem_array according to the address in the linked list and send it to the outside.

[0037] Address recycling step: Send the data address to the free_addr_ctrl module when reading the data address for subsequent scheduling.

[0038] Other features and advantages of the present application will be described in the subsequent specification, and will be partially obvious from the specification, or understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures and processes pointed out in the specification and the drawings. Description of the Drawings

[0039] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 is a schematic structural diagram of a device for implementing a high-performance cache based on a single-port RAM according to an exemplary embodiment of the present application.

[0041] Figure 2 is a block diagram of a system for implementing a high-performance cache based on a single-port RAM according to an exemplary embodiment of the present application.

[0042] Figure 3 is a flowchart framework diagram of a data reading and writing method for implementing a high-performance cache based on a single-port RAM according to an exemplary embodiment of the present application.

[0043] Figure 4 is a schematic flowchart diagram of a data reading and writing method for implementing a high-performance cache based on a single-port RAM according to an exemplary embodiment of the present application. Detailed Embodiments

[0044] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.

[0045] In a high-bandwidth environment, data caching needs to support higher throughput, usually requiring the processing ability of one or two packets per cycle. This poses higher requirements for the design of the memory, especially in scenarios that support read and write operations simultaneously. In related technologies, there are significant differences in area, power consumption, and performance between single-port RAM and pseudo-dual-port RAM, and trade-offs need to be made according to specific application scenarios.

[0046] A single-port RAM can only perform read or write operations in the same clock cycle and cannot perform read and write operations simultaneously. Its characteristics are a relatively high operating frequency (1715N3 135MHz), and relatively small area and power consumption. For example, in the S12 process, the area of a single-port RAM (sram_w48_d256_1pl1) with a capacity of 12,288 bits is 3562.6721 * 2 = 7,125.344. A single-port RAM is suitable for scenarios where read and write operations are not required to be performed simultaneously, but its limitation is that it cannot meet the requirements of high-throughput data processing.

[0047] A pseudo-dual-port RAM can perform read and write operations in the same clock cycle, but the read and write operations cannot access the same address simultaneously. Its operating frequency is relatively low (810 - 1484MHz), and the area and power consumption are relatively large. For example, in the S12 process, the area of a pseudo-dual-port RAM (sram_w48_d256_2p11) with a capacity of 12,288 bits is 4297.7526 * 2 = 8,595.5052, which is 1.2 times that of a single-port RAM. As the memory capacity increases, the gap in area and power consumption will further expand. A pseudo-dual-port RAM is suitable for scenarios where read and write operations need to be performed simultaneously, especially in high-bandwidth data processing, but its relatively high area and power consumption limit its application in resource-constrained scenarios.

[0048] It can be seen from this that a pseudo-dual-port RAM can perform read and write operations in the same clock cycle and is suitable for scenarios of high-throughput data processing. However, directly using a dual-port RAM can meet the high-throughput requirements, but it has a large area, high power consumption, and difficult timing convergence.

[0049] To reduce area and power consumption, patent document CN103594109A proposes a solution to implement the pseudo-dual-port function using a single-port RAM. This solution realizes read and write operations on a single single-port RAM through time-division multiplexing. Although this solution reduces area and power consumption, it cannot achieve continuous read and write, and the data bandwidth improvement is limited.

[0050] Based on the above analysis, referring to Figure 1 the structural schematic diagram of, the present application exemplarily proposes a processing device for implementing a high-performance cache based on a single-port RAM, specifically including:

[0051] The Mem_array module is configured to divide a large-capacity RAM into n single-port RAMs on average to form a mem_array, where n is a natural number greater than 2;

[0052] The free addr ctrl module is configured to store the free addresses of each single-port RAM in an independent address management FIFO for data write operations, and prefetch a number of addresses to accelerate address allocation;

[0053] The bank equalization scheduling module is configured to balance the loads between banks to ensure that the cached data volumes of each bank are basically the same, so as to reduce bank conflicts during reading;

[0054] The VL_PKT_Q module is configured to divide different virtual channels VL according to service requirements, create a doubly linked list within the corresponding virtual channel based on address information and relevant data, and concatenate the corresponding write data addresses in the form of sub-chains;

[0055] The next_point module is configured to update the address information of the doubly linked list.

[0056] Exemplarily, the Mem_array module of the present application divides the large-capacity RAM into n banks, and each bank is implemented by a single-port RAM. Each bank can perform read and write operations independently. Here, n is a natural number greater than 2, and the specific value is obtained by simulating the area, timing, and access conflict probability indicators in the UVM simulation platform. For example: for a depth of 40k (40,960 entries), the simulation results show that when n = 10, the conflict probability approaches 0, and at the same time, the area of the control logic is controllable. The bank division method is to evenly divide the total depth of 40k (40,960 entries) into 10 banks, and the capacity of each bank is 4K. By adopting the uniform division method, the read and write requests are dispersed to 10 banks, and the statistical conflict probability is about 10%. After further verification by the balanced scheduling simulation, it is reduced to nearly 0%, effectively avoiding read and write conflicts; in the S12 process, the total area of 10 1p11 is 35,626.7, which is significantly lower than 42,977.5 of the pseudo dual-port RAM, significantly reducing the area; in addition, the working frequency of a single bank can reach 3135MHz, and the scheduling logic only needs to switch between 10 banks, the critical path delay is controllable, and the timing convergence is flexible.

[0057] Exemplarily, in the free_addr_ctrl module of the present application, in the initialization stage, all free addresses of each single-port RAM are pre-written into the corresponding fifo. In the running stage, several addresses (such as 2 to 4) are pre-fetched from the fifo and cached in the register for quick allocation. By pre-fetching addresses, the delay of address allocation is reduced, and the throughput of data writing is improved.

[0058] By splitting the large-capacity RAM into multiple single-port RAMs, the present application reduces the area and the difficulty of timing convergence of a single piece of RAM; through the bank balanced scheduling module, read and write operations are dynamically allocated to different banks, avoiding read and write conflicts and realizing read and write operations in the same cycle; in addition, the present application increases the flexibility of timing modification through the VL_PKT_Q module to meet the requirements of different service scenarios. Through the free_addr_ctrl module and the next_point module, the address allocation and data read and write efficiency are optimized.

[0059] In some optional implementation manners, the device further includes:

[0060] The MUX module is configured to determine the virtual channel VL of the message destination according to the message information.

[0061] Exemplarily, as described above, this implementation manner improves the parallelism and flexibility of data processing through virtual channel division.

[0062] In an alternative embodiment, the VL_PKT_Q module is configured such that the number of sub-chains is greater than or equal to the mem read latency. This embodiment sets the number of sub-chains to be greater than or equal to the mem read latency, which increases the flexibility of timing modification to meet the requirements of different service scenarios.

[0063] In some alternative embodiments, the bank equalization scheduling module is configured to implement read and write operations within the same cycle through a bank equalization scheduling and conflict avoidance mechanism. If a conflict occurs between the read bank and the write bank, the scheduler selects another idle bank for the write operation. This embodiment realizes read and write operations within the same cycle through an intelligent scheduling mechanism, effectively improving the data bandwidth.

[0064] Correspondingly, as Figure 2 shown, the present application exemplarily provides a system for implementing a high-performance cache based on a single-port RAM, including: as Figure 1 shown, the above-mentioned processing device for implementing a high-performance cache based on a single-port RAM.

[0065] In some alternative embodiments, the device further includes:

[0066] A MUX module, configured to determine the virtual channel VL for the packet destination according to the packet information.

[0067] Exemplarily, as described above, this embodiment improves the parallelism and flexibility of data processing through virtual channel partitioning.

[0068] In an alternative embodiment, the VL_PKT_Q module is configured such that the number of sub-chains is greater than or equal to the mem read latency. This embodiment sets the number of sub-chains to be greater than or equal to the mem read latency, which increases the flexibility of timing modification to meet the requirements of different service scenarios.

[0069] In some alternative embodiments, the bank equalization scheduling module is configured to implement read and write operations within the same cycle through a bank equalization scheduling and conflict avoidance mechanism. If a conflict occurs between the read bank and the write bank, the scheduler selects another idle bank for the write operation. This embodiment realizes read and write operations within the same cycle through an intelligent scheduling mechanism, effectively improving the data bandwidth.

[0070] The above device is exactly the same as the processing device for implementing a high-performance cache based on a single-port RAM provided in the above embodiment. For specific details, reference can be made to the description of the device in the above embodiment, which will not be elaborated here.

[0071] Correspondingly, as Figure 3 shown, the present application exemplarily provides a data read and write method for implementing a high-performance cache based on a single-port RAM, including:

[0072] Data initialization step: Write all the free addresses of each bank into the corresponding bankn_addr_fifo. After initialization, the bank balanced scheduling module schedules the addresses from the corresponding fifo for subsequent data processing;

[0073] Data writing step: When external data is written, according to the header information, the data is allocated to the corresponding virtual channel VL through the MUX module, and a doubly linked list is created within VL based on the scheduled address information and the data;

[0074] Data storage step: While creating the linked list, write the address provided by the corresponding data into the mem_array;

[0075] Linked list information caching step: Cache the next_point information of the linked list for subsequent data reading.

[0076] In some alternative embodiments, creating a doubly linked list within VL based on the scheduled address information and the data specifically includes:

[0077] The data in the first cycle is written to the head pointer of the first sub-chain;

[0078] The data in the second cycle is written to the head pointer of the second sub-chain;

[0079] And so on. There are a total of four sub-chains used to cover the read latency of the mem to ensure data continuity;

[0080] When the data in the fifth cycle arrives, the data is written to the mem_array, and the next_point and tail_point of the first sub-chain are updated;

[0081] When the data in the sixth cycle arrives, the next_point and tail_point of the second sub-chain are updated,

[0082] And so on, to complete the creation and update of the linked list,

[0083] Data output step: According to the addresses in the linked list, read the data cached in the mem_array and send it to the outside;

[0084] Address recycling step: Read the data address and send it to the free_addr_ctrl module simultaneously for subsequent scheduling. Among them, the number of sub-chains needs to satisfy being greater than or equal to the read latency of the memory, so that the address can directly obtain the data address from the sub-chain without waiting for the data of the cached nextpoint to come out of the memory and then read the data, increasing the flexibility of timing modification. When the RAM read latency is equal to the number of sub-chains, the address storage locations in different write cycles do not overlap, thus avoiding data conflicts. Taking the example that four sub-chains are required when the read latency is four beats, write to sub-chain 1 in the 1st beat → read from sub-chain 1 in the 5th beat; write to sub-chain 2 in the 2nd beat → read from sub-chain 2 in the 6th beat; and so on. The read and write operations are completely decoupled in time and space, avoiding read and write conflicts and ensuring the continuity of the read data.

[0085] Correspondingly, as Figure 4 shown, the present application exemplarily provides a method for data reading and writing of a high-performance cache based on a single-port RAM, and the specific steps are as follows:

[0086] Step S401 Data initialization: Write all the free addresses of each bank into the corresponding bankn_addr_fifo. After initialization, the bank balanced scheduling module schedules the addresses from the corresponding fifo for subsequent data processing;

[0087] Step S402 Data writing: When external data is written, according to the header information, the data is allocated to the corresponding virtual channel VL through the MUX module, and a doubly linked list is created within VL according to the scheduled address information and the data;

[0088] Step S403 Data storage: While creating the linked list, write the address provided by the corresponding data into the mem_array;

[0089] Step S404 Linked list information caching: Cache the next_point information of the linked list for subsequent data reading.

[0090] The present application realizes a processing device, system and method with high throughput, low area and low power consumption by using a single-port RAM and a linked list management mechanism. The specific advantages are at least as follows:

[0091] 1. Reduce the area. For a large-capacity cache, using a single-port RAM can significantly reduce the area. Generally, the area of a dual-port RAM is 20% - 50% larger than that of a single-port RAM. By splitting the large-capacity RAM into multiple single-port RAMs and adopting a linked list management mechanism, the area utilization rate is further optimized.

[0092] 2. Save costs. The cost of a single-port RAM is lower than that of a dual-port RAM. By using a single-port RAM to replace a dual-port RAM, the hardware cost can be reduced.

[0093] 3. Easy timing convergence. The operating frequency of a single-port RAM is usually higher than that of a dual-port RAM. At high-speed clocks, a single-port RAM can better meet the timing requirements and improve the overall performance of the system.

[0094] 4. Reduce power consumption. The power consumption of a single-port RAM is significantly lower than that of a dual-port RAM. By using a single-port RAM, the power consumption of the system can be reduced, extending the battery life of the device or reducing the heat dissipation requirements.

[0095] 5. Easy scalability of cache management. Cache management is carried out in a linked list manner, with high flexibility and scalability. By simply increasing the number of corresponding queue linked lists, it can be extended to multiple queue data cache management to adapt to different business requirements.

[0096] It can be understood that the memory circuit structures, module names, and elements described in the above embodiments are only examples. Those skilled in the art can also make easily conceivable combinations and adjustments to the structural features of the above multiple embodiments according to the usage needs, and should not limit the concept of this application to the specific details of the above examples.

[0097] Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A processing device for implementing high-performance cache based on a single-port RAM, characterized in that: include: The mem_array module is configured to divide the large capacity RAM into n single-port RAMs to form mem_array, where n is a natural number greater than 2; free _ addr _ The ctrl module is configured to store the free addresses of each single-port RAM in an independent address management FIFO for data write operations, and pre-fetch several addresses to speed up address allocation; The bank balancing scheduling module is configured to balance the load between banks and ensure that the cache data volume of each bank is basically consistent to reduce bank conflicts during reading; The VL_PKT_Q module is configured to divide different virtual channels VL according to business needs, create a bidirectional linked list in the corresponding virtual channel according to the address information and related data, and connect the corresponding write data addresses in series through sub-chains; The next_point module is configured to update the address information of the doubly linked list.

2. The processing device for implementing high-performance cache based on single-port RAM according to claim 1, characterized in that: The device also includes: The MUX module is configured to determine a virtual channel VL where the message is to be sent according to the message information.

3. The processing device for implementing high-performance cache based on single-port RAM according to claim 1, characterized in that: The bank balancing scheduling module is configured to implement read and write operations within the same beat through bank balancing scheduling and conflict avoidance mechanisms. If a conflict occurs between a read bank and a write bank, other idle banks are selected for write operations.

4. The processing device for implementing high-performance cache based on single-port RAM according to claim 1, characterized in that: The VL_PKT_Q module is configured so that the number of sub-chains is greater than or equal to the mem read delay.

5. The processing device for implementing high-performance cache based on single-port RAM according to claim 2, characterized in that: The device is VL _ PKT _ The Q module and the MUX module realize virtual channel division and improve the parallelism and flexibility of data processing.

6. The processing device for implementing high-performance cache based on single-port RAM according to claim 2, characterized in that: The device is free _ addr _ The ctrl module and next_point module implement dynamic address management and optimize address allocation and data reading and writing efficiency.

7. A system for implementing high-performance cache based on single-port RAM, characterized in that: The system comprises: a processing device for implementing high-performance cache based on a single-port RAM as described in any one of claims 1-6.

8. A data reading and writing method for realizing high-performance cache based on single-port RAM, characterized in that: include: Data initialization step: Write all free addresses of each bank to the corresponding bankn _ addr _ In the fifo, after initialization is completed, the bank balance scheduling module schedules the address from the corresponding fifo for subsequent data processing; Data writing step: When external data is written, the data is allocated to the corresponding virtual channel VL through the MUX module according to the header information, and a bidirectional linked list is created in the VL according to the scheduled address information and data; Data storage steps: while creating the linked list, write the address provided by the corresponding data into mem_array; Link list information caching step: cache the next_point information of the linked list for subsequent data reading.

9. The data reading and writing method according to claim 8, characterized in that: The step of creating a bidirectional linked list in the VL according to the scheduled address information and data specifically includes: The first beat of data is written to the head pointer of the first subchain; The second beat of data is written into the head pointer of the second subchain; By analogy, a total of four sub-chains are used to cover the read delay of mem to ensure data continuity; When the fifth beat of data arrives, the data is written into mem_array, and the next_point and tail_point of the first subchain are updated; When the sixth beat of data arrives, the next_point and tail_point of the second subchain are updated. And so on, the creation and update of the linked list is completed.

10. The data reading and writing method according to claim 9, characterized in that: The method further comprises: Data reading steps: If the sending conditions are met, a read request is generated externally and data needs to be read from the cache. First, get the address of the data from the head information of the first sub-list and read the data. According to the next_point instruction, read the address information cached in next_point and update it to head. The next beat reads the head information of the second subchain, and also reads the address information corresponding to next_point, and updates it to the head of the second subchain. And so on, the data reading and linked list updating are completed; Data output step, according to the address in the linked list, read the data cached in mem_array and send it to the outside; Address recovery step, read the data address and send it to free _ addr _ ctrl module, used for subsequent scheduling.

Citation Information

Patent Citations

  • Memory structure replacing dual-port static memory

    CN103594109A

  • Data storage method and device, electronic equipment and computer readable storage medium

    CN113204317A

  • High-performance read-write linked list cache device and method

    CN113821457A

  • Method for reading and writing single-port SRAM (static random access memory), FIFO (first in first out) module and chip

    CN117116323A

  • Reordering cache-based instruction processing method and device, equipment and medium

    CN117331862A

Cited By

  • Data caching device, method and equipment

    CN120956717A