Cache management device and method, electronic equipment and storage medium
By introducing a management cache device with address segment division and interleaving enable control in SOC, the problems of wasted cache resources and complex interconnection in SOC design are solved, and dynamic switching and efficient utilization of cache mode are realized.
Patent Information
- Application Number
- CN202510259185.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
AI Technical Summary
In SOC design, the prior art has problems such as power consumption, waste of cache resources and complex interconnection, and it is difficult to effectively manage caches to optimize data processing performance and system efficiency.
By introducing source interface modules, registers, system address mappers and interleaver into the device that manages cache, address segment division and interleaving enable control of read and write memory access request information are realized, and the cache mode is dynamically switched to meet the needs of different scenarios.
It realizes the on-demand switching of LLC in different scenarios, optimizes cache utilization and system efficiency, and reduces the problem of resource waste in traditional architectures.
Smart Images

Figure CN120104523A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a method, an apparatus, an electronic device, and a storage medium for managing a cache. Background Art
[0002] Advances in semiconductor technology allow system-on-chip (SOC) to integrate multiple functional modules or subsystems (such as central processing unit (CPU), image processing unit (GPU)), but it also increases the difficulty of bus architecture design. SOC design pursues high performance, low resource usage, and high energy efficiency. Different modules have different transmission requirements, and a flexible multi-layer bus structure helps improve performance. Compatibility is improved by following relevant standard protocols, but in actual applications, there are still problems such as power consumption, cache resource waste, and complex interconnection, which need to be optimized for specific scenarios. Summary of the invention
[0003] At least one embodiment of the present disclosure provides a device for managing cache, including: a source interface module, a register, a system address mapper, and an interleaver. The source interface module is configured to receive read and write memory access request information from a host; the register is configured to determine whether an interleaving enable control signal acting on the system address mapper is valid; the system address mapper is configured to determine whether the read and write memory access request information belongs to the first address segment or the second address segment; the interleaver is configured to distribute the read and write memory access request information to multiple corresponding last-level caches at a first interleaving granularity in response to the interleaving enable signal being valid and the read and write memory access request information belonging to the first address segment, and the interleaver is also configured to distribute the read and write memory access request information to multiple corresponding last-level caches at a second interleaving granularity in response to the interleaving enable signal being valid and the read and write memory access request information belonging to the second address segment; wherein the second address segment is an address segment in the address space where the cache is located that is different from the first address segment.
[0004] For example, in a cache management device provided in an embodiment of the present disclosure, the first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
[0005] For example, in an apparatus for managing cache provided in an embodiment of the present disclosure, the first interleaving granularity is configured based on the size of a cache line, and the second interleaving granularity is configured based on the data bit width of a cache interface.
[0006] For example, in the device for managing cache provided in one embodiment of the present disclosure, the interleaver is also configured to, in response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the first address segment, send the read / write memory access request information to the first last-level cache; and, in response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the second address segment, send the read / write memory access request information to the second last-level cache, wherein the first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
[0007] For example, in the device for managing cache provided in an embodiment of the present disclosure, the device for managing cache further includes an encoder and a decoder. The encoder is configured to perform an encoding operation on the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information; the interleaver is further configured to distribute the read / write memory access request information to the target final-level cache at the first interleaving granularity based on the intermediate interleaving bit information; and the decoder is configured to perform a decoding operation based on the first interleaving bit information to obtain output address information, and provide the output address information to the target final-level cache.
[0008] For example, in the device for managing cache provided in one embodiment of the present disclosure, the encoder is further configured to perform mathematical operations on the first interleaved position information and the first bit position information to obtain the intermediate interleaved position information, wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved position corresponding to the first interleaved position information; and the first interleaved position information is retained through a sideband signal to be provided to the subsequent decoding operation.
[0009] For example, in the device for managing cache provided in one embodiment of the present disclosure, the decoder is further configured to, in response to the first bit position information being valid and the first interleaving bit information being invalid, determine the output address information based on the read / write memory access request information and the first interleaving granularity; or, in response to the first bit position information and the first interleaving bit information being valid, determine the output address information based on the read / write memory access request information and the first interleaving granularity; or, in response to the first bit position information being invalid, determine the output address information based on the read / write memory access request information; wherein the first bit corresponding to the first bit position information is a higher bit of the interleaving bit corresponding to the first interleaving bit information.
[0010] For example, in the device for managing cache provided in one embodiment of the present disclosure, the encoder is further configured to set the second bit information of the read and write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retain the initial value of the second bit information through a sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and constrains the burst data size of the read and write memory access request information to not exceed one half of the address boundary specified by the transmission protocol; and the decoder is further configured to restore the initial value of the second bit information.
[0011] At least one embodiment of the present disclosure provides a device for managing cache, comprising: a source interface module, a system address mapper, and an interleaver. The source interface module is configured to receive read / write memory access request information from a host; the system address mapper is configured to determine whether the read / write memory access request information belongs to the first address segment or the second address segment; the encoder is configured to, in response to the read / write memory access request information belonging to the first address segment, encode the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information; and the interleaver is configured to distribute the read / write memory access request information to the target last-level cache at the first interleaving granularity based on the intermediate interleaving bit information.
[0012] For example, in a cache management device provided in an embodiment of the present disclosure, the cache management device also includes a decoder, which is configured to perform a decoding operation based on the first interleaved bit information to obtain output address information, and provide the output address information to the target last-level cache.
[0013] For example, in the device for managing cache provided in one embodiment of the present disclosure, the device for managing cache also includes a register; the register is configured to determine whether an interleaving enable control signal acting on the system address mapper is valid; the encoder is further configured to, in response to the interleaving enable signal being valid and the read / write memory access request information belonging to the first address segment, perform an encoding operation on the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information.
[0014] For example, in the device for managing cache provided in one embodiment of the present disclosure, the encoder is further configured to perform mathematical operations on the first interleaved position information and the first bit position information to obtain the intermediate interleaved position information, wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved position corresponding to the first interleaved position information, and the first interleaved position information is retained through a sideband signal to be provided to the subsequent decoding operation.
[0015] For example, in the device for managing cache provided in one embodiment of the present disclosure, the decoder is further configured to, in response to the first bit position information being valid and the first interleaving bit information being invalid, determine the output address information based on the read / write memory access request information and the first interleaving granularity; or, in response to the first bit position information and the first interleaving bit information being valid, determine the output address information based on the read / write memory access request information and the first interleaving granularity; or, in response to the first bit position information being invalid, determine the output address information based on the read / write memory access request information; wherein the first bit corresponding to the first bit position information is a higher bit of the interleaving bit corresponding to the first interleaving bit information.
[0016] For example, in the device for managing cache provided in one embodiment of the present disclosure, the encoder is further configured to set the second bit information of the read and write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retain the initial value of the second bit information through a sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and constrains the burst data size of the read and write memory access request information to not exceed one half of the address boundary specified by the transmission protocol.
[0017] For example, in the device for managing cache provided in an embodiment of the present disclosure, the decoder is further configured to restore the initial value of the second bit information.
[0018] At least one embodiment of the present disclosure provides a method for managing a cache, comprising: receiving read and write memory access request information from a host; in response to an interleaving enable signal being valid and the read and write memory access request information being within a range of a first address segment, distributing the read and write memory access request information to a plurality of corresponding last-level caches at a first interleaving granularity; and in response to the interleaving enable signal being valid and the read and write memory access request information being within a range of a second address segment, distributing the read and write memory access request information to a plurality of corresponding last-level caches at a second interleaving granularity, wherein the second address segment is an address segment in the address space where the cache is located that is different from the first address segment.
[0019] For example, in a method for managing cache provided in an embodiment of the present disclosure, the first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
[0020] For example, in a method for managing cache provided in an embodiment of the present disclosure, the first interleaving granularity is configured based on the size of a cache line, and the second interleaving granularity is configured based on the data bit width of a cache interface.
[0021] For example, in a method for managing cache provided in an embodiment of the present disclosure, the method further includes: performing an encoding operation on first interleaving bit information of the read and write memory access request information to obtain intermediate interleaving bit information; based on the intermediate interleaving bit information, distributing the read and write memory access request information to a target final-level cache at the first interleaving granularity; and performing a decoding operation based on the first interleaving bit information to obtain output address information, and providing the output address information to the target final-level cache.
[0022] For example, in the method for managing cache provided in one embodiment of the present disclosure, the encoding operation is performed on the first interleaved bit information of the read and write memory access request information to obtain intermediate interleaved bit information, including: performing mathematical operations on the first interleaved bit information and the first bit position information to obtain the intermediate interleaved bit information, wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved bit corresponding to the first interleaved bit information; and retaining the first interleaved bit information through a sideband signal to provide it to the subsequent decoding operation.
[0023] For example, in the method for managing cache provided in an embodiment of the present disclosure, the decoding operation is performed based on the first interleaved bit information to obtain the output address information, including: in response to the first bit information being valid and the first interleaved bit information being invalid, determining the output address information based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit information and the first interleaved bit information being valid, determining the output address information based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit information being invalid, determining the output address information based on the read / write memory access request information; wherein the first bit corresponding to the first bit information is a higher bit of the interleaved bit corresponding to the first interleaved bit information.
[0024] For example, in the method for managing cache provided in an embodiment of the present disclosure, the encoding operation also includes: setting the second bit information of the read and write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retaining the initial value of the second bit information through a sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and constrains the burst data size of the read and write memory access request information to not exceed one half of the address boundary specified by the transmission protocol; and the decoding operation also includes: restoring the initial value of the second bit information.
[0025] For example, in a method for managing cache provided in an embodiment of the present disclosure, the method further includes: in response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the first address segment, sending the read / write memory access request information to a first last-level cache; and, in response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the second address segment, sending the read / write memory access request information to a second last-level cache, wherein the first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
[0026] At least one embodiment of the present disclosure provides a method for managing a cache, comprising: receiving read and write memory access request information from a host, and in response to the read and write memory access request information being within a range of a first address segment, encoding first interleaving bit information of the read and write memory access request information to obtain intermediate interleaving bit information, wherein the intermediate interleaving bit information is used to select a target final-level cache among the multiple final-level caches; based on the intermediate interleaving bit information, distributing the read and write memory access request information to the target final-level cache at the first interleaving granularity; and performing a decoding operation based on the first interleaving bit information to obtain output address information, and providing the output address information to the target final-level cache.
[0027] For example, in the method for managing cache provided in one embodiment of the present disclosure, in response to the read / write memory access request information being within the range of the first address segment, encoding the first interleaved bit information of the read / write memory access request information to obtain intermediate interleaved bit information, including: in response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment, encoding the first interleaved bit information of the read / write memory access request information to obtain intermediate interleaved bit information.
[0028] At least one embodiment of the present disclosure provides a system on chip, comprising: a plurality of processor cores; and an interconnect bus, wherein the interconnect bus is configured to route access requests of the plurality of processor cores to a cache management device provided by any embodiment of the present disclosure.
[0029] At least one embodiment of the present disclosure provides an electronic device, including the device for managing cache provided by any embodiment of the present disclosure.
[0030] At least one embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory, wherein the memory stores at least one computer program, and when the at least one computer program is executed by the processor, the method for managing cache provided by any embodiment of the present disclosure is implemented.
[0031] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium for non-temporarily storing computer-readable instructions. When the computer-readable instructions are executed by a computer, the method for managing cache provided by any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure.
[0033] Figure 1 A schematic block diagram of a cache management device provided by at least one embodiment of the present disclosure is shown;
[0034] Figure 2 A schematic block diagram of another cache management device provided by at least one embodiment of the present disclosure is shown;
[0035] Figure 3 A flowchart of a method for managing cache provided by at least one embodiment of the present disclosure is shown;
[0036] Figure 4 A flowchart of another method for managing cache provided by at least one embodiment of the present disclosure is shown;
[0037] Figure 5 A schematic diagram of an application of a cache management device provided by at least one embodiment of the present disclosure is shown;
[0038] Figure 6 A schematic diagram of another application of a cache management device provided by at least one embodiment of the present disclosure is shown;
[0039] Figure 7 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown;
[0040] Figure 8 A schematic block diagram showing another electronic device provided by at least one embodiment of the present disclosure; and
[0041] Fig. 9 A schematic diagram of a computer-readable storage medium provided by at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0043] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, similar words such as "one", "one" or "the" do not indicate quantity restrictions, but indicate that there is at least one. Similar words such as "include" or "comprise" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Similar words such as "connect" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0044] For example, tablet SOC or other SOC (as a key component of electronic devices, its performance will affect the user experience of the final product. As semiconductor technology evolves from micron level to nano level, the number of transistors that can be integrated on the chip has increased significantly, which allows SOC to integrate multiple functional modules or subsystems, such as CPU, GPU, memory controller, communication module, etc. However, this high degree of integration also brings new challenges, especially in bus architecture design. In order to achieve efficient communication between these modules, an advanced bus architecture is needed to connect the modules so that data can be quickly transmitted and processed.
[0045] In addition, as users' performance requirements continue to increase, they also hope that electronic devices can have a longer battery life. This requires the SOC bus architecture to not only meet the needs of high-performance data transmission, but also to reduce power consumption as much as possible. For example, low voltage differential signaling (LVDS) technology can be used to achieve high-speed data transmission while reducing power consumption12. In addition, by optimizing the clock management of the bus and dynamically adjusting the clock frequency according to different workloads, power consumption can be further reduced.
[0046] SOC integrates a variety of modules with different functions and performance requirements, and the data transmission rate, data format, operating frequency, etc. of each module are different. For example, high-speed, bursty data transmission is required between the CPU and the memory, while low-speed, periodic data transmission is required between the sensor module and the processor. Therefore, the bus architecture needs to have a flexible topology and protocol. For example, a multi-layer bus structure can be used to connect high-speed modules and low-speed modules to different levels of buses to improve the overall performance and reliability of the system.
[0047] In order to facilitate the interconnection between chips of different types or models and external devices, the SOC bus architecture needs to follow certain standards and specifications. For example, the AMBA (Advanced Microcontroller Bus Architecture) bus standard can be applied to the ARM (Advanced RISC Machine) architecture SOC, which defines a variety of bus protocols and interface specifications, allowing IP cores of different types or models to work together on the same bus, improving the compatibility and scalability of the system.
[0048] However, the inventors of the present disclosure have noticed that although the above-mentioned technology can improve the performance of SOC, there are still some problems that need to be solved in practical applications. For example, in combination with the requirements of different application scenarios, the characteristics of Last Level Cache (LLC) and SRAM should be reasonably allocated and utilized to optimize data processing performance and system efficiency. For example, some scenarios require pre-processing of data obtained from DDR and writing to SRAM, while some scenarios rely on the time and space locality of LLC to improve performance, and some scenarios require the use of SRAM and LLC at the same time. For another example, in a SOC system, the utilization of the Cache directly affects the system performance. For example, some NoC IPs can only implement static interleaving logic for certain fixed address bits, resulting in some sets in the cache always being unable to be used, which seriously wastes Cache resources.
[0049] At least one embodiment of the present disclosure provides a device for managing cache, including: a source interface module, a register, a system address mapper, and an interleaver. The source interface module is configured to receive read and write memory access request information from a host; the register is configured to determine whether an interleaving enable control signal acting on the system address mapper is valid; the system address mapper is configured to determine whether the read and write memory access request information belongs to the first address segment or the second address segment; the interleaver is configured to distribute the read and write memory access request information to multiple corresponding last-level caches at a first interleaving granularity in response to the interleaving enable signal being valid and the read and write memory access request information belonging to the first address segment, and the interleaver is also configured to distribute the read and write memory access request information to multiple corresponding last-level caches at a second interleaving granularity in response to the interleaving enable signal being valid and the read and write memory access request information belonging to the second address segment; wherein the second address segment is an address segment in the address space where the cache is located that is different from the first address segment. The device for managing the cache divides the address segments and enables interleaving control. LLC can realize on-demand switching of multiple modes to meet the needs of multiple scenarios (such as high-performance computing and large-bandwidth data processing).
[0050] At least one embodiment of the present disclosure provides another device for managing cache, including: a source interface module, a system address mapper, an encoder, a decoder, and an interleaver. The source interface module is configured to receive read and write memory access request information from a host; the system address mapper is configured to determine whether the read and write memory access request information belongs to the first address segment or the second address segment; the encoder is configured to, in response to the read and write memory access request information belonging to the first address segment, encode the first interleaved bit information of the read and write memory access request information to obtain intermediate interleaved bit information; and the interleaver is configured to distribute the read and write memory access request information to the target last-level cache at a first interleaved granularity based on the intermediate interleaved bit information; the decoder is configured to perform a decoding operation based on the first interleaved bit information to obtain output address information, and provide the output address information to the target last-level cache. The device for managing cache can solve the problem of low set index utilization in LLC.
[0051] In addition, embodiments of the present disclosure also provide a method for managing cache, an electronic device, and a storage medium corresponding to the above-mentioned device for managing cache.
[0052] Figure 1 A schematic block diagram of a cache management device provided by at least one embodiment of the present disclosure is shown.
[0053] like Figure 1 As shown, the cache management device 100 includes: a source interface module 110 , a register 120 , a system address mapper (SAM, System Address Map) 130 and an interleaver 140 .
[0054] The source interface module 110 is configured to receive read and write memory access request information from the host.
[0055] Register 120 is configured to determine whether an interleaving enable control signal acting on system address mapper 130 is valid.
[0056] The system address mapper 130 is configured to determine whether the read / write memory access request information belongs to the first address segment or the second address segment.
[0057] The interleaver 140 is configured to distribute the read / write memory access request information to a plurality of corresponding last-level caches at a first interleaving granularity in response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment.
[0058] The interleaver 140 is further configured to distribute the read / write memory access request information to a plurality of corresponding last-level caches at a second interleaving granularity in response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the second address segment.
[0059] For example, a read / write memory access request issued by a host (such as a CPU, GPU or other processing unit) may be transmitted to the cache management device 100 via a bus architecture (such as a domain-specific NOC). For example, the read / write memory access request information may be received by the source interface module 110 .
[0060] For example, read and write memory access request information may include key fields such as read request address, write request address, data length or operation type (eg, read / write), which can provide a basis for determining subsequent address segments and interleaving distribution.
[0061] For example, the first address segment may correspond to a cache function address segment. For example, when the read / write memory access request information belongs to the first address segment, the traditional LLC may be used as a cache access mode (cache mode, cache mode). For example, it may be applicable to data access scenarios with spatial locality (such as CPU frequently accessing continuous memory blocks), thereby improving the cache hit rate.
[0062] For example, the first interleaving granularity can be configured based on the cache line size to conform to the characteristic that the cache uses a cache line as the minimum processing unit, thereby improving the cache execution efficiency. For example, the cache line size can be changed according to the needs of the actual application or the limitations of the hardware (such as the limitations of the memory controller), and the embodiments of the present disclosure are not limited to this. Exemplarily, if the cache line size is 128 bytes, the first interleaving granularity can be set to 128B.
[0063] For example, a hardware address comparator or a preconfigured address range register (the embodiments of the present disclosure are not limited to the hardware) may be used to determine whether the read or write memory access request falls within the first address segment.
[0064] For example, when the address of the read / write memory access request completely falls within the first address segment (cache function segment), it can be considered that the read / write memory access request information belongs to the range of the first address segment.
[0065] For example, when the interleaving enable signal is valid (eg, configurable by software via a configuration register), cache line-based interleaving distribution is enabled.
[0066] For example, based on the first interleaving bit information of the read / write memory access request information, the read / write memory access request information may be distributed to a plurality of corresponding last-level caches at a first interleaving granularity.
[0067] Here, the first interleaving bit information of the read / write memory access request information is information indicated by the interleaving bit representing the read / write memory access request information.
[0068] For example, the first interleaving bit information may be used to select a corresponding last-level cache among a plurality of last-level caches (eg, select LLC0 or LLC1).
[0069] For example, the first interleaving bit information may be determined based on the set idex of the read / write memory access request information. Exemplarily, the first interleaving bit information may select a low bit (eg, address 6) in the set idex interval.
[0070] For example, the set_idex range can be determined based on the cache line size.
[0071] For example, multiple independent LLCs (such as LLC0 and LLC1) can be deployed, and each LLC can be independently configured as cache mode (dynamic replacement) or SRAM mode (direct access). For example, when the read and write memory access request information belongs to the range of the first address segment, the request is alternately distributed to different LLCs at the first interleaving granularity to improve the parallel bandwidth.
[0072] For example, before distribution, the interleaving selection bits of the address can be XOR-encoded by the encoder (such as XORing the high-order address with the low-order address) to solve the problem of NoC static interleaving and LLC set index conflicts and fixing certain set index bits to reduce cache group (Set) utilization.
[0073] For example, when the interleaving enable is valid and the read / write memory access request information belongs to the second address segment, it is distributed to the LLC according to the second interleaving granularity.
[0074] For example, the second address segment may correspond to the SRAM direct access address segment. For example, when the read / write memory access request information belongs to the second address segment, LLC may be used as the SRAM access mode (SRAM mode). For example, it may be applicable to application scenarios that require large bandwidth and low latency data preprocessing (such as temporary storage of data before GPU rendering, etc., which is not limited in the embodiments of the present disclosure).
[0075] For example, the second interleaving granularity may be configured based on the data bit width of the cache interface, so that the data block accessed each time may be aligned with the bus bit width, thereby reducing the number of transmissions and improving bandwidth utilization.
[0076] For example, the same LLC physical space is divided into multiple logical segments through address remapping technology. When the read and write memory access request falls on the second address segment, the LLC can enter the SRAM mode, and the cache replacement strategy (such as LRU) can be automatically shielded by hardware, and access can be directly mapped by address.
[0077] For example, in an embodiment of the present disclosure, the second interleaving granularity may be configured according to the data bit width of the cache interface. Exemplarily, if the data bit width of the interface is 64 bytes, the second interleaving granularity may be set to 64B.
[0078] For example, when the interleaving enable signal is valid and the read / write memory access request information belongs to the second address segment, the read / write memory access request information can be distributed to multiple corresponding last-level caches at the second interleaving granularity based on the first interleaving bit information of the read / write memory access request information.
[0079] For example, in scenarios where SRAM and Cache need to be used simultaneously (e.g., some data needs to be frequently accessed after preprocessing), LLC0 can be configured in SRAM mode and LLC1 in Cache mode (dynamic management), and resource isolation and efficient utilization can be achieved through address segment division.
[0080] In at least one embodiment of the present disclosure, through address segment division and interleaving enable control, LLC can realize on-demand switching of multiple modes (for example, cache mode and SRAM mode) to meet the needs of multiple scenarios (such as high-performance computing and large-bandwidth data processing) to further reduce the waste of resources caused by fixed functions in traditional architectures. In addition, adapting the cache mode with the first interleaving granularity can improve the cache execution efficiency; adapting the bus bit width with the second interleaving granularity can maximize the data transmission efficiency. The combination of the two can significantly improve the overall performance of the system.
[0081] In some embodiments of the present disclosure, the interleaver 140 is also configured to send the read-write memory access request information to the first final-level cache in response to the interleaving enable signal being invalid and the read-write memory access request information falling within the range of the first address segment; and, in response to the interleaving enable signal being invalid and the read-write memory access request information falling within the range of the second address segment, send the read-write memory access request information to the second final-level cache.
[0082] For example, in light load scenarios (such as standby, low-power standby such as sensor data acquisition, and deterministic low-latency access), the interleaving enable signal can be turned off to reduce dynamic power consumption.
[0083] For example, when the interleaving enable signal is invalid, circuits such as the dynamic codec module and the multi-way interleaving logic can turn off the clock or cut off the power supply to reduce static and dynamic power consumption.
[0084] For example, when the interleaving enable signal is invalid and the read / write memory access request information is within the range of the first address segment, the read / write memory access request information can be directly routed to the first last-level cache (eg, LLC 1). In this case, LLC 1 operates as a traditional LLC.
[0085] For example, when the interleaving enable signal is invalid and the read / write memory access request information is within the range of the second address segment, the read / write memory access request information can be directly routed to the second last-level cache (eg, LLC 0). In this case, LLC 0 operates as SRAM.
[0086] In a possible implementation, for example, in a standby mode (low power consumption) application scenario, it is determined that the interleave enable signal is turned off (interleave_enable=0). For example, a data write request generated by the Sensor Hub is received, for example, the address is 0x0000_1000, and it is further determined that the address belongs to the range of the second address segment, then the data write request can be directly routed to LLC 0, that is, the data is written to the physical SRAM area, and the low power CPU reads the data from LLC 0.
[0087] In another possible implementation, for example, in an AP domain high performance computing application scenario, it is determined that the interleaving enable signal is turned off (interleave_enable=0), and the CPU needs to frequently access the memory database. For example, the address of the data read request is 0x8000_0000, and it is further determined that the address belongs to the range of the first address segment. Then, the address hits LLC 1, and LLC 1 returns data according to the cache line. If it does not hit, it is pre-fetched from other memories (such as DDR).
[0088] In some embodiments of the present disclosure, by interleaving the enable signal, the system can dynamically select efficient interleaving or low-power static routing according to the load to achieve a balance between performance and power consumption. For example, in high-performance scenarios, LLC can be dynamically reused as cache or SRAM to adapt to diverse needs. For example, in low-power scenarios, LLC can be statically divided to avoid idle resources, further improving the scenario coverage of the SOC cache management solution.
[0089] For example, in the embodiments of the present disclosure, the source interface module 110, the register 120, the system address mapper 130, and the interleaver 140 can be hardware, software, firmware, and any feasible combination thereof. For example, the source interface module 110, the register 120, the system address mapper 130, and the interleaver 140 can be a dedicated or general circuit, chip, or device, or a combination of a processor and a memory. The embodiments of the present disclosure do not limit the specific implementation forms of the above-mentioned modules.
[0090] It should be noted that in the embodiments of the present disclosure, the specific operations that each module of the cache management device 100 is configured to perform and the beneficial effects that can be achieved may correspond to the various steps of the cache management method provided in any embodiment of the present disclosure. Figure 1 The components and structures of the device 100 for managing cache are merely exemplary and non-restrictive. The device 100 for managing cache may further include other components and structures as needed.
[0091] The inventors of the present disclosure also noticed that, for example, in the cache management architecture of a SOC, the conflict between the interleaving logic of the NoC (Network-on-Chip) and the set index of the LLC (Last Level Cache) is a key issue leading to low cache utilization (for example, some sets cannot be used due to fixed address bits).
[0092] In view of this, at least one embodiment of the present disclosure provides a device for managing cache.
[0093] Figure 2 A schematic block diagram of a cache management device provided by at least one embodiment of the present disclosure is shown.
[0094] like Figure 2 As shown, the cache management device 200 includes: a source interface module 210, a system address mapper 220, an encoder 230 and an interleaver 240.
[0095] The source interface module 210 is configured to receive read and write memory access request information from the host.
[0096] The system address mapper 220 is configured to determine whether the read / write memory access request information belongs to the first address segment or the second address segment.
[0097] The encoder 230 is configured to, in response to the read / write memory access request information belonging to the range of the first address segment, perform an encoding operation on the first interleaved bit information of the read / write memory access request information to obtain the intermediate interleaved bit information.
[0098] The interleaver 240 is configured to distribute the read and write memory access request information to the target last-level cache at a first interleaving granularity based on the intermediate interleaving bit information.
[0099] In some embodiments of the present disclosure, the cache management device 200 further includes a decoder 250 .
[0100] The decoder 250 may be configured to perform a decoding operation based on the first interleaved bit information to obtain output address information, and provide the output address information to a target last-level cache.
[0101] The cache management device 200 can use dynamic address encoding and decoding technology to encode the interleaving bits before distributing read and write memory access requests, and restore the real address after distribution, thereby decoupling the interleaving logic and the set index bit of the LLC, and maximizing cache utilization.
[0102] Since some NoC interleaving logics (such as selecting LLC based on the high bits of the address) may fix certain bits of the LLC set index, some cache sets (Sets) cannot be accessed.
[0103] For example, if the interleaving bit of the NoC is address 6 and the set index of the LLC uses address 9:6, when the interleaving bit is fixed to a specific value, the low bit of the set index (such as address 6) may be fixed, resulting in a technical problem that some sets cannot be covered.
[0104] Therefore, through the encoding operation, the set index received by the cache is no longer fixed by bits, so that all sets in the LLC can be effectively used.
[0105] For example, the set index bits of the target address can be extracted from the read or write request to interleave the bits to achieve input address resolution.
[0106] For example, an XOR operation may be performed on the interleaved bit and other bits (eg, slightly higher bits) other than the set index bit (eg, XORing address bit 6 with address bit 17) to generate intermediate interleaved bits (encoded interleaved bits).
[0107] For example, assuming that the interleaved bit of the original address is 0 (binary), and the slightly higher bits other than the set index bit are 1 (binary), after XOR encoding, the middle interleaved bit may become 1 (specifically, it may depend on the XOR rule).
[0108] For example, the encoded interleaving bit is used to select a target LLC (eg, LLC0 or LLC1), while the low bit of setindex (ie, the interleaving bit) is not fixed, so that all Sets can be covered.
[0109] For example, the intermediate interleaving bit information may be used to select a target last-level cache among a plurality of last-level caches.
[0110] Since the interleaving logic of some NoCs directly uses the interleaving bits of the original address to select LLC, this may lead to the technical problem of limited Set distribution. Through the intermediate interleaving bits, the selection of LLC can no longer be strongly bound to the interleaving bits of the original address, thereby improving the utilization of Set.
[0111] For example, during distribution, interleaving may be performed based on the interleaving granularity configured based on the cache line size (eg, 128B), so that data at adjacent addresses may be distributed in different LLCs, making full use of the parallel access capability.
[0112] After receiving the request, LLC needs to access data according to the set index and tag of the original address, but the interleaving bit has changed the interleaving bit information of the address, so decoding is required to restore the real address. The decoding operation can also make the internal logic of LLC (such as cache replacement strategy and consistency protocol) not need to be modified and still work according to the original address.
[0113] For example, when the amount of each burst of read and write accesses is aligned with the interleaving granularity, an inverse operation (such as XOR decoding) can be performed on the intermediate interleaving bits to restore the interleaving bit information of the original address.
[0114] For example, in the aforementioned encoding operation, the interleaved bit 0 is XORed with the slightly higher bit 1 outside the set index bit to generate the intermediate interleaved bit 1. When decoding, the intermediate interleaved bit 1 needs to be XORed with the slightly higher bit 1 outside the current set index bit again to obtain the original interleaved bit 0.
[0115] For example, the decoded interleaved bits can be combined with other bits of the original address (such as Tag, Offset) to generate a final output address for LLC to access data.
[0116] In some embodiments of the present disclosure, the fixed relationship between the NoC interleaving bit and the LLC set index bit can be broken through the encoding operation, so that all Sets can be evenly accessed. For example, the situation where 50% of the Sets may be unusable due to the fixed interleaving bit can achieve 100% utilization after encoding. In addition, the encoding and decoding rules can be configured according to different scenarios (such as the number of XOR bits, address bits involved in encoding, etc., which are not limited by the embodiments of the present disclosure) to flexibly adapt to a variety of LLC topologies. While maintaining the high efficiency of the NoC interleaving logic, the embodiments of the present disclosure can further solve the problem of low LLC setindex utilization and further improve the high performance and low power consumption design of the SOC.
[0117] In some embodiments of the present disclosure, the cache management device 200 further includes a register 260 .
[0118] Register 260 may be configured to determine whether an interleave enable control signal acting on the system address mapper is asserted.
[0119] The encoder 230 may be further configured to, in response to the interleaving enable signal being valid and the read / write memory access request information being within the first address segment, encode the first interleaving bit information of the read / write memory access request information to obtain the intermediate interleaving bit information.
[0120] For example, the operations related to the register 260 determining whether the interleaving enable control signal is valid may refer to the related description of the register 120 in the above embodiment, which will not be described in detail here.
[0121] For example, after the register 260 determines that the interleaving enable control signal is valid, the system address mapper 220 can further determine whether the read / write memory access request information belongs to the first address segment. In the case where the read / write memory access request information belongs to the first address segment, the encoder 230 can perform an encoding operation on the first interleaving bit information of the read / write memory access request information to obtain the intermediate interleaving bit information. In some embodiments of the present disclosure, the encoder 230 can be further configured to perform a mathematical operation on the first interleaving bit information and the first bit information to obtain the intermediate interleaving bit information; and retain the first interleaving bit information through a sideband signal to provide it to a subsequent decoding operation.
[0122] For example, while encoding the first interleaved bit information, the original first interleaved bit information may be retained through a sideband signal, so that a subsequent decoding operation may restore the real address based on the original first interleaved bit information.
[0123] For example, when performing mathematical operations, the input parameters may include first interleaved bit information and first bit position information.
[0124] For example, the first interleaving bit information may include an interleaving bit for selecting LLC in the original address. For example, if two LLCs are interleaved, the interleaving bit may be bit 6, and if four LLCs are interleaved, the interleaving bit may be bit 7:6.
[0125] For example, the first bit information may include a slightly higher address than the set index bit. Exemplarily, the interleaving bit is a lower bit in the set index bit, such as address 6, and the first bit may be address 17.
[0126] For example, an XOR operation may be performed on the first interleaving bit information and the first bit position information to generate intermediate interleaving bit information.
[0127] Exemplarily, if the first interleaved bit information is 0 (binary, address 6=0) and the first bit information is 1 (address 17=1), the XOR result is 0^1=1 (middle interleaved bit information).
[0128] For example, the original first interleaved bit information may be retained through a sideband signal, and the original first interleaved bit information may be transmitted to a subsequent decoder 250 through a sideband signal for decoding operation.
[0129] For example, when the amount of each burst data of read and write access is aligned with the interleaving granularity, the original first interleaving bit information and the intermediate interleaving bit information can be combined for inverse operation during decoding to restore the real address.
[0130] In at least one embodiment of the present disclosure, the interleaving bit is dynamically changed through XOR coding, so that the set idex bit of a certain LLC is dynamically changed, and finally all cache sets in the LLC can be utilized.
[0131] In some embodiments of the present disclosure, the decoder 250 can be further configured to, in response to the first bit information being valid and the first interleaved bit information being invalid, determine the output address information based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit information and the first interleaved bit information being valid, determine the output address information based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit information being invalid, determine the output address information based on the read / write memory access request information.
[0132] Exemplary:
[0133] / / Output address generation logic
[0134] addr_out[7]=addr_in[7]^addr_in
[17] / / XOR operation introduces dynamic adjustment addr_out
[11] =1'b0; / / Force a fixed bit / / Sideband signal logic
[0135] tag_2k = addr_in
[11] ; / / retain original address bits for decoding
[0136] tag_bit7 = addr_in[7];
[0137] / / Dynamic address adjustment
[0138] addr_out = addr_in
[17] ? / / Judge the high flag
[0139] (tag_bit7?addr_in+granu / / Increase granularity offset
[0140] :addr_in-granu) / / Reduce granularity offset
[0141] :=addr_in; / / no adjustment addr_out
[11] =tag_2k; / / restore the original address bit
[0142] For example, if the first bit information is valid, that is, addr_in
[17] is valid (for example, 1), then it is further determined whether the first interleaving bit information is valid, that is, further based on the value of tag_bit7, the offset of the interleaving granularity (granu) is increased or decreased on the basis of the original read and write memory access request information as the output address information (addr_out), so that the data block requested across the granularity can be completely mapped to the continuous LLC.
[0143] For example, if the first bit information is invalid, that is, addr_in
[17] is invalid (for example, is 0), the original read / write memory access request information can be directly used as the output address information (addr_out).
[0144] For example, addr_out
[11] can be set to tag_2k (ie, the original address 11) and retained through the sideband signal so that the key address bits can be correctly restored during decoding.
[0145] For example, the output address information (addr_out) may be calculated according to the difference between the requested data width and the interleaving granularity.
[0146] For example, a decoding operation may be selected based on the first bit information of the read / write memory access request information and the original first interleaved bit information (e.g., transmitted via a sideband signal). When the first bit information is valid and the original first interleaved bit is valid, the output address information (addr_out) may be the input read / write memory access request information plus the interleaved granularity; when the first bit information is valid but the original first interleaved bit information is invalid, the output address information (addr_out) may be the input read / write memory access request information minus the interleaved granularity; when the first bit information is invalid, the output address information (addr_out) is the input read / write memory access request information.
[0147] For example, in some application scenarios of 4K video stream writing, if the data size of each frame is 64 bytes (aligned cache line), the first interleaved bit information can be XOR-encoded and decoded so that the corresponding data can not only be evenly distributed to multiple LLCs, but also the Set utilization of each LLC can be maximized.
[0148] Through the coordinated operation of the encoder 230 and the decoder 250, the disclosed embodiment achieves efficient adaptation to diverse data access scenarios in the SOC, solves the problem of reduced cache utilization caused by fixed interleaving in some architectures, and provides a robust cache management solution for complex application loads.
[0149] In some embodiments of the present disclosure, the encoder 230 can be further configured to set the second bit information of the read and write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retain the initial value of the second bit information through the sideband signal and constrain the burst data size of the read and write memory access request information to not exceed one half of the address boundary specified by the transmission protocol.
[0150] For example, the decoder 250 may be further configured to restore the initial value of the second bit information.
[0151] For example, the second bit information may be determined based on an address boundary specified by a transmission protocol.
[0152] In the SOC cache management process, the address boundary crossing problem is a key challenge in the interleaved codec design. For example, when following the AXI (Advanced eXtensible Interface) bus protocol, the address range of the burst transmission must be strictly constrained. For example, the AXI bus stipulates that a single burst transmission cannot cross the address boundary (such as the 4K address boundary, that is, the address increment must not exceed 4KB) to prevent the data packet from being incorrectly routed to different slave devices.
[0153] The inventors of the present disclosure noticed that during the encoding operation stage, if the first interleaved bit information of the read and write memory access request information is XORed, the address increment of subsequent burst transmissions may cause the boundary address bit specified by the protocol to flip, thereby crossing the address boundary (such as the 4K boundary).
[0154] Therefore, the encoder 230 can force Bit 11 (the second bit) of the input read / write memory access request information to be set to zero during the encoding phase, so that the address will not cross the address boundary (eg, 4K boundary) even if the address is incremented subsequently.
[0155] Exemplary:
[0156] / / Encoding operation logic
[0157] addr_encoded
[11] =1'b0; / / Force Bit11 to zero
[0158] In this way, even if the first address of the burst transfer is modified, its subsequent address increment (such as Increment Burst) will not cross the 4K boundary due to the flipping of Bit 12.
[0159] For example, during the encoding operation phase, the initial value of Bit 11 may also be retained through a sideband signal.
[0160] For example, the Bit 11 value (second bit information) of the original read / write memory access request information may be transmitted to the decoding module via a sideband signal.
[0161] Exemplary:
[0162] sideband_tag_2k = addr_in
[11] ; / / Save the original Bit11
[0163] For example, the decoder 250 may restore the original Bit 11 value during the decoding operation so that the address accessed by the LLC is consistent with the actual request.
[0164] In addition, the inventors of the present disclosure also noticed that the burst length (Burst Size) can determine the amount of data transmitted in a single time, which affects the address increment. For example, if the Burst Size is 8 (transmitting 8 data blocks), the address increment is data width × Burst Length. If the burst transmission length is too large (for example, more than 2K), even if the first address Bit11 is set to zero, the address increment may still exceed the 4K boundary (for example, if the first address is 0, when transmitting 2K+1 bytes, the address will cross the 2048 boundary). Therefore, it is also necessary to design constraint rules to constrain the burst length to not exceed 2K.
[0165] For example, it may be mandatory that the length of burst transmission data sent by the host does not exceed one half (ie, 2K) of the 4K boundary.
[0166] For example, after Bit 11 is set to zero, a 2K address increment (0 to 2047) will not trigger a flip of Bit 12, so that the address is always within the range of 0 to 2K. In addition, short burst transmission (≤2K) can reduce the occupancy time of the bus arbiter, further reducing the long-term exclusive bandwidth of large bursts, thereby improving the concurrent performance of multiple hosts.
[0167] For the decoding operation, for example, the decoder 250 may extract the saved sideband_tag_2k (original Bit11 value) from the sideband signal and write it back to the Bit11 position of the output address to restore the original Bit11 value.
[0168] Exemplary:
[0169] addr_out
[11] = sideband_tag_2k; / / restore original Bit 11
[0170] For example, the original value of Bit 11 can be restored through a decoding operation, so that the final address received by the LLC is consistent with the original request address, thereby improving the correctness of cache access.
[0171] In a possible implementation, for example, in an application scenario of burst writing a 2K data block, the original request address is: 0x0000_1000 (Bit 11=1).
[0172] For example, during the encoding operation phase, the encoder 230 can force the value of Bit 11 to be zero (the address becomes 0x0000_0000); and save the value of Bit 11 through the sideband signal sideband_tag_2k=1. For example, for burst transmission, BurstSize=2K, and the address increment range is 0x0000_0000~0x0000_07FF (Bit 11 is always 0).
[0173] For example, during the decoding operation phase, the decoder 250 can restore Bit 11 to the original value 1, and the final address is, for example, 0x0000_1000 to 0x0000_17FF. All addresses are located in the same 4K block, which meets the AXI transmission protocol requirements.
[0174] In at least one embodiment of the present disclosure, by setting the second bit information to zero and constraining the burst length, the risk of the address crossing the address boundary (e.g., 4K boundary) can be further reduced to meet the requirements of the AXI transmission protocol. In addition, through the sideband signal retention and decoding restoration mechanism, the LLC can access the data according to the original address, further reducing the Tag or Set errors caused by encoding. In addition, short burst transmission can reduce bus congestion, improve multi-host concurrency efficiency, and improve system performance. While achieving LLC interleaving optimization and cache utilization improvement, strictly following the bus protocol constraints provides reliable guarantee for efficient data transmission and resource management of SOC.
[0175] For example, in the embodiments of the present disclosure, the source interface module 210, the system address mapper 220, the encoder 230, the interleaver 240, the decoder 250 and the register 260 can be hardware, software, firmware and any feasible combination thereof. For example, the source interface module 210, the system address mapper 220, the encoder 230, the interleaver 240, the decoder 250 and the register 260 can be a dedicated or general circuit, chip or device, etc., or a combination of a processor and a memory. The embodiments of the present disclosure do not limit the specific implementation form of the above-mentioned modules.
[0176] It should be noted that in the embodiments of the present disclosure, the specific operations that each module of the cache management device 200 is configured to perform and the beneficial effects that can be achieved may correspond to the various steps of the cache management method provided in any embodiment of the present disclosure. Figure 2 The components and structure of the device 200 for managing cache are merely exemplary and non-restrictive. The device 200 for managing cache may further include other components and structures as needed.
[0177] In some embodiments of the present disclosure, the apparatus 100 for managing a cache and the apparatus 200 for managing a cache may be implemented in the same electronic device.
[0178] For example, Figure 1 and Figure 2 The components and structures of the device for managing cache 100 and the device for managing cache 200 shown are merely exemplary and non-restrictive. As needed, the device for managing cache 100 and the device for managing cache 200 may share the same circuit, chip or module, which is not limited in the embodiments of the present disclosure.
[0179] For example, the cache management device 100 and the cache management device 200 may share the same source interface module, system address mapper, encoder, interleaver, decoder, and register, etc., which is not limited in the embodiments of the present disclosure.
[0180] Figure 3 A flowchart of a method for managing cache provided by at least one embodiment of the present disclosure is shown.
[0181] like Figure 3 As shown, the method for managing cache includes steps S300 to S320. The method for managing cache can be applied to a cache management device (eg, cache management device 100) provided in any embodiment of the present disclosure, and the embodiments of the present disclosure are not limited thereto.
[0182] Step S300: receiving a read / write memory access request from a host.
[0183] Step S310: In response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment, the read / write memory access request information is distributed to a plurality of corresponding last-level caches at a first interleaving granularity.
[0184] Step S320: In response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the second address segment, the read / write memory access request information is distributed to a plurality of corresponding last-level caches at a second interleaving granularity.
[0185] For example, the second address segment is an address segment in the address space where the cache is located that is different from the first address segment.
[0186] For example, step S300 may be applied to the source interface module 110 provided in the above embodiment. For example, steps S310 to S320 may be applied to the interleaver 140 provided in the above embodiment.
[0187] For example, in a method for managing cache provided in an embodiment of the present disclosure, the first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
[0188] For example, in a method for managing cache provided in an embodiment of the present disclosure, a first interleaving granularity is configured based on a cache line size, and a second interleaving granularity is configured based on a data bit width of a cache interface.
[0189] For example, in a method for managing cache provided in an embodiment of the present disclosure, the method for managing cache further includes steps S330 to S350.
[0190] Step S330: performing an encoding operation on the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information.
[0191] Step S340: Based on the intermediate interleaving bit information, the read and write memory access request information is distributed to the target last-level cache at the first interleaving granularity.
[0192] Step S350: performing a decoding operation based on the first interleaved bit information to obtain output address information, and providing the output address information to the target last-level cache.
[0193] For example, step S330 may be applied to an encoder of the apparatus 100 for managing a cache. For example, step S340 may be applied to the interleaver 140 provided in the above embodiment. For example, step S350 may be applied to a decoder of the apparatus 100 for managing a cache.
[0194] For example, in the method for managing cache provided in an embodiment of the present disclosure, step S330 may further include steps S331 to S332.
[0195] Step S331: performing a mathematical operation on the first interleaved bit information and the first bit position information to obtain intermediate interleaved bit information.
[0196] For example, the first bit corresponding to the first bit information is a higher bit of the interleaved bit corresponding to the first interleaved bit information.
[0197] Step S332: retain the first interleaved bit information via a sideband signal to provide it for subsequent decoding operations.
[0198] For example, steps S331 to S332 may be applied to an encoder of the apparatus 100 for managing a cache.
[0199] For example, in the method for managing cache provided in an embodiment of the present disclosure, step S350 may further include step S351: in response to the first bit position information being valid and the first interleaved bit information being invalid, determining the output address information based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit position information and the first interleaved bit information being valid, determining the output address information based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit position information being invalid, determining the output address information based on the read / write memory access request information; wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved bit corresponding to the first interleaved bit information.
[0200] For example, step S351 may be applied to a decoder of the apparatus 100 for managing a cache.
[0201] For example, in the method for managing cache provided in an embodiment of the present disclosure, step S330 may further include step S333: setting the second bit information of the read and write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retaining the initial value of the second bit information through the sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and constraining the burst data size of the read and write memory access request information not to exceed one half of the address boundary specified by the transmission protocol.
[0202] For example, step S350 may further include step S352: the decoding operation further includes: restoring the initial value of the second bit information.
[0203] For example, step S333 may be applied to an encoder of the apparatus 100 for managing a cache. For example, step S352 may be applied to a decoder of the apparatus 100 for managing a cache.
[0204] For example, in a method for managing cache provided in an embodiment of the present disclosure, the method for managing cache also includes step S360: in response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the first address segment, sending the read / write memory access request information to the first last-level cache; and, in response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the second address segment, sending the read / write memory access request information to the second last-level cache, wherein the first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
[0205] For example, step S360 may be applied to the interleaver 140 provided in the aforementioned embodiment.
[0206] Figure 4 A flowchart of another method for managing cache provided by at least one embodiment of the present disclosure is shown.
[0207] like Figure 4 As shown, the method for managing cache includes steps S400 to S430. The method for managing cache can be applied to a cache management device (eg, cache management device 200) provided in any embodiment of the present disclosure, and the embodiments of the present disclosure are not limited thereto.
[0208] Step S400: receiving a read / write memory access request from a host.
[0209] Step S410: In response to the read / write memory access request information belonging to the range of the first address segment, encoding operation is performed on the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information.
[0210] Step S420: Based on the intermediate interleaving bit information, the read and write memory access request information is distributed to the target last-level cache at the first interleaving granularity.
[0211] Step S430: performing a decoding operation based on the first interleaved bit information to obtain output address information, and providing the output address information to the target last-level cache.
[0212] For example, step S400 may be applied to the source interface module 210 provided in the aforementioned embodiment.
[0213] For example, step S410 may be applied to the encoder 230 provided in the aforementioned embodiment.
[0214] For example, step S420 may be applied to the interleaver 240 provided in the aforementioned embodiment.
[0215] For example, step S430 may be applied to the decoder 250 provided in the aforementioned embodiment.
[0216] For example, the specific operations in step S400 may also refer to the relevant description of step S300 in the aforementioned embodiment, which will not be repeated here.
[0217] For example, the specific operations involved in steps S410 to S430 may also refer to the relevant descriptions of steps S330 to S350 in the aforementioned embodiment, which will not be repeated here.
[0218] In some embodiments of the present disclosure, for example, step S410 may further include steps S411 to S412.
[0219] Step S411: performing a mathematical operation on the first interleaved bit information and the first bit position information to obtain intermediate interleaved bit information, wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved bit corresponding to the first interleaved bit information.
[0220] Step S412: retain the first interleaved bit information via a sideband signal to provide it for subsequent decoding operations.
[0221] For example, steps S411 to S412 may be applied to the encoder 230 provided in the aforementioned embodiment.
[0222] For example, the specific operations involved in steps S411 to S412 can also refer to the relevant descriptions of steps S331 to S332 in the aforementioned embodiment, which will not be repeated here.
[0223] In some embodiments of the present disclosure, step S410 may further include step S413.
[0224] Step S413: in response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment, encoding the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information.
[0225] For example, step S413 may be applied to the encoder 230 provided in the aforementioned embodiment.
[0226] In some embodiments of the present disclosure, for example, step S430 may further include step S431.
[0227] Step S431: In response to the first bit information being valid and the first interleaved bit information being invalid, the output address information is determined based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit information and the first interleaved bit information being valid, the output address information is determined based on the read / write memory access request information and the first interleaved granularity; or, in response to the first bit information being invalid, the output address information is determined based on the read / write memory access request information.
[0228] For example, step S431 may be applied to the decoder 250 provided in the aforementioned embodiment.
[0229] For example, the specific operations involved in step S431 can refer to the relevant description of step S351 in the above embodiment, which will not be repeated here.
[0230] In some embodiments of the present disclosure, the method for managing cache, the encoding operation involved in step S410 may further include: setting the second bit information of the read and write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retaining the initial value of the second bit information through a sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and constraining the burst data size of the read and write memory access request information not to exceed one half of the address boundary specified by the transmission protocol.
[0231] The decoding operation involved in step S430 may further include: restoring the initial value of the second bit information.
[0232] The further encoding operation in step S410 and the further decoding operation in step S430 involved in this embodiment may refer to the relevant descriptions of the encoding operation and the decoding operation in the aforementioned embodiments, and will not be described in detail here.
[0233] It should be noted that, for the specific operations, functions or beneficial effects of each step in the method for managing cache provided in any embodiment of the present disclosure, for example, reference can be made to the relevant description of the device for managing cache provided in any embodiment of the present disclosure, and will not be repeated here.
[0234] Figure 5 A schematic diagram of an application of a cache management device provided by at least one embodiment of the present disclosure is shown.
[0235] like Figure 5 As shown, Figure 5 The address mapping and interleaving control design of dual LLC physical IP in SOC architecture is demonstrated to realize dynamic switching of LLC resources between cache mode and SRAM mode.
[0236] For example, the dual LLC entity is divided into LLC0 and LLC1, including two physical LLC modules, supporting dynamic configuration of functions. Here, LLC0 and LLC1 may correspond to the second last level cache (LLC 0) and the first last level cache (LLC 1) described in the context, for example.
[0237] For example, the address mapping segment includes a cache function address segment (such as a cache address segment) and a static random access memory direct access address segment (such as an SRAM address segment).
[0238] For example, in the dynamic mode when the interleaving enable signal is valid (interlv_en=1), when the read and write memory access requests belong to the SRAM address segment, the data blocks corresponding to the read and write memory access requests are alternately written into LLC0 and LLC1 with 64B granularity interleaving (to adapt to the large bandwidth requirements of SRAM direct access), and the dual LLC can be used as SRAM.
[0239] For example, in the dynamic mode when the interleave enable signal is valid (interlv_en=1), when the read and write memory access requests belong to the cache address segment, the data blocks corresponding to the read and write memory access requests are alternately written into LLC0 and LLC1 with 128B granularity interleaving (adapting cache line prefetch and spatial locality), and the dual LLC can be used as a Last Level Cache.
[0240] For example, in static mode when the interleaving enable signal is invalid (interlv_en=0), LLC0 is fixed to SRAM mode (flat address mapping), serving low-latency, deterministic access scenarios (such as sensor data processing), and LLC0 is used as an SRAM with a continuous address (flatten). LLC1 is fixed to Cache mode (dynamic management), serving high-performance computing modules (such as CPU / GPU), and LLC1 is used as a Last Level Cache with a continuous address (flatten).
[0241] For example, Figure 5 The others in the address space represent other address segments in the address space.
[0242] Figure 5 For the relevant contents of the operation methods, functions or beneficial effects of the embodiments involved, for example, reference may be made to the relevant description of the embodiments of the device for managing cache provided in any embodiment of the present disclosure, and will not be repeated here.
[0243] Figure 6 A schematic diagram of another application of a cache management device provided by at least one embodiment of the present disclosure is shown.
[0244] Figure 6The improved design of address encoding and decoding logic in NoC (network on chip) is shown. Through dynamic XOR encoding and decoding and sideband signal mechanism, the low LLC cache utilization problem caused by fixed interleaving bits in the original architecture is solved, while ensuring the address boundary constraints of AXI protocol for burst transmission.
[0245] For example, the original architecture problem (see Figure 6 (Left side figure) Fixed interleaving bits lead to Set waste.
[0246] Exemplarily, the NoC selects the target LLC (e.g., Bit12=0→LLC0, Bit12=1→LLC1) based on fixed address bits (e.g., Bit12-11). These interleaved bits overlap with the set index bits (e.g., Bit9-6) of the LLC, causing some Sets to be inaccessible due to the fixed address bits (e.g., if the interleaved bits are fixed to 0, the high bit of the Set index of LLC0 is always 0, and only 50% of the Sets can be used).
[0247] For example, the interleaving bit Bit12 of address 0x8000_1000 is 0, and it is fixedly routed to LLC0, and its Set index bit Bit9 is 0, which results in Sets 0 to 127 of LLC0 being available, and Sets 128 to 255 being forever inaccessible.
[0248] For example, the improved architecture solution proposed in the embodiment of the present disclosure (see Figure 6 ). For example, in the encoding stage (input), the interleaved bit of the original address (such as Bit 7) is XORed with a slightly higher bit (such as Bit 17) other than the set index bit to generate intermediate interleaved bit information (addr_out[7] = addr_in[7]^addr_in
[17] ). The intermediate interleaved bit changes dynamically, breaking the fixed binding with the set index.
[0249] For example, the generated intermediate interleaving bit information may be used to select a target LLC. Here, the target LLC may correspond to the target final level cache described above, that is, the target final level cache selected based on the intermediate interleaving bit information.
[0250] For example, according to the generated intermediate interleaving bit information, the input address may be distributed to the target LLC selected based on the intermediate interleaving bit information at a corresponding interleaving granularity (eg, 128B).
[0251] For example, the original information is retained through the sideband signal, and the original value of Bit 7 and the value of Bit 11 are transmitted to the decoding module through the sideband signal (tag_2k=addr_in
[11] ; tag_bit7=addr_in[7]).
[0252] For example, Bit 11 of the encoded address is forced to be set to zero (addr_out
[11] =1'b0) to prevent burst transmission from crossing the 4K boundary.
[0253] Optionally, for example, for the decoding stage (output), when the amount of each burst of read and write accesses is aligned with the interleaving granularity, the intermediate interleaved bits can be XORed inversely to restore the original interleaved bits, and Bit 11 (addr_out
[11] =tag2k) can be restored in combination with the sideband signal to generate the output address information for LLC access.
[0254] Optionally, for example, for the decoding stage (output), when the amount of each burst of read and write accesses is not aligned with the interleaving granularity, it is possible to determine whether the value of Bit 17 (i.e., addr_in
[17] ) is valid. If Bit 17 is valid, it is further determined whether the value of the original Bit 7 (i.e., tag_bit7, which is retained by the sideband signal) is valid. If the value of Bit 7 is valid, the output address information is equal to the input address information plus the corresponding interleaving granularity (addr_out = addr_in + granu). If the value of Bit 7 is invalid, the output address information is equal to the input address information minus the corresponding interleaving granularity (addr_out = addr_in - granu). If the value of Bit 17 is invalid, the output address information is equal to the input address (addr_out = addr_in)
[0255] For example, for burst transmission constraints, in order to limit Burst Size ≤ 2K, Bit 12 is not flipped after the address is incremented (such as the first address Bit 11 = 0, the 2K data block address range is 0x0000 to 0x07FF, and Bit 11 is always 0).
[0256] Figure 6 For the relevant contents of the operation methods, functions or beneficial effects of the embodiments involved, for example, reference may be made to the relevant description of the embodiments of the device for managing cache provided in any embodiment of the present disclosure, and will not be repeated here.
[0257] Figure 7 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown.
[0258] At least some embodiments of the present disclosure also provide an electronic device, such as Figure 7 As shown, the electronic device 500 includes a cache management device provided by any embodiment of the present disclosure. For example, the cache management device 508 can be implemented as a sender and / or receiver of instructions / data, for example, it can be used for internal memory or external memory.
[0259] For example, processor 501 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units having data processing capabilities and / or program execution capabilities; for example, the central processing unit (CPU) may be a RISC, X86, or ARM architecture, etc.
[0260] The electronic device 500 in the embodiment of the present disclosure may include mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., as well as any devices such as digital TVs, desktop computers, servers, etc., and may also be a combination of any data processing devices and hardware, which is not limited by the embodiments of the present disclosure.
[0261] Figure 7 The electronic device 500 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0262] For example, Figure 7 As shown, in some examples, the processor 501 can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device (not shown) into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the computer system are also stored. Multiple processors 501, ROM 502 and RAM 503 are connected to each other through a communication channel 504. An input / output (I / O) interface 505 is also connected to the communication channel 504.
[0263] For example, the following components can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device including, for example, a magnetic tape, a hard disk, etc., for example, the storage controller of the storage device includes the above-mentioned device for managing the cache; for example, a communication device 509 including a network interface card such as a LAN card, a modem, etc. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data, and perform communication processing via a network such as the Internet. The drive 510 is also connected to the I / O interface 505 as needed. Removable media 511, such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive 510 as needed, so that the computer program read therefrom can be installed into the storage device as needed. Although Figure 7The electronic device 500 is shown to include various devices, but it should be understood that it is not required to implement or include all the devices shown. More or fewer devices may be implemented or included instead.
[0264] For example, the electronic device 500 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a lightning interface, etc. The communication device 509 may communicate with a network and other devices through wireless communication, such as the Internet, an intranet and / or a wireless network such as a cellular phone network, a wireless local area network (LAN) and / or a metropolitan area network (MAN). Wireless communication may use any of a variety of communication standards, protocols, and techniques, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email, instant messaging and / or Short Message Service (SMS), or any other suitable communication protocol.
[0265] It should be noted that in the embodiments of the present disclosure, the specific functions and technical effects of the electronic device 500 can be referred to, for example, the relevant description of the apparatus and method for managing cache in the embodiments of the present disclosure, and will not be repeated here.
[0266] Figure 8 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown.
[0267] At least one embodiment of the present disclosure further provides an electronic device, such as Figure 8 As shown, the electronic device 600 includes at least one processor 610 and at least one memory 620 .
[0268] For example, the memory 620 may be used to store computer-readable instructions (e.g., one or more computer program modules) in a non-temporary manner. The processor 610 may be used to execute the computer-readable instructions, and when the computer-readable instructions are executed by the processor 610, one or more steps in the method for managing cache described above may be executed. The memory 620 and the processor 610 may be interconnected through a bus or link, in a wired, wireless, and / or other form of communication medium, etc., and the embodiments of the present disclosure are not limited to this.
[0269] For example, the processor 610 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units having data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) may be a RISC, X86, or ARM architecture, etc. The processor 610 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 600 to perform desired functions.
[0270] For example, the memory 620 may include any combination of one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disk read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 610 may run one or more computer program modules to implement various functions of the electronic device 600. Various applications and various data, as well as various data used and / or generated by the application, etc. may also be stored in the computer-readable storage medium.
[0271] At least one embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which is used to non-transitorily store computer-readable instructions. When the computer-readable instructions are executed by a computer, the above-mentioned method for managing cache can be implemented.
[0272] Fig. 9 A schematic diagram of a computer-readable storage medium provided for some embodiments of the present disclosure.
[0273] like Fig. 9 As shown, the computer-readable storage medium 700 is used to store computer-readable instructions 710. For example, when the computer-readable instructions 710 are executed by a computer, one or more steps in the method for managing cache described above can be performed.
[0274] For example, the computer-readable storage medium 700 may be applied to the electronic device 500 or the electronic device 600. For example, the relevant description of the non-volatile computer-readable storage medium 700 may also refer to Figure 7 The storage device in the electronic device 500 shown in FIG. Figure 8 The corresponding description of the memory 620 in the electronic device 600 is shown and will not be repeated here.
[0275] It should be noted that, in the embodiments of the present disclosure, the specific functions and technical effects of the computer-readable storage medium 700 can be referred to the above description of the method for managing cache and the device for managing cache, and will not be repeated here.
[0276] At least one embodiment of the present disclosure provides a system on chip, which includes: multiple processor cores and an interconnect bus, wherein the interconnect bus is configured to route access requests of the multiple processor cores to the cache management device provided by any embodiment of the present disclosure.
[0277] This embodiment provides a system on chip, which is characterized by integrating multiple processor cores and interconnection buses. The SOC architecture can optimize data transmission paths, improve computing performance and low-latency data access.
[0278] For example, the SOC includes multiple processor cores, each of which can independently execute tasks and also supports multi-core collaborative working mode, which is suitable for application scenarios with complex computing requirements.
[0279] For example, the interconnect bus can be used to connect each processor core and other components, such as the device for managing cache provided in any embodiment of the present disclosure. The interconnect bus can be responsible for routing access requests from the processor core to the appropriate resource location, such as internal cache or external memory.
[0280] There are a few points to note:
[0281] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.
[0282] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to obtain new embodiments.
[0283] The above description is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A device for managing a cache, comprising: The source interface module is configured to receive the read and write memory access request information of the host; a register configured to determine whether an interleaving enable control signal acting on a system address mapper is valid; The system address mapper is configured to determine whether the read / write memory access request information belongs to the first address segment or the second address segment; as well as an interleaver configured to distribute the read / write memory access request information to a plurality of corresponding last-level caches at a first interleaving granularity in response to an interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment, and In response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the second address segment, distributing the read / write memory access request information to a plurality of corresponding last-level caches at a second interleaving granularity; The second address segment is an address segment in the address space where the cache is located that is different from the first address segment.
2. The device for managing cache according to claim 1, wherein: The first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
3. The device for managing cache according to claim 1, wherein: The first interleaving granularity is configured based on a size of a cache line, and the second interleaving granularity is configured based on a data bit width of a cache interface.
4. The device for managing cache according to claim 1, wherein: The interleaver is further configured to: In response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the first address segment, sending the read / write memory access request information to a first last-level cache; as well as, In response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the second address segment, sending the read / write memory access request information to the second last-level cache, The first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
5. A device for managing a cache, comprising: The source interface module is configured to receive read and write memory access request information from the host; A system address mapper, configured to determine whether the read / write memory access request information belongs to the first address segment or the second address segment; an encoder configured to, in response to the read / write memory access request information belonging to the range of the first address segment, perform an encoding operation on the first interleaved bit information of the read / write memory access request information to obtain intermediate interleaved bit information; as well as The interleaver is configured to distribute the read and write memory access request information to the target last-level cache at the first interleaving granularity based on the intermediate interleaving bit information.
6. The device for managing cache according to claim 5, further comprising: Decoder, The decoder is configured to perform a decoding operation based on the first interleaved bit information to obtain output address information, and provide the output address information to the target last-level cache.
7. The device for managing cache according to claim 5, further comprising: register; The register is configured to determine whether an interleaving enable control signal acting on the system address mapper is valid; The encoder is further configured to, in response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment, perform an encoding operation on the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information.
8. The device for managing cache according to claim 5, wherein: The encoder is further configured to perform a mathematical operation on the first interleaved bit information and the first bit position information to obtain the intermediate interleaved bit information, wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved bit corresponding to the first interleaved bit information, and The first interleaved bit information is retained through a sideband signal to be provided for subsequent decoding operations.
9. The device for managing cache according to claim 6, wherein: The decoder is further configured to, in response to the first bit information being valid and the first interleaving bit information being invalid, determine the output address information based on the read / write memory access request information and the first interleaving granularity; or, In response to the first bit information and the first interleaving bit information being valid, determining the output address information based on the read / write memory access request information and the first interleaving granularity; or, In response to the first bit information being invalid, determining the output address information based on the read / write memory access request information; The first bit corresponding to the first bit information is a higher bit of the interleaved bit corresponding to the first interleaved bit information.
10. The device for managing cache according to claim 5, wherein: The encoder is further configured to set the second bit information of the read / write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retain the initial value of the second bit information through a sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and The burst data size of the read / write memory access request information is constrained not to exceed one half of the address boundary specified by the transmission protocol.
11. The device for managing cache according to claim 6, wherein: The decoder is further configured to restore an initial value of the second bit information.
12. A method for managing a cache, comprising: Receive read and write memory access request information from the host; In response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the first address segment, distributing the read / write memory access request information to a plurality of corresponding last-level caches at a first interleaving granularity; as well as In response to the interleaving enable signal being valid and the read / write memory access request information being within the range of the second address segment, the read / write memory access request information is distributed to a plurality of corresponding last-level caches at a second interleaving granularity, The second address segment is an address segment in the address space where the cache is located that is different from the first address segment.
13. The method for managing cache according to claim 12, wherein: The first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
14. The method for managing cache according to claim 12, further comprising: In response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the first address segment, sending the read / write memory access request information to a first last-level cache; as well as, In response to the interleaving enable signal being invalid and the read / write memory access request information being within the range of the second address segment, sending the read / write memory access request information to the second last-level cache, The first address segment is a cache function address segment in the address space where the cache is located, and the second address segment is a static random access memory direct access address segment in the address space where the cache is located.
15. A method for managing a cache, comprising: Receive read and write memory access request information from the host, In response to the read / write memory access request information belonging to the range of the first address segment, encoding the first interleaving bit information of the read / write memory access request information to obtain intermediate interleaving bit information; Distributing the read and write memory access request information to a target last-level cache at the first interleaving granularity based on the intermediate interleaving bit information; A decoding operation is performed based on the first interleaved bit information to obtain output address information, and the output address information is provided to the target last-level cache.
16. The method for managing cache according to claim 15, wherein: The encoding operation is performed on the first interleaved bit information of the read / write memory access request information to obtain the intermediate interleaved bit information, including: Performing a mathematical operation on the first interleaved position information and the first bit position information to obtain the intermediate interleaved position information, wherein the first bit position corresponding to the first bit position information is a higher bit of the interleaved position corresponding to the first interleaved position information; The first interleaved bit information is retained through a sideband signal to be provided for the subsequent decoding operation.
17. The method for managing cache according to claim 15, wherein: The performing a decoding operation based on the first interleaved bit information to obtain output address information includes: In response to the first bit information being valid and the first interleaving bit information being invalid, determining the output address information based on the read / write memory access request information and the first interleaving granularity; or, In response to the first bit information and the first interleaving bit information being valid, the output address information is determined based on the read / write memory access request information and the first interleaving granularity; or, In response to the first bit information being invalid, determining the output address information based on the read / write memory access request information; The first bit corresponding to the first bit information is a higher bit of the interleaved bit corresponding to the first interleaved bit information.
18. The method for managing cache according to claim 15, wherein: The encoding operation also includes: Setting the second bit information of the read / write memory access request information to zero so that the encoded incremental burst address does not cross the address boundary specified by the transmission protocol, and retaining the initial value of the second bit information through a sideband signal, wherein the second bit information is determined based on the address boundary specified by the transmission protocol, and Constraining the burst data size of the read / write memory access request information to not exceed one half of the address boundary specified by the transmission protocol; and The decoding operation further includes: Restore the initial value of the second bit information.
19. An electronic device, comprising the cache management device according to any one of claims 1-11.
20. An electronic device, comprising: processor; as well as A memory, wherein the memory stores at least one computer program, and when the at least one computer program is executed by the processor, the method for managing cache as described in any one of claims 12 to 18 is implemented.
21. A non-transitory computer-readable storage medium, used for non-temporarily storing computer-readable instructions, and when the computer-readable instructions are executed by a computer, the method for managing cache as described in any one of claims 12-18 is implemented.
Citation Information
Cited By
Cache management apparatus and method, electronic device, and storage medium
WO2026184013A1