Cache, control chip, debugging method and upper computer
By designing a cache structure that maps cache blocks to data blocks one by one in the MCU system, the problems of excessive power consumption and difficult timing convergence caused by high main frequency operation are solved, a balance between transmission performance and power consumption is achieved, and cache utilization and system performance are improved.
Patent Information
- Application Number
- CN202510571748.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-26
AI Technical Summary
In MCU systems, high main frequency operation leads to excessive power consumption and difficulty in timing convergence. Especially when the bus structure runs at the same frequency as the CPU, traditional cache design has problems such as low hit rate, high hardware complexity and unstable operation.
A cache structure is designed in which cache blocks are mapped one-to-one with data blocks. The cache blocks are divided into multiple sub-blocks, which run at different frequencies from the bus and memory. The mapping between cache blocks and data blocks achieves a balance between transmission performance and power consumption, and improves cache utilization.
Through the mapping design of cache blocks and data blocks, a balance between transmission performance and power consumption between the central processing unit and the bus is achieved, which improves cache utilization, reduces power consumption and enhances the overall performance of the system.
Smart Images

Figure CN120705088A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed in this application relate to the field of chip technology, and more specifically, to a cache, a control chip, a debugging method, and a host computer. Background Art
[0002] MCUs (Microcontroller Units), the core of embedded systems, are widely used in scenarios such as smart device control and real-time data processing. With technological advancements and rising market demands, MCU clock speeds are increasing, even evolving towards SoCs (System on Chips). However, requiring complex bus structures to operate at the same frequency as the CPU (Central Processing Unit) can lead to challenges such as excessive power consumption and difficulty closing timing. Summary of the Invention
[0003] According to the embodiments of the present application, the present application proposes a cache, a control chip, a debugging method and a host computer to solve the above problems.
[0004] The first aspect of the present application discloses a cache, which is connected between a central processing unit and a bus, the bus is also connected to a memory, the central processing unit operates at a first frequency, the bus and the memory operate at a second frequency, and the first frequency is greater than the second frequency; the cache includes at least one cache block, any one of which is used to store any one of at least one data block in the memory, the cache block has the same size as the data block, and each cache block is divided into at least two cache sub-blocks, and at least two cache sub-blocks in the cache block are mapped one-to-one with at least two data sub-blocks in the data block.
[0005] In some embodiments, the cache further includes at least one cache control unit; each cache control unit is connected to a corresponding cache block for temporarily storing a base address of the corresponding cache block, wherein the base address of the at least one cache block is different.
[0006] In some embodiments, the size of the cache block is related to the bit width of the bus, wherein the bit width of the bus is n bits, the size of the cache block is m bytes, and m is a positive integer multiple of n / 8.
[0007] The second aspect of the present application discloses a control chip, comprising: a memory; a bus connected to the memory; a central processing unit connected to the bus through a cache; wherein the memory includes at least one data block, and the cache includes at least one cache block, wherein any one of the cache blocks is used to store any one data block, and the data block has the same size as the cache block, wherein each of the cache blocks is divided into at least two cache sub-blocks, and at least two cache sub-blocks in the cache block are mapped one-to-one to at least two data sub-blocks in the data block; the central processing unit is used to: obtain a target access address in response to an operation instruction; compare the target access address with the address of the at least one cache block to perform a corresponding data operation through the cache.
[0008] In some embodiments, the central processing unit is used to: obtain a target access address in response to a read operation instruction; compare the target access address with the address of the at least one cache block; in response to the target access address being the same as the address of one of the at least one cache block and the data in the cache block being not empty, read and return the target data corresponding to the same address in the cache block; in response to the target access address being different from the address of the at least one cache block, obtain the target data corresponding to the target access address returned from the memory to the cache via the bus.
[0009] In some embodiments, the cache is used to: in response to the target access address being different from the address of the at least one cache block, obtain target data corresponding to the target access address from the memory and return the target data to the central processing unit; obtain other data that is in the same loop cycle as the target data from the memory; and write the target data and the other data into one of the at least one cache block in a first-in-first-out manner.
[0010] In some embodiments, the central processing unit is used to: in response to a write operation instruction, obtain a target access address, and write the data corresponding to the write operation instruction into the memory; wherein, in response to the target access address being the same as the address of one of the at least one cache block and the data in the cache block being not empty, the data corresponding to the write operation instruction is also written into the cache block; in response to the target access address being different from the address of the at least one cache block, not performing a write operation on the cache; or, the central processing unit is further used to: in response to the data in the data block being overwritten, clear the data in the cache block mapped by the data block using the clear cache enable port of the cache.
[0011] The third aspect of the present application discloses a debugging method for a host computer, which is connected to the control chip as described in the second aspect. The method includes: in response to a control instruction, the host computer controls the memory and the cache to enter a debugging mode; in response to the debugging instruction, obtains a target address; in response to the target data obtained according to the target address not being a specific value, extracts the target data from the cache for debugging.
[0012] In some embodiments, the cache includes a first cache area and a second cache area, the debug instruction includes a debug enable signal, and the method further includes: in response to the debug enable signal being a first preset value, extracting the target data from the first cache area; in response to the debug enable signal being a second preset value, extracting the target data from the second cache area.
[0013] The fourth aspect of the present application discloses a host computer, comprising a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the debugging method described in the third aspect.
[0014] The beneficial effects of the present application are as follows: the cache is connected between the central processing unit and the bus, the bus is also connected to the memory, the central processing unit operates at a first frequency, the bus and the memory operate at a second frequency, the first frequency is greater than the second frequency, the cache includes at least one cache block, wherein any one cache block is used to store any one of at least one data block in the memory, the cache block and the data block have the same size, and each cache block is divided into at least two cache sub-blocks, at least two cache sub-blocks in the cache block are mapped one-to-one with at least two data sub-blocks in the data block, through arbitrary mapping between cache blocks and data blocks and one-to-one mapping of at least two cache sub-blocks in the cache block with at least two data sub-blocks in the data block, the transmission performance and power consumption between the central processing unit and the bus are balanced, and the utilization of the cache is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The present application will be further described below with reference to the accompanying drawings and implementation methods, in which:
[0016] Figure 1 This is a schematic diagram of a cache group-connected mapping in the related art;
[0017] Figure 2 This is a schematic diagram of the structure of the cache in an embodiment of the present application;
[0018] Figure 3 is a mapping relationship diagram between the cache and the memory in an embodiment of the present application;
[0019] Figure 4 This is a schematic diagram of the internal architecture of a cache according to an embodiment of the present application;
[0020] Figure 5 Schematic diagram of the structure of the control chip of the embodiment of the present application;
[0021] Figure 6 This is a schematic diagram of the prefetch effect of an embodiment of the present application;
[0022] Figure 7 1 is a flowchart of a read operation according to an embodiment of the present application;
[0023] Figure 8 1 is a flowchart of a write operation according to an embodiment of the present application;
[0024] Figure 9 Schematic diagram of the debugging method of the embodiment of the present application;
[0025] Figure 10 This is a flowchart of a debugging method according to an embodiment of the present application;
[0026] Figure 11 It is a structural diagram of the host computer of an embodiment of the present application. DETAILED DESCRIPTION
[0027] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0028] The term "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. In addition, the terms "first", "second", and "third" in this application are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.
[0029] Traditional caches mainly include direct associative, fully associative and set associative design methods. Among them, direct associative means that any address in the main memory can be mapped to any address in the cache. The hardware design is simple, but access conflicts are frequent and the hit rate is low. Fully associative means that main memory addresses at fixed intervals are mapped to the same cache address. Although the hit rate is high, the hardware is complex. Set associative is between fully associative and direct associative, such as Figure 1 As shown, Figure 1 This is a mapping diagram of cache group connection in the related technology. The main memory of the same color is mapped to the cache of the same color. The main memory of the same color and the cache of the same color can be mapped arbitrarily. However, the group connection requires at least two SRAMs (Static Random-Access Memory), one for storing TAG information and the other for storing the data itself. The SRAM needs to be initialized when performing related operations. The startup time is long, and there are still problems such as instability and complex operation.
[0030] To this end, the present application proposes a cache, a control chip, a debugging method and a host computer to solve the above problems.
[0031] In order to enable those skilled in the art to better understand the technical solution of the present application, the technical solution of the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0032] See also Figure 2-Figure 3 , Figure 2 is a schematic diagram of the structure of the cache in an embodiment of the present application, Figure 3 : is a mapping diagram of the cache and memory in the embodiment of the present application. Figure 2 As shown, cache 100 is connected between central processing unit 200 and bus 300, and bus 300 is also connected to memory 400. Central processing unit 200 operates at a first frequency, and bus 300 and memory 400 operate at a second frequency, where the first frequency is greater than the second frequency. In some examples, cache 100 can be a cache, central processing unit 200 is a CPU, bus 300 can be a system bus, and memory 400 can be flash memory or static random access memory (SRAM). For example, central processing unit 200 can operate at 200 MHz, and bus 300 and memory 400 can operate at 100 MHz. Cache 100 can be used to cache data required by central processing unit 200, thereby reducing power consumption while maintaining performance.
[0033] Cache 100 includes at least one cache block 110, and memory 400 includes at least one data block 410. Any cache block 110 is used to store any one of the at least one data blocks 410. Cache block 110 and data block 410 are of the same size, and the data within data block 410 is mapped correspondingly to the data within cache block 110. Each data block 410 is divided into at least two data sub-blocks 401, and each cache block 110 is divided into at least two cache sub-blocks 101. The at least two cache sub-blocks 101 are mapped one-to-one with at least two data sub-blocks 401. The number of data sub-blocks 401 in a data block 410 is equal to the number of cache sub-blocks 101 in a cache block 110. In other words, any data block 410 can be stored in any cache block 110, and the mapping between data sub-blocks 401 in a data block 410 and cache sub-blocks 110 in a cache block 110 is fixed.
[0034] In some examples, such as Figure 3 As shown, at least one cache block 110 includes cache block 110a, cache block 110b..., and at least one data block 410 includes cache block 410a...cache block 410n. For example, data block 410n includes four data sub-blocks, namely data sub-block 4011, data sub-block 4012, data sub-block 4013 and data sub-block 4014, and the data sub-blocks are used to store different data; cache block 110a includes four cache sub-blocks, namely cache sub-block 1011, cache sub-block 1012, cache sub-block 1013 and cache sub-block 1014; cache block 110b includes four cache sub-blocks, namely cache sub-block 1015, cache sub-block 1016, cache sub-block 1017 and cache sub-block 1018.
[0035] Data block 410n can be mapped to any cache block 110 in cache 100, and the data sub-block 401 in data block 410 is mapped one-to-one with the cache sub-block 110 in cache block 110. For example, the data of data block 410n can be stored in cache block 110a or cache block 110b at will. If data block 410n is mapped to cache block 110a, then data sub-block 4011, data sub-block 4012, data sub-block 4013 and data sub-block 4014 are mapped one-to-one with cache sub-block 1011, cache sub-block 1012, cache sub-block 1013 and cache sub-block 1014. Alternatively, if data block 410n is mapped to cache block 110b, data sub-blocks 4011, 4012, 4013, and 4014 are mapped one-to-one with cache sub-blocks 1015, 1016, 1017, and 1018.
[0036] In this embodiment, the cache 100 is connected between the central processing unit 200 and the bus 300, and the bus 300 is also connected to the memory 400. The central processing unit 200 operates at a first frequency, and the bus 300 and the memory 400 operate at a second frequency. The first frequency is greater than the second frequency. The cache 100 includes at least one cache block 110, wherein any one cache block 110 is used to store any one data block 410 of at least one data block 410 in the memory 400, and the size of the cache block 110 is the same as that of the data block 410. In the same manner, each cache block 110 is divided into at least two cache sub-blocks 101, and at least two cache sub-blocks 101 in the cache block 110 are mapped one-to-one with at least two data sub-blocks 401 in the data block 410. By arbitrarily mapping between the cache block 110 and the data block 410 and mapping at least two cache sub-blocks 101 in the cache block 110 one-to-one with at least two data sub-blocks 401 in the data block 410, the transmission performance and power consumption between the central processing unit 200 and the bus 300 are balanced, and the utilization rate of the cache 100 is improved.
[0037] In some embodiments, as Figure 2 As shown, the cache 100 further includes at least one cache control unit 120. Each cache control unit 120 is connected to a corresponding cache block 110 and is used to temporarily store the base address of the corresponding cache block 100, wherein the base address of at least one cache block 110 is different.
[0038] For ease of understanding, the internal architecture of the cache 100 is illustrated. In some examples, the size of the memory 400 is 4MB, the size of the cache 100 is 2KB, and the address width of the bus 300 is 32 bits. Figure 4 As shown, Figure 4 This is a schematic diagram of the internal architecture of a cache in an embodiment of the present application, wherein the size of a cache block 110 is 32 bytes. The 2KB cache 100 can be divided into 64 cache blocks 110, namely cache block 0...cache block 63. Each cache block 110 can be divided into 8 cache sub-blocks for temporarily storing data corresponding to 8 addresses. The data size corresponding to each address is 32 bits. And, as Figure 4 As shown, the 64 cache blocks may correspond to 64 cache control units (cache_ctrl_unit), namely control unit 0 . . . control unit 63. The 64 cache control units may be used to temporarily store the base addresses (baseAddress) of the corresponding 64 cache blocks.
[0039] Further, if Figure 4As shown, the total capacity of memory 400 is 4MB. The main memory physical address range is 22 bits ([21:0]), and the bus 300 is 32 bits wide ([31:0]). During transmission, the actual effective address only requires 22 bits, so the upper address bits [31:22] are the same, for example, they can be fixed to 0. That is, the upper address bits [31:22] do not need to be temporarily stored in the cache control unit 120. The size of cache block 110 is 32 bytes, and the internal offset (Offset) occupies 5 bits ([4:0]). When accessing cache 100, it can be directly located by the lower 5 bits of the address, without the need for temporary storage in the cache control unit 120. There are 64 cache blocks 100, and the corresponding group index (Index) occupies 6 bits ([10:5]). The corresponding controller can be selected directly by address decoding, and no temporary storage in the cache control unit 120 is required. That is, the cache control unit 120 only needs to temporarily store 11 bits of data corresponding to the address [21:11], thereby saving hardware costs.
[0040] In some embodiments, the size of the cache block 110 is related to the bit width of the bus 300 , where the bit width of the bus 300 is x bits, the size of the cache block 410 is y bytes, and y is a positive integer multiple of x / 8.
[0041] For example, x=32, y=4n; or x=64, y=8n, where n is a positive integer. In some examples, the bit width of bus 300 is 32 bits, and the size of cache block 110 can be 4*n bytes; or the bit width of bus 300 is 64 bits, and the size of cache block 110 can be 8*n bytes.
[0042] In some examples, the size of memory 400 is 8MB, the size of cache 100 is 4KB, and the address width of bus 300 is 64 bits. If n = 8, the size of cache block 110 can be 64 bytes, that is, cache 100 can be divided into 64 cache blocks 110, where each cache block 110 can include data corresponding to 16 addresses, and the data corresponding to each address is 32 bits in size. The main memory physical address range is 23 bits ([22:0]), and the bus 300 is 64 bits wide ([63:0]). During transmission, the actual effective address only requires 23 bits, so the high-order address bits [63:23] are the same, for example, they can be fixed to 0, that is, the high-order address bits [63:23] do not need to be temporarily stored in cache control unit 120. The size of cache block 110 is 64 bytes, and the internal offset (Offset) occupies 6 bits ([5:0]). When accessing the cache, the low-order 6 bits of the address can be directly used to locate the address, without the need for temporary storage in cache control unit 120. The 64 cache blocks 100 have a corresponding group index (Index) of 6 bits ([11:6]). The corresponding controller can be directly selected by address decoding without requiring temporary storage in the cache control unit 120. In other words, the cache control unit 120 only needs to temporarily store the data corresponding to the 11 bits of address [22:12].
[0043] Alternatively, in other examples, the size of memory 400 is 4MB, the size of cache 100 is 1KB, and the address width of bus 300 is 64 bits. If n=8, the size of cache block 110 can be 64 bytes, that is, cache 100 can be divided into 16 cache blocks 110, where each cache block 110 can include data corresponding to 8 addresses, and the data size corresponding to each address is 64 bits. The main memory physical address range is 22 bits ([21:0]), and the bus 300 bit width is 64 bits ([63:0]). During transmission, the actual effective address only requires 22 bits, so the high-order address bits [63:22] are the same, for example, they can be fixed to 0, that is, the high-order address bits [63:22] do not need to be temporarily stored in cache control unit 120. The size of cache block 110 is 64 bytes, and the internal offset (Offset) occupies 6 bits ([5:0]). When accessing the cache, the low-order 6 bits of the address can be directly used to locate the address, without the need for temporary storage in cache control unit 120. The 16 cache blocks 100 have a corresponding group index of 4 bits ([9:6]). The corresponding controller can be directly selected by address decoding without requiring temporary storage in the cache control unit 120. In other words, the cache control unit 120 only needs to temporarily store the data corresponding to the 12 bits of address [21:10].
[0044] See also Figure 5 , Figure 51 is a schematic diagram of the structure of a control chip according to an embodiment of the present application. The control chip 1000 may include the aforementioned cache 100, a central processing unit 200, a bus 300, and a memory 400. The bus 300 is connected to the memory 400, and the central processing unit 200 is connected to the bus 300 through the cache 100. In some examples, the central processing unit 200 may operate at a first frequency, and the bus 300 and the memory 400 may operate at a second frequency, where the first frequency is greater than the second frequency.
[0045] Cache 100 includes at least one cache block 110, and memory 400 includes at least one data block 410. Any cache block 110 is used to store any one of the at least one data blocks 410. Cache block 110 and data block 410 are of the same size, and the data within data block 410 is mapped correspondingly to the data within cache block 110. Each data block 410 is divided into at least two data sub-blocks 401, and each cache block 110 is divided into at least two cache sub-blocks 101. The at least two cache sub-blocks 101 are mapped one-to-one with at least two data sub-blocks 401. The number of data sub-blocks 401 in a data block 410 is equal to the number of cache sub-blocks 101 in a cache block 110. In other words, any data block 410 can be stored in any cache block 110, and the mapping between data sub-blocks 401 in a data block 410 and cache sub-blocks 110 in a cache block 110 is fixed.
[0046] The central processing unit 200 is configured to obtain a target access address in response to an operation instruction, compare the target access address with the address of at least one cache block 110, and perform a corresponding data operation through the cache 100. For example, the central processing unit 200 obtains the target access address in response to the operation instruction, accesses the cache 100, compares the target access address with the addresses stored in at least one cache block 110, and performs a corresponding data operation through the cache 100 based on the comparison result, wherein the data operation includes a read operation, a write operation, a prefetch operation, an update operation, a clear operation, etc.
[0047] In this embodiment, in response to an operation instruction, a target access address is obtained, and the target access address is compared with the address of at least one cache block 110 to perform a corresponding data operation through the cache 100, wherein any cache block 110 is used to store any data block 410 of at least one data block 410 in the memory 400, and the cache block 110 has the same size as the data block 410. By arbitrarily mapping between the cache block 110 and the data block 410 and mapping at least two cache sub-blocks 101 in the cache block 110 to at least two data sub-blocks 401 in the data block 410 one to one, the utilization rate of the cache 100 is improved, thereby improving the performance of the chip 1000.
[0048] In some embodiments, as Figure 2-Figure 4 As shown, in the cache 100 of the control chip 1000, each cache block 110 corresponds to a cache 100 control unit 120, and the cache 100 control unit 120 corresponding to the cache block 110 is used to temporarily store the address of the cache block 110. In actual applications, the relevant content can be referred to the above embodiment and will not be repeated here.
[0049] In some embodiments, in the cache 100 of the control chip 1000, the size of the cache block 110 can be related to the bit width of the bus 300. For example, x = 32, y = 4n; or x = 64, y = 8n, where n is a positive integer. In some examples, if the bit width of the bus 300 is 32 bits, the size of the cache block 110 can be 4*n bytes; or if the bit width of the bus 300 is 64 bits, the size of the cache block 110 can be 8*n bytes. The relevant details can be referred to in the above embodiments and will not be repeated here.
[0050] In some embodiments, the central processing unit 200 is used to: obtain a target access address in response to a read operation instruction; compare the target access address with the address of at least one cache block 110; in response to the target access address being the same as the address of one cache block 110 in the at least one cache block 110 and the data in the cache block 110 is not empty, read and return the target data corresponding to the same address in the cache block 110.
[0051] With the above Figure 4 For example, at least one cache block 110 includes cache block 0 through cache block 63. Each cache block 110 is divided into eight cache sub-blocks for temporarily storing data corresponding to eight addresses, each of which has a data size of 32 bits. Furthermore, the 64 cache blocks correspond to 64 cache control units, namely control unit 0 through control unit 63. Each of the 64 cache control units can temporarily store the base addresses of the corresponding 64 cache blocks, with each cache control unit storing a different address.
[0052] In some examples, a target access address is obtained in response to a read operation instruction, and the target access address is compared with the address of at least one cache block 110, for example, the target access address is simultaneously compared with the addresses temporarily stored in 64 cache control units. In response to the target access address being the same as the address of one of the at least one cache block 110, target data corresponding to the same address in the cache block 110 is read and returned. For example, if the target access address is the same as an address temporarily stored in one of the 64 cache control units, and the data corresponding to the address is not empty, that is, a cache hit, the data corresponding to the same address is read from the cache 100 and returned to the central processing unit 200.
[0053] That is to say, when accessing cache 100, if the access hits, only one cache block 110 will be hit, that is, only one comparison result among the 64 cache control units is 1, and only one 64-bit one-hot code (one hot) can be directly decoded by a one-hot code to binary module to obtain the high-order address [10:5] of cache 100, and the low-order address [4:0] can be obtained by splicing the low-order bits of bus 300.
[0054] Furthermore, in some embodiments, in response to the target access address being different from the address of at least one cache block 110 , target data corresponding to the target access address is retrieved from the memory 400 to the cache 100 via the bus 300 .
[0055] Continue with the above Figure 4 For example, in some examples, in response to a read operation instruction, a target access address is obtained, and the target access address is compared with the address of at least one cache block 110, that is, the target access address is simultaneously compared with the addresses temporarily stored in 64 cache control units. In response to the target access address being different from the address of at least one cache block 110, for example, the target access address is different from the addresses temporarily stored in the 64 cache control units, that is, a cache miss, the target data corresponding to the target access address is obtained from the memory 400, wherein the target data corresponding to the target access address is first returned from the memory 400 to the cache 100 via the bus 300, and then returned from the cache 100 to the central processing unit 200, wherein the bus 300 can use a wrap burst mechanism for data transmission.
[0056] In some embodiments, the cache 100 is used to: in response to the target access address being different from the address of at least one cache block 110, obtain target data corresponding to the target access address from the memory 400 and return the target data to the central processing unit 200; obtain other data that is in the same loop cycle as the target data from the memory 400; and write the target data and the other data into one of the at least one cache block 110 in a first-in-first-out manner.
[0057] In some examples, in response to the target access address being different from the address of at least one cache block 110, target data corresponding to the target access address is retrieved from the memory 400 and returned to the central processing unit 200. Furthermore, other data within the same loop cycle as the target data is pre-fetched from the memory 400 into the cache 100. That is, after the central processing unit 200 retrieves the target data returned from the memory 400 into the cache 100, it can continue to pre-fetch other data within the same loop cycle as the target data from the memory 400 via the bus 300 and temporarily store it in the cache block where the target data is currently located, thereby completing writing to the cache block.
[0058] For example, the bus 300 may use 8-beat burst transmission with loopback transmission to access continuous data blocks. In some examples, Figure 6 As shown, Figure 6 This is a schematic diagram of the prefetch effect of an embodiment of the present application. For example, in response to a miss, the target data n3 corresponding to the target access address is returned from the memory 400 to the cache block N in the cache 100 via the bus 300. Furthermore, other data in the same loop cycle as the target data, namely data n0, data n1, data n2, data n4, data n5, data n6, and data n7, are continuously obtained via the bus 300 to fill the cache block N and implement data prefetching. After returning the target data n3, data n4, data n5, data n6, data n7, data n0, data n1, and data n2 are returned in sequence.
[0059] In this embodiment, the cache 100 is also used to pre-fetch other data that is in the same loop cycle as the target data from the memory 400 to the cache 100 through the bus 300, that is, the cache 100 implements both data reading and pre-fetching functions in one transmission, thereby improving transmission efficiency and reducing hardware expenses while fully utilizing bandwidth.
[0060] In some embodiments, the cache 100 is configured to write target data and other data into one of the at least one cache block 110 in a First In, First Out (FIFO) manner.
[0061] Continue with the above Figure 4Taking an example to illustrate, in some examples, in response to the target access address being different from the address of at least one cache block 110, for example, the target access address being different from the address temporarily stored by the 64 cache control units, that is, a miss, the target data corresponding to the target access address is obtained from the memory 400, wherein the target data corresponding to the target access address is first returned from the memory 400 to one of the 64 cache control units through the bus 300. For example, cache block 0 can be selected for replacement storage based on a first-in-first-out manner. Accordingly, if a miss occurs again, cache block 1 can be selected for replacement storage based on a first-in-first-out manner.
[0062] Furthermore, in some embodiments, the cache 100 control unit 120 includes a counter, and the counter is used to count the data written into a cache block 110 to determine whether the cache block 110 is full.
[0063] After the central processing unit 200 obtains the target data returned from the memory 400 to the cache 100, it can continue to pre-fetch other data in the same loop cycle as the target data from the memory 400 through the bus 300, and temporarily store them in the cache block where the target data is currently located. Among them, the counter can be used to count to determine whether the target data and other data in the same loop cycle as the target data are completely written into a cache block.
[0064] As can be understood, caching the address in the cache control unit 120 allows the cache 100 to obtain a fixed initial value after the power-on reset is released, thus eliminating the need for initialization of the cache 100. Furthermore, initially, each address and address counter in the cache control unit 120 is 0, indicating a miss state. Only after data is read from the memory 400 is the cache address updated and the address counter incremented. This eliminates the possibility of false hits after power-on, and eliminates the need for initialization, saving startup time and simplifying software operations.
[0065] For ease of understanding, the process of the read operation in the embodiment of the present application is described with an example. Figure 7 As shown, Figure 7This is a flowchart of a read operation according to an embodiment of the present application. In response to a read operation instruction, a target access address is obtained, and the target access address is compared with the address of at least one cache block 110. If the target access address matches the address of one of the at least one cache block 110 and is not empty, the cache block 110 is hit, and the target data corresponding to the same address in the cache block 110 is read and returned. If there is a miss, the cache control unit 120 obtains the target address, for example, by latching the address through a latch, and forwards the read operation instruction to the memory 400, thereby obtaining the target data corresponding to the target access address returned from the memory 400 to the cache 100 through the bus 300, and pre-fetching other data in the same loop cycle as the target data from the memory 400, and temporarily storing the target data and the remaining data in the corresponding cache block 110 in the cache 100, and determining whether the corresponding cache block 110 is full to end an access operation.
[0066] In some embodiments, the central processing unit 200 is used to: in response to a write operation instruction, obtain a target access address, and write the data corresponding to the write operation instruction into the memory 400; wherein, in response to the target access address being the same as the address of a cache block 110 in at least one cache block 110 and the data in the cache block 110 being not empty, the data corresponding to the write operation instruction is also written into the cache block 110.
[0067] Continue with the above Figure 4 For example, at least one cache block 110 includes cache block 0...cache block 63, and the 64 cache blocks correspond to 64 cache control units, namely control unit 0...control unit 63. In some examples, in response to a write operation instruction, a target access address is obtained, and data corresponding to the write operation instruction is written to the memory 400. The target access address is simultaneously compared with the addresses of the 64 cache blocks 110, that is, the target access address is simultaneously compared with the addresses temporarily stored in the 64 cache control units. In response to the target access address being the same as the address of a cache block 110 in the at least one cache block 110 and the data in the cache block 110 is not empty, for example, the target access address is the same as an address temporarily stored by control unit 2 in the 64 cache control units, and the data corresponding to the address is not empty, the data corresponding to the write operation instruction is written to cache block 2, that is, the data content temporarily stored in cache block 2 is updated. Alternatively, in response to the target access address being different from the address temporarily stored in the at least one cache block 110, that is, a miss, no processing is performed on the cache 100.
[0068] In some embodiments, the cache 100 includes a clear cache enable port, and the CPU 200 is further configured to: in response to data in the data block 410 being overwritten, clear the data in the cache block 110 mapped by the data block 410 using the clear cache enable port.
[0069] In response to the data in the data block 410 being overwritten, the data in the cache block 110 mapped by the data block 410 is cleared using the clear cache enable port. For example, when the data block 410n is mapped to the cache block 110n in the cache 100 and the data in the data block 410n is overwritten by another master, the central processing unit 200 can clear the temporarily stored data in the cache block 110n through the clear cache enable port, thereby maintaining the consistency between the cache 100 and the memory 400.
[0070] For ease of understanding, the process of the write operation in the embodiment of the present application is described with an example. Figure 8 As shown, Figure 8 This is a flowchart of a write operation according to an embodiment of the present application. In response to a write operation instruction, a target access address is obtained, the write operation instruction is forwarded to the memory 400, and the data corresponding to the write operation instruction is written into the memory 400. At the same time, the target access address is compared with the address of at least one cache block 110. In response to the target access address matching the address of one of the at least one cache block 110 and the data in the cache block 110 is not empty, that is, a hit, the data corresponding to the write operation instruction is also written into the cache block 110. In response to the target access address not being identical to the address temporarily stored in the at least one cache block 110, that is, a miss, no processing is performed on the cache 100.
[0071] See also Figure 9 , Figure 9 1 is a flow chart of the debugging method of the embodiment of the present application. The method can be applied to a host computer with computing functions, which is connected to the control chip 1000 as described above. It should be noted that if there is substantially the same result, the method of the present application is not based on Figure 9 The process sequence shown is limited.
[0072] In some possible implementations, the method may be implemented by a processor calling computer-readable instructions stored in a memory, such as Figure 9 As shown, the method may include the following steps:
[0073] S91: In response to the control instruction, the host computer controls the memory and cache to enter the debugging mode.
[0074] In response to the control instruction, the host computer controls the memory 400 and cache 100 to enter debug mode. For example, in response to the control instruction, the host computer controls the memory 400 and cache 100 to enter debug mode simultaneously. At this time, the CPU 200's operation of reading the memory 400 will be intercepted by the cache 100, and the CPU 200 will no longer read the memory 400.
[0075] S92: Responding to the debug instruction, obtaining a target address.
[0076] In response to the debugging instruction sent by the host computer, the target address to be accessed is obtained.
[0077] S93: In response to the target data obtained according to the target address not being a specific value, extract the target data from the cache for debugging.
[0078] In response to the target data obtained according to the target address not being a specific value, the target data is extracted from the cache for debugging, for example, the target address is obtained to access the cache 100 to determine whether there is a hit. If the target address is the same as the address stored in the cache 100, it is a hit, and the data corresponding to the address is returned from the cache to the central processing unit 200; if the target address is different from the address stored in the cache 100, it is a miss, and a specific value is replied to the central processing unit 200, thereby determining whether the abnormal data comes from the cache 100 based on whether the replied target data is a specific value, so as to perform fault elimination debugging.
[0079] In this embodiment, in response to a control instruction, the host computer controls the memory and cache to enter a debugging mode, and in response to the debugging instruction, obtains a target address. In response to the target data obtained according to the target address not being a specific value, the target data is extracted from the cache for debugging, that is, the internal data of the cache 100 is read out through the debug interface of the central processing unit 200, which is conducive to eliminating fault debugging operations.
[0080] In some embodiments, the cache 100 includes a first cache area 101 and a second cache area 102 . The first cache area 101 may be a data cache area (Data Cache, dCache), and the second cache area 102 may be an instruction cache area (Instruction Cache, iCache).
[0081] The debug instruction includes a debug enable signal. In response to the debug enable signal being a first preset value, target data is extracted from the first cache area 101 , or in response to the debug enable signal being a second preset value, target data is extracted from the second cache area 102 .
[0082] In some examples, such as Figure 10 As shown, Figure 10This is a flow chart of a debugging method according to an embodiment of the present application. In response to a debugging enable signal being a first preset value, target data is extracted from the first cache area. For example, the debug enable signal iCache_debug_en can be configured by the host computer to be 1, indicating that the target data is extracted from the first cache area 101, and the debug enable signal iCache_debug_en is 0, indicating that the target data is extracted from the second cache area 102; the debug enable signal iCache_debug_en can also be configured by the host computer to be 0, indicating that the target data is extracted from the first cache area 101, and the debug enable signal iCache_debug_en is 1, indicating that the target data is extracted from the second cache area 102.
[0083] In this embodiment, in response to the debug enable signal being a first preset value, the target data is extracted from the first cache area 101, or in response to the debug enable signal being a second preset value, the target data is extracted from the second cache area 102, so that it can be determined whether the abnormal data comes from the first cache area 101 or the second cache area 102.
[0084] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0085] See also Figure 11 , Figure 11 1 is a schematic diagram of the structure of the host computer of an embodiment of the present application. Host computer 1100 includes a memory 1101 and a processor 1102 coupled to each other. Processor 1102 is configured to execute program instructions stored in memory 1101 to implement the steps of the debugging method embodiment described above. In a specific implementation scenario, the host computer may include, but is not limited to, a microcomputer or a server, and is not limited here.
[0086] Specifically, the processor 1102 is used to control itself and the memory 1102 to implement the steps of the above-mentioned debugging method embodiment. The processor 1102 can also be called a CPU (Central Processing Unit), and the processor 1102 may be an integrated circuit chip with signal processing capabilities. The processor can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 1102 can be implemented by an integrated circuit chip.
[0087] The non-volatile computer-readable storage medium of the embodiment of the present application can be used to store a computer program. When the computer program is executed by the processor 1102, for example, when executed by the processor in the above embodiment, it is used to implement the steps of the above embodiment for the debugging method.
[0088] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0089] In the several embodiments provided in this application, it should be understood that the disclosed methods and related devices can be implemented in other ways. For example, the above-described related device implementation methods are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication disconnection shown or discussed can be through some interfaces, indirect coupling or communication disconnection of devices or units, which can be electrical, mechanical or other forms.
[0090] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0091] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0092] It is easy for a person skilled in the art to know that many modifications and variations can be made to the apparatus and method while maintaining the teaching content of the present application.Therefore, the above disclosure should be considered as being limited only by the scope of the appended claims.
Claims
1. A cache, characterized in that: connected between a central processing unit and a bus, the bus being further connected to a memory, the central processing unit operating at a first frequency, the bus and the memory operating at a second frequency, the first frequency being greater than the second frequency; The cache includes at least one cache block, wherein any one of the cache blocks is used to store any one of the at least one data block in the memory, and the cache block has the same size as the data block; Each of the cache blocks is divided into at least two cache sub-blocks, and at least two cache sub-blocks in the cache block are mapped one-to-one with at least two data sub-blocks in the data block.
2. The cache according to claim 1, wherein: The cache further includes at least one cache control unit; Each of the cache control units is connected to a corresponding cache block and is used to temporarily store a base address of the corresponding cache block, wherein the base address of at least one cache block is different.
3. The cache according to claim 1, wherein: The size of the cache block is related to the bit width of the bus, wherein the bit width of the bus is x bits, the size of the cache block is y bytes, and y is a positive integer multiple of x / 8.
4. A control chip, characterized in that: include: Memory; bus, connected to the memory; a central processing unit connected to the bus via a cache; The memory includes at least one data block, the cache includes at least one cache block, any one of the cache blocks is used to store any one data block, the data block and the cache block have the same size, each of the cache blocks is divided into at least two cache sub-blocks, and at least two cache sub-blocks in the cache block are mapped one-to-one with at least two data sub-blocks in the data block; The central processing unit is used for: In response to the operation instruction, obtaining a target access address; The target access address is compared with the address of the at least one cache block to perform a corresponding data operation through the cache.
5. The chip according to claim 4, characterized in that The central processing unit is used for: In response to a read operation instruction, obtaining a target access address; comparing the target access address with the address of the at least one cache block; In response to the target access address being the same as the address of one of the at least one cache block and the data in the cache block being not empty, reading and returning target data corresponding to the same address in the cache block; In response to the target access address being different from the address of the at least one cache block, target data corresponding to the target access address is retrieved from the memory to the cache via the bus.
6. The chip according to claim 5, characterized in that The cache is used to: In response to the target access address being different from the address of the at least one cache block, obtaining target data corresponding to the target access address from the memory, and returning the target data to the central processing unit; Acquire other data that is in the same loop cycle as the target data from the memory; The target data and the other data are written into one of the at least one cache block in a first-in-first-out manner.
7. The chip according to claim 4, characterized in that The central processing unit is used for: In response to a write operation instruction, obtaining a target access address, and writing data corresponding to the write operation instruction into the memory; wherein, in response to the target access address being the same as the address of one of the at least one cache block and the data in the cache block being not empty, the data corresponding to the write operation instruction is also written into the cache block; In response to the target access address being different from the address of the at least one cache block, not performing a write operation on the cache; or, The central processing unit is also used for: In response to the data in the data block being overwritten, the data in the cache block mapped by the data block is cleared using the clear cache enable port of the cache.
8. A debugging method, characterized in that: Applied to a host computer, the host computer is connected to the control chip according to any one of claims 4 to 7, and the method includes: In response to a control instruction, the host computer controls the memory and the cache to enter a debugging mode; In response to a debug instruction, obtaining a target address; In response to the target data obtained according to the target address not being a specific value, the target data is extracted from the cache for debugging.
9. The method according to claim 8, characterized in that The cache includes a first cache area and a second cache area, the debug instruction includes a debug enable signal, and the method further includes: In response to the debug enable signal being a first preset value, extracting the target data from the first buffer area; In response to the debug enable signal being a second preset value, the target data is extracted from the second buffer area.
10. A host computer, characterized in that: It comprises a memory and a processor coupled to each other, wherein the processor is used to execute program instructions stored in the memory to implement the debugging method according to any one of claims 8 to 9.
Citation Information
Patent Citations
High-performance instruction cache system and method
CN104424132A
Information block transfer management in a multiprocessor computer system employing private caches for individual center processor units and a shared cache
US6006309A
Cited By
Data reading and writing method and system, computer equipment and storage medium
CN121070285A
Data reading and writing method, system, computer device and storage medium
CN121070285B