Cache device, data handling method and chip
By configuring the first and second address spaces in the cache device and implementing the DMA function using the Cache Miss mechanism, the problem of integrating DMA modules in the MCU or SOC increases cost and power consumption, and a low-cost and low-power DMA solution is realized.
Patent Information
- Application Number
- CN202510899479.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The prior art integrates DMA modules in MCUs or SOCs significantly increases cost and power consumption, and a low-cost and low-power DMA solution is urgently needed.
By configuring the first and second address spaces in the cache device, the DMA function is realized using the Cache Miss mechanism, and the target cache missing is triggered to obtain data through the bus and transport it to the designated storage space, multiplexing the processing logic of cache missing to avoid adding DMA modules.
While implementing DMA function, it reduces the cost and power consumption of the chip, improves system efficiency and reduces the CPU burden.
Smart Images

Figure CN120407473A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data transmission technologies, and particularly to a cache device, a data transfer method, and a chip. Background Art
[0002] Direct Memory Access (DMA) is a technology in a computer system that allows certain hardware subsystems to directly access the main memory, independent of the Central Processing Unit (CPU). DMA can perform data transfer without occupying CPU resources, improving system efficiency, reducing the CPU burden, increasing data transfer speed, and optimizing system performance. However, the DMA module itself has a relatively large area, and integrating a DMA module on an MCU (Microcontroller Unit) or SOC (System on Chip) that pursues extreme area and cost will significantly increase costs and power consumption.
[0003] Therefore, there is an urgent need for a low-cost and low-power DMA solution. Summary of the Invention
[0004] In view of this, this application provides a cache device, a data transfer method, and a chip, which reduce the cost and power consumption of the chip while implementing the DMA function.
[0005] In a first aspect, this application provides a cache device, on which a first address space and a second address space are configured; the cache device is configured to return data by the cache device when a read instruction for the first address space is received; The cache device is further configured to trigger a target cache miss when a read instruction for the second address space is received, so as to obtain target data in a second storage area through a bus on the chip, and transfer the target data in the second storage area to the specified storage space.
[0006] In an optional implementation manner, a first function register is included in the device; an address of a second storage area, an address of a specified storage space, and a data transfer length are stored in the first function register.
[0007] In an optional implementation manner, a space judgment module is further included in the device; the space judgment module is configured to receive a read instruction and judge the memory access address in the read instruction; if the memory access address is located in the first address space and the cache is not hit, an original cache miss is triggered; if the memory access address is located in the second address space, the target cache miss is triggered.
[0008] In an alternative embodiment, the device further includes a replacement control module; the replacement control module is configured to receive a cache line load request sent after the space determination module triggers a target cache miss; The replacement control module is configured to, after receiving the cache line load request, read data of the transfer data length according to the second storage area address and write it into the specified storage space address.
[0009] In an alternative embodiment, the replacement control module further includes a current data read address register, a current data storage address register, and a remaining data length counter.
[0010] Second, a data transfer method is provided. The method is applied to the above cache device, and the method includes: When a read instruction for the second address space is received, a target cache miss is triggered, and target data in the second storage area is obtained through the bus on the chip; Transfer the target data in the second storage area to a specified storage space.
[0011] In an alternative embodiment, the method further includes: Before the target data in the second storage area starts to be transferred, return a target identifier to the processor in the chip.
[0012] In an alternative embodiment, the method further includes: When the transfer of data of the transfer data length from the second storage area is completed, send an interrupt notification to the processor in the chip.
[0013] In an alternative embodiment, the method further includes: receiving a configuration instruction sent by the processor; the configuration instruction includes the address of the second storage area, the address in the specified storage space, and the transfer data length; Configure the first function register in the cache device according to the configuration instruction.
[0014] Third, a chip is provided, and the above cache device is provided in the chip.
[0015] Fourth, a computer-readable storage medium is provided. Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause the cache device to execute the above data transfer method.
[0016] The technical solution provided by this application may include the following beneficial effects: In the present application, the cache device can be configured to return data when a read instruction for a first address space is received, that is, when a read instruction for the first address space is received and the cache misses, an original cache miss (i.e., a normal cache miss) is triggered. At this time, according to the processing flow of the original cache miss, data in the first storage area is obtained through the on-chip bus and transferred to the cache device, and then the data is returned to the CPU. If the cache hits, the cached data in the cache device is directly returned to the CPU; in addition to retaining the above normal cache miss function, the present application also reuses the processing function and logic of the cache miss, that is, the cache device is configured to trigger a target cache miss when a read instruction for a second storage area is received, so as to obtain the target data in the second storage area through the on-chip bus and transfer the target data in the second storage area to a specified storage space; since the cache miss mechanism is triggered, the target data in the second storage area can be transferred through the on-chip bus without the intervention of the CPU. At this time, the cache device can directly transfer the obtained data to the specified storage space, thereby realizing the DMA function. The above solution can realize the DMA function without adding a DMA module, thereby reducing the cost and power consumption of the chip while realizing the DMA function. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 Shows a low-cost and low-power MCU system architecture diagram.
[0019] Figure 2 Shows a structural schematic diagram of a cache device.
[0020] Figure 3 Shows a Cache function architecture diagram in an embodiment of the present application.
[0021] Figure 4 Shows a Cache function architecture diagram related to an embodiment of the present application.
[0022] Figure 5 Shows a method flow diagram of a data transfer method related to an embodiment of the present application.
[0023] Figure 6The flowchart of the method of the cache device involved in the embodiments of the present application is shown.
[0024] Figure 7 It is a schematic structural diagram of a chip provided by an alternative embodiment of the present application. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0026] In the description of the embodiments of the present application, the term "corresponding" may indicate a direct or indirect corresponding relationship between two parties, may also indicate an associated relationship between two parties, or may be a relationship such as indication and being indicated, configuration and being configured, etc.
[0027] Figure 1 A low-cost and low-power MCU system architecture diagram is shown. As Figure 1 shown, during the processing of the CPU, data can be obtained through the Cache. On the other hand, a DMA module is also provided in the MCU. The DMA can perform data transmission without occupying CPU resources, improving system efficiency and reducing the burden on the CPU. However, generally, the area of the DMA module itself is relatively large. Therefore, in an MCU that pursues extreme area and cost, integrating the DMA module will increase the complexity, area, and power consumption of the circuit.
[0028] Figure 2 A schematic structural diagram of a cache device is shown. As Figure 2 shown, the cache device is configured with a first address space and a second address space; the first address space and the second address space can be for the space mapping relationship of the virtual space; the cache device is configured to return data when receiving a read instruction for the first address space. First, obtain the data in the first storage area through the bus on the chip and transfer it to the cache device, and then return the data to the CPU; The cache device is further configured to trigger a target cache miss when receiving a read instruction for the second address space, so as to obtain the target data in the second storage area through the bus on the chip and transfer the target data in the second storage area to the specified storage space.
[0029] Optionally, as Figure 2As shown, both the first storage area and the second storage area are located in an external storage outside the chip. When a read instruction for the first address space is received, it can be determined whether the cache hits; if the cache hits, the data already cached in the cache device is directly returned, and if the cache misses and the original cache miss is triggered, the data in the first storage area of the external storage can be obtained through the external storage controller, transported to the cache device, and then the data is returned to the CPU for use; when the chip triggers a target cache miss, the data in the second storage area of the external storage can be obtained through the external storage controller and transported to a specified storage space outside the cache device, such as an off-chip storage space or Figure 2 the on-chip system SRAM (Static Random-Access Memory) in .
[0030] Optionally, in the embodiment of the present application, the cache device may be a Cache, and a specified address space (including the first address space and the second address space) is configured on the Cache. The specified address space can be used to trigger the Cache Miss mechanism of the Cache. Once the cache device receives a read instruction for the specified address space, if the data in the specified address space has not been cached, the Cache Miss mechanism of the Cache will be triggered.
[0031] "Cache Miss" (cache miss) means that when accessing the cache (Cache), the required data is not in the Cache and must be obtained from the next-level cache or a slower next-level storage space, resulting in additional access latency and performance overhead. In the Cache Miss mechanism, the data acquisition process is completely automatically completed by the cache controller in the hardware. The CPU does not need to execute additional instructions or microcode to coordinate this process and only needs to wait for the data to be returned. Therefore, the Cache Miss mechanism can be used to implement the DMA function to perform data transfer without occupying CPU resources, improve system efficiency, and reduce the CPU burden.
[0032] The data transfer requirements mainly include instruction transfer and data transfer. Among them, instruction transfer is to load from the external storage to the Cache storage space; data transfer refers to loading from the external storage to a specified storage space, such as the on-chip system SRAM or the off-chip storage space.
[0033] However, in the Cache Miss mechanism, when data is fetched from the next-level cache or a slower next-level storage space (off-chip storage in the embodiments of the present application), the data is directly stored in the storage space of the Cache. Therefore, in the embodiments of the present application, the Cache needs to be configured such that when fetching data from the second storage area outside the chip, the data is directly fetched and stored in the designated storage space (optionally, the designated storage space in the embodiments of the present application can be the on-chip system SRAM), so as to implement a function similar to DMA data fetching (for ease of description, in the embodiments of the present application, this function similar to DMA data fetching is simply referred to as the DMA function. It should be noted that the cache device in the embodiments of the present application does not have a DMA module).
[0034] Figure 3 Fig. shows a Cache function architecture diagram in the embodiments of the present application. As Figure 3 shown, the Cache function architecture includes a Cache configuration register, a Cacheable space judgment module, a Cache storage space, and a replacement control module. Optionally, the Cache in the embodiments of the present application can be an instruction cache or a data cache. Among them, the Cache configuration register is used to configure the basic functions of the Cache; The Cacheable space judgment module is configured to judge whether the address accessed by the current memory access operation is a cacheable space. If it is judged that the current memory access operation accesses a cacheable space, the memory access address is sent to the main memory Cache address mapping module. If it is judged that the current memory access operation does not access a cacheable space, the memory access operation directly accesses the corresponding non-cacheable space through the bus interface, such as the configuration space of system peripherals, etc.; Optionally, the space judgment module further includes a main memory Cache address mapping conversion module, which is configured to judge whether the currently accessed cacheable space has been cached in the Cache storage space. If so, a Cache Hit (cache hit) occurs, and data is retrieved from the Cache storage space and returned to the CPU; if not, a Cache Miss occurs, and the access address is sent to the replacement control module.
[0035] The replacement control module is configured to retrieve the missing data from the next-level cache or a slower next-level storage space when a Cache Miss occurs. Generally, the data / instructions of the entire Cacheline (cache line, the length is generally 16Byte, 32Byte, etc.) where the missing data is located are retrieved, and then written into the Cache storage space according to the Cache replacement policy, and then the data is returned to the CPU.
[0036] Figure 4shows the Cache functional architecture diagram involved in the embodiments of the present application. As Figure 4 shown, further, in order to implement the above DMA function on the basis of the original Cache function, a first function register needs to be set in the device; the address of the second storage area, the address of the specified storage space, and the data transfer length are stored in the first function register.
[0037] That is, in the embodiments of the present application, relevant configurations of the DMA function need to be added to the DMA configuration register (i.e., the first function register) of the Cache, which need to include the starting address of the data in the external storage (i.e., the address in the second storage area), the starting address of the target SRAM storage space (the address in the specified storage space), and the length of the original data to be transferred (i.e., the data transfer length).
[0038] In the embodiments of the present application, the Cache further includes a second space address register for storing the starting address and the ending address of the second address space. In an optional implementation manner of the embodiments of the present application, the first function register may include the above-mentioned second space address register to store the starting address and the ending address of the second address space.
[0039] Optionally, in the DMA configuration register, a DMA function switch and a DMA interrupt status register may also be configured. Among them, enabling the DMA function can be independent of the original Cache function and is not affected by the Cache enable control. The DMA interrupt status register can be used to indicate whether the DMA data transfer is completed.
[0040] Further, in order to stably trigger the Cache Miss function of the Cache, a space judgment module is further included in the device; the space judgment module is configured to receive a read instruction and judge the memory access address in the read instruction; if the memory access address is in the first address space and the cache is not hit, the original cache miss is triggered; if the memory access address is in the second address space, the target cache miss is triggered.
[0041] In the embodiments of the present application, when the DMA function configuration is completed and the DMA function is enabled, the space judgment module monitors the memory access address sent by the CPU. If it is in the Cacheable space configured by the DMA function, a Cache Miss is generated, and a Cacheline load request is sent to the replacement control module for Cacheline loading.
[0042] Specifically, in the embodiments of the present application, the replacement control module is specifically configured to receive the cache line load request sent after the space judgment module triggers the target cache miss; The replacement control module is configured to, after receiving the cache line load request, read data of the transfer data length according to the second storage area address and write it into the specified storage space address.
[0043] Therefore, the original replacement control module reads back the Cacheline data in the off-chip storage that is missing and writes it into the Cache storage space according to the Cache replacement policy. After adding the DMA function, the replacement control module can, according to the DMA data transfer configuration, start from a preset address in the external storage space, load data of a preset length, and sequentially write it into the target address space in the on-chip SRAM, so as to implement the DMA function through the Cache.
[0044] In order to implement the functions of the above replacement control module, in addition to the original Cacheline load control logic and Cacheline buffer (buffer for caching Cacheline), etc., the replacement control module may further include a current data read address register, a current data storage address register, and a remaining data length counter. Through this design, compared with the area of the DMA module, the increase in the logic area of the replacement control module is basically negligible.
[0045] After the DMA function is configured, the CPU executes a Load (read) instruction to read a data from the original data start address in the second storage space, triggering a Cache miss and starting the DMA data transfer process; thereafter, the Cache can return the data corresponding to the Load address to the CPU as the execution result of this Load instruction, without blocking the CPU instruction execution. The Load instruction only functions as an operation to start the DMA function.
[0046] And in the embodiments of the present application, an interrupt of the Cache is further added to the system to indicate the completion of the DMA data transfer: after the DMA transfers all the data, it sends an interrupt to the CPU, and the CPU processes the Cache interrupt and starts to perform subsequent processing on the data transferred and completed in the SRAM.
[0047] In the embodiments of the present application, the above logic can implement specific data transfer processing in the following two ways: Method 1: There is a Cacheline buffer in the cache replacement control module, which can be reused. Each time data is transferred, it is processed according to the Cacheline length. For example, if the Cacheline length is 16 bytes of data, 16 bytes of data are first read from the external storage in sequence and stored in the Cacheline Buffer each time, and then the 16-byte data is taken out from the Cacheline buffer and written into the target space of the on-chip target SRAM through the bus master interface of the cache.
[0048] Among them, the Cacheline Buffer is a dedicated buffer in the cache replacement control module for receiving the entire Cacheline data returned from the next-level cache or slower next-level storage. When a cache miss occurs, the buffer first receives the complete Cacheline (such as 16 bytes or 32 bytes, etc.), and then writes it into the cache storage space. Therefore, in Method 1, the buffer can be reused for the DMA function. At this time, the replacement control module can first read from the external storage according to the Cacheline size, the obtained data returns to the Cacheline Buffer, and then the entire Cacheline data is written into the target space of the on-chip target SRAM in sequence through the bus master interface of the cache.
[0049] Method 2: The replacement control module retrieves one data each time, such as 1 Word, 4 bytes, and directly writes it into the SRAM target space without passing through the Cacheline Buffer.
[0050] In Method 2, the replacement control module can directly bypass the Cacheline Buffer. In the cache replacement control module, one unit granularity of data (such as 4 bytes) is read from the external storage each time and directly written into the corresponding address in the SRAM through the bus master interface of the cache without first caching the entire Cacheline as a whole.
[0051] For the above two transfer methods, selection can be made according to the actual length of the data to be transferred. One of them can be selected, or the two methods can be used in combination. Specifically, in an alternative implementation manner of the embodiment of the present application, it is possible to directly select to transfer data only according to Method 1 or transfer data only according to Method 2. Generally speaking, when transferring data according to Method 1, the pressure of writing data into the target space of the on-chip target SRAM can be significantly reduced. However, during the process of transferring data according to Method 1, the size of the data to be transferred is not exactly an integer multiple of the Cacheline length. For example, after transferring data of a certain Cacheline length, if the remaining data to be transferred is 8 Byte, at this time, there is no need to use the Cacheline Buffer, and directly switch to Method 2 to transfer the remaining 8 Byte data to the target space of the on-chip target SRAM.
[0052] Through the above configuration, the cache device in the embodiment of the present application can have the following logic: First, the cache device has a conventional Cache function, such as the processing of Cache Hit and Cache Miss, which is carried out according to normal logic. When there is no Cache Miss in the conventional Cache function, the DMA function can proceed normally according to the above logic.
[0053] However, when currently processing a conventional Cache Miss, even if the cache device receives a Load instruction to start DMA at this time, the DMA function is not started.
[0054] However, if a Cache Miss corresponding to the DMA function occurs, that is, data transfer is in progress and at this time a Cache Miss also occurs in the conventional Cache function, then after the DMA completes the transfer of the data of the current Cacheline (reading and writing by Cacheline) or a single data (reading and writing by a single data), the DMA function is paused, and the replacement control module saves the on-site state of the DMA transfer, and turns to preferentially process the conventional Cache Miss to ensure the performance of the CPU.
[0055] This function saves the state of the current DMA transfer through the DMA intermediate state save register of the replacement control module, then starts to process the conventional Cache Miss. After the processing is completed, the state in the DMA intermediate state save register is restored to the control register of the replacement control module to restore the on-site state of the DMA transfer and continue the DMA data transfer.
[0056] If the Cacheable space of the DMA function overlaps with the Cacheable space of the original cache, when the CPU sends a Load instruction to start the DMA function, the space accessed by the Load instruction should not fall within the Cacheable space of the original function; if the data space read by the DMA is a subset of the original Cacheable space, the normal cache function can be considered to be paused first, that is, after pausing the normal cache function, start the DMA transfer through the Load instruction, and then start the cache, so that while not affecting the DMA transfer, the CPU performance is also hardly affected.
[0057] After the DMA function is started, once a DMA Cache Miss occurs, the DMA data loading is started. After that, before the data transfer is completed, all subsequent access operations to the second address space are processed according to the normal Load instruction until the current DMA transfer is completed, and then the DMA function is reconfigured and enabled.
[0058] After the DMA is started, it is automatically disabled after the data transfer is completed. If you want to reuse the DMA function, you need to reconfigure and enable the DMA function.
[0059] Moreover, the solution shown in the embodiments of the present application is applicable regardless of whether the MCU is a von Neumann bus architecture or a Harvard bus architecture.
[0060] In summary, in the present application, the cache device can be configured to return data when a read instruction for the first address space is received, that is, when a read instruction for the first address space is received and the cache miss occurs, the original cache miss (i.e., the normal cache miss) is triggered. At this time, according to the processing flow of the original cache miss, the data in the first storage area is obtained through the on-chip bus and transferred to the cache device, and then the data is returned to the CPU. If the cache hit occurs, the cached data in the cache device is directly returned to the CPU; in addition to retaining the above normal cache miss function, the present application also reuses the processing function and logic of the cache miss, that is, the cache device is configured to trigger a target cache miss when a read instruction for the second storage area is received, so as to obtain the target data in the second storage area through the on-chip bus and transfer the target data in the second storage area to the specified storage space; since the cache miss mechanism is triggered, the target data in the second storage area can be transferred through the on-chip bus without the intervention of the CPU. At this time, the cache device can directly transfer the obtained data to the specified storage space, so as to implement the DMA function. The above solution can implement the DMA function without adding a DMA module, thereby reducing the cost and power consumption of the chip while implementing the DMA function.
[0061] Correspondingly, the present application also provides a data transfer method, which can be applied to a cache device as Figure 2 shown. Figure 5 FIG. shows a flowchart of a data transfer method according to an embodiment of the present application. As Figure 5 shown, the method includes: Step 501, when a read instruction for the second address space is received, trigger a target cache miss, and obtain the target data in the second storage area through the on-chip bus.
[0062] Optionally, receive a configuration instruction sent by the processor; the configuration instruction includes the address of the second storage area, the address in the specified storage space, and the length of the transferred data; configure the first function register in the cache device according to the configuration instruction.
[0063] In the embodiment of the present application, in the cache device, the original function register corresponding to the Cache can be configured in advance. For example, the original cache miss function corresponding to the Cache is configured first, so that the Cache is configured to trigger the original cache miss once a read instruction for the first address space is received and the cache miss occurs; according to the processing mechanism of the Cache Miss, once it is triggered, data will be obtained from the lower-level storage.
[0064] On the other hand, in this cache device, the first function register corresponding to the Cache can also be pre-configured so that the Cache is configured to trigger a cache miss corresponding to the DMA once a read instruction for the second address space is received. According to the cache miss handling mechanism, once triggered, data will be fetched from the lower-level storage. In the embodiments of the present application, the address of the second storage area read by the target cache miss function, the specified storage space (i.e., the address in the specified storage space) where the data is to be stored, and the length of data transferred each time can be configured as needed.
[0065] Furthermore, when a read instruction for the second address space is received, it is necessary to determine whether the DMA function flag in the first function register of the cache device is enabled. If it is enabled, a cache miss is triggered to trigger the DMA function; if it is not enabled, it is processed according to the normal memory access operation, and the data in the accessed address space is returned.
[0066] Step 502: Transfer the target data in the second storage area to the specified storage space.
[0067] In the embodiments of the present application, the target cache miss triggered by the second address space is not completely the same as the conventional cache miss function. When the conventional cache miss function is triggered, data is fetched from the lower-level storage and filled into the cache storage space corresponding to the missing cache line; while for the cache miss triggered by the second address space, after the data is fetched from the second storage area outside the chip, the data fetched from the second storage area is directly transferred to the specified storage space (this function can be implemented by configuring the replacement control module).
[0068] Therefore, in the logic of the target cache miss triggered by the second address space, the read instruction for the second address space can be used as a trigger instruction to trigger a cache miss to implement the DMA data transfer function. Thus, in the embodiments of the present application, after the cache miss corresponding to the DMA function is triggered, the data in the second storage area outside the chip will be transferred to the specified storage space according to the address stored in the first function register of the cache device. Before the data in the second storage area starts to be transferred, the cache device can return a target flag to the processor in the chip; this target flag is the return result of the read instruction for the second address space, thereby informing the processor that the DMA function has started to be executed.
[0069] Optionally, in the embodiments of the present application, the read instruction used to trigger the DMA function is a single instruction. At this time, the target flag can be the data corresponding to this single read instruction.
[0070] Alternatively, in the embodiments of the present application, the target identifier may also be pre-configured data for informing the CPU that the cache device has triggered the DMA function at this time.
[0071] Further, after the target data transfer in the second storage area is completed, the cache device needs to send an interrupt notification to the processor in the chip.
[0072] In a possible implementation manner, during the process of transferring the target data in the second storage area to the specified storage space, the target data in the second storage area can be obtained from the external storage according to the size of the CachelineBuffer in the Cache replacement control module, and the read data is returned to the Cacheline Buffer, and then the data in the Cacheline Buffer is written into the specified storage space (such as SRAM) in sequence through the bus interface of the Cache.
[0073] In another possible implementation manner, during the process of transferring the target data in the second storage area to the specified storage space, whenever data of a unit granularity (such as 1 Word, 4 Bytes) is obtained, the data of the unit granularity is directly written into the corresponding address in the specified storage space through the bus interface of the Cache, without first caching the data of the entire Cacheline.
[0074] The above two data transfer methods can be seen in detail Figure 2 in the corresponding Embodiment 1 and Embodiment 2, which will not be elaborated here.
[0075] Please refer to Figure 6 , which shows the method flow chart of the cache device involved in the embodiments of the present application. As Figure 6 shown, in the embodiments of the present application, before the Cache (i.e., the cache device) works normally, the CPU needs to configure the function registers of the Cache, including the registers for realizing the normal function of the Cache (such as the original cache miss) and the DMA function registers. After the Cache is configured, the CPU can send a read instruction to the Cache. When the Cache receives the read instruction, it can judge the memory access address of the read instruction to determine whether it belongs to the first address space or the second address space; If the memory access address in the read instruction is in the first address space and the cache hits, the data is directly read from the cache device and returned to the CPU.
[0076] If the memory access address in the read instruction is located in the first address space and the cache miss occurs, the original cache miss is triggered, that is, the normal Cache Miss process. At this time, the cache sequentially reads the target Cacheline data in the first storage area and writes the read data into the corresponding storage space of the cache. The CPU can then obtain the data from the storage space of the cache and process it.
[0077] If the memory access address in the read instruction is located in the second address space, the target cache miss is triggered, that is, the Cache Miss for implementing the DMA function. At the same time, the target identifier is returned to the CPU to inform the CPU that the DMA function has been started at this time. Then the cache sequentially reads the target data in the second storage area, but does not write it into the storage space of the cache. Instead, the target data from the second storage area is directly written into the specified storage space, and an interrupt notification can be sent to the CPU after this process is executed.
[0078] Therefore Figure 6 In the method flow of the cache device shown, by reusing the Cache Miss function, the cache can be reused to implement the DMA function with only simple configuration on the premise of realizing its original function. Compared with the solution of adding a DMA module to implement the DMA function, both the cost and power consumption are reduced.
[0079] Furthermore, in the cache device shown as Figure 6 the first address space and the second address space can be two independent address spaces (that is, there is no overlapping part between the first address space and the second address space), or there is at least partial overlap between the first address space and the second address space.
[0080] First, if the first address space and the second address space are two independent address spaces, for example, the range of the first address space is (1 - 1000), and the range of the second address space is (1001 - 2000). At this time, if the address of the read instruction falls within the first address space, the normal cache function is triggered, that is, if the cache hits, the data is directly returned, and if the cache misses, the original cache miss is triggered; if the address of the read instruction falls within the second address space, the target cache miss is triggered, thus implementing the DMA function through the above logic.
[0081] Furthermore, if there is at least partial overlap between the first address space and the second address space. For example, the range of the first address space is (1 - 1000), and the range of the second address space is (800 - 1200), then the range (800 - 1000) belongs to both the first address space and the second address space. Therefore, within the overlapping range, the triggering conditions for both the regular Cache function and the target cache miss are satisfied.
[0082] In a possible implementation, the Cache can be set to not trigger the Cache function and the target cache miss simultaneously. In this embodiment of the present application, if there is at least partial overlap between the first address space and the second address space, the CPU can pre-control either the original function register or the first function register in the cache device; so that when the CPU issues a read instruction, either the normal Cache function or the target cache miss function is in a paused state. At this time, even if the read instruction is located in the overlapping part of the first address space and the second address space, the unpaused function can be triggered.
[0083] Optionally, the second address space may also be a subset of the first address space. For example, the range of the first address space is (1 - 1000), and the range of the second address space is (500 - 700). At this time, the address in the read instruction issued by the CPU needs to be located in the second address space, and the normal Cache function needs to be paused to trigger the DMA function.
[0084] In another possible implementation, the Cache can also be set to allow the normal Cache function and the target cache miss to be triggered simultaneously. However, since the normal Cache function includes Cache Hit and Cache Miss, they need to be described separately.
[0085] If there is partial overlap between the first address space and the second address space, and the read instruction issued by the CPU is exactly located in the overlapping part of the first address space and the second address space, and there is data in the overlapping part at this time, that is, Cache Hit can be triggered. At this time, on the one hand, the Cache runs according to the Cache Hit process, that is, directly returns the data in the overlapping part to the CPU; on the other hand, the DMA function is also triggered to transfer the target data in the second storage area to the specified storage space (such as SRAM).
[0086] If there is a partial overlap between the first address space and the second address space, and the read instruction issued by the CPU is exactly located in the overlapping part of the first address space and the second address space, and there is no data in the overlapping part at this time, that is, when a Cache Miss is triggered, in order not to affect the processing performance of the CPU, the Cache will first execute the normal Cache Miss process, obtain the data in the first storage area through the bus on the chip and transfer it to the cache device, so as to return the data to the CPU; then after the normal Cache Miss processing process is completed, the Cache will execute the target cache miss, that is, trigger the DMA function, and transfer the target data in the second storage area to the specified storage space (such as SRAM).
[0087] In summary, in the present application, the cache device can be configured to return data when receiving a read instruction for the first address space, that is, when receiving a read instruction for the first address space and the cache is not hit, trigger the original cache miss (that is, the normal cache miss). At this time, according to the processing process of the original cache miss, obtain the data in the first storage area through the bus on the chip and transfer it to the cache device, and then return the data to the CPU. If the cache is hit, directly return the cached data in the cache device to the CPU; and in addition to retaining the above normal cache miss function, the present application also reuses the processing function and logic of the cache miss, that is, the cache device is configured to trigger the target cache miss when receiving a read instruction for the second storage area, so as to obtain the target data in the second storage area through the bus on the chip and transfer the target data in the second storage area to the specified storage space; since the cache miss mechanism is triggered, without the intervention of the CPU, the target data in the second storage area can be transferred through the bus on the chip. At this time, the cache device can directly transfer the obtained data to the specified storage space, thus realizing the DMA function. The above solution can realize the DMA function without adding a DMA module, thereby reducing the cost and power consumption of the chip while realizing the DMA function.
[0088] The embodiment of the present application also provides a chip, which may include the cache device shown in the embodiment of the present application to implement the DMA function in the chip. Please refer to Figure 7 , Figure 7 is a schematic structural diagram of a chip provided by an alternative embodiment of the present application, as Figure 7As shown, the chip includes: a processor 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the chip. In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple chips can be connected, and each device provides part of the necessary operations (e.g., as a server array, a set of blade servers, or a multi-processor system). Figure 7 In Figure 7 , a processor 10 is taken as an example.
[0089] The processor 10 can be a central processing unit, a network processor, or a combination thereof. The above chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof, such as an MCU or an SOC.
[0090] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0091] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the chip, etc. In addition, the memory 20 can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the chip through a network.
[0092] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory; the memory 20 can also include a combination of the above types of memories.
[0093] The chip also includes a communication interface 30 for the chip to communicate with other devices or communication networks.
[0094] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0095] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A cache device, characterized in that, A first address space and a second address space are configured on the cache device; the cache device is configured to return data when a read instruction for the first address space is received; The cache device is further configured to trigger a target cache miss when a read instruction for the second address space is received, so as to obtain target data in a second storage area through a bus on the chip, and transfer the target data in the second storage area to a specified storage space.
2. The device according to claim 1, wherein The device includes a first function register; the second storage area address, the address in the specified storage space, and the data transfer length are stored in the first function register.
3. The device according to claim 2, wherein The device further includes a space judgment module; the space judgment module is configured to receive a read instruction and judge the memory access address in the read instruction; if the memory access address is in the first address space and the cache is not hit, an original cache miss is triggered; if the memory access address is in the second address space, the target cache miss is triggered.
4. The device according to claim 3, characterized in that, The device further includes a replacement control module; the replacement control module is configured to receive a cache line load request sent after the space judgment module triggers a target cache miss; The replacement control module is configured to, after receiving the cache line load request, read data of the data transfer length according to the second storage area address and write it into the specified storage space address.
5. The device according to claim 4, characterized in that The replacement control module further includes a current data read address register, a current data storage address register, and a remaining data length counter.
6. A data transfer method, characterized in that, The method is applied to the cache device according to any one of claims 1 to 5, and the method includes: When a read instruction for the second address space is received, triggering a target cache miss and obtaining target data in a second storage area through the bus on the chip; Transferring the target data in the second storage area to a specified storage space.
7. The method according to claim 6, characterized in that, The method further includes: Before the target data in the second storage area starts to be transferred, returning a target identifier to a processor in the chip.
8. The method according to claim 6, characterized in that The method further includes: When the transfer of data of the data transfer length from the second storage area is completed, sending an interrupt notification to the processor in the chip.
9. The method according to any one of claims 6 to 8, characterized in that The method further includes: Receiving a configuration instruction sent by the processor; the configuration instruction includes the address of the second storage area, the address in the specified storage space, and the data transfer length; Configuring the first function register in the cache device according to the configuration instruction.
10. A chip, characterized in that, The cache device according to any one of claims 1 to 5 is provided in the chip.
Citation Information
Patent Citations
Direct memory access unit and control component
CN113031849A
Data writing method and device, equipment and medium
CN119415048A
Methods and apparatus for improving throughput of cache-based embedded processors by switching tasks in response to a cache miss
CN1547701A
Method and apparatus for providing test mode access to an instruction cache and microcode ROM
US20030061545A1
Method and apparatus for efficient cache refilling by the use of forced cache misses
US5553264A