Data processing method, readable storage medium, program product and electronic equipment

By setting a buffer in the processor to store the first memory data and sending the second memory request at the same time, the problem of inefficient execution caused by continuous cache misses is solved, and the data acquisition speed and execution efficiency of the processor are improved.

CN120429019APending Publication Date: 2025-08-05ARM TECH CHINA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510519927.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The processor experiences continuous cache misses during the execution of instructions, resulting in reduced execution efficiency, especially when continuous instructions in the pipeline need to obtain data from memory, which increases the processor's waiting time.

Method used

After obtaining the first memory data, the processor stores it in the buffer and simultaneously sends a second memory data request in the process, reducing the waiting time and thus improving the execution efficiency of the processor.

Benefits of technology

By reducing the time loss of continuous cache misses, the processor's data acquisition speed in continuous cache misses is improved, and the processor's execution efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429019A_ABST
    Figure CN120429019A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a data processing method, a readable storage medium, a program product and electronic equipment. The data processing method is applied to the electronic equipment, the electronic equipment comprises a processor, if continuous cache miss occurs in the process that the processor executes a first instruction and a second instruction which are continuous, the processor sends a first message for obtaining first data of the first instruction to a memory, and in the process that the memory returns the first memory data corresponding to the first message, the processor can also send a second message for obtaining the second data of the second instruction to the memory. And after the processor acquires the first memory data, acquiring second memory data corresponding to the second message. Therefore, the processor only needs to set one buffer area to store the memory data acquired from the memory, so that the time loss of the cache miss of the processor is reduced, and the speed of acquiring the data from the memory by the processor under the continuous cache miss is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, a readable storage medium, a program product, and an electronic device. Background Art

[0002] Currently, when a processor in an electronic device executes an instruction, if the data required to execute the instruction is in the processor's cache, it is considered a cache hit, and the processor can directly retrieve the instruction from the cache. Conversely, if the data required to execute the instruction is not stored in the cache, it is called a cache miss. In this case, the processor needs to retrieve the data from the memory, which takes a long time.

[0003] In some cases, when the processor is executing instructions, multiple instructions may miss the cache consecutively. The processor needs to obtain the data of one instruction from the memory, and then obtain the data of the next instruction from the memory, which increases the time the processor takes to execute instructions, thereby reducing the efficiency of the processor in executing instructions. Summary of the Invention

[0004] Embodiments of the present application provide a data processing method, a readable storage medium, a program product, and an electronic device.

[0005] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to an electronic device, wherein the electronic device includes a processor and a first memory, wherein the processor includes a first buffer and a cache. The method includes: the processor of the electronic device detects a first request to obtain first data corresponding to a first instruction from the first memory, and the first request includes a first address of the first data. The processor sends a first address to the first memory, obtains first memory data from the first memory based on the first address, the first memory data includes the first data, and stores the first memory data in the first buffer. Before the processor obtains the first memory data, the processor detects a second request to obtain second data corresponding to a second instruction from the first memory, and the second request includes the second address of the second data. The processor sends a second address to the first memory, and after the processor stores the first memory data in the first buffer into the cache, the processor obtains second memory data from the first memory based on the second address, the second memory data includes the second data, and stores the second memory data in the first buffer.

[0006] In some embodiments of the present application, when the processor of the electronic device is executing the first instruction, if it detects a first request to obtain the first data of the first instruction from the first memory, the processor may send the first address of the first data to the first memory. The first memory may then return the first memory data at the first address to the processor. It is understandable that when the first memory returns data to the processor, it returns it at the granularity of one memory data, and the first memory data is a memory data including the first data. After the processor obtains the first memory data, it may store the first memory data in a buffer. Before the processor obtains all the first memory data, if the processor detects a second request to obtain the second data of the second instruction from the first memory, the processor may send the second address of the second data in the second request to the first memory. In this way, while the first memory returns the first memory data to the processor, the processor may send the second address to the first memory and the first memory may prepare to return the data corresponding to the second address. This can reduce the time for the first memory to return the second data to the processor, thereby improving the efficiency of the processor in executing instructions.

[0007] Secondly, compared with some implementation schemes, by setting multiple first buffers in the processor, the processor can simultaneously obtain corresponding data of multiple addresses from the first memory, but this solution requires the processor to be configured with more register resources, resulting in higher processor costs.

[0008] In this solution, after the processor has retrieved all of the first memory data and stored the first memory data from the first buffer into the cache, the processor can retrieve the second memory data including the second data. In this way, only one first buffer needs to be set up in the processor, eliminating the need to consume additional register space to set up the first buffer, thus ensuring low processor costs.

[0009] In a possible implementation of the first aspect above, the processor sends a first address to the first memory, and obtains first memory data from the first memory based on the first address, including: the processor sends the first address to the first memory, obtains the first data from the first memory, and then obtains data other than the first data from the first memory.

[0010] In some embodiments of the present application, since the processor retrieves data from the first memory at the granularity of a single memory data, when retrieving the first memory data, in order to ensure the execution of the first instruction, the processor may prioritize retrieving the first data from the first memory and then retrieve the data other than the first data from the first memory data. In this way, the processor can retrieve the first data earlier, so as to facilitate the execution of the first instruction.

[0011] In a possible implementation of the first aspect, the processor detecting a first request to retrieve first data corresponding to a first instruction from a first memory includes: the processor executing the first instruction in a first stage of a pipeline, and the processor detecting that the first data corresponding to the first instruction is not located in a cache of the processor. The processor detecting the first request to retrieve the first data corresponding to the first instruction from the first memory.

[0012] In some embodiments of the present application, the processor may execute the first instruction based on a pipeline, and may detect in the first stage of the pipeline that the first data required by the first instruction is not located in the cache of the processor, and the processor may detect the first request.

[0013] In a possible implementation of the first aspect, the second instruction is executed after the first instruction. The processor detects a second request to retrieve second data corresponding to the second instruction from the first memory before completing the acquisition of the first memory data, including: after the processor acquires the first data from the first memory, the first stage of the processor's pipeline completes the execution of the first instruction, and before the processor completes the acquisition of the first memory data, the processor executes the second instruction on the first stage of the pipeline, and the processor detects that the second data is not in the processor's cache. The processor detects the second request to retrieve the second data corresponding to the second instruction from the first memory.

[0014] In some embodiments of the present application, when the processor is executing the first instruction based on the pipeline, the pipeline enters a pause phase because the processor needs to obtain the first data from the first memory when the first instruction is executed in the first stage. The pipeline will not continue to run until the processor obtains the first data from the first memory and the first stage of the pipeline completes the execution of the first instruction. However, the processor still needs to continue to obtain the first memory data from the first memory. If the second instruction runs to the first stage of the pipeline before the processor obtains all the first memory data, and it is detected in the first stage that the second data required by the second instruction is not in the processor's cache, the processor can detect the second request.

[0015] In a possible implementation of the first aspect, the second instruction is executed after the first instruction. The processor detects a second request to retrieve second data corresponding to the second instruction from the first memory before completing the acquisition of the first memory data, including: after the processor sends the first address to the first memory, the processor executes the second instruction based on the second stage of the pipeline, and the processor detects in the second stage that the second data of the second instruction is not located in the processor's cache. The processor detects the second request to retrieve the second data corresponding to the second instruction from the first memory.

[0016] In some embodiments of the present application, a processor pipeline includes a second stage, and the second stage precedes the first stage. The second stage can also detect whether the data corresponding to the instruction running in the stage is located in the cache. Because a cache miss occurs when the processor runs the first instruction in the first stage of the pipeline, the processor needs to obtain the first data corresponding to the first instruction from the first memory. Therefore, the pipeline is in a paused state. If the second instruction running in the second stage detects that the data required by the second instruction is not in the cache, the processor can also detect a second request to obtain the second data from the first memory.

[0017] In a possible implementation of the first aspect, the processor further includes a memory subsystem, and the cache and the first buffer are arranged in the memory subsystem.

[0018] In some embodiments of the present application, a processor of an electronic device obtains data from a first memory through a memory controller, and temporarily stores the obtained data in a first buffer. After obtaining memory data of a granularity, the memory controller can store the data in the first buffer in a cache so that the memory controller can obtain data from the memory next time.

[0019] In a possible implementation of the first aspect, the size of the first buffer is the same as the size of the space occupied by the first memory data.

[0020] In some embodiments of the present application, the size of the first buffer is equal to the size of the space occupied by memory data of one granularity in the first memory. The first memory data is memory data of one granularity in the first memory. Therefore, the size of the first buffer can be equal to the size of the space occupied by the first memory data. In this way, the space occupied by the first buffer can be reduced, ensuring miniaturization of the processor.

[0021] In a second aspect, the present application provides an electronic device comprising: a memory for storing instructions; and at least one processor for executing the instructions to cause the device to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the second aspect can be referenced to the beneficial effects of the method provided in any embodiment of the first aspect and will not be further elaborated here.

[0022] In a third aspect, the present application provides a computer-readable storage medium storing instructions that, when executed by a device, cause a computer to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achieved in the third aspect can be referenced to the beneficial effects of the method provided in any embodiment of the first aspect and are not further elaborated here.

[0023] In a fourth aspect, the present application provides a computer program product that, when executed on a device, causes the device to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achieved in the fourth aspect can be referenced to the beneficial effects of the method provided in any embodiment of the first aspect and will not be further elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of an electronic device entering a screen-on state is shown;

[0025] Figure 2 The operation process of a five-stage pipeline is shown;

[0026] Figure 3 A schematic structural diagram of an electronic device is shown;

[0027] Figure 4A It shows an interactive flow chart of MSS getting data from memory;

[0028] Figure 4B A schematic diagram of storing data in a memory is shown;

[0029] Figure 5 Another interactive flow chart of MSS getting data from memory is shown;

[0030] Figure 6 According to an embodiment of the present application, an interactive flow chart of a data processing method is shown;

[0031] Figure 7 According to some embodiments of the present application, an interactive flow chart of a processor obtaining data from a first memory is shown;

[0032] Figure 8 According to some embodiments of the present application, a schematic diagram of a processor acquiring data is shown;

[0033] Figure 9 According to some embodiments of the present application, a schematic structural diagram of an electronic device is shown. DETAILED DESCRIPTION

[0034] Illustrative embodiments of the present application include, but are not limited to, a data processing method, a readable storage medium, a program product, and an electronic device.

[0035] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0036] As described in the background art, if cache misses occur for multiple consecutive instructions during the process of instruction execution by a processor of an electronic device, the efficiency of instruction execution by the processor of the electronic device will be greatly reduced.

[0037] To facilitate understanding, some of the terms and related technologies involved in this application are explained below.

[0038] 1. Pipeline.

[0039] Processors can generally be divided into three-stage pipelines and five-stage pipelines, depending on their design and structure. A five-stage pipeline divides the instruction execution process into stages such as instruction fetch, decode, execute, memory access, and writeback. Each stage is handled by a dedicated hardware unit. Once an instruction enters the pipeline, its different stages will sequentially occupy different hardware units, forming a pipeline process. Therefore, in the same clock cycle, a five-stage pipeline can execute up to five instructions simultaneously. Therefore, by processing each stage of an instruction in parallel, the pipeline of a processor (such as a central processing unit (CPU)) can complete more computing tasks within the same clock cycle, thereby improving computing efficiency. However, in a pipeline, if the instruction that is executed earlier in the order stops at the corresponding stage, the instruction that is executed later in the order cannot continue, resulting in reduced pipeline execution efficiency. For example, instruction 1 stops at the memory access stage. At this time, instruction 2, which is executed after instruction 1 in the execution order, cannot enter the memory access stage during the execution stage, and instruction 3 cannot enter the execution stage during the decoding stage. Therefore, the entire pipeline enters a pause state, which reduces the efficiency of the pipeline's instruction execution.

[0040] The following describes the execution process of a load instruction at each stage of a five-stage pipeline, taking the processor executing a load instruction as an example.

[0041] Instruction Fetch: The processor fetches the next instruction from the instruction cache (e.g., a load instruction). This stage involves incrementing the program counter to point to the next instruction and reading that instruction from memory.

[0042] Decode: The processor decodes the instruction, determines the operation to be performed, and reads the operands from the register bank. For load instructions, this stage may also include resolving the memory address and preparing for the memory access operation.

[0043] Execute: During the execute phase, for load instructions, the ALU may be involved in calculating the effective address (if the instruction contains indirect addressing). However, the actual memory access operation is usually performed in the next phase.

[0044] Memory Access: During the Memory Access phase, the processor can read the data required by the load instruction from memory. This typically involves accessing a data cache or main memory using the effective address calculated during the execute phase. In some embodiments, if the data required by the load instruction is located in a cache within the processor, the processor may prioritize obtaining the data required by the load instruction from the cache.

[0045] Write Back: Finally, in the Write Back phase, the data read from memory is written to the specified register. This phase completes the execution of the load instruction.

[0046] 2. Cache.

[0047] Cache refers to storage that allows for high-speed data exchange. It exchanges data with the processor before the main memory, resulting in a very fast data exchange rate. Cache configuration is one of the key factors in achieving high performance in computer systems.

[0048] Cache is based on the principle of program locality, which means that programs often access certain data or instructions repeatedly during execution. By storing this data or instructions in cache, the processor can access them faster, reducing the number of memory accesses and thus improving system performance.

[0049] In computer systems, caches are usually divided into multiple levels to provide different levels of performance optimization. Common cache hierarchies include:

[0050] Level 1 cache (L1 cache): The cache closest to the processor, typically divided into two parts: data cache and instruction cache. L1 cache has very fast access speeds but a relatively small capacity.

[0051] Second level cache (L2 cache): Slightly larger than L1 cache and slightly slower to access, but still much faster than RAM. L2 cache is usually integrated into the CPU or located on a separate chip.

[0052] Level 3 cache (L3 cache): This has a larger capacity and relatively slower access speed, but is still faster than RAM. L3 cache is commonly used in multi-core CPUs to improve data exchange efficiency between processors by sharing the cache.

[0053] For example, the access speeds of caches such as L1 cache, L2 cache, and L3 cache decrease in sequence, but their capacities increase in sequence. When the processor needs to read data, it first searches the L1 cache, then the L2 cache, and finally the L3 cache. If the data exists in any of the first-level caches (i.e., a cache hit), the processor can read the data directly from the cache without accessing memory. This can significantly increase the speed of data reads.

[0054] However, if the data is not in the cache (a cache miss), the processor needs to access the memory to read the data. The memory access speed is relatively slow, so this will increase the waiting time of the processor.

[0055] Since the access speed in the L1 cache is the fastest, if the processor misses the L1 cache, it needs to search the L2 cache for the data, which increases the clock cycle for the processor to access the data. Similarly, if the processor misses the L2 cache, it can continue to access the data in the next level of cache until it accesses the data from the main memory.

[0056] 3. Memory subsystem (MSS).

[0057] The memory subsystem is a core component of computer hardware architecture, providing high-speed data transfer and temporary storage between the processor and external storage devices (such as hard drives and solid-state drives). The memory subsystem typically includes cache, a memory controller, a memory management unit, buffers, and related bus interfaces.

[0058] 4. Advanced microcontroller bus architecture advanced eXtensible interface (AMBA AXI).

[0059] The AMBA AXI protocol (hereinafter referred to as the AXI protocol or AXI bus) is an on-chip bus designed for high performance, high bandwidth, and low latency. This protocol has several notable features, including separation of address / control and data phases, support for unaligned data transfers, requiring only the first address for burst transfers, separate read and write data channels, and support for out-of-order access. These features make the AXI protocol widely used in high-performance system designs.

[0060] The following takes loading instructions as an example to introduce the process of electronic devices executing instructions.

[0061] For example, Figure 1 A schematic diagram of an electronic device entering a screen-on state is shown.

[0062] It should be noted that the electronic equipment in the embodiments of the present application may also be referred to as a terminal (termina1), a user terminal, a mobile terminal, a user equipment (UE), a terminal device, a mobile station, a mobile terminal (mobile terminal, MT), etc. The terminal device may be a mobile phone, a smart TV, a wearable device, a tablet computer (Pad), a computer with a wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a terminal in industrial control (industrial control), a terminal in self-driving (self driving), a terminal in a smart grid (smart grid), a wireless terminal in transportation safety (transportation safety), a terminal in a smart city (smart city), a terminal in a smart home (smart home), and other terminal devices. The following will be described using a mobile phone as an example. However, it will be understood that the technical solution described in the present application is applicable to various electronic devices that need to load data when the above-mentioned processor executes instructions, and is not limited to mobile phones.

[0063] For example, the process of executing the loading instruction by the processor of the electronic device 100 is described below by taking the electronic device 100 waking up the display screen as an example.

[0064] like Figure 1 As shown, after the electronic device 100 detects the operation of waking up the display screen, it enters the screen-on state and then displays the desktop M1 and the icons of various application programs on the desktop M1.

[0065] In this process, when the electronic device 100 needs to display the desktop M1, it needs to read data related to the image displayed on the desktop M1 from the memory. The data may include pixel information, color information, size information, etc. of the image.

[0066] Then, the processor of the electronic device 100 loads these data from the memory into the register or cache of the processor by executing the load instruction, so as to perform subsequent processing and rendering on the image-related data of the desktop M1, thereby displaying the desktop M1.

[0067] For example, Figure 2 The operation process of a five-stage pipeline is shown.

[0068] The following describes the process of electronic devices executing instructions.

[0069] For example, Figure 2 As shown, the five-stage pipeline includes instruction fetch, decode, execute, memory access, and write back stages.

[0070] In the clock cycle T0, the load instruction Load0 enters the instruction fetch stage. In the clock cycle T0, the processor of the electronic device retrieves the address of the load instruction Load0 from the instruction cache and reads the load instruction Load0 from the memory according to the address of the load instruction Load.

[0071] During clock cycle T1, load instruction Load0 enters the decode phase, and load instruction Load1 enters the instruction fetch phase. During this clock cycle, the processor decodes load instruction Load0, resolves its memory address and prepares for memory access, and then fetches load instruction Load1.

[0072] During clock cycle T2, load instruction Load0 enters the execute phase, load instruction Load1 enters the decode phase, and load instruction Load2 enters the instruction fetch phase. During this clock cycle, the processor can calculate the effective address of load instruction Load0, decode load instruction Load1, and fetch load instruction Load2.

[0073] During clock cycle T3, Load0 enters the memory access phase, Load1 enters the execution phase, Load2 enters the decoding phase, and Load3 enters the instruction fetch phase. During this clock cycle, the processor can retrieve the data to be loaded by Load0 based on its address, execute Load1, decode Load2, and fetch Load3.

[0074] It can be understood that during the memory access phase of the load instruction Load0, the processor can first confirm whether the data required by the load instruction Load0 is located in the cache based on the address of the load instruction. If the data corresponding to the load instruction Load0 is located in the cache, it means a cache hit, and the processor can quickly obtain the data corresponding to the load instruction Load0 from the cache, thereby completing the memory access of the load instruction Load0. However, if the data of the load instruction Load0 is not in the cache, it means a cache miss, and the processor needs to obtain the data of the load instruction Load0 from the memory based on the address of the load instruction Load0, which takes a long clock cycle.

[0075] For example, continue to refer to Figure 2 If a cache miss occurs in the load instruction Load0, the load instruction Load0 will be in the memory access stage from the clock cycle T3 to the clock cycle Tn-1, causing the pipeline to pause for n-3 clock cycles.

[0076] In the clock cycle of Tn, after the processor obtains the data of the load instruction Load0 from the memory, the load instruction Load0 enters the write-back stage. At this time, the load instruction Load1 enters the memory access stage. If the data of the load instruction Load1 also has a cache miss, the processor still needs to spend a longer clock cycle to obtain the data of the load instruction Load1 from the memory. For example, it takes Tn+m clock cycles for the processor to obtain the data of the load instruction Load1. In other words, the processor has experienced cache misses continuously.

[0077] In this case, the processor first obtains the data of the load instruction Load0 from the memory, and then obtains the data of the load instruction Load1 from the memory, thereby further reducing the efficiency of the processor in executing instructions.

[0078] Next, the structure of the electronic device is introduced.

[0079] For example, Figure 3 A structural diagram of an electronic device is shown.

[0080] like Figure 3 As shown, the electronic device 100 includes a processor 10 and a memory 20, wherein the processor 10 includes an MSS 11, an MSS

[0081] 11 includes a memory management unit 111 , a memory controller 112 , a buffer 113 and a cache 114 .

[0082] When the processor 10 executes an instruction, it generally obtains the data required by the instruction from the cache 114 of the MSS 11. If the data required by the instruction is not in the cache, that is, the processor 10 has a cache miss, the processor 10 needs to access the data in the memory 20. During this process, the MSS 11 can initiate a memory access request through the memory management unit 111. The memory management unit 111 can convert the virtual address of the data required by the instruction into a physical address and determine the access type (e.g., read or write) and size (e.g., byte, word, double word, etc.). Then, the memory management unit 111 can send the memory access request to the memory controller 112.

[0083] Memory controller 112 can receive memory access requests from memory management unit 111 via a bus interface (e.g., an AXI bus) and is responsible for processing memory access requests. For example, memory controller 112 can interact with memory 20 via AXI bus 30 to execute memory access requests. Memory controller 112 can store data retrieved from memory 20 in buffer 113. After retrieving the data, memory controller 112 can store the data in buffer 113 in cache 114.

[0084] In some embodiments of the present application, in order to better illustrate the scenario in which the processor 10 obtains data from the memory 20 through the MSS 11, some basic parameters may be set.

[0085] For example, the number of bytes of the granularity at which the memory controller 112 in the MSS11 obtains data from the memory 20 through the AXI bus 30 can be cacheline_size. It can be understood that in the memory 20, cacheable data is usually stored with the size of cacheline_size as a granularity, and when the external processor 10 obtains data from the memory 20, it also usually obtains data with the size of a granularity. That is, when the processor 10 obtains data from the memory 20, a complete access needs to obtain data of the size of cacheline_size (not read in one clock cycle). The size of cacheline_size is defined by the architecture of the processor 10. In some embodiments, cacheline_size is usually 32 bytes or 64 bytes. It can be understood that one byte is 8 bits of data, and 32 bytes are 256 bits of data. In some embodiments of the present application, the process of MSS11 obtaining data from the memory 20 can be introduced by taking the cacheline_size of 32 bytes as an example.

[0086] In some embodiments of the present application, the maximum number of bytes that can be accessed at one time by each load instruction executed by processor 10 is load_byte_n, which is at least 1 byte. For example, load_byte_n can be 4 bytes, 8 bytes, 16 bytes, etc., and the embodiments of the present application do not limit the size of load_byte_n. In some embodiments of the present application, assuming load_byte_n is 4 bytes, the process of MSS 11 retrieving the data required by the load instruction from memory 20 when a cache miss occurs during the execution of a load instruction by processor 10 is described.

[0087] In some embodiments of the present application, the number of bytes that the AXI bus 30 can transmit in one clock cycle can be trans_data_size. In some embodiments, trans_data_size is generally the same as load_byte_n. In some embodiments of the present application, trans_data_size is also 4 bytes.

[0088] Therefore, when a cache miss occurs in the processor 10 and the processor retrieves data from the memory 20, the minimum number of cycles required to complete a complete access to the memory 20 is trans_t. Here, trans_t = cacheline_size / trans_data_size. That is, the size of the data cacheline_size that the processor 10 needs to retrieve from the memory 20 to complete a complete access is divided by the size of the data trans_data_size that the AXI bus 30 can transmit in one clock cycle. It can be understood that when the cacheline_size is 32 bytes and the trans_data_size is 4 bytes, trans_t = 8. In other words, the processor 10 needs 8 clock cycles to transfer the data required for the load instruction from the memory 20 to the buffer 113. In addition, the processor 10 also needs to spend alloc_t clock cycles to store the data in the buffer 113 in the cache 114. In some embodiments, alloc_t can be 4 clock cycles.

[0089] Next, a process is described in which the MSS 11 obtains the data of the load instructions Load0 and Load1 from the memory 20 after a cache miss occurs during the execution of the load instructions Load0 and Load1 by the processor 10 .

[0090] For example, Figure 4A An interactive flow chart of MSS 11 acquiring data from memory 20 is shown.

[0091] like Figure 4A As shown, the process includes:

[0092] S401 , MSS11 sends the first address of the data corresponding to Load0 to the memory 20 .

[0093] For example, in some embodiments, after detecting that a cache miss occurs when the processor 10 executes the load instruction Load0, MSS11 can obtain the first address of the data corresponding to the load instruction Load0 through the memory management unit 111 and parse the first address. Then, the memory management unit 111 can generate an access request and send the access request to the memory controller 112 for the memory controller 112 to process the access request. The access request can, for example, include the first address. The memory controller 112 can send the access request to the memory 20 through the AXI bus 30 to obtain the data corresponding to the load instruction Load0 from the memory 20.

[0094] For example, referring to Figure 4BFor example, the first address of the data corresponding to the load instruction Load0 starts from byte 0 of the memory 20 and ends at byte 3, that is, the data corresponding to the load instruction Load0 is 4 bytes in total, load_byte_n=4.

[0095] In some embodiments of the present application, the size of the first address is, for example, smaller than trans_data_size. That is, the AXI bus 30 can transmit the first address to the memory 20 within one clock cycle at the fastest. Therefore, the clock cycle address_t occupied by the memory controller 112 in the MSS 11 sending the first address to the memory 20 via the AXI bus 30 can be at most one clock cycle.

[0096] S402 , the memory 20 returns the first memory data corresponding to the first address to the MSS 11 .

[0097] For example, in some embodiments of the present application, after receiving the first address, memory 20 may determine, based on the first address, the first memory data to be returned to MSS 11. In some embodiments, after receiving a transaction that returns the data at the first address, memory 20 needs to spend a certain response time, response_t, to respond to the transaction. The response time of memory 20 may be 5 to 10 clock cycles. In some embodiments of the present application, the response time of memory 20 may be 8 clock cycles. In other words, the response time from memory 20 receiving the transaction that returns the data at the first address to providing the data is response_t = 8.

[0098] It is understood that the initial position of the first address is 0 bytes, that is, the first memory data is the data in the first cacheline_size in the memory 20. In some embodiments of the present application, the cacheline_size is 32, that is, the first memory data is the first 32 bytes of data in the memory 20, that is, the data from bytes 0 to 31. It is understood that the processor 10 generally obtains data from the memory 20 at a granularity of one cacheline_size. Therefore, even if the number of bytes of data corresponding to the first address is 4, the processor 10 still needs to obtain 32 bytes of data from the memory 20.

[0099] It is understood that the AXI bus 30 can transmit 4 bytes of data in one clock cycle, so the number of clock cycles required to transmit 32 bytes of data is trans_t=8. That is, the memory 20 takes 8 clock cycles to return the first memory data to the MSS 11, and the MSS 11 stores it in the corresponding buffer 113.

[0100] For example, take the load instruction Load0 as an example, refer to Figure 4B , the data required by the load instruction Load0 is data from 0 bytes to 3 bytes, that is, the memory 20 can send the data required by the load instruction Load0 to the processor 10 in the first clock cycle. Therefore, the load instruction Load0 actually only pauses for address_t+response_t+1 clock cycles in the pipeline. Afterwards, the load instruction Load1 enters the execution stage, but because the memory 20 still needs to send the data from the 4th byte to the 31st byte to MSS11, the buffer 113 in MSS11 is still in an occupied state at this time. Therefore, if a cache miss occurs in the execution stage of the load instruction Load1, MSS11 cannot immediately send the second address of the data corresponding to the load instruction Load1 to the memory 20 because the previous transaction has not been completed (the transaction to obtain the first memory data), that is, the load instruction Load1 enters a paused state in the pipeline.

[0101] S403 , the MSS 11 stores the first memory data in the cache 114 .

[0102] For example, in some embodiments of the present application, after acquiring the first memory data, MSS11 may store the first memory data in cache 114 .

[0103] For example, MSS 11 also needs to store the first memory data from buffer 113 into cache 114, a process that takes alloc_t clock cycles. At this point, MSS 11 retrieves the first memory data from memory 20 and stores it in cache 114, taking a total of address_t + response_t + trans_t + allloc_t clock cycles. Afterward, MSS 11 can execute the next message to retrieve data from memory 20.

[0104] S404 , MSS11 sends the second address of the data corresponding to Load1 to the memory 20 .

[0105] For example, similar to the process of S401 , the clock cycle occupied by MSS11 sending the second address to the memory 20 is also, for example, address_t, and the response time of the memory 20 is also response_t.

[0106] S405 , the memory 20 returns the second memory data corresponding to the second address to the MSS 11 .

[0107] For example, the second address starts at byte 36 of memory 20. That is, the data in load instruction Load1 is from bytes 36 to 39 in memory 20, a total of 4 bytes. Therefore, similar to the process in S402, the clock cycles occupied by memory 20 returning the second memory data to MSS 11 are also, for example, trans_t. However, memory 20 can return the data from load instruction Load1 to MSS 11 in two clock cycles. That is, it takes one clock cycle to return bytes 32 to 35 to MSS 11, and then another clock cycle to return bytes 36 to 39 to MSS 11. Therefore, the clock cycle in which the processor 10 actually obtains the data of the load instruction Load1 is: (trans_t-1)+alloc_t+address_t+response_t+2, that is, the clock cycle (trans_t-1) in which the memory 20 continues to return the first memory data to MSS11 after the processor 10 obtains the data of the first address, the clock cycle alloc_t in which the memory controller 112 in MSS11 stores the first memory data in the buffer 112 in the cache 114, the clock cycle address_t in which the memory 20 obtains the second address, the clock cycle response_t in which the memory 20 responds to the data returned at the second address, and the two clock cycles in which the memory 20 returns the data of the second address.

[0108] S406 , the MSS 11 stores the second memory data in the cache 114 .

[0109] Exemplarily, similar to the process of S403 , MSS11 spends alloc_t clock cycles to store the second memory data in the cache 114 .

[0110] Afterwards, MSS 11 may execute the next message to obtain data from memory 20 .

[0111] It can be understood that when cache misses occur when the processor 10 executes the load instruction Load0 and the load instruction Load1, the processor 10 needs to pause for a total of (address_t+1+response_t)+((trans_t-1)+alloc_t+address_t+response_t+2)=2×address_t+2×response_t+trans_t+alloc_t+2 clock cycles. When address_t=1, response_t=8, trans_t=8, and alloc_t=4, the pipeline of the processor 10 needs to pause for 32 clock cycles. It can be understood that the more consecutive cache misses occur, the longer the clock cycle the processor 10 is paused, resulting in low efficiency in executing instructions by the processor 10.

[0112] It can be understood that the above 32 clock cycles are the most ideal situation. If the data of the load instruction Load0 and the load instruction Load1 are both at the end of different cacheline_sizes in the memory 20, the memory 20 will need to spend at most two trans_t clock cycles to return the data of the load instruction Load0 and the load instruction Load1 to the processor 10, that is, the processor 10 needs to pause for a total of 2×address_t+2×response_t+2×trans_t+alloc_t=38 clock cycles at most.

[0113] In other embodiments, if the processor 10 encounters continuous cache misses, in order to improve the efficiency of the processor 10 in obtaining data from the memory 20, the processor 10 may also continuously send multiple requests to obtain data to the memory 20, and allocate more buffers 113 to store the data obtained from the memory 20 by each request.

[0114] For example, Figure 5 Another interactive flow chart of MSS 11 acquiring data from memory 20 is shown.

[0115] like Figure 5 As shown, the process includes:

[0116] S501 , MSS11 sends the first address of the data corresponding to Load0 to the memory 20 .

[0117] For example, similar to the process of S401 , the clock cycle occupied by MSS11 sending the first address to the memory 20 is also, for example, address_t.

[0118] S502 , MSS11 sends the second address of the data corresponding to Load1 to the memory 20 .

[0119] For example, in some embodiments of the present application, MSS 11 may send the second address corresponding to the data of load instruction Load2 to memory 20 without waiting for load instruction Load0 to obtain the corresponding data.

[0120] For example, refer to Figure 2 In some embodiments, the execution stage in the pipeline of the processor 10 can also determine whether the corresponding instruction has a cache miss. For example, when the load instruction Load0 obtains the corresponding data in the memory access stage, the processor 10 has a cache miss, and the processor 10 needs to obtain the data of the load instruction Load0 from the memory 20. The load instruction Load0 is paused in the memory access stage, causing the pipeline to enter a stagnant state. At this time, the load instruction Load1 is in the execution stage. If the processor 10 detects that the load instruction Load1 also has a cache miss in the execution stage of the pipeline, the MSS11 can send the second address of the data corresponding to Load1 to the memory 20 without waiting for all the first memory data to be placed in the cache 114.

[0121] Therefore, the clock cycle taken by memory 20 to return the first memory data to MSS11 overlaps with the clock cycle taken by MSS11 to send the second address to memory 20, and the time taken by memory 20 to respond to the data returned from the first address overlaps with the event of responding to the data from the second address. In this way, the clock cycle in which the processor 10 is paused when cache misses occur continuously can be reduced, thereby improving the efficiency of the processor 10 in executing instructions.

[0122] S503 , the memory 20 returns the first memory data corresponding to the first address and the second memory data corresponding to the second address to the MSS 11 .

[0123] For example, in some embodiments of the present application, after receiving the second address, the memory 20 may return the second memory data corresponding to the second address to the MSS 11. The memory 20 needs to return 64 bytes of data to the MSS 11, that is, the memory 20 returns the first memory data and the second memory data, which takes a total of two trans_t clock cycles.

[0124] Furthermore, since MSS11 continuously sends two requests for obtaining data to memory 20, in order to ensure that MSS11 can receive the first memory data and the second memory data, more buffers 113 need to be set in MSS11 to receive the first memory data and the second memory data.

[0125] It can be understood that in the process of memory 20 returning data to MSS11, an address identifier (address read ID, ARID) can be added to the transaction of the first address of the data corresponding to Load0 sent by MSS11. For example, the address identifier of the first address of the data corresponding to Load0 is R0, and the address identifier of the first address of the data corresponding to Load1 is R1. Memory 20 can process transactions corresponding to R0 and R1 at the same time. For example, in the process of simultaneously processing transactions of R0 and R1, memory 20 returns data to MSS11 not by returning the first memory data and then returning the second memory data, but the 4 bytes of data returned by memory 20 to MSS11 in each clock cycle may be data in the first memory data or data in the second data, and a total of 64 bytes of data can be returned. After receiving the data, MSS11 can store the first memory data and the second memory data in different buffers 113 respectively.

[0126] Therefore, through the above processing method, the memory 20 can also return the data of the load instruction Load1 to the MSS11 in advance. Figure 4BAfter memory 20 obtains the first address and completes its response to the first address, it can send the data required by load instruction Load0 to processor 10 in the first clock cycle, that is, return the data from bytes 0 to 3 in memory 20 (the first four bytes of data in the first memory data). Therefore, load instruction Load0 actually only pauses in the pipeline for address_t+response_t+1 clock cycles. Moreover, while returning the data required by load instruction Load0 to MSS11, memory 20 receives the second address of load instruction Load1 and completes its response to the second address. In other words, the clock cycle address_t in which memory 20 receives the second address and the clock cycle response_t in which memory 20 responds to the second address coincides with the clock cycle in which processor 10 sends the first address of load instruction Load0 to memory 20, the clock cycle in which memory 20 responds to the first address, and the clock cycle in which memory 20 returns the data required by load instruction Load0 to MSS11, that is, the clock cycle in which memory 20 receives the second address coincides with the clock cycle in which processor 10 sends the first address of load instruction Load0 to memory 20, the clock cycle in which memory 20 responds to the first address, and the clock cycle in which memory 20 returns the data required by load instruction Load0 to MSS11, that is, the clock cycle in which memory 20 returns the data required by load instruction Load0 to MSS11 coincides with address_t+response_t+1. Therefore, in the most ideal state, after memory 20 returns the data required by load instruction Load0 to MSS11, it can return the second memory data corresponding to the second address. For example, in the second clock cycle when memory 20 returns data to MSS11, it returns the data from bytes 32 to 35 to MSS11 (corresponding to the first 4 bytes of the second memory data). In the third clock cycle, memory 20 returns the data from bytes 36 to 39 to MSS11 (corresponding to the data required by load instruction Load1 in the second memory data). In other words, within the fastest three clock cycles, processor 10 can obtain the data of load instructions Load0 and load instructions Load1.

[0127] It can be understood that cache misses occurred continuously during the execution of the load instructions Load0 and Load1 by the processor 10, but a total of one address_t and one response_t clock cycle was spent, and the memory 20 transmitted data to the processor 10 for 3 clock cycles, that is, the processor 10 paused for a total of address_t+response_t+3 clock cycles. When address_t=1 and response_t=8, the pipeline of the processor 10 needs to be paused for at least 12 clock cycles.

[0128] It can be understood that the above-mentioned address_t+response_t+3 clock cycles are the most ideal situation. If the data of the load instruction Load0 and the load instruction Load1 are both at the end of different cacheline_sizes in the memory 20, the memory 20 will need to spend at most two clock cycles of trans_t to return the data of the load instruction Load0 and the load instruction Load1 to the processor 10, that is, the processor 10 needs to pause for a total of at most address_t+response_t+2×trans_t clock cycles. When trans_t=8, the pipeline of the processor 10 may need to be paused for at most 25 clock cycles.

[0129] It can be understood that although the processor 10 obtains the data of the load instruction Load0 and the load instruction Load1, the memory 20 will still continue to return the first memory data and the other data in the second memory data to the MSS 11.

[0130] In some embodiments, during the memory access phase, the pipeline of processor 10 can determine whether a cache miss will occur when processor 10 retrieves instructions for that phase. For example, if a cache miss occurs during the memory access phase for load instruction Load0, the pipeline will enter a pause before processor 10 retrieves the data required for load instruction Load0. After the aforementioned address_t+response_t+1 clock cycles, processor 10 retrieves the data required for load instruction Load0, and load instruction Load0 enters the next pipeline phase or exits the pipeline. Then, load instruction Load1 enters the memory access phase. Processor 10 determines that the data required for load instruction Load1 is also not in the cache, meaning that a cache miss also occurred during the execution of load instruction Load1 by processor 10. Therefore, processor 10 can send the second address of the data corresponding to load instruction Load1 to memory 20 via MSS11 as soon as address_t+response_t+1 clock cycles later. It can be understood that during this process, MSS11 of processor 10 has not yet fully retrieved the first memory data corresponding to the first address from memory 20. After memory 20 obtains the second address, it can respond to the second address and then return the second memory data corresponding to the second address to MSS 11. At this time, if memory 20 has not yet returned all the first memory data to MSS 11, the data returned by memory 20 to MSS 11 within one clock cycle may be data from the first memory data or data from the second memory data. After memory 20 returns the data required by load instruction Load1 to MSS 11, load instruction Laod1 can complete execution in the memory access stage and the pipeline can continue to run.

[0131] S504 , the MSS 11 stores the first memory data and the second memory data in the cache 114 .

[0132] For example, within two clock cycles of trans_t, MSS11 stores the first memory data and the second memory data in different buffers 113. Afterwards, MSS11 may continue to spend two clock cycles of alloc_t to store the data in buffer 113 in cache 114.

[0133] It can be understood that the processor 10 has a cache miss situation in the process of executing the load instruction Load0 and the load instruction Load1, but the shortest total time it takes is address_t+response_t+3 clock cycles, which is relatively Figure 4A In the embodiment, the clock cycles spent are 2×address_t+2×response_t+trans_t+alloc_t+2, which reduces the clock cycles of address_t+response_t+trans_t+alloc_t-1. However, an additional buffer 113 needs to be added to the MSS11 of the processor 10. It can be understood that if the processor 10 needs to continuously send more transaction requests to the memory 20 to obtain data, more buffers 113 need to be set up in the processor 10 to store the data corresponding to each transaction request. If the processor 10 increases the number of buffers 113 in order to reduce the processing time of cache misses, that is, the processor 10 needs to configure more register resources as buffers 113, thereby increasing the cost of the processor 10.

[0134] As mentioned above, when the processor encounters continuous cache misses during instruction execution, the clock cycle for the processor to obtain data from the memory is longer, resulting in reduced efficiency in executing instructions by the processor.

[0135] To solve the problem of low efficiency in executing instructions in an electronic device, the present application proposes a data processing method. The electronic device includes a processor and a first memory, wherein the processor includes a first buffer (e.g., a buffer). The method includes:

[0136] The processor detects a first request to obtain first data corresponding to a first instruction from a first memory, where the first request includes a first address of the first data.

[0137] The processor sends a first address to the first memory, obtains first memory data from the first memory based on the first address, the first memory data including the first data, and stores the first memory data in a first buffer.

[0138] Before the processor completes acquiring the first memory data, it detects a second request for acquiring second data corresponding to a second instruction from the first memory, where the second request includes a second address of the second data.

[0139] The processor sends a second address to the first memory, and after storing the first memory data in the first buffer into the cache, obtains second memory data from the first memory based on the second address, the second memory data including the second data, and stores the second memory data in the first buffer.

[0140] Through the above solution, before the first processor obtains all the data of the first memory data, if the first processor detects that the second instruction also has a cache miss, the first processor can send a second request for obtaining the second data corresponding to the second instruction to the first memory, thereby saving the time for the first processor to send the second request to the first memory and reducing the time for the processor to obtain data from the first memory when the processor has continuous cache misses. In addition, after the first processor sends the data in the first buffer (for example, stores the data in the first buffer in the processor's cache), it obtains the second memory data. In this way, the first processor of the electronic device does not need to set up more first buffers to store the data obtained from the first memory, thereby ensuring the size of the cache space in the first processor and reducing the occurrence of cache misses in the processor.

[0141] Next, the data processing method in the implementation of this application is introduced.

[0142] Figure 6 According to an embodiment of the present application, an interactive flow chart of a data processing method is shown.

[0143] like Figure 6 As shown, the process includes:

[0144] S601: A processor detects a first request for obtaining first data corresponding to a first instruction from a first memory, where the first request includes a first address of the first data.

[0145] In some embodiments of the present application, during a pipeline-based execution of a first instruction by a processor of an electronic device, when the first instruction reaches the first stage of the pipeline (e.g., a memory access stage), the processor detects that first data required by the first instruction is not located in the processor's cache, i.e., a cache miss occurs. At this point, the processor may detect a first request to retrieve the first data corresponding to the first instruction from a first memory, the first request including a first address of the first data, so that the processor can retrieve the first data corresponding to the first instruction from the first memory, thereby completing execution of the first instruction.

[0146] For example, the processor includes an MSS, and the MSS can detect a request to obtain data from a first memory. After a cache miss occurs when the processor executes a first instruction, the MSS can detect a first request to obtain first data corresponding to the first instruction from the first memory. The first request includes a first address of the first data, so that the MSS obtains the first data from the first memory based on the first request.

[0147] S602: The processor sends a first address to the first memory, obtains first memory data from the first memory based on the first address, the first memory data including first data, and stores the first memory data in a first buffer.

[0148] For example, in some embodiments of the present application, the MSS in the processor may send a first request to the first memory to obtain the first data of the first instruction, where the request includes the first address of the first data.

[0149] For example, the first instruction is a load instruction Load0. The MSS can send a first request to the first memory via the AXI bus to obtain the first data corresponding to the load instruction Load0. The first request includes the first address of the first data. The AXI bus can add an ARID to the first request. For example, the ARID corresponding to the first request is R0. The AXI bus can process the corresponding request based on the ARID.

[0150] It can be understood that since the cacheable data in the first memory is usually stored in a granularity of cacheline_size. After the first memory obtains the first address of the first data, it can determine the position of the first data in the first memory, and use the data of cacheline_size corresponding to the position as the first memory data. The first memory can return the first memory data to the MSS through the AXI bus, and the MSS can store the first memory data in a buffer (as an example of a first buffer zone). In some embodiments, the size of the storage space of the buffer is the same as the size of the storage space occupied by the first memory data. That is, a buffer can store data of one granularity in the first memory.

[0151] In some embodiments of the present application, the size of the first data is 4 bytes, and the cacheline_size is 32 bytes. After the first memory obtains the address of the 4 bytes of data of the first data, the first memory data corresponding to the 4 bytes of data can be determined. Figure 4BFor example, if the 4-byte address of the first data starts at byte 0 of the first memory and ends at byte 3, the first memory data corresponding to the first data starts at byte 0 and ends at byte 31. In other words, the first memory data needs to return 32 bytes of data to the MSS via the AXI bus.

[0152] In some embodiments of the present application, when the first memory returns data through the AXI bus, it can start from the initial position of the first memory data and gradually return 32 bytes of data to the MSS. However, this method of returning data may cause the processor to obtain the first data at a relatively slow speed. For example, the AXI bus can return 4 bytes of data in one clock cycle. If the first data is located in the first 4 bytes of the first memory data, the AXI bus only needs to spend one clock cycle to return the first data to the processor. However, if the first data is located in the last 4 bytes of the first memory data, the AXI bus needs to spend 8 clock cycles to return the first data to the processor, which will cause the first instruction corresponding to the first data to be paused in the pipeline for a longer time.

[0153] Therefore, in some embodiments of the present application, the AXI bus supports WRAP burst type (a transmission mode) to transmit data. Through this transmission mode, the first memory can preferentially return the 4 bytes of data corresponding to the first data through the AXI bus, and then return the remaining data in the first memory data to the processor. That is, the processor sends the first address to the first memory, obtains the first data from the first memory, and then obtains the data in the first memory data other than the first data from the first memory. In this way, no matter where the first data is in the first memory data, the first memory only needs to spend one clock cycle to return the first data to the processor, thereby greatly improving the speed at which the processor obtains the first data.

[0154] S603: Before the processor obtains the first memory data, it detects a second request to obtain second data corresponding to the second instruction from the first memory, where the second request includes a second address of the second data.

[0155] For example, in some embodiments of the present application, during the process of obtaining first memory data from the first memory based on the first address, the processor also receives a second request to obtain second data corresponding to the second instruction from the first memory. The execution order of the second instruction follows the execution order of the first instruction, wherein the processor detects the second request to obtain the second data corresponding to the second instruction from the first memory before the first memory data is obtained, which may include: after the processor obtains the first data from the first memory, the first stage on the pipeline of the processor completes the execution of the first instruction, and before the processor completes the execution of the first memory data, the processor runs the second instruction on the first stage of the pipeline, and the processor detects that the second data is not located in the cache of the processor. The processor detects the second request to obtain the second data corresponding to the second instruction from the first memory.

[0156] For example, still taking the processor through the pipeline instruction as an example, refer to Figure 2 , corresponding to the clock cycle of T3, a cache miss occurs when the processor executes the load instruction Load0 (as an example of the first instruction). The MSS in the processor can send the first address of the first data corresponding to the load instruction Load0 to the first memory. The first memory returns the first memory data to the MSS according to the first address. For example, the first memory transmits data in a WRAP burst type manner, then it actually only takes 1 clock cycle for the first memory to return the first data to the MSS. In some embodiments of the present application, it also only takes 1 clock cycle for the MSS to send the first address to the first memory. The time for the first memory to respond and return the first memory data is 8 clock cycles, then it actually only takes 10 clock cycles for the processor to obtain the first data. That is to say, after 10 clock cycles, the processor obtains the first data from the first memory, and the memory access stage on the pipeline of the processor (as an example of the first stage) completes the execution of the load instruction Load0, (corresponding to Figure 2(clock cycle Tn in ), the load instruction Load0 enters the write-back phase. At this time, the first memory still needs to return data other than the first data in the first memory data to the processor. Before the processor completes acquiring the first memory data, the load instruction Load1 (as an example of the second instruction) enters the memory access phase, that is, the execution order of the load instruction Load1 is after the execution order of the load instruction Load0. In the process of running the load instruction Load1 in the memory access phase, the processor also detects that the second data of the load instruction Load1 is not in the cache, that is, the processor also has a cache miss, that is, the processor has a cache miss continuously, and the processor can detect a second request to acquire the second data of the load instruction Load1 (as an embodiment of the second instruction) from the first memory. At this time, the first memory is still returning data other than the first data in the first memory data to the processor. That is, before the processor completes acquiring the first memory data, it detects a second request to acquire the second data corresponding to the second instruction from the first memory.

[0157] In other embodiments, the processor detects a second request to obtain second data corresponding to a second instruction from the first memory before obtaining the first memory data. Alternatively, after the processor sends the first address to the first memory, the processor runs the second instruction based on the second stage of the pipeline, and the processor detects in the second stage that the second data of the second instruction is not located in the processor's cache, and the processor detects a second request to obtain the second data corresponding to the second instruction from the first memory.

[0158] For example, continue to refer to Figure 2 , in the clock cycle T3, after the load instruction Load0 encounters a cache miss in the memory access stage, the pipeline enters a pause. The processor detects the first request to obtain the first data of the load instruction Load0 from the first memory. After the processor sends the first address to the first memory, the processor can also determine whether the load instruction Load1 will encounter a cache miss based on the execution stage of the pipeline (as an example of the second stage). If it is determined that the processor detects that the second data required by the load instruction Load1 is not in the cache during the execution stage, it can be determined that the processor has a cache miss, that is, the processor has a cache miss continuously, then the processor will also detect the second request to obtain the second data of the load instruction Load1 from the first memory.

[0159] S604, the processor sends a second address to the first memory, and after storing the first memory data in the first buffer into the cache, obtains second memory data from the first memory based on the second address, the second memory data including the second data, and stores the second memory data in the first buffer.

[0160] In some embodiments of the present application, if a processor detects a second request before it has completely retrieved the data from the first memory, the processor may send the second address in the second request to the first memory via the AXI bus. However, in the process of the AXI bus adding an ARID to the second request, the ARID of the second request may be set to be the same as the ARID of the first request, i.e., both are R0. It is understood that if the AXI bus sets the ARID of the second request to R1, then for the AXI bus, R0 and R1 are different requests, and the AXI bus does not need to execute the R0 request and the R1 request in sequence. For example, the AXI bus can return 4 bytes of data in one clock cycle. The 4 bytes of data that the AXI bus can return in one clock cycle can be the 4 bytes of data corresponding to the R0 request or the 4 bytes of data corresponding to the R1 request. In this case, the processor needs to set up two buffers to store the data corresponding to the R0 request and the data corresponding to the R1 request, respectively. This requires the processor to allocate an additional portion of space to set up as a buffer, which in turn requires the processor to configure more register resources, thereby increasing the cost of the processor.

[0161] Therefore, in some embodiments of the present application, the AXI bus can set the ARID of the second request to R0. Then, when the AXI bus returns the request corresponding to R0, it returns it in order. That is, the AXI bus first returns the first memory data corresponding to the first request, and then returns the second memory data corresponding to the second request. In this way, the processor can set only one buffer (as an example of a first buffer zone). After the processor obtains the first memory data and stores the first memory data from the buffer to the cache, the processor will obtain the second memory data from the first memory and store the second memory data in the buffer.

[0162] It is understood that, through the above solution, only one buffer can be set up in the processor, thereby ensuring the low cost of the processor. In addition, the processor can continuously send requests to the first memory to obtain data. In the event that the processor repeatedly experiences cache misses, the time it takes for the processor to obtain data from the first memory can be reduced.

[0163] The following describes a process in which the processor obtains data from the first memory.

[0164] For example, Figure 7 According to some embodiments of the present application, an interactive flow chart of a processor obtaining data from a first memory is shown.

[0165] like Figure 7 As shown, the process includes:

[0166] S701: The MSS of the processor sends a first address of first data corresponding to a first instruction to a first memory.

[0167] For example, in some embodiments of the present application, if a cache miss occurs during the process of the processor executing the first instruction based on the pipeline, the MSS of the processor may send the first address of the first data corresponding to the first instruction to the first memory so as to obtain the first data corresponding to the first instruction from the memory, so that the processor continues to execute the first instruction based on the pipeline.

[0168] For example, the MSS obtains the first data from the first memory through the AXI bus, and sends a first request to the first memory through the AXI bus to obtain the first data. The first request includes the first address of the first data. The AXI bus can assign an ARID to the first request, for example, the ARID is R0.

[0169] For example, in some embodiments of the present application, the MSS may send the first address to the first memory in the clock cycle of address_t. In some embodiments, address_t may be 1, meaning that the MSS takes one clock cycle to send the first address to the first memory. The first memory may respond with data corresponding to the first address in the response duration response_t of 8 clock cycles.

[0170] For example, Figure 8 According to some embodiments of the present application, a schematic diagram of a processor acquiring data is shown.

[0171] like Figure 8 As shown, in some embodiments of the present application, t1-t33 respectively correspond to 1 clock cycle, for example, t1 is 1 clock cycle, t2 is 1 clock cycle, ..., t33 is 1 clock cycle.

[0172] In the clock cycle t1, the memory controller in the MSS of the processor may execute S701, thereby sending the first address to the first memory. For example, it takes 1 clock cycle for the memory controller in the MSS to send the first address to the first memory.

[0173] S702: The MSS of the processor sends a second address of second data corresponding to a second instruction to the first memory.

[0174] For example, in some embodiments of the present application, the second instruction is an instruction that is executed after the first instruction in the pipeline. Because a cache miss occurs when the processor executes the first instruction, the processor needs to retrieve the first data corresponding to the first instruction from the first memory. Therefore, the first instruction enters a paused state, and the pipeline also enters a paused state accordingly. In this case, the stage of the pipeline corresponding to the second instruction can also determine that a cache miss will occur during the execution of the second instruction. Therefore, the processor's MSS can send the second address of the second data corresponding to the second instruction to the first memory.

[0175] For example, the MSS sends a second request to obtain the second data to the first memory through the AXI bus. The second request includes the second address of the second data. The AXI bus can assign the same ARID as the first request to the second request. For example, the ARID of the second request is also R0. In this way, the AXI bus can return the data corresponding to the first request and the data corresponding to the second request in sequence.

[0176] For example, referring to Figure 8 In the clock cycle of t2, the memory controller in the MSS can execute S702. That is, the memory controller in the MSS also takes one clock cycle to send the second address to the first memory. This clock cycle can overlap with the response duration response_t of the first memory responding to the R0 request.

[0177] S703: The first memory returns first memory data corresponding to the first address to the MSS of the processor.

[0178] In some embodiments of the present application, after receiving the first address, the first memory may determine the first memory data of the first data corresponding to the first address, and then return the first memory data corresponding to the first address to the MSS.

[0179] For example, in some embodiments of the present application, the AXI bus can transmit 4 bytes of data in one clock cycle, and the first memory data has a total of 32 bytes. Therefore, the AXI bus needs to spend 8 clock cycles to return the first memory data to the MSS. After the memory controller in the MSS receives the first memory data, it can store the first memory data in the buffer.

[0180] In some embodiments of the present application, the first memory transmits data in a WRAP burst type manner, and the first memory data is 4 bytes, that is, within one clock cycle, the AXI bus can return the first data to the processor. Figure 8After the first memory takes 8 clock cycles to respond to the R0 request, it can return the first memory data to the MSS in the t10 clock cycle. For example, in the t10 clock cycle, the first memory can return 4 bytes of data to the MSS. These 4 bytes of data are the first data in the first memory data. Therefore, in the t10 clock cycle, the processor can obtain the first data and continue to execute the first instruction.

[0181] For example, because the AXI bus sets the ARID of the request to obtain the second data to R0, the first memory returns the second memory data including the second data only after returning the first memory data. For example, at clock cycle t21, the first memory can return the first memory data to the MSS. That is, the first memory takes 8 clock cycles to return the first memory data to the MSS.

[0182] For example, in some embodiments of the present application, only certain stages of the processor pipeline can determine whether a cache miss will occur during the processor's execution of the instruction. For example, Figure 2 In the five-stage pipeline of the processor, during the process of executing the load instruction in the memory access phase, it can be determined whether the processor will encounter a cache miss when executing the corresponding load instruction. Therefore, after the first instruction is executed to the memory access phase and a cache miss occurs, the pipeline cannot proceed. In the embodiment of the present application, at the clock cycle t10, the processor can obtain the first data of the first instruction, and the pipeline can continue to execute the first instruction (for example, the load instruction Load0). And at the clock cycle t11 (for example, the corresponding Figure 2 In the clock cycle of Tn in the memory), the second instruction (for example, the load instruction Load1) enters the memory access stage, during which the processor again encounters a cache miss. Therefore, the MSS of the processor can also send a second address to the first memory in the clock cycle of t11. That is, only after the processor obtains the first data of the first instruction can it determine whether a cache miss will occur when executing the second instruction. After the processor determines that a cache miss occurs in the second instruction, the MSS can send the second address of the second data corresponding to the second instruction to the first memory before obtaining the first memory data. That is, in some embodiments, the process of S702 can be executed after the process of S703 starts, refer to Figure 8 At clock cycle t11, the MSS can execute the process of S702. That is, the clock cycle in which the MSS sends the second address to the first memory overlaps with the clock cycle in which the first memory returns the first memory data to the MSS. The clock cycle in which the first memory responds to and processes the second address can also overlap with the clock cycle in which the first memory returns the first memory data to the MSS.

[0183] S704 , the MSS of the processor stores the first memory data in the cache of the processor.

[0184] In some embodiments of the present application, after the buffer in the MSS stores the first memory data, the MSS may store the first memory data from the buffer into the cache. For example, the clock cycles alloc_t taken by the MSS to store the first memory data from the buffer into the cache may be 4 clock cycles. That is, during the clock cycles t18 to t21, the MSS may store the first memory data into the cache.

[0185] It can be understood that if the MSS sends the second address to the first memory during clock cycle t11, the clock cycle for the first memory to respond and process the second address is from clock cycle t12 to clock cycle t19. In other words, the clock cycle for the first memory to respond and process the second address overlaps with the partial clock cycle for the first memory to return the first memory data to the MSS (clock cycle t12 to clock cycle t17), and the partial clock cycle for the MSS to store the first memory data from the buffer to the cache (clock cycle t18 and clock cycle t19). This can reduce the time it takes for the processor to retrieve data from the first memory when cache misses occur continuously.

[0186] S705: The first memory returns second memory data corresponding to the second address to the MSS of the processor.

[0187] For example, in some embodiments of the present application, after the MSS stores the first memory data in the buffer into the cache and the first memory responds to the processor's second address, the first memory can return the second memory data corresponding to the second address to the MSS, and the MSS can store the second memory data in the buffer. Therefore, in some embodiments of the present application, only one buffer is required in the processor, eliminating the need to increase buffer space in the processor. This avoids increasing the processor area and ensures processor miniaturization.

[0188] For example, referring to Figure 8 The first memory begins returning the second memory data to the MSS at clock cycle t22 and completes returning the first memory data at clock cycle t29, taking a total of eight clock cycles. It is understood that the first memory can transfer data to the MSS using a WRAP burst. That is, at clock cycle t22, the first memory can transfer the second data to the processor, and the processor can then continue executing the second instruction.

[0189] In this way, when cache misses occur in the process of executing the first instruction and the second instruction, the pipeline of the processor is paused for a total of address_t+response_t+trans_t+alloc_t+1=22 clock cycles.

[0190] S706 , the MSS of the processor stores the second memory data in the cache of the processor.

[0191] For example, in an embodiment of the present application, after the processor obtains the second data, the MSS still needs to spend alloc_t clock cycles to store the second memory data in the cache to ensure that the buffer stores the data obtained from the first memory next time.

[0192] Through the above solution, the processor does not need to add an additional buffer and can increase the speed of obtaining corresponding instructions from the first memory to avoid a long waiting time when the processor encounters continuous cache misses, thereby improving the efficiency of the processor in executing instructions.

[0193] The electronic devices involved in the above embodiments are introduced below.

[0194] For example, Figure 9 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown.

[0195] The electronic device 100 can be used to implement the data processing methods provided in the aforementioned embodiments.

[0196] like Figure 9 As shown, the electronic device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output (I / O) device 105. Among them:

[0197] The processor 101 may include one or more processing units, such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor or a field programmable gate array (FPGA), a neural network processing unit (NPU), etc. The processing module or processing circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU may be configured with an MSS and obtain data from the system memory 102 via an AXI bus.

[0198] The system memory 102 is a volatile memory, such as random-access memory (RAM) or double data rate synchronous dynamic random access memory (DDR SDRAM). The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 can be used to store data corresponding to instructions, such as the first data corresponding to the first instruction and the second data corresponding to the second instruction in the above-mentioned embodiments. The system memory 102 can also be used to store instructions for the data processing methods provided in the above-mentioned embodiments.

[0199] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card. In other embodiments, the non-volatile memory 103 may be used to store instructions for the data processing methods provided in the aforementioned embodiments.

[0200] In particular, the system memory 102 and the non-volatile memory 103 may respectively include a temporary copy and a permanent copy of the instruction 107. The instruction 107 may include instructions that, when executed by at least one of the processors 101, enable the electronic device 100 to implement the data processing methods provided in various embodiments of the present application.

[0201] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, thereby communicating with any other suitable device via one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the electronic device 100, for example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the electronic device 100 can communicate with other devices via the communication interface 104.

[0202] The input / output (I / O) device 105 may include input devices such as a keyboard and a mouse, and output devices such as a display. A user may interact with the electronic device 100 through the input / output (I / O) device 105 .

[0203] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface with other modules of the electronic device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide interfaces to the system memory 102 and the non-volatile memory 103.

[0204] In some embodiments, at least one of the processors 101 may be packaged together with the logic of one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).

[0205] I understand. Figure 9 The structure of the electronic device 100 shown is only an example. In other embodiments, the electronic device 100 may include more or fewer components than shown, or may combine or separate some components, or arrange the components differently. The components shown may be implemented in hardware, software, or a combination of software and hardware.

[0206] It is understood that the electronic device 100 can be any electronic device, including but not limited to a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook computer, etc., and the embodiments of the present application are not limited thereto.

[0207] An embodiment of the present application further provides a program product, which, when executed on an electronic device, can enable the electronic device to implement the methods provided in the aforementioned embodiments.

[0208] An embodiment of the present application further provides a readable storage medium, in which one or more programs are stored. When the one or more programs are executed by an electronic device, the electronic device implements the methods provided in the aforementioned embodiments.

[0209] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0210] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application specific integrated circuit, or a microprocessor.

[0211] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0212] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried by or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed over a network or via other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to floppy disks, optical disks, optical discs, compact disc-read only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable storage for transmitting information via the Internet in the form of electrical, optical, acoustical, or other propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Thus, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0213] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0214] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.

[0215] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0216] While the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.

Claims

1. A data processing method, applied to an electronic device, characterized in that: The electronic device includes a processor and a first memory, wherein the processor includes a cache and a first buffer zone, and the method includes: The processor detects a first request for obtaining first data corresponding to a first instruction from the first memory, where the first request includes a first address of the first data; The processor sends the first address to the first memory, obtains first memory data from the first memory based on the first address, the first memory data including first data, and stores the first memory data in the first buffer; Before acquiring all the data in the first memory, the processor detects a second request for acquiring second data corresponding to a second instruction from the first memory, where the second request includes a second address of the second data; The processor sends the second address to the first memory, and after the processor stores the first memory data in the first buffer into the cache, obtains second memory data from the first memory based on the second address, the second memory data including the second data, and stores the second memory data in the first buffer.

2. The method according to claim 1, characterized in that The processor sends the first address to a first memory, and obtains first memory data from the first memory based on the first address, including: The processor sends the first address to the first memory, obtains the first data from the first memory, and then obtains data other than the first data in the first memory from the first memory.

3. The method according to claim 1 or 2, characterized in that The processor detects a first request to obtain first data corresponding to a first instruction from the first memory, including: The processor executes the first instruction in a first stage of a pipeline, and the processor detects that first data corresponding to the first instruction is not located in a cache of the processor; The processor detects a first request to obtain first data corresponding to a first instruction from the first memory.

4. The method according to claim 3, characterized in that The execution order of the second instruction is after the execution order of the first instruction; The processor detects, before completing the acquisition of the first memory data, a second request for acquiring second data corresponding to a second instruction from the first memory, including: After the processor obtains the first data from the first memory, the first stage of the pipeline of the processor completes executing the first instruction, and before the processor completes obtaining the first memory data, the processor executes the second instruction on the first stage of the pipeline, and the processor detects that the second data is not located in the cache of the processor; The processor detects a second request to obtain second data corresponding to a second instruction from the first memory.

5. The method according to claim 3, characterized in that The execution order of the second instruction is after the execution order of the first instruction; The processor detects, before completing the acquisition of the first memory data, a second request for acquiring second data corresponding to a second instruction from the first memory, including: After the processor sends the first address to the first memory, the processor executes the second instruction based on the second stage of the pipeline, and the processor detects in the second stage that the second data of the second instruction is not located in the cache of the processor; The processor detects a second request to obtain second data corresponding to a second instruction from the first memory.

6. The method according to claim 1, characterized in that The processor further includes a memory subsystem, and the cache and the first buffer are arranged in the memory subsystem.

7. The method according to claim 1, characterized in that The size of the first buffer is the same as the size of the space occupied by the first memory data.

8. An electronic device, characterized in that: include: a memory for storing instructions; At least one processor is configured to execute the instructions so that the electronic device implements the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is run on a device, the device is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent caching method and device based on self-adaptive caching strategy

    CN121579388A