Providing Fast Memory Abandonment in a Processor-Based Device
By introducing a fast memory waste mechanism in processor-based devices, the waste value maintenance problem caused by traditional ISA operations is solved, and more efficient hardware resource utilization and performance improvement is achieved.
Patent Information
- Application Number
- CN202080084223.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-06
- Filing Date
- 2020-11-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-05
AI Technical Summary
Traditional ISA operations result in maintenance of waste values in system memory, resulting in unnecessary hardware resource consumption and performance degradation.
Introducing a fast memory depreciation mechanism in a processor-based device, by providing memory loading instructions that can indicate a final memory loading operation, positioning an intermediate memory entry and performing a final memory loading operation, setting an abandon indicator to free the entry for reuse.
Effectively eliminate the need to maintain waste value, reduce the consumption of hardware resources, and improve the system performance level.
Smart Images

Figure CN114746839B_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to memory access and maintenance in processor-based devices, and more particularly, to optimizing performance by avoiding unnecessary memory store operations to the system memory. Background Art
[0002] Implementing an instruction set architecture (ISA) of a processor-based device is basically oriented around the use of memory, where memory store instructions provided by the ISA are used to write values to the system memory, and memory load instructions provided by the ISA are used to read back values from the system memory. One use of such memory store instructions and memory load instructions is to temporarily save the values of registers and then restore them later to allow these registers to be used for other purposes within the processor-based device. As a result of this memory usage, once a memory load operation that reads a value from the system memory and restores the value to a register is executed, the originally saved register value may no longer be needed to remain in the system memory. Such a value can be considered "discarded" because no subsequent instruction will need to reference the memory location in the system memory to obtain the register value.
[0003] However, the rules of traditional ISA operations consider such values written to memory as "permanent", such that each value stored to the system memory remains available to any subsequent memory load instruction that reads the memory location until a subsequent memory store instruction writes a new value to the memory location. Therefore, a processor-based device is required to maintain the values in the system memory even in cases where no instruction will attempt to read the value again before it is overwritten by a subsequent memory store operation.
[0004] In addition, some traditional ISAs support a feature called "store-to-load forwarding", where a memory load operation after an earlier memory store operation that references the same memory location can be executed before the memory store operation has written its value to the system memory. With store-to-load forwarding, the value to be written by the memory store operation can be obtained from an intermediate memory such as a store buffer before it reaches the system memory and can be used to execute the memory load operation. In this case, if no subsequent memory load instruction will access the value after the memory load operation, the value can be considered discarded even before it reaches the system memory. As a result, requiring a processor-based device to hold discarded values leads to consumption of unnecessary hardware resources such as store buffers and increases the number of such hardware resources required to achieve a desired level of system performance.
[0005] Therefore, a more efficient mechanism is needed to eliminate the need to maintain discarded values. Summary of the Invention
[0006] Exemplary embodiments disclosed herein include providing fast memory discard in a processor-based device. In this regard, in one exemplary embodiment, an instruction set architecture (ISA) of a processor-based device is implemented to provide a memory load instruction that can indicate a final memory load operation from a given memory address (i.e., can indicate that after performing the memory load operation represented by a memory load instruction, it is no longer necessary to maintain the value stored at the memory address). In some exemplary embodiments, the memory load instruction may include a custom opcode, while some exemplary embodiments may provide that the memory load instruction includes an existing opcode and a custom final read indicator (e.g., a bit indicator). After a memory load instruction is received in an execution pipeline of a processing element (PE) of a processor-based device, an entry corresponding to the memory address of the memory load instruction is located in an intermediate memory external to the system memory of the processor-based device and is used to perform the final memory load operation. In some exemplary embodiments, the intermediate memory may be a buffer (e.g., as a non-limiting example, a store buffer, a write-back buffer, a pre-commit buffer, or a memory controller buffer), or may be a cache (e.g., as a non-limiting example, a data cache, a unified cache, or a level 1 (L1), level 2 (L2), level 3 (L3), or level 4 (L4) cache).
[0007] After performing the final memory load operation using the entry, a discard logic circuit (e.g., as a non-limiting example, located in a load comparator or a cache controller) sets a discard indicator for the entry to indicate that the entry can be reused before the content of the entry is written to the system memory. For example, the discard indicator may be a validity indicator (in embodiments where the intermediate memory is a buffer or a cache) or a dirty indicator (in embodiments where the intermediate memory is a cache). Then, traditional buffer and / or cache maintenance operations performed by the processor-based device can release the entry for reuse before the content of the entry is written to the system memory. In some embodiments, the processor-based device may also cancel a pending memory store operation initiated by a memory store instruction prior to the memory load instruction.
[0008] In another exemplary embodiment, a processor-based device is provided. The processor-based device includes a system memory and also includes a processing element (PE) that includes an execution pipeline and one or more load comparators. The execution pipeline includes one or more load stages. The processor-based device also includes an intermediate memory that is external to the system memory and includes a plurality of entries and corresponding plurality of discard indicators. The processor-based device is configured to receive a memory load instruction including a memory address using the execution pipeline of the PE, the memory load instruction indicating a final memory load operation from the memory address. The processor-based device is also configured to locate, by a load comparator among the one or more load comparators of the PE, an entry corresponding to the memory address among the plurality of entries in the intermediate memory. The processor-based device is also configured to perform the final memory load operation using the entry. The processor-based device is also configured to set a value of a discard indicator of the entry using discard logic circuitry of the processor-based device, wherein the discard indicator of the entry indicates that the entry can be reused before the content of the entry is written to the system memory.
[0009] In another exemplary embodiment, a method for providing fast memory discard in a processor-based device is provided. The method includes receiving a memory load instruction including a memory address using an execution pipeline of a processing element (PE) of the processor-based device, the memory load instruction indicating a final memory load operation from the memory address. The method also includes locating, using a load comparator of the PE, an entry corresponding to the memory address among a plurality of entries in an intermediate memory external to the system memory of the processor-based device. The method also includes performing the final memory load operation using the entry. The method also includes setting a value of a discard indicator of the entry using discard logic circuitry of the processor-based device, wherein the discard indicator of the entry indicates that the entry can be reused before the content of the entry is written to the system memory.
[0010] In another exemplary embodiment, a non-transitory computer-readable medium is provided. A computer-readable memory stores thereon an instruction program that includes a plurality of computer-executable instructions for execution by a processor. The plurality of computer-executable instructions includes a memory load instruction that includes a memory address and indicates a final memory load operation from the memory address.
[0011] Those skilled in the art will understand the scope of the present disclosure and recognize additional embodiments thereof after reading the following detailed description of the preferred embodiments in conjunction with the accompanying drawings. Description of the Drawings
[0012] The accompanying drawings are incorporated in and form a part of this specification, which illustrate several embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0013] Figure 1 is a schematic diagram of an exemplary processor-based device that includes a processing element (PE) for providing fast memory discard;
[0014] Figure 2A and Figure 2B is a block diagram of an exemplary memory load instruction that corresponds to Figure 1 and is used to indicate a final memory load operation;
[0015] Figure 3A and Figure 3B are respectively block diagrams of an exemplary buffer and cache embodiment of an intermediate memory that can be used to perform a fast memory discard operation; Figure 1
[0016] Figure 4A and Figure 4B is a flowchart of an exemplary operation for providing fast memory discard by a processor-based device via Figure 1 ; and
[0017] Figure 5 is a block diagram of an exemplary processor-based device configured to provide fast memory discard (such as a processor-based device of Figure 1 ). Detailed Description
[0018] Exemplary embodiments disclosed herein include providing fast memory discard in a processor-based device. In this regard, in one exemplary embodiment, an instruction set architecture (ISA) of a processor-based device is implemented to provide a memory load instruction that can indicate a final memory load operation from a given memory address (i.e., can indicate that after performing the memory load operation represented by a memory load instruction, it is no longer necessary to maintain the value stored at that memory address). In some exemplary embodiments, the memory load instruction may include a custom opcode, while some exemplary embodiments may provide that the memory load instruction includes an existing opcode and a custom final read indicator (e.g., a bit indicator). After a memory load instruction is received in an execution pipeline of a processing element (PE) of a processor-based device, an entry corresponding to the memory address of the memory load instruction is located in an intermediate memory external to the system memory of the processor-based device and is used to perform the final memory load operation. In some exemplary embodiments, the intermediate memory may be a buffer (e.g., as a non-limiting example, a store buffer, a write-back buffer, a pre-commit buffer, or a memory controller buffer), or may be a cache (e.g., as a non-limiting example, a data cache, a unified cache, or a level 1 (L1), level 2 (L2), level 3 (L3), or level 4 (L4) cache).
[0019] After performing the final memory load operation using the entry, discard logic circuitry (e.g., as a non-limiting example, located in a load comparator or a cache controller) sets the value of a discard indicator for the entry to indicate that the entry can be reused before the content of the entry is written to the system memory. For example, the discard indicator may be a validity indicator (in embodiments where the intermediate memory is a buffer or a cache) or a dirty indicator (in embodiments where the intermediate memory is a cache). Then, traditional buffer and / or cache maintenance operations performed by the processor-based device can release the entry for reuse before the content of the entry is written to the system memory. In some embodiments, the processor-based device may also cancel a pending memory store operation initiated by a memory store instruction prior to the memory load instruction.
[0020] In this regard, Figure 1 An exemplary processor-based device 100 is shown that provides a processing element (PE) 102 for providing fast memory discard. The PE 102 may include a central processing unit (CPU) having one or more processor cores, or may include a single processor core that includes a logical execution unit and associated caches and functional units. Additionally, in some embodiments, the PE 102 may be one of a plurality of similarly configured PEs (not shown) of the processor-based device 100.
[0021] In Figure 1 the example of, PE 102 is communicatively coupled to an interconnect bus 104, which in some embodiments may include additional components (e.g., by way of non-limiting example, bus controller circuitry and / or arbitration circuitry) not shown in Figure 1 for clarity. PE 102 is also communicatively coupled to a memory controller 106 via the interconnect bus 104. The memory controller 106 includes a memory controller buffer 107 and controls access to system memory 108 and manages the data stream to and from system memory 108. System memory 108 provides addressable memory for data storage of the processor-based device 100 and may thus include, by way of non-limiting example, synchronous dynamic random access memory (SDRAM). Figure 1 The PE 102 in the example of
[0022] Figure 1 is also communicatively coupled to a level 2 (L2) cache 110 and a level 3 (L3) cache 112 via the interconnect bus 104, where each cache represents a layer in a hierarchical cache structure used by the processor-based device 100 to cache frequently accessed data for faster retrieval (as compared to retrieving data from system memory 108). Figure 1 Figure 1 The PE 102 of
[0023] includes an execution pipeline 114 that includes circuitry configured to execute a stream of computer-executable instructions, such as exemplary memory store instructions (“MEM STORE”) 116 and subsequent memory load instructions (“MEM LOAD”) 118. In Figure 1 the example ofthe execution pipeline 114 includes a fetch stage 120 for fetching instructions for execution, a decode stage 122 for converting the fetched instructions into control signals for instruction execution, an execution stage 124 for actually performing the instruction execution, and a memory access stage 126 for performing memory access operations (e.g., memory load operations and / or memory store operations) resulting from the instruction execution. It should be understood that in some embodiments, the execution pipeline 114 may include fewer or more stages than those shown in Figure 1The PE 102 also includes a data cache 130, which is managed by a cache controller 132 and can be used to cache local copies of frequently accessed data within the PE 102 for faster access during the memory access stage 126 of the execution pipeline 114.
[0024] In traditional operation, the execution stage 124 of the execution pipeline can access the GPRF 128 to retrieve operands and / or store the results of arithmetic or logical operations. The results of memory store operations that will ultimately be committed to the system memory 108 can be temporarily stored in a store buffer 134 before optionally being cached in the data cache 130. Then, data values from the store buffer 134 and / or the data cache 130 can be moved to a write-back buffer 136 and subsequently to a pre-commit buffer 138 before being written to the system memory 108.
[0025] As described above, some traditional ISAs support store-to-load forwarding, which enables a memory load instruction such as the memory load instruction 118 that references a memory address 140 and is after (i.e., before the memory load instruction 118 in program order) an earlier memory store instruction 116 that references the same memory address 140, to be executed before the memory store instruction 116 has written its value to the system memory 108 (i.e., before the pending memory store operation initiated by the memory store instruction 116 is complete). This can be achieved by retrieving the value written by the memory store instruction 116 from, for example, the store buffer 134, the write-back buffer 136, or the pre-commit buffer 138 before sending the value to the system memory 108. Thus, in Figure 1 the example of, the memory access stage 126 of the execution pipeline includes one or more load stages ("LOAD"). Before attempting to access the system memory 108 to retrieve data for the memory load instruction 118, the (multiple) load stages 142 can use load comparators 144(0)-144(2) (corresponding to the store buffer 134, the write-back buffer 136, and the pre-commit buffer 138 respectively) to determine whether the corresponding buffer contains an entry corresponding to the memory address 140 of the memory load instruction 118. If so, the entry can be used to perform the memory load operation without waiting for the result of the memory store instruction 116 to be committed to the system memory 108.
[0026] Figure 1The processor-based device 100 can include any one of known digital logic elements, semiconductor circuits, processing cores, and / or memory structures, as well as other elements or combinations thereof. The embodiments described herein are not limited to any particular arrangement of elements, and the disclosed techniques can be easily extended to various structures and layouts on semiconductor sockets or packages. It should be understood that some embodiments of the processor-based device 100 may include more or fewer elements than Figure 1 shown. For example, the PE 102 may also include more or fewer memory devices to execute pipeline stages, and / or controller circuits. Specifically, in some embodiments, in addition to or instead of Figure 1 the cache shown, the processor-based device 100 may also include buffers and caches (e.g., as non-limiting examples, L1 cache, L4 cache, and / or unified cache).
[0027] As described above, one use of instructions such as the memory store instruction 116 and the memory load instruction 118 is to temporarily save the values of registers and then restore them later (e.g., within the GPRF 128) to allow those registers to be used for other purposes within the processor-based device 100. Thus, once the memory load instruction 118 has been executed, the value that was read may no longer be needed to continue residing in the system memory 108. Such a value can be considered "discarded" because no subsequent instruction needs to reference that storage location in the system memory 108 to obtain its value. However, traditional embodiments of processor-based devices may need to maintain such a value, even in cases where no memory load operation will attempt to read that value again before it is overwritten by a subsequent memory store operation. Additionally, in the above-described store-to-load forwarding scenario, if a value read from, for example, the store buffer 134, the write-back buffer 136, or the pre-commit buffer 138 is not accessed again by a subsequent memory load operation, that value can be considered discarded before it reaches the system memory 108. As a result, entries within the store buffer 134, the write-back buffer 136, and the pre-commit buffer 138 may be wasted by being used to store discarded data, and system resources can be wasted in maintaining discarded data as the discarded data moves through the buffers to the system memory 108.
[0028] In this regard, the processor-based device 100 is configured to provide fast memory discard. Specifically, the memory load instruction 118 provided by the ISA of the processor-based device 100 indicates a final memory load operation from the memory address 140 (i.e., indicates that the memory load instruction after the memory load instruction 118 will not access the current content stored at the memory address 140 again). Thus, after the execution pipeline 114 of the processing element 102 receives the memory load instruction 118, the entry corresponding to the memory address 140 is located within the intermediate memory 146 of the processor-based device 100. As discussed in more detail below, the intermediate memory 146 may include one or more buffers of the processor-based device 100 (e.g., as non-limiting examples, a store buffer 134, a write-back buffer 136, a pre-commit buffer 138, and / or a memory controller buffer 107) and / or caches (e.g., as non-limiting examples, a data cache 130, an L2 cache 110, and / or an L3 cache 112).
[0029] This entry is used to perform the final memory load operation (e.g., as a non-limiting example, using traditional store-to-load forwarding to read from a buffer or by accessing cache data from the cache), and then the value of the discard indicator for this entry is set to indicate that this entry can be reused before the content of this entry is written to the system memory 108. Since the intermediate memory 146 may include one or more buffers and / or caches within the processor-based device 100, the logic for setting the value of the discard indicator can be implemented by the discard logic circuit ("LOG CIR") 148 in one or more of the cache controller 132, load comparators 144(0)-144(2), L2 cache 110 (or its cache controller), L3 cache 113 (or its cache controller), and / or memory controller buffer 107. As discussed below with reference to Figure 3A and Figure 3B In embodiments where the intermediate memory 146 is a buffer or a cache, the discard indicator may be a validity indicator, and its value is set to indicate that this entry is no longer valid (e.g., as a non-limiting example, by setting the value of the validity indicator to "false" or zero (0)). Embodiments where the intermediate memory 146 is a cache may provide that the discard indicator is a dirty indicator, and its value is set to indicate that the content of this entry has not been modified. Thus, in the latter embodiment, setting the value of the discard indicator may include setting the value of the dirty indicator to "false" or zero (0), such that the content of the cache line will not be written to the system memory 108.
[0030] In some embodiments, after the value of the discard indicator for an entry is set, the processor-based device 100 may detect that the discard indicator indicates that the entry can be reused, and may release the entry for reuse before the content of the entry is written to the system memory 108. For example, in embodiments where the intermediate memory 146 is a buffer or cache and the discard indicator is a validity indicator for a buffer entry, the processor-based device 100 may detect that the buffer or cache entry is no longer valid, and may reuse the buffer or cache entry. Similarly, in embodiments where the intermediate memory 146 is a cache and the discard indicator is a dirty indicator for a cache entry, the processor-based device 100 may detect that the cache entry is not dirty (i.e., does not contain modified data), and may avoid writing the content of the cache entry to the system memory 108.
[0031] According to some embodiments in which the PE 102 is one of a plurality of PEs of the processor-based device 100, after detecting that the discard indicator indicates that the entry can be reused, the processor-based device 100 may be configured to cancel a coherence operation corresponding to the memory address 140 from the first PE (e.g., PE 102) to one or more other PEs of the plurality of PEs of the processor-based device 100. For example, conventionally, the PE 102 may be configured to update other PEs of the processor-based device 100 to indicate that, as a result of a memory store operation to the memory address 140, the cache entry in the cache has changed state (e.g., as a non-limiting example, has been modified or has been declared invalid). Thus, conventionally, the PE 102 may be configured to perform a coherence operation to notify other PEs of the processor-based device 100 of the state change. However, performing the final memory store operation described herein may render such a coherence operation unnecessary, and thus the PE 102 may cancel such a coherence operation corresponding to the memory address 140 when the final memory store operation is performed.
[0032] Embodiments of the processor-based device 100 may also be configured to cancel a memory store instruction 116 to a memory address 140 (i.e., a memory store instruction 116 that writes data to an entry within an intermediate memory 146 for performing a final memory load operation) after setting the value of the discard indicator for the entry and before the result of the memory store instruction 116 in the memory store is committed to the system memory 108. This can save system resources that would otherwise be consumed in processing the memory store instruction 116, even if the memory load instruction 118 has made the processor-based device 100 aware that the content to be stored at the memory address 140 will not be accessed by any subsequent memory load instruction. Some embodiments of the processor-based device 100 may provide further security by being configured to overwrite the content of the entry after performing the memory load operation indicated by the memory load instruction 118.
[0033] To illustrate an exemplary memory load operation corresponding to a memory load instruction 118 for indicating a final memory load operation Figure 1 provided are Figure 2A and Figure 2B . Figure 2A Illustrated is a memory load instruction 200 that is functionally corresponding to the Figure 1 memory load instruction 118. In the Figure 2A example, the memory load instruction 200 includes a custom opcode 202 (i.e., an opcode specifically provided by the underlying ISA for explicitly indicating memory discard). In contrast, Figure 2B illustrates a memory load instruction 204 that includes an existing opcode 206 and a custom final read indicator 208. The existing opcode 206 corresponds to an opcode provided by the ISA for a traditional memory load operation, and the custom final read indicator 208 includes an additional indicator (e.g., a bit indicator) that can be set to indicate that the memory load operation to be performed is a final memory load operation to a specified memory address.
[0034] Figure 3A and Figure 3B respectively illustrate exemplary buffer and cache embodiments of the intermediate memory 146 on which a fast memory discard operation may be performed. In the Figure 1 example, the intermediate memory 146 includes a buffer 300, which may correspond to, for example, as a non-limiting example, Figure 3A Figure 1 a storage buffer 134, a write-back buffer 136, a pre-commit buffer 138, and / or a memory controller buffer 107. The buffer 300 includes a plurality of entries 302(0)-302(B), and each entry is associated with a corresponding validity indicator 304(0)-304(B) indicating whether the associated entry 302(0)-302(B) is in a valid state. As described above, for embodiments where the intermediate memory 146 is a buffer such as buffer 300, each of the validity indicators 304(0)-304(B) can be considered a discard indicator 306 for the associated entry 302(0)-302(B).
[0035] In Figure 3B , the intermediate memory 146 includes a cache 308, which may correspond to, for example, by way of non-limiting example Figure 1 a data cache 130, an L2 cache 110, and / or an L3 cache 112. As Figure 3B shown, the cache 308 includes a plurality of entries 310(0)-310(C), and each entry is associated with a corresponding dirty indicator 312(0)-312(C) indicating whether the associated entry 310(0)-310(C) contains dirty (i.e., modified) data. Each of the plurality of entries 310(0)-310(C) is also associated with a corresponding validity indicator 314(0)-314(C) indicating whether the associated entry 310(0)-310(C) is in a valid state. Thus, for some embodiments where the intermediate memory 146 is a cache such as cache 308, each of the dirty indicators 312(0)-312(C) can be considered a discard indicator 306 for the associated entry 310(0)-310(C), while in some embodiments each of the validity indicators 314(0)-314(C) can be used as a discard indicator 306 for the associated entry 310(0)-310(C).
[0036] Figure 4A and Figure 4B illustrates an exemplary operation 400 for providing fast memory discard by a Figure 1 processor-based device 100. For clarity, reference is made to the elements of Figure 4A and Figure 4B when describing Figure 1 and Figure 3A and Figure 3B According to some embodiments, Figure 4AOperation 400 therein begins with the execution pipeline 114 of the PE 102 of the processor-based device 100 receiving a memory load instruction 118 including a memory address 140, and the memory load instruction 118 indicates a final memory load operation from the memory address 140 (block 402). The processor-based device 100 locates an entry corresponding to the memory address 140 in a plurality of entries 302(0)-302(B), 310(0)-310(C) of the intermediate memory 146 external to the system memory 108 of the processor-based device 100 (e.g., Figure 3A and Figure 3B one of the entries 302(0) or 310(0)) (block 404). Then, the processor-based device 100 performs the final memory load operation using the entries 302(0), 310(0) (block 406).
[0037] Next, the processor-based device 100 sets the value of the discard indicator 306 of the entries 302(0), 310(0), where the discard indicator 306 of the entries 302(0), 310(0) indicates that the entries 302(0), 310(0) can be reused before the contents of the entries 302(0), 310(0) are written to the system memory 108 (block 408). In some embodiments (e.g., where the intermediate memory 146 is a buffer or a cache), the operation for setting the value of the discard indicator 306 in block 408 may include, for example, setting the value of the validity indicator 304(0) of the entry 302(0) to indicate that the entry 302(0) is no longer valid (block 410). Some embodiments (e.g., where the intermediate memory 146 is a cache) may provide that the operation for setting the value of the discard indicator 306 in block 408 may include, for example, setting the value of the dirty indicator 312(0) of the entry 310(0) to indicate that the contents of the entry 310(0) have not been modified (block 412). In some embodiments, the processing may continue in Figure 4B in.
[0038] Now refer to Figure 4B, shows further operations that can be performed by the processor - based device 100. In some embodiments, next, the processor - based device 100 can cancel a pending memory store operation initiated by the memory - stored instruction 116 to the memory address 140, where the memory - stored instruction 116 is before the memory load instruction 118 in program order (block 414). Some embodiments may provide that the processor - based device 100 detects that the discard indicators 306 of entries 302(0), 310(0) indicate that entries 302(0), 310(0) can be reused (block 416). In response to detecting that the discard indicators 306 indicate that entries 302(0), 310(0) can be reused, the processor - based device 100 can release entries 302(0), 310(0) for reuse before the contents of entries 302(0), 310(0) are written to the system memory 108 (block 418).
[0039] According to some embodiments, the processor - based device 100 can detect that the discard indicators 306 of entries 302(0), 310(0) indicate that entries 302(0), 310(0) can be reused (block 420). In response to detecting that the discard indicators 306 indicate that entries 302(0), 310(0) can be reused, in embodiments where the PE 102 is one of a plurality of PEs, the processor - based device 100 can cancel coherence operations corresponding to the memory address 140 from a first PE (e.g., PE 102) to one or more other PEs of the plurality of PEs of the processor - based device 100 (block 422). Some embodiments may provide that the processor - based device 100 can overwrite the contents of entries 302(0), 310(0) (block 424).
[0040] Figure 5 is an exemplary block diagram of a processor - based device 500 of a processor - based device 100 that provides fast memory discard such as Figure 1 The processor - based device 500 can be one or more circuits included in an electronic board, such as a printed circuit board (PCB), a server, a personal computer, a desktop computer, a laptop computer, a personal digital assistant (PDA), a computing board, a mobile device, or any other device, and can represent, for example, a server or a user's computer. In this example, the processor - based device 500 includes a processor 502. The processor 502 represents one or more general - purpose processing circuits, such as a microprocessor, a central processing unit, etc., and can correspond to Figure 1The PE 102. The processor 502 is configured to execute the processing logic in the instructions for performing the operations and steps discussed herein. In this example, the processor 502 includes an instruction cache 504 for temporary, fast access memory storage of instructions and an instruction processing circuit 510. Instructions fetched or prefetched from memory, such as from the system memory 508 via the system bus 506, are stored in the instruction cache 504. The instruction processing circuit 510 is configured to process the instructions fetched into the instruction cache 504 and process the instructions for execution.
[0041] The processor 502 and the system memory 508 are coupled to the system bus 506 (corresponding to Figure 1 the interconnect bus 104), and can be mutually coupled to peripheral devices included in the processor-based device 500. As is well known, the processor 502 communicates with these other devices by exchanging address, control, and data information on the system bus 506. For example, the processor 502 can transmit a bus transaction request to the memory controller 512 in the system memory 508, which is an example of a peripheral device. Although not shown in Figure 5 it, multiple system buses 506 can be provided, where each system bus constitutes a different architecture. In this example, the memory controller 512 is configured to provide a memory access request to the memory array 514 in the system memory 508. The memory array 514 consists of an array of storage bit cells for storing data. The system memory 508 can be read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM), etc.), and static memory (e.g., flash memory, static random access memory (SRAM), etc.), as non-limiting examples.
[0042] Other devices can be connected to the system bus 506. As Figure 5As shown, these devices may include, by way of example, system memory 508, one or more input devices 516, one or more output devices 518, a modem 524, and one or more display controllers 520. The (multiple) input devices 516 may include any type of input device, including but not limited to input keys, switches, voice processors, etc. The (multiple) output devices 518 may include any type of output device, including but not limited to audio, video, other visual indicators, etc. The modem 524 may be any device configured to allow data to be exchanged to and from a network 526. The network 526 may be any type of network, including but not limited to wired or wireless networks, private or public networks, local area networks (LANs), wireless local area networks (WLANs), wide area networks (WANs), Bluetooth™ networks, and the Internet. The modem 524 may be configured to support any type of communication protocol required. The processor 502 may also be configured to access the (multiple) display controllers 520 via the system bus 506 to control the information sent to one or more displays 522. The (multiple) displays 522 may include any type of display, including but not limited to cathode ray tubes (CRTs), liquid crystal displays (LCDs), plasma displays, etc.
[0043] Figure 5 The processor-based device 500 in may include an instruction set 528, which may be encoded with an explicit consumer naming model based on arrival for execution by the processor 502 for any application as desired according to the instructions. The instructions 528 may be stored in the system memory 508, the processor 502, and / or the instruction cache 504, by way of example of a non-transitory computer-readable medium 530. The instructions 528 may also reside, in whole or at least in part, within the system memory 508 and / or within the processor 502 during their execution. The instructions 528 may also be sent or received over the network 526 via the modem 524, such that the network 526 includes a computer-readable medium 530.
[0044] Although the computer-readable medium 530 is shown as a single medium in the exemplary embodiment, the term "computer-readable medium" should be regarded as including a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) storing one or more instruction sets 528. The term "computer-readable medium" should also be regarded as including any medium that is capable of storing, encoding, or carrying an instruction set for execution by a processing device and that causes the processing device to perform any one or more of the methods of the embodiments disclosed herein. Thus, the term "computer-readable medium" should be regarded as including but not limited to solid-state memory, optical media, and magnetic media.
[0045] The embodiments disclosed herein include various steps. The steps of the embodiments disclosed herein may be formed by hardware components or may be implemented in machine-executable instructions that may be used to cause a general or special-purpose processor programmed with the instructions to perform these steps. Alternatively, these steps may be performed by a combination of hardware and software.
[0046] The embodiments disclosed herein may be provided as a computer program product or software that may include a machine-readable medium (or computer-readable medium) having instructions stored thereon that may be used to program a computer system (or other electronic device) to perform a process according to the embodiments disclosed herein. The machine-readable medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, the machine-readable medium includes: machine-readable storage media (e.g., ROM, random access memory (“RAM”), magnetic disk storage media, optical storage media, flash devices, etc.).
[0047] Unless otherwise specifically stated and apparent from the foregoing discussion, it is to be understood that throughout the specification, discussions using terms such as “processing,” “computing,” “determining,” “displaying,” etc., refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers of the computer system into other data similarly represented as physical quantities within a memory of the computer system or registers or other such information storage, transmission, or display devices.
[0048] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. According to the teachings herein, various systems may be used with the program, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The structure required for various systems will be apparent from the foregoing description. Further, the embodiments described herein are not described with reference to any particular programming language. It should be understood that various programming languages may be used to implement the teachings of the embodiments described herein.
[0049] Those skilled in the art will further recognize that the various illustrative logical blocks, modules, circuits, and algorithms described in connection with the embodiments disclosed herein can be implemented as electronic hardware, instructions stored in memory or another computer-readable medium and executed by a processor or other processing device, or a combination of both. The components of the distributed antenna system described herein can be used in any circuit, hardware component, integrated circuit (IC), or IC chip, by way of example. The memory disclosed herein can be of any type and size and can be configured to store any type of information required. To clearly illustrate this interchangeability, the various illustrative components, boxes, modules, circuits, and steps have been described generally above in terms of their functionality. How this functionality is implemented depends on the particular application, design choices, and / or design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in different ways for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of the present embodiments.
[0050] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or executed using a processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Additionally, a controller can be a processor. The processor can be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with DSP cores, or any other such configuration).
[0051] The embodiments disclosed herein can be implemented in hardware and in instructions stored in hardware, and can reside in, for example, RAM, flash memory, ROM, electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable medium known in the art. The exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium can be integrated into the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a remote station. Alternatively, the processor and the storage medium can reside as discrete components in a remote station, a base station, or a server.
[0052] It should also be noted that the operational steps described in any exemplary embodiment herein are provided for example and discussion. The described operations can be performed in many different sequences other than the shown sequence. Additionally, the operations described in a single operational step can actually be performed in multiple different steps. Further, one or more operational steps discussed in the exemplary embodiments can be combined. Those skilled in the art will also understand that any of a variety of technologies and techniques can be used to represent information and signals. For example, the data, instructions, commands, information, signals, bits, symbols, and chips referred to throughout the above description can be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0053] Unless otherwise expressly stated, any method presented herein is not to be construed as requiring that its steps be performed in a particular order. Accordingly, where a method claim does not actually recite an order to be followed by its steps, or where the claims or specification do not otherwise expressly state that the steps will be limited to a particular order, no particular order is intended to be implied.
[0054] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the spirit or scope of the invention. Since modifications, combinations, sub - combinations and variations of the disclosed embodiments incorporating the spirit and substance of the invention may occur to those skilled in the art, the invention should be construed to include all such within the scope of the appended claims and their equivalents.
Claims
1. A processor-based device, comprising: A system memory; A processing element PE, including an execution pipeline; and An intermediate memory, which is external to the system memory and includes a plurality of entries and corresponding a plurality of discard indicators; The processor-based device is configured to: Use the execution pipeline of the PE to receive a memory load instruction including a memory address, the memory load instruction indicating a final memory load operation from the memory address; Locate an entry corresponding to the memory address among the plurality of entries in the intermediate memory; Execute the final memory load operation using the entry; And Use the discard logic circuit of the processor-based device to set the value of the discard indicator of the entry, wherein the discard indicator of the entry indicates that the entry can be reused before the content of the entry is written to the system memory.
2. The processor-based device according to claim 1, wherein: The intermediate memory includes one or more of a buffer and a cache; The plurality of discard indicators include a plurality of validity indicators corresponding to the plurality of entries; And The processor-based device is configured to: set the value of the discard indicator of the entry by being configured to set the value of the validity indicator of the entry to indicate that the entry is no longer valid.
3. The processor-based device according to claim 1, wherein: The intermediate memory includes a cache; The plurality of discard indicators include a plurality of dirty indicators corresponding to the plurality of entries; And The processor-based device is configured to: set the value of the discard indicator of the entry by being configured to set the value of the dirty indicator of the entry to indicate that the content of the entry has not been modified.
4. The processor-based device according to claim 1, wherein the processor-based device is further configured to: Detect that the discard indicator of the entry indicates that the entry can be reused; and In response to detecting that the discard indicator indicates that the entry can be reused, release the entry for reuse before the content of the entry is written to the system memory.
5. The processor-based device according to claim 1, wherein the processor-based device is further configured to: after setting the value of the discard indicator of the entry, cancel a pending memory store operation initiated by a memory store instruction to the memory address, the memory store instruction being in program order before the memory load instruction.
6. The processor-based device according to claim 1, wherein: The PE includes a first PE among a plurality of PEs of the processor-based device; and The processor-based device is further configured to: Detect that the discard indicator of the entry indicates that the entry can be reused; and In response to detecting that the discard indicator indicates that the entry is discarded, cancel the coherence operation corresponding to the memory address from the first PE to one or more other PEs among the multiple PEs of the processor-based device.
7. The processor-based device according to claim 1, wherein the processor-based device is further configured to: after performing the final memory load operation, overwrite the content of the entry.
8. The processor-based device according to claim 1, wherein the memory load instruction includes a custom opcode of the instruction set architecture (ISA) of the processor-based device.
9. The processor-based device according to claim 1, wherein the memory load instruction includes a custom final read indicator and an existing opcode of the ISA of the processor-based device.
10. A method for providing fast memory discard in a processor-based device, comprising: Using an execution pipeline of a processing element (PE) of the processor-based device to receive a memory load instruction including a memory address, the memory load instruction indicating a final memory load operation from the memory address; Locating an entry corresponding to the memory address in multiple entries of an intermediate memory external to the system memory of the processor-based device; Performing the final memory load operation using the entry; And Using the discard logic circuit of the processor-based device to set the value of a discard indicator of the entry, wherein the discard indicator of the entry indicates that the entry can be reused before the content of the entry is written to the system memory.
11. The method according to claim 10, wherein: The intermediate memory includes one or more of a buffer and a cache; The discard indicator includes a validity indicator of the entry; And Setting the discard indicator of the entry includes setting the value of the validity indicator of the entry to indicate that the entry is no longer valid.
12. The method according to claim 10, wherein: The intermediate memory includes a cache; The discard indicator includes a dirty indicator of the entry; and Setting the discard indicator of the entry includes setting the value of the dirty indicator of the entry to indicate that the content of the entry has not been modified.
13. The method according to claim 10, further comprising: Detecting that the discard indicator of the entry indicates that the entry can be reused; And In response to detecting that the discard indicator indicates that the entry can be reused, releasing the entry for reuse before the content of the entry is written to the system memory.
14. The method according to claim 10, further comprising: After setting the value of the discard indicator of the entry, cancel a pending memory store operation initiated by a memory store instruction to the memory address, the memory store instruction being in program order before the memory load instruction.
15. The method according to claim 10, wherein: The PE includes a first PE among the multiple PEs of the processor-based device; And The method further comprises: detecting that the discard indicator of the entry indicates that the entry can be reused; and in response to detecting that the discard indicator of the entry indicates that the entry can be reused, canceling a coherence operation corresponding to the memory address from the first PE to one or more other PEs among the plurality of PEs of the processor-based device.
16. The method according to claim 10, further comprising: After performing the final memory load operation, overwrite the content of the entry.
17. The method according to claim 10, wherein the memory load instruction comprises a custom opcode of an instruction set architecture ISA of the processor-based device.
18. The method according to claim 10, wherein the memory load instruction comprises an existing opcode of the ISA of the processor-based device and a custom final read indicator.
Citation Information
Patent Citations
Operand cache control techniques
US20170075810A1