Instruction execution device and method, processor and computer equipment
Execute atomic operation instructions by multiplexing cache access pipelines, solving the problem of low queue resource utilization and achieving efficient execution of atomic operation instructions.
Patent Information
- Application Number
- CN202410111237.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, atomic operation instructions need to be split into multiple instructions in the processor of the RISC-V architecture, resulting in low queue resource utilization and excessive queue resources occupied.
By multiplexing the cache access pipeline, atomic computing instructions are directly executed, including loading storage units, tag pipelines, data pipelines, logical computing units and data writing units, reducing resource consumption and improving execution efficiency.
It realizes efficient execution of atomic operation instructions, reduces resource consumption during instruction execution, and improves queue utilization.
Smart Images

Figure CN120371391A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of chip technology, and particularly to an instruction execution device, method, processor, and computer device. Background Art
[0002] An atomic memory operation (AMO) instruction is an instruction for processing atomic operations in a processor of the RISC-V architecture. An atomic operation refers to an operation that cannot be split or interrupted by other instructions or operations, and is usually used in a multi-core processor system.
[0003] In the related art, it is necessary to split an atomic instruction into three instructions, and atomic instructions are generally stored in a queue in the processor. Excessive instructions will occupy limited queue resources, resulting in low queue utilization. Summary of the Invention
[0004] The embodiments of the present application provide an instruction execution device, method, processor, and computer device, which can improve the execution efficiency of atomic arithmetic instructions and reduce resource consumption. The technical solutions are as follows:
[0005] On the one hand, the embodiments of the present application provide an instruction execution device, which includes a load / store unit, a tag pipeline, a data pipeline, a logic operation unit, and a data writing unit. The tag pipeline and the data pipeline are used to access the first-level cache. The tag pipeline includes m tag reading paths, and different tag reading paths correspond to different cache lines. The data pipeline includes m data reading paths, and different data reading paths correspond to different cache lines;
[0006] The load / store unit is used to transmit an atomic arithmetic instruction to the tag pipeline and the data pipeline. The atomic arithmetic instruction includes a target memory access address and a first data to be operated;
[0007] The tag pipeline is used to output a cache hit signal based on the target address tag corresponding to the target memory access address and the address tags corresponding to the cache lines in each tag reading path. The cache hit signal is used to indicate the cache hit status in the first-level cache;
[0008] The data pipeline is used to output the original cache line data corresponding to the target memory access address based on the cache hit signal;
[0009] The logic operation unit is used to perform a logic operation on the first data to be operated and a second data to be operated in the original cache line data to obtain a logic operation result;
[0010] The data writing unit is configured to write the target cache line data corresponding to the logical operation result into the target data memory in the first-level cache.
[0011] On the other hand, an embodiment of the present application provides an instruction execution method, which is used for the instruction execution device described in the above aspect. The instruction execution device includes a load / store unit, a tag pipeline, a data pipeline, a logical operation unit, and a data writing unit;
[0012] The method includes:
[0013] Transmit an atomic operation instruction to the tag pipeline and the data pipeline through the load / store unit. The atomic operation instruction includes a target memory access address and a first data to be operated on;
[0014] Output a cache hit signal through the tag pipeline based on the target address tag corresponding to the target memory access address and the address tags corresponding to the cache lines in each tag read path. The cache hit signal is used to indicate the cache hit status in the first-level cache;
[0015] Output the original cache line data corresponding to the target memory access address through the data pipeline based on the cache hit signal;
[0016] Perform a logical operation on the first data to be operated on and the second data to be operated on in the original cache line data through the logical operation unit to obtain a logical operation result;
[0017] Write the target cache line data corresponding to the logical operation result into the target data memory in the first-level cache through the data writing unit.
[0018] In some embodiments, the tag pipeline includes a tag reading unit, and the tag read path includes a tag memory, a first register, and a comparator. The tag memory stores the address tags corresponding to the cache lines;
[0019] The tag reading unit responds to the atomic operation instruction and sends a tag reading instruction to the tag memory in each tag read path;
[0020] The tag memory responds to the tag reading instruction, transmits the address tag to the first register, and the first register performs a pipelining process on the address tag and then transmits it to the comparator;
[0021] The comparator compares the address tag and the target address tag and outputs a tag comparison result;
[0022] The label pipeline outputs the cache hit signal based on the label comparison results corresponding to each label reading path.
[0023] In some embodiments, the data pipeline includes a data reading unit and a first multiplexer. The data reading path includes a data memory and a second register. The data memory stores the cache line data corresponding to the cache line. The input end of the first multiplexer is connected to the output end of the second register;
[0024] The data reading unit responds to the atomic operation instruction and sends a data reading instruction to the data memories in each data reading path;
[0025] The data memory responds to the data reading instruction, transmits the cache line data to the second register, and the second register performs pipelining processing on the cache line data and then transmits it to the first multiplexer;
[0026] The first multiplexer outputs the original cache line data corresponding to the target memory access address when the cache hit signal indicates a cache hit.
[0027] In some embodiments, the device further includes a third register. The input end of the third register is connected to the output end of the label pipeline, and the output end of the third register is connected to the strobe end of the first multiplexer;
[0028] The third register receives the cache hit signal transmitted by the label pipeline and transmits the cache hit signal to the first multiplexer.
[0029] In some embodiments, the data pipeline further includes a path prediction unit. The input end of the path prediction unit is connected to the output end of the load / store unit, and the output end of the path prediction unit is connected to the input end of the data reading unit;
[0030] The path prediction unit responds to the atomic operation instruction, predicts the data reading path corresponding to the target memory access address, and transmits a path prediction result to the data reading unit;
[0031] The data reading unit sends the data reading instruction to the data memory in the data reading path corresponding to the path prediction result.
[0032] In some embodiments, when the data reading path corresponding to the path prediction result is different from the data reading path corresponding to the cache hit signal, the path prediction unit re-predicts the data reading path corresponding to the target memory access address and transmits a new path prediction result to the data reading unit.
[0033] In some embodiments, the apparatus further includes a request generation unit and a backfill unit. The input end of the request generation unit is connected to the output end of the tag pipeline;
[0034] The request generation unit receives the cache hit signal transmitted by the tag pipeline;
[0035] When the cache hit signal indicates a cache miss, the request generation unit generates a data miss request, writes the data miss request into the miss request queue, and transmits the data miss request to the secondary cache through the miss request queue;
[0036] The backfill unit outputs the original cache line data corresponding to the target memory access address based on the request feedback corresponding to the data miss request.
[0037] In some embodiments, the apparatus further includes a second multiplexer. The input end of the second multiplexer is connected to the output end of the first multiplexer, the input end of the second multiplexer is connected to the output end of the backfill unit, and the output end of the second multiplexer is connected to the input end of the logic operation unit;
[0038] When the cache hit signal indicates a cache hit, the second multiplexer receives the original cache line data corresponding to the target memory access address output by the first multiplexer, and transmits the original cache line data to the logic operation unit;
[0039] When the cache hit signal indicates a cache miss, the second multiplexer receives the original cache line data corresponding to the target memory access address output by the backfill unit, and transmits the original cache line data to the logic operation unit.
[0040] In some embodiments, the apparatus further includes a data merging unit and a fourth register. The input end of the data merging unit is connected to the output end of the second multiplexer, the input end of the data merging unit is connected to the output end of the logic operation unit, and the output end of the data merging unit is connected to the input end of the fourth register;
[0041] The data merging unit receives the original cache line data output by the second multiplexer and the logic operation result output by the logic operation unit;
[0042] The data merging unit performs data merging based on the original cache line data and the logic operation result to obtain the target cache line data, and the fourth register performs pipelining processing on the target cache line data and then transmits it to the data writing unit.
[0043] In some embodiments, when the original cache line data is located in the level-1 cache, the data writing unit determines the target data memory corresponding to the original cache line data in the level-1 cache, and writes the target cache line data into the target data memory; when the original cache line data is located in the level-2 cache, the data writing unit determines the data memory with the longest data update time distance from the current time in the level-1 cache as the target data memory based on the data update conditions of the data memories corresponding to the data pipelines, and writes the target cache line data into the target data memory.
[0044] In some embodiments, the apparatus further includes a status pipeline, and the status pipeline includes m status reading paths, and different status reading paths correspond to different cache lines;
[0045] The status pipeline responds to the atomic operation instruction, and outputs the cache line status corresponding to each cache line through each status reading path, and the cache line status is used to characterize the data read / write permission of the cache line corresponding thereto.
[0046] In some embodiments, the status pipeline further includes a status reading unit and a third multiplexer, and the status reading path includes a status memory and a fifth register;
[0047] The input end of the third multiplexer is connected to the output end of the fifth register, and the selection end of the third multiplexer is connected to the output end of the tag pipeline;
[0048] The status reading unit responds to the atomic operation instruction, and sends a status reading instruction to the status memory in each status reading path;
[0049] The status memory responds to the status reading instruction, transmits the cache line status to the fifth register, and the fifth register performs a pipelining process on the cache line status and then transmits it to the third multiplexer;
[0050] The third multiplexer does not output the cache line status when the cache hit signal indicates cache miss;
[0051] The third multiplexer does not output the target cache line status when the cache hit signal indicates cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address has data read / write permission.
[0052] In some embodiments, when the cache hit signal indicates a cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address does not have data read / write permission, the third multiplexer transmits the target cache line status to the request generation unit;
[0053] Based on the target cache line status, the request generation unit generates a data miss request, writes the data miss request into the miss request queue, and transmits the data miss request to the secondary cache through the miss request queue; or,
[0054] The request generation unit generates a permission acquisition request based on the target cache line status, and the permission acquisition request is used to request to obtain the data read / write permission of the target cache line in the first-level cache.
[0055] On the other hand, an embodiment of the present application provides a processor, and the processor includes an instruction execution device as described in the above aspect.
[0056] On the other hand, an embodiment of the present application provides a computer device, and the computer device includes a processor and a memory as described in the above aspect, and the processor is connected to the memory through a bus.
[0057] On the other hand, an embodiment of the present application provides a computer device, and the computer device includes a processor, a memory, and an instruction execution device as described in the above aspect, wherein the processor is connected to the instruction execution device, and the processor is connected to the memory through a bus.
[0058] In the embodiment of the present application, by multiplexing the cache access pipeline, the load / store unit sends an atomic operation instruction to the tag pipeline and the data pipeline. Thus, based on the target address tag corresponding to the target memory access address in the atomic operation instruction and the address tags corresponding to the cache lines in each tag read path, the tag pipeline outputs a cache hit signal, and the data pipeline outputs the original cache line data corresponding to the target memory access address according to the cache hit signal. Then, the logic operation unit performs a logic operation based on the first data to be operated in the atomic operation instruction and the second data to be operated in the original cache line data to obtain a logic operation result. Furthermore, the data write unit writes the target cache line data corresponding to the logic operation result into the target data memory in the first-level cache. By adopting the solution provided by the embodiment of the present application, executing the atomic operation instruction by multiplexing the cache access pipeline reduces the resource consumption in the instruction execution process and improves the execution efficiency of the atomic operation instruction. Description of the Drawings
[0059] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0060] Figure 1 Fig. shows a schematic structural diagram of an instruction execution device provided by an exemplary embodiment of the present application;
[0061] Figure 2 Fig. shows a schematic structural diagram of an instruction execution device provided by another exemplary embodiment of the present application;
[0062] Figure 3 Fig. shows a schematic structural diagram of an instruction execution device provided by yet another exemplary embodiment of the present application;
[0063] Figure 4 Fig. shows a schematic diagram of an instruction merging process provided by an exemplary embodiment of the present application;
[0064] Figure 5 Fig. shows a schematic structural diagram of an instruction execution device provided by yet another exemplary embodiment of the present application;
[0065] Figure 6 Fig. shows a schematic data flow diagram under the condition of cache hit provided by an exemplary embodiment of the present application;
[0066] Figure 7 Fig. shows a schematic data flow diagram under the condition of cache miss provided by an exemplary embodiment of the present application;
[0067] Figure 8 Fig. shows a flowchart of an instruction execution method provided by an exemplary embodiment of the present application;
[0068] Figure 9 Fig. shows a block diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed Embodiments
[0069] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0070] It should be understood that the "several" mentioned in this article refers to one or more, and "multiple" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0071] Atomic Memory Operation (AMO) is an instruction for processing atomic operations in RISC-V processors. Atomic operations are operations that cannot be divided or interrupted by other instructions or operations, and are usually used in multi-core processor systems. Optionally, AMO instructions can include atomic addition instructions, atomic exchange instructions, atomic logic instructions, and atomic magnitude instructions.
[0072] In the related art, in order to execute an AMO instruction, taking the AMO instruction as an atomic addition instruction as an example, one instruction needs to be split into three instructions to implement, including a data load instruction, which is used to load the data in the memory access address into the processor; an addition instruction, which is used to perform the addition operation; and a data write instruction, which is used to write the addition calculation result into the original memory access address. Instructions are generally stored in a queue in the processor, and too many instructions will occupy limited queue resources, resulting in low queue utilization.
[0073] Therefore, in an embodiment of the present application, by delegating the implementation of the AMO instruction to the cache access pipeline, the logic operation unit directly executes the instruction operation without splitting the instruction, thereby improving the instruction execution efficiency; and by reusing the cache access pipeline, the resource consumption during the instruction execution process is reduced.
[0074] Please refer to Figure 1 , which shows a schematic diagram of the structure of an instruction execution device provided by an exemplary embodiment of the present application. Figure 1 As shown, the instruction execution device 100 provided in the embodiment of the present application includes a load storage unit 110, a tag pipeline 120, a data pipeline 130, a logic operation unit 140 and a data writing unit 150.
[0075] Optionally, the tag pipeline and the data pipeline are cache access pipelines for accessing the first-level cache (L1 Cache). The tag pipeline includes m tag reading paths, and different tag reading paths correspond to different cache lines; the data pipeline includes m data reading paths, and different data reading paths correspond to different cache lines.
[0076] Optionally, a cache line refers to the smallest unit in a cache for storing data. Different cache lines correspond to different tag pipelines and data pipelines in the cache access pipeline.
[0077] Optionally, each cache line in the cache corresponds to an address tag (Tag). For example, the first n bits of the storage address corresponding to the cache line can be set as the address tag corresponding to the cache line, that is, the address tags corresponding to each cache line are unique. Then, in the process of reading data from the cache, it is possible to first determine whether the cache line exists in the cache according to the address tag, and then decide whether to perform data reading, thereby improving the data reading efficiency in the cache.
[0078] In a possible design, each cache line in the first-level cache can be numbered, so that according to the cache line number, the tag read path (Tag Way) and the data read path (Data Way) are in one-to-one correspondence, that is, the number of the tag read path of each cache line in the tag pipeline is the same as the number of the data read path in the data pipeline. For example, Tag Way0 and Data Way0 correspond to cache line0, Tag Way1 and Data Way1 correspond to cache line1, and so on.
[0079] A load store unit (LSU) 110 is used to transmit atomic operation instructions to the tag pipeline and the data pipeline. The atomic operation instructions include a target memory access address and a first data to be operated on.
[0080] Optionally, in the process of reusing the cache access pipeline to execute atomic operation instructions, the instruction execution device can first initiate an atomic operation instruction through the load store unit, that is, the load store unit transmits the atomic operation instructions to the tag pipeline and the data pipeline respectively.
[0081] Optionally, the atomic operation instruction can be an atomic addition instruction, an atomic exchange instruction, an atomic logic instruction, and an atomic fetch maximum / minimum value instruction. The atomic operation instruction can include a target memory access address and a first data to be operated on.
[0082] The tag pipeline 120 is used to output a cache hit signal based on the target address tag corresponding to the target memory access address and the address tags corresponding to the cache lines in each tag read path. The cache hit signal is used to indicate the cache hit status in the first-level cache.
[0083] Optionally, in order to determine whether a cache line corresponding to the target memory access address is stored in the first-level cache, it is necessary to compare the target address tag corresponding to the target memory access address with the address tags corresponding to each cache line. Thus, if there is an address tag that is the same as the target address tag, it is determined that a cache line corresponding to the target memory access address is stored in the first-level cache; if there is no address tag that is the same as the target address tag, it is determined that there is no cache line corresponding to the target memory access address in the first-level cache.
[0084] In a possible design, after receiving the atomic operation instruction transmitted by the load / store unit, the tag pipeline needs to compare the target address tag corresponding to the target memory access address with the address tags corresponding to the cache lines in each tag read channel, so as to output a cache hit signal, which is used to indicate the cache hit status in the first-level cache.
[0085] The data pipeline 130 is used to output the original cache line data corresponding to the target memory access address based on the cache hit signal.
[0086] In a possible design, in response to the atomic operation instruction, the data pipeline reads data in each data read path respectively. In order to improve the data transmission efficiency and reduce the waste of wiring resources, the data pipeline can first determine the cache hit status according to the cache hit signal output by the tag pipeline, and then decide whether to output data.
[0087] Optionally, when the cache hit signal indicates a cache hit, that is, there is a cache line corresponding to the target memory access address in the first-level cache, the data pipeline outputs the original cache line data corresponding to the target memory access address; when the cache hit signal indicates a cache miss, that is, there is no cache line corresponding to the target memory access address in the first-level cache, the data pipeline does not output data.
[0088] The logical operation (Arith Logic Unit, ALU) unit 140 is used to perform a logical operation on the first data to be operated and the second data to be operated in the original cache line data to obtain a logical operation result.
[0089] Optionally, after determining the original cache line data corresponding to the target memory access address, the logical operation unit can perform a logical operation on the first data to be operated and the second data to be operated in the original cache line data, so as to obtain a logical operation result.
[0090] In a possible design, corresponding to the instruction category of the atomic operation instruction, the operation categories that the logical operation unit can execute include but are not limited to addition operation, exchange operation, logical operation, and taking the maximum or minimum value.
[0091] The Data Write unit 150 is configured to write the target cache line data corresponding to the logical operation result into the target data memory in the first-level cache.
[0092] In a possible design, after completing the logical operation corresponding to the atomic operation instruction and obtaining the logical operation result, the target cache line data corresponding to the logical operation result can be written into the target data memory in the first-level cache through the data write unit, thereby realizing the update of the cache data.
[0093] In summary, in the embodiment of the present application, by reusing the cache access pipeline, the load / store unit sends the atomic operation instruction to the tag pipeline and the data pipeline. Thus, the tag pipeline outputs a cache hit signal based on the target address tag corresponding to the target memory access address in the atomic operation instruction and the address tags corresponding to the cache lines in each tag read path, and the data pipeline outputs the original cache line data corresponding to the target memory access address according to the cache hit signal. Then, the logical operation unit performs a logical operation based on the first data to be operated in the atomic operation instruction and the second data to be operated in the original cache line data to obtain the logical operation result. Furthermore, the data write unit writes the target cache line data corresponding to the logical operation result into the target data memory in the first-level cache. By adopting the solution provided in the embodiment of the present application, the atomic operation instruction is executed by reusing the cache access pipeline, reducing the resource consumption during the instruction execution process and improving the execution efficiency of the atomic operation instruction.
[0094] Next, the specific processing flow in the tag pipeline and the data pipeline will be described.
[0095] Please refer to Figure 2 , which shows a schematic structural diagram of an instruction execution device provided by another exemplary embodiment of the present application. In this embodiment, it is described by taking the tag pipeline including 2 tag read paths and the data pipeline including 2 data read paths as an example.
[0096] In a possible design, the tag pipeline includes a Tag Read unit, and each tag read path includes a tag memory, a first register, and a Compare. The tag memory stores the address tags corresponding to the cache lines.
[0097] Exemplarily, the output end of the tag memory in the tag read path is connected to the input end of the first register, and the output end of the first register is connected to the input end of the comparator. Schematically, as Figure 2As shown, the output end of the tag reading unit 201 is connected to the input ends of tag memory 0 and tag memory 1. The output end of tag memory 0 is connected to the input end of the first register 0. The output end of the first register 0 is connected to the input end of comparator 0. The output end of tag memory 1 is connected to the input end of the first register 1. The output end of the first register 1 is connected to the input end of comparator 1.
[0098] Among them, placing the first register between the tag memory and the comparator is to have the first register perform clock gating on the data to optimize the timing issue during data transmission. By placing the first register, there is no need to consider the delay issue during wiring, thus achieving the randomness of wiring, which is beneficial to alleviating the problem of wiring congestion.
[0099] Optionally, after the load storage unit transmits the atomic operation instruction to the tag pipeline, the tag reading unit in the tag pipeline responds to the atomic operation instruction and sends tag reading instructions to the tag memories in each tag reading path in parallel. Furthermore, the tag memories in each tag reading path respond to the tag reading instructions, transmit the address tags to the first register, and the first register performs clock gating on the address tags and then transmits them to the comparator. Thus, the comparator compares the address tags and the target address tags and outputs the tag comparison result. Finally, the tag pipeline outputs the cache hit signal according to the tag comparison results corresponding to each tag reading path.
[0100] Optionally, when the target address tag is the same as the address tag read from the tag memory, it indicates a cache hit; when the target address tag is different from the address tag read from the tag memory, it indicates a cache miss.
[0101] In a possible design, since the cache hit status is divided into two cases: cache hit and cache miss, and there is only a cache hit in one tag reading path in the case of a cache hit, the cache hit (way_sel) signal can be represented by a one-hot signal of binary numbers. Exemplarily, taking the example that the tag pipeline includes two tag reading paths, in the case of a cache miss, the cache hit signal can be represented as 00; in the case where the target address tag and the address tag corresponding to the cache line in tag reading path 0 (Tag Way0) hit, the cache hit signal can be represented as 01; in the case where the target address tag and the address tag corresponding to the cache line in tag reading path 1 (Tag Way1) hit, the cache hit signal can be represented as 10.
[0102] In a possible design, a data read unit and a first multiplexer (Mux) are included in a data pipeline. A data memory and a second register are included in a data read path, and cache line data corresponding to a cache line is stored in the data memory.
[0103] Wherein, an output end of the data memory is connected to an input end of the second register, and an input end of the first multiplexer is connected to an output end of the second register. Schematically, as Figure 2 shown, an output end of the data read unit 202 is connected to input ends of a data memory 0 and a data memory 1. An output end of the data memory 0 is connected to an input end of the second register 0. An input end of the first multiplexer is connected to an output end of the second register 0. An output end of the data memory 1 is connected to an input end of the second register 1. An input end of the first multiplexer is connected to an output end of the second register 1.
[0104] Wherein, the function of placing the second register between the data memory and the first multiplexer is the same as that of placing the first register between the tag memory and the comparator, which will not be elaborated here.
[0105] Optionally, after the load store unit transmits an atomic operation instruction to the data pipeline, the data read unit responds to the atomic operation instruction and sends a data read instruction to the data memories in each data read path. Thus, the data memories respond to the data read instruction, transmit the cache line data to the second register, and the second register performs a pipelining process on the cache line data and then transmits it to the first multiplexer. Furthermore, when the cache hit signal indicates a cache hit, the first multiplexer outputs the original cache line data corresponding to the target memory access address.
[0106] Exemplarily, when the cache hit signal indicates that the target memory access address hits the tag read path 1 (Tag Way1), the first multiplexer outputs the cache line data output in the data read path 1 (Data Way1).
[0107] In a possible design, in order to reduce the latency problem in the data transmission process, a third register can be placed between the tag pipeline and the first multiplexer, that is, an input end of the third register is connected to an output end of the tag pipeline, and an output end of the third register is connected to a strobe end of the first multiplexer. Optionally, the third register is used to receive the cache hit signal transmitted by the tag pipeline and transmit the cache hit signal to the first multiplexer.
[0108] In a possible design, in order to improve data reading efficiency and reduce resource waste during data reading, a Way Predict unit can also be placed in the data pipeline. The Way Predict unit performs way prediction, so that the process of parallelly reading data from each cache line is changed to only reading the cache line data in the predicted way.
[0109] Optionally, the input end of the Way Predict unit is connected to the output end of the load / store unit, and the output end of the Way Predict unit is connected to the input end of the data reading unit. Schematically, as Figure 2 shown, the input end of the Way Predict unit 203 is connected to the output end of the load / store unit, and the output end of the Way Predict unit 203 is connected to the input end of the data reading unit 202.
[0110] Optionally, in response to an atomic operation instruction, the Way Predict unit predicts the data reading way corresponding to the target memory access address and transmits the way prediction result to the data reading unit, so that the data reading unit sends a data reading instruction to the data memory in the data reading way corresponding to the way prediction result. Finally, the first multiplexer can only receive the cache line data output by the second register in the predicted way.
[0111] Optionally, when the data reading way corresponding to the way prediction result is the same as the data reading way corresponding to the target cache line indicated by the cache hit signal in the case of cache hit, the first multiplexer can directly output the original cache line data in the target cache line, thus realizing efficiency optimization in the data reading process.
[0112] Optionally, when the data reading way corresponding to the way prediction result is different from the data reading way corresponding to the cache hit signal, the Way Predict unit needs to re-predict the data reading way corresponding to the target memory access address and transmit a new way prediction result to the data reading unit until the data reading way corresponding to the way prediction result is the same as the data reading way corresponding to the target cache line indicated by the cache hit signal in the case of cache hit.
[0113] In the above embodiments, in the tag pipeline, the address tags corresponding to the cache lines in each tag reading way are parallelly read, and the address tags are compared with the target address tag by a comparator, so as to output a cache hit signal. At the same time, in the data pipeline, the cache line data in each data reading way are parallelly read, and based on the cache hit signal, the original cache line data corresponding to the target memory access address are output, improving the reading efficiency of the cache line data.
[0114] In addition, by placing a path prediction unit in the data pipeline, the path prediction unit first performs path prediction, and then the data reading unit reads the cache line data, optimizing the data reading process and reducing the resource consumption during data reading.
[0115] In a possible design, considering that in the case of cache miss, that is, when the cache line corresponding to the target memory access address does not exist in the first-level cache, it is impossible to complete the execution of the atomic operation instruction only relying on the first-level cache. Therefore, it is also necessary to further send a data reading request to the second-level cache to obtain the original cache line data corresponding to the target memory access address.
[0116] Please refer to Figure 3 , which shows a schematic structural diagram of an instruction execution device provided by another exemplary embodiment of the present application. In this embodiment, it is described by taking an example that there are 2 tag reading paths in the tag pipeline and 2 data reading paths in the data pipeline.
[0117] Optionally, the instruction execution device may further include a request generation unit (MQ request gen) and a refill unit (Refill Unit, RFU). Among them, the input end of the request generation unit is connected to the output end of the tag pipeline.
[0118] Optionally, the request generation unit is configured to receive the cache hit signal transmitted by the tag pipeline, and generate a data miss request when the cache hit signal indicates a cache miss, so as to read the original cache line data corresponding to the target memory access address from the second-level cache (L2 Cache) based on the data miss request.
[0119] In a possible design, in order to improve the timing of reading data from the second-level cache, a miss queue (Miss Queue, MQ) may also be set. Among them, the input end of the miss queue is connected to the output end of the request generation unit, and the output end of the miss queue is connected to the input end of the second-level cache.
[0120] Schematically, as Figure 3 shown, the input end of the miss queue 302 is connected to the output end of the request generation unit 301, and the output end of the miss queue 302 is connected to the input end of the second-level cache 303.
[0121] Optionally, after the request generation unit generates a data miss request, the data miss request can be written into the miss queue, and the data miss request is transmitted to the second-level cache through the miss queue. Thus, the refill unit receives the request feedback corresponding to the data miss request output by the second-level cache, and the refill unit outputs the original cache line data corresponding to the target memory access address according to the request feedback.
[0122] In a possible design, considering that the data sources of the original cache line data corresponding to the target memory access address are different for both cache hit and cache miss cases. In the case of a cache hit, the original cache line data is output by the first multiplexer; in the case of a cache miss, the original cache line data is output by the refill unit. Therefore, in order to improve data transmission efficiency, a second multiplexer can also be placed in the instruction execution device.
[0123] Among them, the input end of the second multiplexer is connected to the output end of the first multiplexer, the input end of the second multiplexer is connected to the output end of the refill unit, and the output end of the second multiplexer is connected to the input end of the logical operation unit. Schematically, as Figure 3 shown, the input end of the refill unit 304 is connected to the output end of the secondary cache 303, and the output end of the refill unit 304 is connected to the input end of the second multiplexer.
[0124] Optionally, when the cache hit signal indicates a cache hit, the second multiplexer receives the original cache line data corresponding to the target memory access address output by the first multiplexer and transmits the original cache line data to the logical operation unit.
[0125] Optionally, when the cache hit signal indicates a cache miss, the second multiplexer receives the original cache line data corresponding to the target memory access address output by the refill unit and transmits the original cache line data to the logical operation unit.
[0126] Optionally, since a cache line contains stored data corresponding to multiple memory access addresses, and in the logical operation process, only the second data to be operated in the original cache line data is logically operated with the first data to be operated. Therefore, after obtaining the logical operation result, when replacing the second data to be operated with the logical operation result, a data merging operation also needs to be performed.
[0127] In a possible design, the instruction execution device also includes a data merge unit and a fourth register. Among them, the input end of the data merge unit is connected to the output end of the second multiplexer, the input end of the data merge unit is connected to the output end of the logical operation unit, the output end of the data merge unit is connected to the input end of the fourth register, and the output end of the fourth register is connected to the input end of the data writing unit.
[0128] Schematically, as Figure 3 shown, the input end of the data merge unit 306 is connected to the output end of the second multiplexer, the input end of the data merge unit 306 is connected to the output end of the logical operation unit 305, the output end of the data merge unit 306 is connected to the input end of the fourth register, and the output end of the fourth register is connected to the input end of the data writing unit 307.
[0129] Among them, the role of the fourth register placed between the data merging unit and the data writing unit is the same as that of the first register placed between the tag memory and the comparator, which will not be elaborated here.
[0130] Optionally, in addition to receiving the logical operation result output by the logical operation unit, the data merging unit also needs to receive the original cache line data output by the second multiplexer, and then perform data merging based on the original cache line data and the logical operation result to obtain the target cache line data, and the fourth register performs pipelining processing on the target cache line data and then transmits it to the data writing unit.
[0131] Schematically, as Figure 4 shown, taking the atomic operation instruction being an atomic operation performed on W2 as an example, a cache line includes 8 word data, W2 is the second data to be operated in the original cache line data. After the logical operation unit performs a logical operation on the first data to be operated and the second data to be operated, the logical operation result can be used as the new W2 data, and thus the target cache line data is obtained through data merging.
[0132] Optionally, considering the two cases of cache hit and cache miss, the data sources of the original cache line data corresponding to the target memory access address are different. Therefore, in the data writing process, the writing method of the target cache line data also needs to be divided into two cases.
[0133] Optionally, when the original cache line data is located in the first-level cache, the data writing unit can directly determine the target data memory corresponding to the original cache line data in the first-level cache and write the target cache line data into the target data memory.
[0134] Optionally, when the original cache line data is located in the second-level cache, since each data memory in the first-level cache stores other cache line data, in order to reduce the impact on the data storage in the first-level cache, the data writing unit can determine the data memory with the longest data update time distance from the current time in the first-level cache as the target data memory according to the data update conditions of the data memories corresponding to each data pipeline, so as to clear the cache line data originally stored in the target data memory and write the target cache line data into the target data memory.
[0135] In the above embodiments, when the cache hit signal indicates a cache hit, the original cache line data is directly output through the first multiplexer, and logical operations and data merging are performed to obtain the target cache line data; when the cache hit signal indicates a cache miss, the request generation unit is required to generate a data miss request, so as to obtain the original cache line data from the secondary cache and output it via the backfill unit. By providing two data acquisition methods, the smooth completion of atomic operation instructions is ensured while reusing the cache access pipeline, and the instruction execution process is optimized.
[0136] In a possible design, during the cache access process, in addition to considering whether the cache is hit, it is also necessary to determine the data read and write permissions corresponding to each cache line. That is, only when the cache is hit and the original cache line data corresponding to the target memory access address has read and write permissions, the first multiplexer can effectively output the original cache line data corresponding to the target memory access address.
[0137] Please refer to Figure 5 , which shows a schematic structural diagram of an instruction execution device provided by another exemplary embodiment of the present application.
[0138] In a possible design, in addition to the tag pipeline for reading address tags and the data pipeline for reading cache line data, the instruction execution device also needs to include a status pipeline, which includes m status reading paths, and different status reading paths correspond to different cache lines.
[0139] Schematically, as Figure 5 shown, taking the instruction execution device including two tag reading paths (Tag Way0 and TagWay1), two status reading paths (Status Way0 and Status Way1), and two data reading paths (DataWay0 and Data Way1) as an example.
[0140] Optionally, in response to the atomic operation instruction transmitted by the load / store unit, the status pipeline outputs the cache line status corresponding to each cache line through each status reading path, and the cache line status is used to represent the data read and write permissions corresponding to the cache line. For example, Status Way0 outputs the data read and write permissions corresponding to Cache Line0, and Status Way1 outputs the data read and write permissions corresponding to CacheLine1.
[0141] In a possible design, the status pipeline may further include a Status Read unit and a third multiplexer. The status read path includes a status memory and a fifth register. The input end of the third multiplexer is connected to the output end of the fifth register, and the strobe end of the third multiplexer is connected to the output end of the tag pipeline.
[0142] Schematically, as Figure 5 shown, the output end of the status read unit 501 is connected to the input ends of status memory 0 and status memory 1. The output end of status memory 0 is connected to the input end of the fifth register 0. The output end of the fifth register 0 is connected to the input end of the third multiplexer. The output end of status memory 1 is connected to the input end of the fifth register 1. The output end of the fifth register 1 is connected to the input end of the third multiplexer.
[0143] Among them, the function of placing the fifth register between the status memory and the third multiplexer is the same as that of placing the first register between the tag memory and the comparator, which will not be elaborated here.
[0144] Optionally, after the load / store unit transmits an atomic operation instruction to the status pipeline, the status read unit responds to the atomic operation instruction and sends a status read instruction to the status memories in each status read path. As a result, the status memories respond to the status read instruction, transmit the cache line status to the fifth register, and the fifth register beats the cache line status and then transmits it to the third multiplexer.
[0145] Optionally, considering that there are two cases of cache hit and cache miss, and there are also two cases of having read / write permission and not having read / write permission. And it is meaningless to output the cache line status in the case of cache miss, and it is also meaningless to output the cache line status in the case of cache hit and having read / write permission. Therefore, in order to reduce resource consumption, the third multiplexer can determine whether to output data according to the cache hit signal and the target cache line status.
[0146] Optionally, when the cache hit signal indicates a cache miss, the third multiplexer does not output the cache line status; when the cache hit signal indicates a cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address has data read / write permission, the third multiplexer can also not output the target cache line status.
[0147] Optionally, when the cache hit signal indicates a cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address does not have data read / write permission, that is, when the first multiplexer cannot output valid data based on the cache hit signal, the third multiplexer can transmit the target cache line status (status_out) to the request generation unit.
[0148] Schematically, as Figure 5 shown, the output terminal of the third multiplexer is connected to the input terminal of the request generation unit 502.
[0149] Optionally, when the cache hit signal indicates a cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address does not have data read / write permission, there can be two data transmission methods. One is to request to obtain the data read / write permission of the target cache line and continue to read the original cache line data corresponding to the target memory access address from the first-level cache; the other is to directly request to read the original cache line data corresponding to the target memory access address from the second-level cache. Therefore, the request generation unit can have two request generation methods.
[0150] In a possible design, the request generation unit can generate a data miss request according to the target cache line status, write the data miss request into the miss request queue, and transmit the data miss request to the second-level cache through the miss request queue, so that the refill unit can transmit the original cache line data corresponding to the target memory access address to the third multiplexer according to the request feedback transmitted by the second-level cache.
[0151] In another possible design, the request generation unit can generate a permission acquisition request according to the target cache line status. This permission acquisition request is used to request to obtain the data read / write permission of the target cache line in the first-level cache. Furthermore, when the target cache line has data read / write permission, the data memory corresponding to the target memory access address in the data pipeline can output the cache line data to the second register, so that the second multiplexer can respond to the cache hit signal and transmit the original cache line data corresponding to the target memory access address to the third multiplexer.
[0152] In the above embodiments, during the execution of the atomic operation instruction, in addition to considering whether the cache is hit, it is also judged through the status pipeline whether the target cache line has data read / write permission when the cache is hit. Thus, only when the cache is hit and has data read / write permission, the first multiplexer will output the original cache line data corresponding to the target memory access address, ensuring the effectiveness of data output.
[0153] And in the case of cache hit but without data read / write permission, a data miss request is generated by the request generation unit, so as to directly obtain the original cache line data from the secondary cache; or, a permission acquisition request is generated by the request generation unit, and after obtaining the permission, the original cache line data is output through the first multiplexer, ensuring the smooth execution of the atomic operation instruction and optimizing the execution process of the atomic operation instruction.
[0154] Please refer to Figure 6 , which shows a data flow diagram in the case of cache hit provided by an exemplary embodiment of the present application. Taking the example that there are two paths (way0 and way1) in the pipeline and the cache hits way0 for illustration.
[0155] First, the load / store unit transmits the atomic operation instruction to the tag pipeline, the status pipeline, and the data pipeline respectively. The tag pipeline responds to the atomic operation instruction, reads the address tags in tag way0 and tag way1 respectively, and compares the tags through the comparator. In the case where the target address tag hits way0, the tag pipeline outputs a cache hit signal 01, and transmits it to mux2 through the register. At the same time, the data pipeline responds to the atomic operation instruction, reads the cache line data in data way0 and data way1 respectively, and transmits it to mux2 through the register. Thus, mux2 outputs the cache line data in data way0 as the original cache line data corresponding to the atomic operation instruction according to the cache hit signal 01, and then performs logical operations through the arithmetic logic unit (ALU), and performs data merging through the data merge unit (merge), and finally transmits it to the data write unit (data write) through the register, and data write writes the target cache line data into data way0.
[0156] Please refer to Figure 7 , which shows a data flow diagram in the case of cache miss provided by an exemplary embodiment of the present application. Taking the example that there are two paths (way0 and way1) in the pipeline for illustration.
[0157] First, the load store unit transmits atomic operation instructions to the tag pipeline, the status pipeline, and the data pipeline respectively. The tag pipeline responds to the atomic operation instructions, reads the address tags in tag way0 and tag way1 respectively, and compares the tags through a comparator. In the case of a cache miss, the tag pipeline outputs a cache hit signal 00 and transmits it to the request generation unit (MQ request gen). The request generation unit generates a data miss request, thereby obtaining the original cache line data corresponding to the target memory access address in the secondary cache (L2cache), performing logical operations through the arithmetic logic unit (ALU), performing data merging through the data merge unit (merge), and finally transmitting the data to the data write unit through the register. When data way1 is a memory that has not been updated for a long time, the data write unit writes the target cache line data into data way1.
[0158] Please refer to Figure 8 , which shows a flowchart of an instruction execution method provided by an exemplary embodiment of the present application. This method is used for the instruction execution device provided by each of the above embodiments. The method includes:
[0159] Step 801, transmit atomic operation instructions to the tag pipeline and the data pipeline through the load store unit. The atomic operation instructions include the target memory access address and the first data to be operated.
[0160] Step 802, output a cache hit signal through the tag pipeline based on the target address tag corresponding to the target memory access address and the address tags corresponding to the cache lines in each tag reading path. The cache hit signal is used to indicate the cache hit status in the first-level cache.
[0161] Step 803, output the original cache line data corresponding to the target memory access address through the data pipeline based on the cache hit signal.
[0162] Step 804, perform a logical operation on the first data to be operated and the second data to be operated in the original cache line data through the arithmetic logic unit to obtain a logical operation result.
[0163] Step 805, write the target cache line data corresponding to the logical operation result into the target data memory in the first-level cache through the data write unit.
[0164] In some embodiments, the tag pipeline includes a tag reading unit, and the tag reading path includes a tag memory, a first register, and a comparator. The tag memory stores the address tags corresponding to the cache lines.
[0165] Optionally, the tag reading unit responds to the atomic operation instruction and sends a tag reading instruction to the tag memory in each tag reading path; the tag memory responds to the tag reading instruction, transmits the address tag to the first register, and the first register performs a pipelining process on the address tag and then transmits it to the comparator; the comparator compares the address tag and the target address tag and outputs a tag comparison result; the tag pipeline outputs a cache hit signal based on the tag comparison results corresponding to each tag reading path.
[0166] In some embodiments, the data pipeline includes a data reading unit and a first multiplexer, the data reading path includes a data memory and a second register, the data memory stores cache line data corresponding to a cache line, and the input end of the first multiplexer is connected to the output end of the second register.
[0167] Optionally, the data reading unit responds to the atomic operation instruction and sends a data reading instruction to the data memory in each data reading path; the data memory responds to the data reading instruction, transmits the cache line data to the second register, and the second register performs a pipelining process on the cache line data and then transmits it to the first multiplexer; the first multiplexer outputs the original cache line data corresponding to the target memory access address when the cache hit signal indicates a cache hit.
[0168] In some embodiments, the device further includes a third register, the input end of the third register is connected to the output end of the tag pipeline, and the output end of the third register is connected to the strobe end of the first multiplexer.
[0169] Optionally, the third register receives the cache hit signal transmitted by the tag pipeline and transmits the cache hit signal to the first multiplexer.
[0170] In some embodiments, the data pipeline further includes a path prediction unit, the input end of the path prediction unit is connected to the output end of the load / store unit, and the output end of the path prediction unit is connected to the input end of the data reading unit.
[0171] Optionally, the path prediction unit responds to the atomic operation instruction, predicts the data reading path corresponding to the target memory access address, and transmits a path prediction result to the data reading unit; the data reading unit sends a data reading instruction to the data memory in the data reading path corresponding to the path prediction result.
[0172] In some embodiments, when the data reading path corresponding to the path prediction result is different from the data reading path corresponding to the cache hit signal, the path prediction unit re-predicts the data reading path corresponding to the target memory access address and transmits a new path prediction result to the data reading unit.
[0173] In some embodiments, the apparatus further includes a request generation unit and a backfill unit, and an input end of the request generation unit is connected to an output end of the tag pipeline.
[0174] Optionally, the request generation unit receives a cache hit signal transmitted by the tag pipeline; when the cache hit signal indicates a cache miss, the request generation unit generates a data miss request, writes the data miss request into a miss request queue, and transmits the data miss request to a secondary cache through the miss request queue; the backfill unit outputs original cache line data corresponding to a target memory access address based on a request feedback corresponding to the data miss request.
[0175] In some embodiments, the apparatus further includes a second multiplexer, an input end of the second multiplexer is connected to an output end of the first multiplexer, the input end of the second multiplexer is connected to an output end of the backfill unit, and an output end of the second multiplexer is connected to an input end of the logic operation unit.
[0176] Optionally, when the cache hit signal indicates a cache hit, the second multiplexer receives original cache line data corresponding to a target memory access address output by the first multiplexer, and transmits the original cache line data to the logic operation unit.
[0177] Optionally, when the cache hit signal indicates a cache miss, the second multiplexer receives original cache line data corresponding to a target memory access address output by the backfill unit, and transmits the original cache line data to the logic operation unit.
[0178] In some embodiments, the apparatus further includes a data merging unit and a fourth register, an input end of the data merging unit is connected to an output end of the second multiplexer, the input end of the data merging unit is connected to an output end of the logic operation unit, and an output end of the data merging unit is connected to an input end of the fourth register.
[0179] Optionally, the data merging unit receives original cache line data output by the second multiplexer and a logic operation result output by the logic operation unit; performs data merging based on the original cache line data and the logic operation result to obtain target cache line data, and the fourth register performs pipelining processing on the target cache line data and then transmits the data to a data writing unit.
[0180] In some embodiments, when the original cache line data is located in the first-level cache, the data writing unit determines the target data memory corresponding to the original cache line data in the first-level cache, and writes the target cache line data into the target data memory; when the original cache line data is located in the second-level cache, the data writing unit determines the data memory with the longest data update time distance from the current time in the first-level cache as the target data memory based on the data update conditions of the data memories corresponding to each data pipeline, and writes the target cache line data into the target data memory.
[0181] In some embodiments, the apparatus further includes a status pipeline, and the status pipeline includes m status reading paths, and different status reading paths correspond to different cache lines.
[0182] Optionally, the status pipeline responds to an atomic operation instruction, and outputs the cache line status corresponding to each cache line through each status reading path, and the cache line status is used to represent the data read / write permission corresponding to the cache line.
[0183] In some embodiments, the status pipeline further includes a status reading unit and a third multiplexer, and the status reading path includes a status memory and a fifth register; the input end of the third multiplexer is connected to the output end of the fifth register, and the selection end of the third multiplexer is connected to the output end of the tag pipeline.
[0184] Optionally, the status reading unit responds to an atomic operation instruction, and sends a status reading instruction to the status memory in each status reading path; the status memory responds to the status reading instruction, transmits the cache line status to the fifth register, and the fifth register performs a beating process on the cache line status and then transmits it to the third multiplexer; the third multiplexer does not output the cache line status when the cache hit signal indicates a cache miss; the third multiplexer does not output the target cache line status when the cache hit signal indicates a cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address has data read / write permission.
[0185] In some embodiments, when the cache hit signal indicates a cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address does not have data read / write permission, the third multiplexer transmits the target cache line status to the request generation unit; the request generation unit generates a data miss request based on the target cache line status, writes the data miss request into the miss request queue, and transmits the data miss request to the second-level cache through the miss request queue; or, the request generation unit generates a permission acquisition request based on the target cache line status, and the permission acquisition request is used to request to obtain the data read / write permission corresponding to the target cache line in the first-level cache.
[0186] In summary, in the embodiments of the present application, by reusing the cache access pipeline, the load / store unit sends atomic operation instructions to the tag pipeline and the data pipeline. Thus, based on the target address tag corresponding to the target memory access address in the atomic operation instruction and the address tags corresponding to the cache lines in each tag read path, the tag pipeline outputs a cache hit signal, and the data pipeline outputs the original cache line data corresponding to the target memory access address according to the cache hit signal. Then, the logical operation unit performs a logical operation based on the first data to be operated in the atomic operation instruction and the second data to be operated in the original cache line data to obtain a logical operation result. Furthermore, the data write unit writes the target cache line data corresponding to the logical operation result into the target data memory in the first-level cache. By adopting the solution provided by the embodiments of the present application, the atomic operation instructions are executed by reusing the cache access pipeline, reducing the resource consumption during the instruction execution process and improving the execution efficiency of the atomic operation instructions.
[0187] In some embodiments, the instruction execution device in the embodiments of the present application may be integrated in the processor or independently provided outside the processor.
[0188] Please refer to Figure 9 , which shows a structural block diagram of a computer device 900 provided by an exemplary embodiment of the present application. Among them, the computer device 900 may be a portable mobile terminal, such as: a smart phone, a tablet computer, a Moving Picture Experts Group Audio Layer III (MP3) player, a Moving Picture Experts Group Audio Layer IV (MP4) player. The computer device 900 may also be referred to by other names such as user equipment, portable terminal, workstation, server, etc.
[0189] Generally, the computer device 900 includes: a processor 901 and a memory 902.
[0190] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0191] In some embodiments, the processor 901 may be integrated with the instruction execution device provided in the above embodiments, or the processor 901 may also be connected to an independently provided instruction execution device. When the processor 901 has a need to execute atomic operation instructions, the atomic operation instructions can be executed through this instruction execution device.
[0192] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be tangible and non-transitory. The memory 902 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices.
[0193] In some embodiments, the computer device 900 may also optionally include a peripheral device interface 903 and at least one peripheral device.
[0194] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the computer device 900, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0195] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0196] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An instruction execution device, characterized in that, The device includes a load store unit, a tag pipeline, a data pipeline, a logic operation unit, and a data write unit. The tag pipeline and the data pipeline are used to access the level-1 cache. The tag pipeline includes m tag read paths, and different tag read paths correspond to different cache lines. The data pipeline includes m data read paths, and different data read paths correspond to different cache lines; The load store unit is used to transmit atomic operation instructions to the tag pipeline and the data pipeline. The atomic operation instructions include a target memory access address and a first data to be operated; The tag pipeline is used to output a cache hit signal based on the target address tag corresponding to the target memory access address and the address tags corresponding to the cache lines in each tag read path. The cache hit signal is used to indicate the cache hit status in the level-1 cache; The data pipeline is used to output the original cache line data corresponding to the target memory access address based on the cache hit signal; The logic operation unit is used to perform a logic operation on the first data to be operated and a second data to be operated in the original cache line data to obtain a logic operation result; The data write unit is used to write the target cache line data corresponding to the logic operation result into the target data memory in the level-1 cache.
2. The device according to claim 1, wherein The tag pipeline includes a tag read unit. The tag read path includes a tag memory, a first register, and a comparator. The tag memory stores the address tags corresponding to the cache lines; The tag read unit is used to respond to the atomic operation instructions and send tag read instructions to the tag memories in each tag read path; The tag memory is used to respond to the tag read instructions, transmit the address tags to the first register, and after the first register performs a pipelining process on the address tags, transmit them to the comparator; The comparator is used to compare the address tags and the target address tags and output a tag comparison result; The tag pipeline is used to output the cache hit signal based on the tag comparison results corresponding to each tag read path.
3. The device according to claim 2, wherein The data pipeline includes a data read unit and a first multiplexer. The data read path includes a data memory and a second register. The data memory stores the cache line data corresponding to the cache lines. The input end of the first multiplexer is connected to the output end of the second register; The data read unit is used to respond to the atomic operation instructions and send data read instructions to the data memories in each data read path; The data memory is used to respond to the data read instructions, transmit the cache line data to the second register, and after the second register performs a pipelining process on the cache line data, transmit them to the first multiplexer; The first multiplexer is used to output the original cache line data corresponding to the target memory access address when the cache hit signal indicates a cache hit.
4. The device according to claim 3, characterized in that, The device further includes a third register, the input end of the third register is connected to the output end of the tag pipeline, and the output end of the third register is connected to the strobe end of the first multiplexer; The third register is configured to receive the cache hit signal transmitted by the tag pipeline and transmit the cache hit signal to the first multiplexer.
5. The device according to claim 3, wherein The data pipeline further includes a path prediction unit, the input end of the path prediction unit is connected to the output end of the load / store unit, and the output end of the path prediction unit is connected to the input end of the data reading unit; The path prediction unit is configured to respond to the atomic operation instruction, predict the data reading path corresponding to the target memory access address, and transmit the path prediction result to the data reading unit; The data reading unit is configured to send the data reading instruction to the data memory in the data reading path corresponding to the path prediction result.
6. The device according to claim 5, characterized in that, The path prediction unit is further configured to: In the case that the data reading path corresponding to the path prediction result is different from the data reading path corresponding to the cache hit signal, re-predict the data reading path corresponding to the target memory access address, and transmit the new path prediction result to the data reading unit.
7. The device according to claim 3, characterized in that, The device further includes a request generation unit and a fill-back unit, the input end of the request generation unit is connected to the output end of the tag pipeline; The request generation unit is configured to receive the cache hit signal transmitted by the tag pipeline; The request generation unit is further configured to generate a data miss request in the case that the cache hit signal indicates a cache miss, write the data miss request into the miss request queue, and transmit the data miss request to the secondary cache through the miss request queue; The fill-back unit is configured to output the original cache line data corresponding to the target memory access address based on the request feedback corresponding to the data miss request.
8. The device according to claim 7, characterized in that, The device further includes a second multiplexer, the input end of the second multiplexer is connected to the output end of the first multiplexer, the input end of the second multiplexer is connected to the output end of the fill-back unit, and the output end of the second multiplexer is connected to the input end of the logic operation unit; The second multiplexer is configured to receive the original cache line data corresponding to the target memory access address output by the first multiplexer and transmit the original cache line data to the logic operation unit in the case that the cache hit signal indicates a cache hit; The second multiplexer is further configured to receive the original cache line data corresponding to the target memory access address output by the fill-back unit and transmit the original cache line data to the logic operation unit in the case that the cache hit signal indicates a cache miss.
9. The device according to claim 8, characterized in that, The device further includes a data merging unit and a fourth register, the input end of the data merging unit is connected to the output end of the second multiplexer, the input end of the data merging unit is connected to the output end of the logic operation unit, and the output end of the data merging unit is connected to the input end of the fourth register; The data merging unit is configured to receive the original cache line data output by the second multiplexer and the logical operation result output by the logical operation unit; The data merging unit is further configured to perform data merging based on the original cache line data and the logical operation result to obtain the target cache line data, and the fourth register performs pipelining processing on the target cache line data and then transmits it to the data writing unit.
10. The device according to claim 9, characterized in that, The data writing unit is further configured to: When the original cache line data is located in the first-level cache, determine the target data memory corresponding to the original cache line data in the first-level cache, and write the target cache line data into the target data memory; When the original cache line data is located in the second-level cache, based on the data update conditions of the data memories corresponding to the data pipelines, determine the data memory with the longest time distance from the data update moment to the current moment in the first-level cache as the target data memory, and write the target cache line data into the target data memory.
11. The device according to any one of claims 1 to 10, characterized in that The apparatus further includes a status pipeline, and the status pipeline includes m status reading paths, and different status reading paths correspond to different cache lines; The status pipeline is configured to respond to the atomic operation instruction, and output the cache line status corresponding to each cache line through each status reading path, and the cache line status is used to represent the data read / write permission corresponding to the cache line.
12. The device according to claim 11, wherein, The status pipeline further includes a status reading unit and a third multiplexer, and the status reading path includes a status memory and a fifth register; The input end of the third multiplexer is connected to the output end of the fifth register, and the strobe end of the third multiplexer is connected to the output end of the tag pipeline; The status reading unit is configured to respond to the atomic operation instruction and send a status reading instruction to the status memory in each status reading path; The status memory is configured to respond to the status reading instruction, transmit the cache line status to the fifth register, and the fifth register performs pipelining processing on the cache line status and then transmits it to the third multiplexer; The third multiplexer is configured to not output the cache line status when the cache hit signal indicates cache miss; The third multiplexer is further configured to not output the target cache line status when the cache hit signal indicates cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address has data read / write permission.
13. The apparatus according to claim 12, wherein The third multiplexer is further configured to transmit the target cache line status to the request generation unit when the cache hit signal indicates cache hit and the target cache line status indicates that the target cache line corresponding to the target memory access address does not have data read / write permission. The request generation unit is configured to generate a data miss request based on the target cache line state, write the data miss request into a miss request queue, and transmit the data miss request to a secondary cache through the miss request queue; or, The request generation unit is configured to generate a permission acquisition request based on the target cache line state, where the permission acquisition request is used to request to acquire the data read / write permission of the target cache line in the first-level cache.
14. An instruction execution method, characterized in that, The method is used for an instruction execution device according to any one of claims 1 to 13, where the instruction execution device includes a load storage unit, a tag pipeline, a data pipeline, a logical operation unit, and a data writing unit; The method includes: Transmitting an atomic operation instruction to the tag pipeline and the data pipeline through the load storage unit, where the atomic operation instruction includes a target memory access address and a first data to be operated; Outputting a cache hit signal through the tag pipeline based on the target address tag corresponding to the target memory access address and the address tags corresponding to the cache lines in each tag reading path, where the cache hit signal is used to indicate the cache hit state in the first-level cache; Outputting the original cache line data corresponding to the target memory access address through the data pipeline based on the cache hit signal; Performing a logical operation on the first data to be operated and a second data to be operated in the original cache line data through the logical operation unit to obtain a logical operation result; Writing the target cache line data corresponding to the logical operation result into a target data memory in the first-level cache through the data writing unit.
15. A processor, characterized in that, The processor includes an instruction execution device according to any one of claims 1 to 13.
16. A computer device, characterized in that, The computer device includes a processor and a memory according to claim 15, where the processor is connected to the memory through a bus.
17. A computer device, characterized in that, The computer device includes a processor, a memory, and an instruction execution device according to any one of claims 1 to 13, where the processor is connected to the instruction execution device, and the processor is connected to the memory through a bus.