Instruction processing assembly, instruction processing method, processor, and related apparatus
By introducing a data load queue circuit into the processor, the problem of frequently re-issuing memory access instructions when cache is missing is solved, and the utilization rate of the pipeline and the efficiency of the processor are improved.
Patent Information
- Application Number
- PCT/CN2024/120291
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-09-23
- Publication Date
- 2025-07-03
AI Technical Summary
Existing processors frequently re-issues memory access instructions when cache is missing, resulting in waste of pipeline resources and increased power consumption, reducing the efficiency of the processor.
The data loading queue circuit is introduced into the processor. The main pipeline circuit writes the data address of the memory access instruction to the data loading queue circuit when the cache query fails. The data loading queue circuit queries and outputs data from the backfill data, reducing the number of resents of the memory access instruction.
It effectively reduces the number of resents of memory access instructions, improves the utilization rate of the mainstream pipeline, and improves the efficiency of processor processing memory access instructions.
Smart Images

Figure CN2024120291_03072025_PF_FP_ABST
Abstract
Description
Instruction processing component, instruction processing method, processor and related device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 25, 2023, with application number 2023117973888 and application name “Instruction processing component, instruction processing method, processor and computer device”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of chip technology, and in particular to instruction processing. Background Art
[0003] In the processor, when the processor receives a memory access request, it accesses the multi-level cache read request information. If the data is missing in the multi-level cache, the processor will search the memory.
[0004] In the related art, when the processor detects that the data required for the memory access request is missing in the cache during the memory access process, it will directly put the memory access instruction back into the transmission queue for re-transmission until the transmitted memory access instruction obtains complete data.
[0005] However, in the above scheme, since the instruction will be directly reissued when the cache is missed, a memory access instruction will be issued multiple times. A memory access instruction being reissued multiple times will cause the pipeline that subsequently processes the memory access instruction to process the memory access instruction multiple times, occupying pipeline resources and reducing the effective utilization rate of the pipeline. In addition, when the memory access instruction execution device of the pipeline processes the memory access instruction, the power consumption of reading the memory is very high, which will result in a waste of power consumption of reading the memory.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide an instruction processing component, an instruction processing method, a processor, and a computer device, which can improve the efficiency of the processor in processing memory access instructions. The technical solution is as follows.
[0008] In one aspect, an instruction processing component is provided, the instruction processing component comprising: a main pipeline circuit, a data buffer, and a data load queue circuit;
[0009] The main pipeline circuit is electrically connected to the data buffer, and the main pipeline circuit is electrically connected to the data loading queue circuit;
[0010] The main pipeline circuit is used to receive a memory access instruction and send an access request to the data buffer based on the memory access instruction, wherein the access request is used to request to query the data buffer for data corresponding to the memory access instruction;
[0011] The main pipeline circuit is further configured to write the data address corresponding to the memory access instruction into the data loading queue circuit when the query of the data corresponding to the memory access instruction from the data buffer fails;
[0012] The data loading queue circuit is used to receive backfill data, query corresponding data from the backfill data based on the data address written from the mainstream pipeline circuit, and output the queried data; the backfill data is the data filled into the data buffer.
[0013] In another aspect, a method for processing an instruction is provided. The method is performed by the instruction processing component described above, and the method includes:
[0014] Receiving a memory access instruction through the main pipeline circuit, and sending an access request to the data buffer based on the memory access instruction, wherein the access request is used to request to query the data buffer for data corresponding to the memory access instruction;
[0015] When the query of the data corresponding to the memory access instruction from the data buffer fails, the data address corresponding to the memory access instruction is written into the data loading queue circuit through the main pipeline circuit;
[0016] The backfill data is received through the data loading queue circuit, and based on the data address written from the mainstream pipeline circuit, the corresponding data is queried from the backfill data and the queried data is output; the backfill data is the data filled into the data buffer.
[0017] On the other hand, a processor is provided, comprising at least one instruction processing component as described above.
[0018] On the other hand, a computer device is provided. The computer device includes at least one processor, and the processor includes at least one instruction processing component as described above.
[0019] On the other hand, an embodiment of the present application provides a storage medium, which is used to store a computer program, and the computer program is used to execute the method of the above aspect.
[0020] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when executed on a computer, enables the computer to execute the above method.
[0021] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:
[0022] A data loading queue circuit is set in the processor. When the mainstream pipeline circuit fails to query the data corresponding to the memory access instruction from the data cache, the data address corresponding to the memory access instruction is written into the data loading queue circuit. The data loading queue circuit receives the backfill data, queries the corresponding data from the backfill data based on the written data address, and outputs the queried data. That is to say, in the above scheme, when the mainstream pipeline circuit fails to query from the data cache during the process of processing the memory access instruction, the memory access instruction is not directly reissued. Instead, the data loading queue circuit queries the corresponding data from the backfill data of the data cache and outputs it. This can effectively reduce the number of reissues of the memory access instruction, improve the utilization rate of the mainstream pipeline, and thereby improve the efficiency of the processor in processing the memory access instruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] FIG1 is a schematic diagram of the structure of a processor;
[0024] FIG2 is a schematic diagram of a conventional instruction processing component work pipeline;
[0025] FIG3 is a schematic diagram of a work pipeline of an instruction processing component provided by an exemplary embodiment of the present application;
[0026] FIG4 is a schematic diagram of the structure of an instruction processing component provided by an exemplary embodiment of the present application;
[0027] FIG5 is a schematic diagram of the structure of an instruction processing component provided by an exemplary embodiment of the present application;
[0028] FIG6 is a schematic diagram of the structure of an instruction processing component provided by an exemplary embodiment of the present application;
[0029] FIG7 is a schematic diagram of the structure of an instruction processing component provided by an exemplary embodiment of the present application;
[0030] FIG8 is a schematic diagram of the structure of an instruction processing component provided by an exemplary embodiment of the present application;
[0031] FIG9 is a schematic diagram of the structure of an instruction processing component provided by an exemplary embodiment of the present application;
[0032] FIG10 is a schematic diagram of the operation of an instruction processing component provided by an exemplary embodiment of the present application;
[0033] FIG11 is a schematic diagram of a pipeline structure according to an exemplary embodiment of the present application;
[0034] FIG12 is a schematic diagram of a micro-architecture structure involved in an exemplary embodiment of the present application;
[0035] FIG13 is a schematic diagram of the operation of a refill split unit according to an exemplary embodiment of the present application;
[0036] FIG14 is a schematic diagram of a right shifter operation according to an exemplary embodiment of the present application;
[0037] FIG15 is a flowchart of an instruction processing method according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0039] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter without departing from the scope of this disclosure. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0040] The following first introduces some concepts involved in this application:
[0041] 1) Main pipe: The main pipeline is used to execute instructions and complete the processor's computing tasks. In a processor, the processing of an instruction is divided into multiple stages, and different stages are executed by different circuits in the main pipeline.
[0042] 2) MSHR: miss status handling register; a register used to record status information when a cache miss occurs in the processor. It is usually used in the processor's control logic to take appropriate measures to handle cache misses when a miss occurs.
[0043] 3) LDQ: load queue, data load queue; a data structure used inside the processor to record the status of load instructions being executed and the data involved in these instructions.
[0044] 4) Dff: D-flip-flop, D trigger; a basic timing element in digital circuits, used to store the state of binary data.
[0045] 5) Mux: multiplexer, a digital circuit component used to select one of multiple input signals and pass it to a single output. It functions like a switch, which can select an output from multiple inputs.
[0046] 6) Round robin (RR): This is an arbitrator design used to coordinate access to shared resources among multiple devices. In this arbitration scheme, each device has an opportunity to gain access to the shared resource within a round, ensuring fairness and uniformity. The basic idea is to sequentially assign arbitration signals to each device, ensuring that each device has a chance to obtain arbitration rights within a round. If a device does not obtain arbitration rights in the current round, it waits for the next round.
[0047] 7) LOD: leading one detect; a technique for detecting the first 1 at the most significant bit in data. In digital signal processing, communication systems, and hardware design, LOD is often used to identify the highest non-zero bit of data to determine the range of the data or perform other related operations.
[0048] FIG1 shows a schematic diagram of the structure of a processor, as shown in FIG1 :
[0049] Processor 100 consists of a controller 101, an arithmetic unit 102, and a memory 103. Controller 101 includes an instruction fetch unit 101a, which is responsible for fetching instructions from memory via bus 110. Arithmetic unit 102 is used to perform operations, such as addition and subtraction, on the instructions fetched from memory by controller 101. Memory 103 is used to store data used in instructions. Instruction fetch unit 101a includes an instruction counter 101a1 and an instruction register 101a2. Instruction counter 101a1 is a register used to store the address of the currently executing instruction. When a computer executes a program, it executes instructions one by one. The function of instruction counter 101a1 is to record the address of the instruction to be executed. Each time an instruction is executed, the address in instruction counter 101a1 is updated to point to the next instruction to be executed, thereby achieving sequential execution of the program's instructions. The instruction register 101a2 is a special register in the computer, which is used to store the instruction currently being executed. The main function of the instruction fetch unit 101a is to read instructions from the memory and pass them to the instruction decoder for parsing and execution. When the computer runs a program, the instruction fetch unit 101a reads instructions from the memory through the bus 110, and then stores the instructions in the instruction register 101a2. The instruction register 101a2 then passes the instruction opcode and operand to the instruction decoder, and the instruction decoder performs the corresponding operation according to the opcode.
[0050] Multi-level cache technology is widely used in modern processors. Direct access to memory is slow, typically taking up to several hundred clock cycles. However, processors can perform computations very quickly. To keep pace with this speed, a multi-level cache is placed between the processor and memory. Cache closer to the processor is faster, but has a smaller capacity. Due to limited cache capacity, the processor may miss required data during cache access. In such cases, the processor must search the lower-level cache until the missing data is found.
[0051] When a cache miss occurs, the processor needs to perform a series of related processes to adapt. Since the cache miss does not provide the data required by the memory access instruction, the memory access instruction must be reissued. If the data is found in the L2 cache after an L1 cache miss, the data can be returned quickly. However, if the data is missing from the L2 cache and subsequent levels of cache, the required data must be searched in memory, which is very time-consuming. How the processor handles cache miss scenarios is a key technical point in processor design.
[0052] Please refer to Figure 2, which shows a schematic diagram of an existing instruction processing component workflow. As shown in Figure 2, the existing solution, as shown in the figure above, is that a memory access instruction is emitted from the issue queue and processed by the four-stage main pipeline (from the first stage main pipeline to the fourth stage main pipeline). If no exceptions occur during the memory access process, the complete memory access data required by the memory access instruction can be obtained at the last stage and written directly. However, if an exception occurs, such as a cache miss detected in the third stage main pipeline, a miss replay is performed directly, placing the memory access instruction back into the issue queue for execution.
[0053] However, in Figure 2 above, if a cache miss is detected in the third stage of the main pipeline, the memory access instruction will be reissued directly, which will cause the memory access instruction to be issued multiple times. Since pipeline resources are relatively important, multiple reissues of a memory access instruction will cause the pipeline to process the memory access instruction multiple times, reducing the effective utilization rate of the pipeline.
[0054] If a miss occurs in the lower-level cache, the number of times the memory access instruction needs to be reissued increases greatly. During the execution of the memory access instruction, the data cache and data tag memory need to be read. The power consumption of reading the memory is very high, which greatly wastes power.
[0055] Please refer to Figure 3, which shows a schematic diagram of the instruction processing component workflow provided by an exemplary embodiment of the present application. The instruction processing component includes: a main pipeline circuit 302, a data buffer 303, and a data load queue circuit 304. The data buffer 303 can be a first-level cache in the processor or a cache at another level, which is not limited in this application.
[0056] In Figure 3, one / multiple output ports of the instruction emission circuit 301 are connected to one / multiple input ports of the mainstream pipeline circuit 302, and the memory access instruction is transmitted from the above-mentioned output port of the instruction emission circuit 301 to the above-mentioned input port of the mainstream pipeline circuit 302, so that the memory access instruction can be input into the mainstream pipeline circuit 302.
[0057] The main pipeline circuit 302 is electrically connected to the data buffer 303 , and the main pipeline circuit 302 is electrically connected to the data load queue circuit 304 .
[0058] One / multiple output ports of the mainstream pipeline circuit 302 and one / multiple input ports of the data buffer 303 are connected, and the access request is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the data buffer 303, so that the access request can be input into the data buffer 303.
[0059] One / multiple output ports of the data buffer 303 are connected to one / multiple input ports of the mainstream pipeline circuit 302, and the query result is transmitted from the above-mentioned output port of the data buffer 303 to the above-mentioned input port of the mainstream pipeline circuit 302, so that the query result is input into the mainstream pipeline circuit 302.
[0060] One / more output ports of the mainstream pipeline circuit 302 are connected to one / more input ports of the data loading queue circuit 304, and the data address corresponding to the memory access instruction is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the data loading queue circuit 304, so that the data address corresponding to the memory access instruction is input into the data loading queue circuit 304.
[0061] The main pipeline circuit 302 is used to receive a memory access instruction and send an access request to the data buffer 303 based on the memory access instruction. The access request is used to request to query the data buffer 303 for data corresponding to the memory access instruction.
[0062] Among them, the above-mentioned memory access instructions can also be called memory access instructions; in the process of processing memory access instructions, the mainstream pipeline circuit 302 needs to obtain the data corresponding to the memory access instructions and write them out; in this process, the mainstream pipeline circuit 302 can give priority to sending an access request to the first-level cache (that is, the above-mentioned data cache) to query the data corresponding to the memory access instruction from the data cache 303.
[0063] The main pipeline circuit 302 is further configured to write the data address corresponding to the memory access instruction into the data loading queue circuit 304 when the query of the data corresponding to the memory access instruction from the data buffer 303 fails.
[0064] The above-mentioned failure to query the data corresponding to the memory access instruction from the data buffer 303 may refer to the absence of all or part of the data corresponding to the memory access instruction in the data buffer 303; that is, when the data buffer 303 does not contain all the data corresponding to the memory access instruction, it is considered that the query of the data corresponding to the memory access instruction from the data buffer 303 has failed. When the data buffer 303 only contains a part of the total data corresponding to the memory access instruction, it is also considered that the query of the data corresponding to the memory access instruction from the data buffer 303 has failed. In an embodiment of the present application, when the query of the data corresponding to the memory access instruction from the data buffer 303 fails, the mainstream pipeline circuit 302 does not immediately trigger the retransmission of the memory access instruction, but instead writes the data address corresponding to the memory access instruction into the data load queue circuit 304, and the data load queue circuit 304 continues to obtain the corresponding data. In some embodiments, the above-mentioned memory access instruction can be used to obtain data corresponding to one or more data addresses. In this case, the mainstream pipeline circuit 302 can write the one or more data addresses into the data load queue circuit 304.
[0065] The data loading queue circuit 304 is used to receive backfill data, query corresponding data from the backfill data based on the written data address, and output the queried data; the backfill data is the data filled into the data buffer 303.
[0066] Among them, the above-mentioned backfill data can refer to data that is not queried from the data cache, but is subsequently queried in the lower-level cache or memory of the data cache 303, and is written to the data cache 303; for example, at a certain moment 1, the mainstream pipeline circuit 302 queries the data cache 303 for data corresponding to a data address A. The data cache 303 does not store the data corresponding to the data address A at moment 1. Later, at moment 2, the data corresponding to the data address A is queried from the lower-level cache or memory and the query is successful. At this time, the data corresponding to the data address A can be written to the data cache 303. The data corresponding to the data address A written to the data cache 303 here can be called backfill data; optionally, when writing the data corresponding to the data address A to the data cache 303, the data block where the data corresponding to the data address A is located can be written to the data cache 303. At this time, the data block where the data corresponding to the data address A is located can be called the above-mentioned backfill data. In an embodiment of the present application, in addition to being written to the data cache 303, the backfill data will also be written to the data loading queue circuit 304. For example, the line connected from the lower-level cache or memory to the data buffer can also be extended to the data loading queue circuit 304 , so that when backfill data is written into the data buffer 303 , the backfill data will also be sent to the data loading queue circuit 304 .
[0067] To sum up, the scheme shown in the embodiment of the present application is to set a data loading queue circuit 304 in the processor. When the mainstream pipeline circuit 302 fails to query the data corresponding to the memory access instruction from the data cache 303, the data address corresponding to the memory access instruction is written into the data loading queue circuit 304. The data loading queue circuit 304 receives the backfill data, queries the corresponding data from the backfill data based on the written data address, and outputs the queried data. That is to say, in the above scheme, when the mainstream pipeline circuit 302 fails to query from the data cache 303 during the process of processing the memory access instruction, the memory access instruction is not directly reissued. Instead, the data loading queue circuit 304 queries the corresponding data from the backfill data of the data cache 303 and outputs it. This can effectively reduce the number of reissues of the memory access instruction, improve the utilization rate of the mainstream pipeline, and thereby improve the efficiency of the processor in processing the memory access instruction.
[0068] Based on the instruction processing component shown in Figure 3, please refer to Figure 4, which shows a schematic diagram of the structure of the instruction processing component provided by an exemplary embodiment of the present application. The data loading queue circuit 304 may include: at least two data loading queue sub-circuits 3041, and an output sub-circuit 305; at least two data loading queue sub-circuits 3041 are electrically connected to the mainstream pipeline circuit 302 respectively; at least two data loading queue sub-circuits 3041 are electrically connected to the output sub-circuit 305 respectively.
[0069] One or more output ports of the main pipeline circuit 302 are respectively connected to one or more input ports of the at least two data loading queue sub-circuits 3041. The multiple data addresses are transmitted from the output ports of the main pipeline 302 to the input ports of the at least two data loading queue sub-circuits 3041, so that the multiple data addresses are split into two parts and respectively input into the at least two data loading queue sub-circuits 3041.
[0070] One / more output ports of each of the at least two data loading sub-circuits 3041 are connected to one / more input ports of the output sub-circuit 305, and the data queried by each of the at least two data loading queue sub-circuits 3041 are transmitted from the above-mentioned output ports of each of the at least two data loading queue sub-circuits 3041 to the above-mentioned input ports of the output sub-circuit, so that the queried data is input into the output sub-circuit 305.
[0071] The main pipeline circuit 302 is used to disperse and write multiple data addresses into at least two data loading queue sub-circuits 3041.
[0072] Exemplarily, the mainstream pipeline circuit 302 splits the data address into at least two data loading queue sub-circuits 3041 according to the logic of parallel writing through a preset algorithm or rule, ensuring uniform address distribution and improving the parallelism of data query by the data loading queue circuit 304.
[0073] At least two data loading queue sub-circuits 3041 are configured to concurrently execute the steps of receiving backfill data and searching for corresponding data from the backfill data based on the written data address.
[0074] The output sub-circuit 305 is used to obtain and output the data queried by the at least two data loading queue sub-circuits 3041 .
[0075] In an embodiment of the present application, multiple data addresses can be dispersed into at least two data loading queue sub-circuits 3041, so that the step of querying data from the backfill data is executed in parallel by at least two data loading queue sub-circuits 3041, thereby improving the concurrency of data query and improving the efficiency of obtaining data corresponding to memory access instructions that failed to query from the backfill data.
[0076] Please refer to Figure 5, which shows a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application. The data loading queue sub-circuit 3041 in Figure 4 includes: an address entry queue 3041a, an address comparison device 3041b, a data extraction device 3041c and an output selection circuit 3041d.
[0077] One / multiple output ports of the address entry queue 3041a are connected to one / multiple input ports of the address comparison device 3041b, and the data address in the address entry is passed from the above-mentioned output port of the address entry queue 3041a to the above-mentioned input port of the address comparison device 3041b, so that the data address in the address entry is input into the address comparison device 3041b.
[0078] One or more output ports of the address entry queue 3041a are connected to one or more input ports of the output selection circuit 3041d. The backfill field is transferred from the output port of the address entry queue 3041a to the input port of the output selection circuit 3041d, so that the backfill field is input into the output selection circuit 3041d.
[0079] One / more output ports of the address comparison device 3041b are connected to one / more input ports of the data extraction device 3041c.
[0080] One / multiple input ports of the address comparison device 3041b are connected to one / multiple output ports of the mainstream pipeline circuit 302, and the data address of the backfill data is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the address comparison device 3041b, so that the data address of the backfill data is input into the address comparison device 3041b.
[0081] The address entry queue 3041a includes multiple address entries, each of which is used to store a written data address.
[0082] The above address entry contains information such as number, data address and backfill field.
[0083] The data extraction device 3041c includes a data entry queue 3041c1, a data selector 3041c2, and a data outputter 3041c3. The data entry queue 3041c1 includes a plurality of data entries.
[0084] One / multiple input ports of the data entry queue 3041c1 and one / multiple output ports of the data selector 3041c2 are connected, and the data corresponding to the data address in the address entry is passed from the above-mentioned output port of the data selector 3041c2 to the above-mentioned input port of the data entry queue 3041c1, so that the above-mentioned data is input into the data entry queue 3041c1.
[0085] One / multiple output ports of the data entry queue 3041c1 are connected to one / multiple input ports of the data outputter 3041c3, and the data entries in the data entry queue 3041c1 are transferred from the above-mentioned output ports of the data entry queue 3041c1 to the above-mentioned input ports of the data outputter 3041c3, so that the above-mentioned data entries are input into the data outputter 3041c3.
[0086] One / multiple input ports of the data selector 3041c2 are connected to one / multiple output ports of the mainstream pipeline circuit 302, and the backfill data is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the data selector 3041c2, so that the backfill data is input into the data selector 3041c2.
[0087] One / multiple input ports of the data outputter 3041c3 are connected to one / multiple outputs of the output selection circuit 3041d, and the number of the address entry is passed from the above-mentioned output port of the output selection circuit 3041d to the above-mentioned input port of the data outputter 3041c3, so that the number of the address entry is input into the data outputter 3041c3.
[0088] The address comparison device 3041b is used to compare the data address in the address entry with the data address of the backfill data, and set the backfill field in the address entry to a specified value when the data address in the address entry matches the data address of the backfill data.
[0089] The above-mentioned address matching means that when the value of the data address in the address entry is compared with the value of the data address of the backfill data, the values of the two addresses are the same.
[0090] The specified value is a specific value or status that is set for the backfill field when the address is matched. The specified value can be defined by the developer during development and design. It can be a binary number, a status flag, or other content that needs to be set when matching. For example, the specified value can be 1 or 0.
[0091] The data selector 3041c2 is used to extract data corresponding to the data address in the address entry from the backfill data when the data address in the address entry matches the data address of the backfill data, and store the extracted data in the data bits of the data entry corresponding to the address entry.
[0092] The output selection circuit 3041d is used to select the number of an address entry whose backfill field is a specified value, and pass the number of the address entry to the data output component.
[0093] The above numbers are used to uniquely identify or recognize different address entries. Each address entry in the processor may be pre-set with a fixed number.
[0094] Alternatively, the numbers of the address entries may be generated by a processor. For example, in some embodiments, a self-increasing number may be used as the number, and each time a new address entry is added, the number automatically increases; for example, 1, 2, 3, etc. In some embodiments, a hash function may be used to generate the number, and the hash function may convert the data into a hash value of a fixed length, and use these hash values as the number; typically, the hash value is represented by a fixed-length string or number. In some embodiments, the number may also be generated based on business rules, which may include specific prefixes, dates, locations, and other information. For example, 20231127 represents November 27, 2023.
[0095] In an embodiment of the present application, the output selection circuit 3041d can search each address entry in turn according to the specified value of the backfill field, find the address entry whose backfill field is the specified value, then query the number of the address entry, and pass the number to the data output component.
[0096] The data outputter 3041c3 is used to transmit the data stored in the data entry corresponding to the address entry number in the data entry queue 3041c1 to the output sub-circuit 305 according to the address entry number transmitted by the output selection circuit 3041d.
[0097] In the solution shown in the embodiment of the present application, an address comparison device is set to compare the data address of the backfill data with the data address of the data that failed to be queried from the data cache 303, thereby triggering the data selector 3041c2 to accurately extract the data that failed to be queried before from the backfill data, and output the queried data through the output selection circuit 3041d and the data outputter 3041c3, thereby realizing a circuit structure for data extraction based on address comparison, ensuring that the data corresponding to the data address of the memory access instruction that failed to be queried before can be obtained from the backfill data through the hardware circuit, thereby ensuring the execution efficiency of data extraction.
[0098] Please refer to Figure 6, which shows a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application. Based on the fact that each data entry in the data entry queue 3041c1 in Figure 5 also includes a valid signal bit 3041c1a, the data extraction device 3041c also includes an initialization circuit 3041c4.
[0099] One / multiple input ports of the valid signal bit 3041c1a are connected to one / multiple output ports of the initialization circuit 3041c4, and the valid signal is transmitted from the above-mentioned output port of the initialization circuit 3041c4 to the above-mentioned input port of the valid signal bit 3041c1a, so that the valid signal is input into the valid signal bit 3041c1a.
[0100] One / multiple input ports of the initialization circuit 3041c4 are connected to one / multiple output ports of the mainstream pipeline circuit 302, and the forward data is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the initialization circuit 3041c4, so that the forward data is input into the initialization circuit 3041c4.
[0101] The mainstream pipeline circuit 302 is also used to write the forward data into the initialization circuit 3041c4; the forward data is data composed of the results of the query from the data buffer; the forward data includes the query result data and the valid signal, and the valid signal is used to indicate whether the query result data is valid.
[0102] Optionally, the valid signal bit may be 1 or 0, where 1 indicates that the query result data is valid and 0 indicates that the query result data is invalid; or, 0 indicates that the query result data is valid and 1 indicates that the query result data is invalid.
[0103] Initialization circuit 3041c4 is used to write the query result data into the data entry queue 3041c1, the data bit of the data entry corresponding to the address entry where the data address of the query result data is located, and write the valid signal into the valid signal bit 3041c1a of the data entry corresponding to the address entry where the data address of the query result data is located.
[0104] The initialization circuit 3041c4 obtains data from the main pipeline circuit 302 and writes a valid signal into the valid signal bit 3041c1a of the data entry corresponding to the address entry where the data address of the query result data is located.
[0105] Exemplarily, when the data address in the address entry matches the data address of the backfilled data, it indicates that the data address in the address entry has been backfilled, and the initialization circuit 3041c4 writes 0 to the valid signal bit of the data entry corresponding to the address entry, and the valid signal indicates invalid; when the data address in the address entry does not match the data address of the backfilled data, it indicates that the data address in the address entry has not been backfilled, and the initialization circuit 3041c4 writes 1 to the valid signal bit of the data entry corresponding to the address entry, and the valid signal indicates valid.
[0106] The data selector 3041c2 is used to extract the data corresponding to the data address in the address entry from the backfill data of the input data extraction device when the data address in the address entry matches the data address of the backfill data and the valid signal indication in the valid signal bit 3041c1a of the data entry corresponding to the address entry is invalid, and store the extracted data in the data bit of the data entry corresponding to the address entry.
[0107] In the scheme shown in the embodiment of the present application, the forward data is also written into the data entry queue 3041c1. When obtaining the data corresponding to the data address of the memory access instruction that failed to be queried previously from the backfill data, there is no need to repeatedly query from the backfill data for some of the data that has been queried, thereby further improving the efficiency of obtaining the data corresponding to the data address of the memory access instruction that failed to be queried previously from the backfill data.
[0108] Please refer to Figure 7, which shows a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application. Based on the data loading queue sub-circuit 3041 in Figure 6, which includes M data extraction devices 3041c, the data loading queue sub-circuit 3041 also includes: a forward splitting unit 3041e, and a backfill splitting unit 3041f.
[0109] One / multiple input ports of the forward splitting unit 3041e are connected to one / multiple output ports of the mainstream pipeline circuit 302, and the forward data is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the forward splitting unit 3041e, so that the forward data is input into the forward splitting unit 3041e.
[0110] One / multiple output ports of the forward splitting unit 3041e are connected to one / multiple input ports of the initialization circuit 3041c4, and the split forward data is transmitted from the above-mentioned output port of the forward splitting unit 3041e to the above-mentioned input port of the initialization circuit 3041c4, so that the split forward data is input into the initialization circuit 3041c4.
[0111] One / multiple input ports of the backfill splitting unit 3041f are connected to one / multiple output ports of the mainstream pipeline circuit 302, and the backfill data is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the backfill splitting unit 3041f, so that the backfill data is input into the backfill splitting unit 3041f.
[0112] One / multiple output ports of the backfill splitting unit 3041f are connected to one / multiple input ports of the data selector 3041c2, and the split backfill data is transmitted from the above-mentioned output port of the backfill splitting unit 3041f to the above-mentioned input port of the data selector 3041c2, so that the split backfill data is input into the data selector 3041c2.
[0113] The forward splitting unit 3041e is used to disperse and write the forward data of each M data address into the initialization circuit 3041c4 in the M data extraction devices.
[0114] For example, the data addresses are A1, A2, A3, A4, A5, A6, A7, A8, A9, and M is 3. The forward splitting unit 3041e writes the data addresses into the initialization circuits in the three data extraction devices in a dispersed manner. The resulting groups may be group 1: A1, A2, A3, group 2: A4, A5, A6, and group 3: A7, A8, A9.
[0115] The backfill splitting unit 3041f is used to disperse and write the backfill data of every M data addresses into the M data extraction devices 3041c.
[0116] For example, the data addresses are B1, B2, B3, B4, B5, B6, B7, B8, and B9, and M is 3. The backfill split unit 3041f writes the data addresses into three data extraction devices in a dispersed manner. The resulting groups may be group 1: B1, B2, and B3; group 2: B4, B5, and B6; and group 3: B7, B8, and B9.
[0117] In the scheme shown in the embodiment of the present application, M data extraction devices 3041c can be set in a data loading queue sub-circuit 3041, and the forward data and backfill data input into the data loading queue sub-circuit 3041 will be dispersed to the M data extraction devices 3041c, so that the M data extraction devices 3041c can execute the steps of querying data from the backfill data in parallel, thereby improving the concurrency of data query and improving the efficiency of obtaining data corresponding to the memory access instruction that failed the query from the backfill data.
[0118] Please refer to Figure 8, which shows a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application. Based on Figure 7, the data loading queue sub-circuit 3041 also includes a first beat register 3041g, and the data extraction device 3041c also includes a second beat register 3041c5.
[0119] One / multiple input ports of the first beat register 3041g are connected to one / multiple output ports of the mainstream pipeline circuit 302, and the data address corresponding to the memory access instruction is transmitted from the above-mentioned output port of the mainstream pipeline circuit 302 to the above-mentioned input port of the first beat register 3041g, so that the data address corresponding to the memory access instruction is input into the first beat register 3041g.
[0120] One / more output ports of the first beat register 3041g are connected to one / more input ports of the address comparison device 3041b, and the data address corresponding to the memory access instruction after the beat processing is transferred from the above-mentioned output port of the first beat register 3041g to the above-mentioned input port of the address comparison device 3041b, so that the data address corresponding to the memory access instruction after the beat processing is input into the address comparison device 3041b.
[0121] One / multiple input ports of the second beat register 3041c5 are connected to one / multiple output ports of the backfill splitting unit 3041f, and the split backfill data is transferred from the above-mentioned output port of the backfill splitting unit 3041f to the above-mentioned input port of the second beat register 3041c5, so that the split backfill data is input into the second beat register 3041c5.
[0122] One / multiple output ports of the second beat register 3041c5 are connected to one / multiple output ports of the data selector 3041c2, and the backfill data after the beat processing is transmitted from the above-mentioned output port of the second beat register 3041c5 to the above-mentioned output port of the data selector 3041c2, so that the backfill data after the beat processing is input into the data selector 3041c2.
[0123] The main pipeline circuit 302 is used to write the data address corresponding to the memory access instruction into the first beat register 3041g for beat processing.
[0124] The backfill splitting unit is used to disperse and write M consecutive bytes of backfill data into the second beat register 3041c5 in M data extraction devices for beat processing.
[0125] In the solution shown in the embodiment of the present application, the address and backfill data of the backfill data can be delayed at the entrance of the data extraction device 3041c to avoid the problem of too long processing cycle caused by too long processing flow in a single processing process, thereby avoiding affecting the main frequency design of the processor.
[0126] Please refer to Figure 9, which shows a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application. The output sub-circuit 305 in Figure 4 is used to obtain and output the data queried by at least two data loading queue sub-circuits 3041 in a polling manner.
[0127] The input port of the output sub-circuit 305 is connected to the output port of the data loading queue circuit 304 respectively; the data queried by the at least two data loading queue sub-circuits 3041 are respectively transmitted from the above-mentioned output ports of the data loading queue circuit 304 to the above-mentioned input ports of the output sub-circuit 305, so that the above-mentioned queried data is input into the output sub-circuit 305.
[0128] The output subcircuit 305 is configured to obtain the data queried by the at least two data loading queue subcircuits 3041 in a round-robin manner, shift the data queried by the at least two data loading queue subcircuits 3041, and then output the data;
[0129] The displacement amount of the data shift is the lower three bits of the address of the data queried by the at least two data loading queue sub-circuits 3041 .
[0130] In some embodiments, the output sub-circuit 305 may include a polling arbitration sub-circuit 3051 and a polling arbitration sub-circuit 3052 .
[0131] The input port of the polling arbitration sub-circuit 3051 is connected to the output port of the data loading queue sub-circuit 3041 respectively; the data queried by the at least two data loading queue sub-circuits 3041 are transmitted from the output ports of the data loading queue sub-circuit 3041 to the input port of the polling arbitration sub-circuit 3051 respectively, so that the queried data is input into the polling arbitration sub-circuit 3051;
[0132] The input port of the polling arbitration sub-circuit 3052 and the output port of the data loading queue sub-circuit 3041 are respectively connected; the data queried by at least two data loading queue sub-circuits 3041 are respectively transmitted from the above-mentioned output ports of the data loading queue sub-circuit 3041 to the above-mentioned input port of the polling arbitration sub-circuit 3052, so that the above-mentioned queried data is input into the polling arbitration sub-circuit 3052.
[0133] The polling arbitration sub-circuit 3051 is configured to obtain the address data outputted by the at least two data loading queue sub-circuits 3041 in a polling manner, and obtain one of the address data outputted by the at least two data loading queue sub-circuits 3041;
[0134] The polling arbitration sub-circuit 3052 is configured to obtain the address data outputted by the at least two data loading queue sub-circuits 3041 in a polling manner, and obtain one of the address data outputted by the at least two data loading queue sub-circuits 3041 .
[0135] By polling, the backfill data received and the written data addresses in at least two data loading queue subcircuits can be processed simultaneously, improving processing efficiency in a near-parallel manner. By shifting data, valid data can be concentrated in the low bits, improving data processing efficiency.
[0136] Please refer to FIG. 10 , which shows a schematic diagram of the operation of the instruction processing component provided by an exemplary embodiment of the present application. The data buffer 303 includes a missing status processing register 3031 .
[0137] Among them, one / multiple input ports of the missing status processing register 3031 are connected to one / multiple output ports of the mainstream pipeline circuit 302. When there is no data corresponding to the access request in the data buffer 303 and the missing status processing register 3031 is not full, the access request is passed from the above-mentioned output port of the mainstream pipeline circuit 302 to the missing status processing register 3031, so that the access request is input into the missing status processing register 3031.
[0138] The data buffer 303 is used to buffer the access request into an entry in the missing status processing register 3031 when the data corresponding to the access request does not exist in the data buffer 303 and the missing status processing register 3031 is not full.
[0139] The above entry may include information such as the address of the missing data and the access type (read or write).
[0140] The data buffer 303 is further used to send the access request in the missing status processing register 3031 to the lower-level cache or memory to query the data corresponding to the access request;
[0141] The access request may include a read or write request, a request address, and the like.
[0142] The data buffer 303 is further configured to remove the access request from the missing status processing register when receiving data corresponding to the access request returned by the lower-level cache or the memory.
[0143] In an embodiment of the present application, a missing status processing register 3031 is also set in the data cache 303. For a certain access request, when the access request fails to query in the data cache 303, there is no need to resend the corresponding memory access instruction, but it is temporarily stored in the missing status processing register 3031. Subsequently, the data cache 303 sends the access request to the lower-level cache or memory for data query, thereby reducing the number of retransmissions of the memory access instruction and improving the processing efficiency of the mainstream pipeline.
[0144] The data buffer 303 is used to discard the access request when the data corresponding to the access request does not exist in the data buffer 303 and the access request already exists in the missing status processing register 3031.
[0145] In an embodiment of the present application, for access requests that appear repeatedly and fail to be queried in the data buffer 303, it is not necessary to temporarily store them in the missing status processing register 3031 each time. Instead, when there is an access request in the missing status processing register 3031 and a new identical access request arrives, the two access requests are merged (one of them is discarded), thereby reducing the repeated processing of the access requests and only improving the efficiency of data requests.
[0146] The main pipeline circuit 302 is further configured to trigger retransmission of the memory access instruction when the data corresponding to the access request does not exist in the data buffer 303 and the missing status processing register 3031 is full.
[0147] In an embodiment of the present application, for a certain access request, when the access request fails to query in the data cache 303, and the missing status processing register 3031 is full at this time, the retransmission of the memory access instruction can be triggered, thereby avoiding the situation where the memory access instruction cannot be correctly executed subsequently due to the overflow of the missing status processing register 3031.
[0148] Based on the scheme shown in any one of Figures 3 to 10 above, the present application can provide a load queue design that supports parallel backfill data detection. When a cache miss occurs, but the corresponding miss status processing register is not full, the current request can enter the miss status processing register, and the current memory access instruction does not need to be reissued. The backfill data of the L1 cache is subsequently continuously detected, and parallel detection and capture are achieved by dividing the bank and the high and low queues, which greatly improves the performance of the processor's memory access instructions.
[0149] Based on any of the solutions shown in FIG3 to FIG10 , the technical solution of this embodiment is described in detail as follows:
[0150] 1. Product side
[0151] In the processor field, multi-level cache technology is widely used. The load queue design proposed in this embodiment that supports parallel backfill data detection improves the effective utilization rate of the pipeline and reduces chip power consumption, effectively improving the competitiveness of the product.
[0152] 2. Technical side
[0153] Please refer to Figure 11, which shows a schematic diagram of the pipeline structure involved in an exemplary embodiment of the present application. As shown in Figure 11, in the pipeline stage s1 (main pipeline level 1), an access request (L1 request) is issued to the L1 dcache (L1 data cache). The MSHR in the L1 dcache is used to handle access requests with cache misses. If a cache miss occurs during the current request to access the L1, the request directly enters the MSHR. The L1 dcache selects an access request from the MSHR and transmits it to the lower-level L2 dcache (L2 request).
[0154] MSHR usually has multiple entries. When an access request for a memory access instruction that is cache-missed enters MSHR, it will occupy an entry. When the access request for the memory access instruction is sent to the lower-level cache and the return data is received, the corresponding entry will be released.
[0155] MSHR also has a function of merging access requests. If the address corresponding to the access request in the MSHR entry matches the address of the access request currently missing from the cache, the access request currently missing from the cache will no longer apply for entry for storage, but will share the returned data with the matched access request.
[0156] Thanks to the existence of the MSHR, even if a memory access instruction encounters a cache miss in the L1 dcache, the access request will be subsequently sent to the lower-level cache as long as it can enter the MSHR. Therefore, when a cache miss occurs and the MSHR is not full, the access request flows backward normally. In the s4 stage (the 4-stage main pipeline), it is determined whether the corresponding data is available from the dcache. If it is available, it is directly written out. Otherwise, it enters the LDQ and waits. At this time, the forward data passed from the store pipeline needs to be written to the LDQ, which is the forward_in stage in the figure.
[0157] Subsequently, the LDQ needs to continuously monitor the backfill data returned from the lower-level cache. This refill data is sent to the L1 dcache to backfill the missing cache line and, in parallel, to the LDQ. If the backfill address matches the data address of the access request in the LDQ, the backfill data is merged into the entry corresponding to the access request. Once the backfill data in the LDQ is ready, it is written out.
[0158] As shown in Figure 12, it shows a schematic diagram of the micro-architecture structure involved in an exemplary embodiment of the present application. As shown in Figure 12, LDQ is divided into two queues, namely L_LDQ (low LDQ) and H_LDQ (high LDQ). Assuming that the queue depth of LDQ is N, L_LDQ stores entries numbered 0 to N / 2-1, and H_LDQ stores entries numbered N / 2 to N-1. The backfill data will be copied twice and enter the high and low queues respectively. After a simple split in each queue, it will directly enter dff2 for beating. Since the backfill data bit width is relatively wide, up to 512 bits, this method of dividing the queues into high and low queues can effectively reduce the congestion of the connection, and the method of beating at the entrance of the high and low queues is also more friendly to the timing. Only L_LDQ is explained below, and H_LDQ is isomorphic to it.
[0159] The widest data type for memory access instructions is double word, or 64 bits. L_LDQ is divided into eight banks, each handling 1 byte (8 bits) of data. Refill data first enters the refill split unit for byte-by-byte processing.
[0160] As shown in Figure 13, it shows a schematic diagram of the refill split unit involved in an exemplary embodiment of the present application. As shown in Figure 13, the schematic diagram of splitting by byte is shown above. A small vertical bar in the above figure represents 1 byte. The input is 512 bits, that is, 64 bytes of data. The splitting method is to extract the bytes at the same position in every 8 bytes together. For example, the green square is the 0th byte in the 8-byte data. After extracting them together, they are sent to bank 0. This processing method processes data of different bytes in different banks. The processing between different banks is completely parallel and does not interfere with each other in wiring. It is highly layout and wiring friendly in chip physical implementation.
[0161] Forward_in is the data passed forward through the store pipeline. It consists of two parts: one is the forward data, which is 64 bits wide, and the other is the data valid signal, which is 8 bits. If valid[i] is high, it means that the i-th byte of the forward data is valid. Since the data passed forward from the store pipeline is updated, the forward data has the highest priority during the merging process with the dcache backfill data. That is, if the i-th byte of the forward data is valid, the backfill data cannot overwrite it. Forward_in needs to enter the forward split unit for splitting. The splitting logic is the same as the refill split unit. After splitting, the forward valid signal and data are written to the assigned entry respectively.
[0162] Refill_addr_in is sent to L_LDQ after being clocked in dff1. The backfill address is cmp(compare)ed with the addr(address) of each entry in L_LDQ. If the addresses match, it means that the backfill address matches the address of the current entry and needs to be captured. The capture process selects 8 bits from the 64-bit input data based on the address information. This selection function is implemented by mux. The capture enable signal enable = compare_sucess&&(!forward), which means that when the address matches and the current entry has no data forwarded, the backfill data can be captured.
[0163] When a successful backfill is detected, the corresponding entry's refilled field is set to 1. The LOD circuit then selects an entry with a refilled field of 1 for writing. The LOD outputs the number of the first refilled entry found, and uses this number to select the data field of the entry through mux1.
[0164] Finally, the high and low queues each select a refilled entry and perform round-robin arbitration through RR to obtain the data that needs to be written.
[0165] Figure 14 illustrates a schematic diagram of the right shifter operation in accordance with an exemplary embodiment of the present application. As shown in Figure 14 , the data output via round-robin arbitration also requires a data shift by the lower three bits of the address. Figure 14 illustrates two shifting scenarios for different data types. For a half word data type, the data at addresses 4 and 5 must be right-shifted by 4 bytes to concentrate the valid data in the lower bits. For a byte data type, the data at address 7 must be right-shifted by 7 bytes.
[0166] Please refer to Figure 15, which shows a flow chart of an instruction processing method involved in an exemplary embodiment of the present application. As shown in Figure 15, the method can be performed by an instruction processing component, which can be the instruction processing component described in any of Figures 3 to 10. As shown in Figure 15, the method can include the following steps:
[0167] Step 1501: Receive a memory access instruction through the main pipeline circuit, and send an access request to the data buffer based on the memory access instruction. The access request is used to request to query the data corresponding to the memory access instruction from the data buffer.
[0168] Step 1502: When the query of the data corresponding to the memory access instruction from the data buffer fails, the data address corresponding to the memory access instruction is written into the data loading queue circuit through the main pipeline circuit.
[0169] Step 1503: Receive backfill data through the data loading queue circuit, query corresponding data from the backfill data based on the data address written from the mainstream pipeline circuit, and output the queried data; the backfill data is the data filled into the data buffer.
[0170] In some embodiments, the data loading queue circuit includes: at least two data loading queue sub-circuits and an output sub-circuit; the at least two data loading queue sub-circuits are electrically connected to the main pipeline circuit respectively; at least two of the data loading queue sub-circuits are electrically connected to the output sub-circuit respectively;
[0171] Writing the data address corresponding to the memory access instruction into the data loading queue circuit includes:
[0172] Writing multiple data addresses into at least two of the data loading queue sub-circuits in a dispersed manner through the main pipeline circuit;
[0173] The receiving backfill data through the data loading queue circuit, querying corresponding data from the backfill data based on the written data address, and outputting the query data includes:
[0174] The steps of receiving backfill data and querying corresponding data from the backfill data based on the written data address are performed in parallel by at least two of the data loading queue sub-circuits;
[0175] The output sub-circuit obtains and outputs the data queried by each of at least two data loading queue sub-circuits.
[0176] In some embodiments, the data load queue subcircuit includes: an address entry queue, an address comparison device, a data extraction device, and an output selection circuit;
[0177] The address entry queue includes a plurality of address entries, each of which is used to store a written data address;
[0178] The data extraction device comprises a data entry queue, a data selector and a data outputter, wherein the data entry queue comprises a plurality of data entries;
[0179] The step of receiving backfill data in parallel by at least two of the data loading queue sub-circuits and querying corresponding data from the backfill data based on the written data address includes:
[0180] comparing the data address in the address entry with the data address of the backfill data by the address comparison device, and setting the backfill field in the address entry to a specified value when the data address in the address entry matches the data address of the backfill data;
[0181] extracting, by the data selector, data corresponding to the data address in the address entry from the backfill data when the data address in the address entry matches the data address of the backfill data, and storing the extracted data in a data bit of the data entry corresponding to the address entry;
[0182] selecting, by the output selection circuit, a number of an address entry whose backfill field is the specified value, and passing the number of the address entry to the data output component;
[0183] The data outputter transmits the data stored in the data entry corresponding to the address entry number in the data entry queue to the output sub-circuit according to the address entry number transmitted from the output selection circuit.
[0184] In some embodiments, each data entry in the data entry queue further comprises a valid signal bit; the data extraction device further comprises an initialization circuit;
[0185] The method further comprises:
[0186] Writing forward data into the initialization circuit through the main pipeline circuit; the forward data is data composed of the result of querying the data buffer; the forward data includes the query result data and a valid signal, the valid signal is used to indicate whether the query result data is valid;
[0187] Writing the query result data into the data entry queue, through the initialization circuit, a data bit of the data entry corresponding to the address entry where the data address of the query result data is located, and writing the valid signal into a valid signal bit of the data entry corresponding to the address entry where the data address of the query result data is located;
[0188] The step of extracting, by the data selector, data corresponding to the data address in the address entry from the backfill data when the data address in the address entry matches the data address of the backfill data, and storing the extracted data in a data bit of the data entry corresponding to the address entry, comprises:
[0189] Through the data selector, when the data address in the address entry matches the data address of the backfill data, and the valid signal indication in the valid signal bit of the data entry corresponding to the address entry is invalid, the data corresponding to the data address in the address entry is extracted from the backfill data input to the data extraction device, and the extracted data is stored in the data bit of the data entry corresponding to the address entry.
[0190] In some embodiments, the data loading queue sub-circuit includes M data extraction devices, and the data loading queue sub-circuit further includes: a forward split unit and a backfill split unit;
[0191] The method further comprises:
[0192] The forward data of each of the M data addresses is written dispersedly into the initialization circuits in the M data extraction devices through the forward splitting unit;
[0193] The backfill data of every M data addresses are written into the M data extraction devices in a dispersed manner through the backfill splitting unit.
[0194] In some embodiments, the data loading queue subcircuit further includes a first beat register, and the data extraction device further includes a second beat register;
[0195] The step of distributing and writing a plurality of data addresses into at least two of the data loading queue sub-circuits through the main pipeline circuit includes:
[0196] Writing the data address corresponding to the memory access instruction into the first beat register for beat processing through the main pipeline circuit;
[0197] The method further comprises:
[0198] Through the backfill splitting unit, the continuous M bytes of backfill data are scattered and written into the second beat registers in the M data extraction devices for beat processing.
[0199] In some embodiments, obtaining and outputting the data queried by each of the at least two data loading queue subcircuits through the output subcircuit includes:
[0200] The output sub-circuit obtains and outputs the data queried by at least two data loading queue sub-circuits in a polling manner.
[0201] In some embodiments, the step of obtaining and outputting the data queried by the at least two data loading queue subcircuits respectively in a polling manner through the output subcircuit includes:
[0202] Obtaining, by the output subcircuit, data queried by each of the at least two data loading queue subcircuits in a polling manner, and performing data shifting on the data queried by each of the at least two data loading queue subcircuits before outputting the data;
[0203] The displacement of the data shift is the lower three bits of the address of the data queried by each of the at least two data loading queue sub-circuits.
[0204] In some embodiments, the data buffer includes a miss status processing register; the method further includes:
[0205] caching the access request into an entry in the missing status processing register by the data buffer when the data corresponding to the access request does not exist in the data buffer and the missing status processing register is not full;
[0206] Sending the access request in the missing status processing register to a lower-level cache or memory through the data buffer to query data corresponding to the access request;
[0207] When the data corresponding to the access request is received from the lower-level cache or the memory through the data buffer, the access request is removed from the missing status processing register.
[0208] In some embodiments, the method further comprises:
[0209] By using the data buffer, when the data corresponding to the access request does not exist in the data buffer and the access request already exists in the missing status processing register, the access request is discarded.
[0210] In some embodiments, the method further comprises:
[0211] Through the mainstream pipeline circuit, when the data corresponding to the access request does not exist in the data buffer and the missing status processing register is full, the retransmission of the memory access instruction is triggered.
[0212] The solutions shown in the above embodiments of the present application can be applied to a processor. Specifically, the present application further provides a processor comprising at least one instruction processing component as shown in any one of Figures 3 to 10 above.
[0213] In a possible implementation, the present application further provides a computer device, which includes at least one processor, and the processor includes at least one instruction processing component as shown in any one of Figures 3 to 10 above.
[0214] In addition, an embodiment of the present application further provides a storage medium, which is used to store a computer program, and the computer program is used to execute the method provided by the above embodiment.
[0215] An embodiment of the present application further provides a computer program product including a computer program, which, when executed on a computer, enables the computer to execute the method provided in the above embodiment.
[0216] Those skilled in the art will understand that all or part of the steps of implementing the above embodiments may be accomplished by hardware, or by programs instructing related hardware to accomplish the steps. The programs may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0217] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An instruction processing component, the instruction processing component comprising: Main pipeline circuit, data buffer, and data loading queue circuit; The main pipeline circuit is electrically connected to the data buffer, and the main pipeline circuit is electrically connected to the data loading queue circuit; The main pipeline circuit is configured to receive a memory access instruction, and send an access request to the data buffer based on the memory access instruction, where the access request is used to request to query data corresponding to the memory access instruction from the data buffer; The main pipeline circuit is further configured to write the data address corresponding to the memory access instruction into the data loading queue circuit when the query for the data corresponding to the memory access instruction from the data buffer fails; The data loading queue circuit is configured to receive backfill data, query corresponding data from the backfill data based on the data address written from the main pipeline circuit, and output the queried data; the backfill data is data filled into the data buffer.
2. The instruction processing component according to claim 1, wherein the data loading queue circuit comprises: At least two data loading queue sub - circuits, and an output sub - circuit; The at least two data loading queue sub - circuits are respectively electrically connected to the main pipeline circuit; At least two of the data loading queue sub - circuits are respectively electrically connected to the output sub - circuit; The main pipeline circuit is configured to disperse and write multiple data addresses into at least two of the data loading queue sub - circuits; At least two of the data loading queue sub - circuits are configured to receive backfill data in parallel, and query corresponding data from the backfill data based on the written data addresses; The output sub - circuit is configured to obtain the data queried by each of the at least two data loading queue sub - circuits and output it.
3. The instruction processing component according to claim 2, wherein the data loading queue sub-circuit comprises: Address entry queue, address comparison device, data extraction device, and output selection circuit; The address entry queue contains multiple address entries, and each address entry is used to store a written data address; The data extraction device contains a data entry queue, a data selector, and a data outputter, and the data entry queue contains multiple data entries; The address comparison device is configured to compare the data address in the address entry with the data address of the backfill data, and set the backfill field in the address entry to a specified value when the data address in the address entry matches the data address of the backfill data; The data selector is configured to extract the data corresponding to the data address in the address entry from the backfill data when the data address in the address entry matches the data address of the backfill data, and store the extracted data into the data bit of the data entry corresponding to the address entry; The output selection circuit is configured to select the number of an address entry whose backfill field is the specified value, and transfer the number of the address entry to the data output component; The data outputter is configured to transfer the data stored in the data entry corresponding to the number of the address entry in the data entry queue to the output sub - circuit according to the number of the address entry transferred by the output selection circuit.
4. The instruction processing component according to claim 3, wherein each data entry in the data entry queue further includes a valid signal bit; and the data extraction device further includes an initialization circuit; The main pipeline circuit is further configured to write the forwarded data into the initialization circuit; the forwarded data is composed of the result queried from the data buffer; the forwarded data includes the query result data and a valid signal, and the valid signal is used to indicate whether the query result data is valid; The initialization circuit is configured to write the query result data into the data bit of the data entry corresponding to the address entry where the data address of the query result data is located in the data entry queue, and write the valid signal into the valid signal bit of the data entry corresponding to the address entry where the data address of the query result data is located; The data selector is configured to, when the data address in the address entry matches the data address of the backfill data, and the valid signal in the valid signal bit of the data entry corresponding to the address entry indicates invalid, extract the data corresponding to the data address in the address entry from the backfill data input to the data extraction device, and store the extracted data into the data bit of the data entry corresponding to the address entry.
5. The instruction processing component according to claim 4, wherein the data loading queue sub-circuit includes M of the data extraction devices, and the data loading queue sub-circuit further includes: A forwarding splitting unit and a backfill splitting unit; The forwarding splitting unit is configured to disperse and write the forwarded data of every M data addresses into the initialization circuits in M data extraction devices; The backfill splitting unit is configured to disperse and write the backfill data of every M data addresses into M data extraction devices.
6. The instruction processing component according to claim 5, wherein the data loading queue sub-circuit further includes a first pipeline register, and the data extraction device further includes a second pipeline register; The main pipeline circuit is configured to write the data address corresponding to the memory access instruction into the first pipeline register for pipelining processing; The backfill splitting unit is configured to disperse and write the backfill data of consecutive M bytes into the second pipeline registers in M data extraction devices for pipelining processing.
7. The instruction processing component according to claim 2, The output sub-circuit is configured to obtain and output the data queried by at least two data loading queue sub-circuits in a polling manner.
8. The instruction processing component according to claim 7, The output sub-circuit is configured to obtain the data queried by at least two data loading queue sub-circuits in a polling manner, perform data shifting on the data queried by at least two data loading queue sub-circuits respectively, and then output the data; Among them, The displacement amount of the data shifting is the lower three bits of the addresses of the data queried by at least two data loading queue sub-circuits respectively.
9. The instruction processing component according to any one of claims 1 to 8, wherein the data buffer includes a miss status processing register; The data buffer is configured to, when the data corresponding to the access request does not exist in the data buffer, and the When the missing state processing register is not full, cache the access request into one of the entries in the missing state processing register; The data buffer is further configured to send the access request in the missing state processing register to a lower-level cache or memory to query the data corresponding to the access request; The data buffer is further configured to remove the access request from the missing state processing register when receiving the data corresponding to the access request returned by the lower-level cache or memory.
10. The instruction processing component according to claim 9, The data buffer is configured to discard the access request when the data corresponding to the access request does not exist in the data buffer and the access request already exists in the missing state processing register.
11. The instruction processing component according to claim 9, The main pipeline circuit is further configured to trigger retransmission of the memory access instruction when the data corresponding to the access request does not exist in the data buffer and the missing state processing register is full.
12. An instruction processing method, which is executed by an instruction processing component, and the instruction processing component is the instruction processing component according to any one of claims 1 to 11. The method includes: Receiving a memory access instruction through the main pipeline circuit, and sending an access request to the data buffer based on the memory access instruction. The access request is used to request to query the data corresponding to the memory access instruction from the data buffer; Writing the data address corresponding to the memory access instruction into the data loading queue circuit through the main pipeline circuit when failing to query the data corresponding to the memory access instruction from the data buffer; Receiving backfill data through the data loading queue circuit, querying the corresponding data from the backfill data based on the data address written from the main pipeline circuit, and outputting the queried data; the backfill data is the data filled into the data buffer.
13. The method according to claim 12, wherein the data loading queue circuit comprises: At least two data loading queue sub-circuits, and an output sub-circuit; The at least two data loading queue sub-circuits are electrically connected to the main pipeline circuit respectively; At least two of the data loading queue sub-circuits are electrically connected to the output sub-circuit respectively; The writing the data address corresponding to the memory access instruction into the data loading queue circuit includes: Writing a plurality of data addresses into at least two of the data loading queue sub-circuits dispersedly through the main pipeline circuit; The receiving backfill data through the data loading queue circuit, querying the corresponding data from the backfill data based on the written data address, and outputting the queried data includes: Executing the steps of receiving backfill data and querying the corresponding data from the backfill data based on the written data address in parallel through at least two of the data loading queue sub-circuits; Obtaining and outputting the data queried by each of the at least two data loading queue sub-circuits through the output sub-circuit.
14. A processor, which includes at least one instruction processing component according to any one of claims 1 to 11.
15. A computer device, the computer device comprising at least one processor, the processor including at least one instruction processing component as described in any one of claims 1 to 11.
16. A storage medium, the storage medium being used for storing a computer program, the computer program being used for executing the method described in claim 12 or 13.
17. A computer program product including a computer program, when it runs on a computer, causing the computer to execute the method described in claim 12 or 13.
Citation Information
Patent Citations
High-performance low-power-consumption embedded processor based on command dual-transmission
CN101526895A
Non-blocking cache missing processing method and device
CN111142941A
Cache replacement system and method based on instruction stream and memory access mode learning
CN113986774A
Data processing unit, data processing method, and electronic device
CN116225979A