Instruction processing component, instruction processing method, processor and computer equipment

By designing an instruction processing component that includes the main pipeline circuit and the data loading queue circuit in the processor, the problem of excessive instruction resend when data is missing in the multi-level cache is solved, and the processor's memory access instruction processing efficiency and pipeline utilization are improved.

CN120216034APending Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311797388.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the processor, when data is missing in the multi-level cache, directly resending the memory access instruction will cause the instruction to be transmitted multiple times, occupying pipeline resources, reducing the effective usage of the pipeline, and increasing the power consumption of the read memory.

Method used

An instruction processing component is designed, including a mainstream pipeline circuit, a data buffer and a data load queue circuit. When the main pipeline circuit fails to query the data corresponding to the memory access instruction from the data buffer, it writes the data address to the data load queue circuit. The data load queue circuit receives the backfill data and queries the corresponding data based on the written data address.

Benefits of technology

It effectively reduces the number of resents of memory access instructions, improves the utilization rate of the mainstream pipeline, and thus improves the efficiency of processor processing memory access instructions, and reduces the power consumption of read memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216034A_ABST
    Figure CN120216034A_ABST
Patent Text Reader

Abstract

The invention discloses an instruction processing component, an instruction processing method, a processor and computer equipment, and relates to the technical field of chips. The instruction processing component comprises an instruction transmitting circuit, a main pipeline circuit, a data buffer and a data loading queue circuit, the data cache is a first-level cache in the processor; the main pipeline circuit is used for writing a data address corresponding to the memory access instruction into the data loading queue circuit under the condition that the data corresponding to the memory access instruction fails to be inquired from the data buffer; and the data loading queue circuit is used for receiving the backfill data, querying corresponding data from the backfill data based on the written data address, and outputting the queried data. According to the scheme, the retransmission times of the memory access instruction can be effectively reduced, the utilization rate of the main assembly line is improved, and the memory access instruction processing efficiency of the processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chip technology, and particularly to an instruction processing component, an instruction processing method, a processor, and a computer device. Background Art

[0002] In a processor, when the processor receives a memory access request, it accesses multiple levels of caches to read the request information. If the data is missing in all levels of caches, then the processor has to search the memory.

[0003] In the related art, during the memory access process of the processor, when detecting a cache miss situation, the memory access instruction is directly put back into the issue queue for execution until the issued instruction obtains the complete memory access data.

[0004] However, in the above solution, due to directly reissuing the instruction, a situation where an instruction is issued multiple times may occur. An instruction being reissued multiple times will occupy pipeline resources, reducing the effective utilization rate of the pipeline. And in the pipeline memory access instruction execution device, the power consumption of reading the memory is very large, resulting in waste of the power consumption of reading the memory. Summary of the Invention

[0005] Embodiments of this application provide an instruction processing component, an instruction processing method, a processor, and a computer device, which can improve the efficiency of the processor in processing memory access instructions. The technical solution is as follows.

[0006] On the one hand, an instruction processing component is provided. The instruction processing component includes: an instruction issue circuit, a main pipeline circuit, a data buffer, and a data load queue circuit; the data buffer is the first-level cache in the processor;

[0007] The main pipeline circuit is electrically connected to the data buffer, and the main pipeline circuit is electrically connected to the data load queue circuit;

[0008] The main pipeline circuit is configured to receive a memory access instruction, and send an access request to the data buffer based on the memory access instruction. The access request is used to request to query the data corresponding to the memory access instruction from the data buffer;

[0009] The main pipeline circuit is further configured to, in the case of failing to query the data corresponding to the memory access instruction from the data buffer, write the data address corresponding to the memory access instruction into the data load queue circuit;

[0010] The data load queue circuit is configured to receive the fill data, query the corresponding data from the fill data based on the written data address, and output the queried data; the fill data is the data filled into the data buffer.

[0011] On the other hand, a method for processing instructions is provided. The method is executed by the instruction processing component as described above, and the method includes:

[0012] Through the main pipeline circuit, receive a memory access instruction, and send an access request to the data buffer based on the memory access instruction. The access request is used to request to query the data corresponding to the memory access instruction from the data buffer;

[0013] Through the main pipeline circuit, in the case where the query for the data corresponding to the memory access instruction from the data buffer fails, write the data address corresponding to the memory access instruction into the data load queue circuit;

[0014] Through the data load queue circuit, receive the backfill data, query the corresponding data from the backfill data based on the written data address, and output the queried data; the backfill data is the data filled into the data buffer.

[0015] In some embodiments, the data load queue circuit includes: at least two data load queue sub - circuits, and an output sub - circuit; the at least two data load queue sub - circuits are electrically connected to the main pipeline circuit respectively; at least two of the data load queue sub - circuits are electrically connected to the output sub - circuit respectively;

[0016] The step of writing the data address corresponding to the memory access instruction into the data load queue circuit includes:

[0017] Through the main pipeline circuit, disperse and write multiple data addresses into at least two of the data load queue sub - circuits;

[0018] The step of, through the data load queue circuit, receiving the backfill data, querying the corresponding data from the backfill data based on the written data address, and outputting the queried data includes:

[0019] Through at least two of the data load queue sub - circuits, perform in parallel the steps of receiving the backfill data and querying the corresponding data from the backfill data based on the written data address;

[0020] Through the output sub - circuit, obtain the data queried by each of the at least two data load queue sub - circuits and output.

[0021] In some embodiments, the data load queue sub - circuit includes: an address entry queue, an address comparison device, a data extraction device, and an output selection circuit;

[0022] The address entry queue contains multiple address entries, and each address entry is used to store a written data address;

[0023] The data extraction device includes a data entry queue, a data selector, and a data outputter. The data entry queue contains multiple data entries;

[0024] The step of receiving backfill data in parallel through at least two of the data loading queue sub - circuits and querying corresponding data from the backfill data based on the written data address includes:

[0025] Through the address comparison device, compare the data address in the address entry with the data address of the backfill data, and when the data address in the address entry matches the data address of the backfill data, set the backfill field in the address entry to a specified value;

[0026] Through the data selector, when the data address in the address entry matches the data address of the backfill data, extract the data corresponding to the data address in the address entry from the backfill data, and store the extracted data into the data bit of the data entry corresponding to the address entry;

[0027] Through the output selection circuit, select the number of an address entry whose backfill field is the specified value, and transfer the number of the address entry to the data output component;

[0028] Through the data outputter, according to the number of the address entry transferred by the output selection circuit, transfer the data stored in the data entry corresponding to the number of the address entry in the data entry queue to the output sub - circuit.

[0029] In some embodiments, each data entry in the data entry queue further includes a valid signal bit; the data extraction device further includes an initialization circuit;

[0030] The method further includes:

[0031] Write the forwarded data into the initialization circuit through the main pipeline circuit; the forwarded data is composed of the result of querying from the data buffer; the forwarded data includes query result data and a valid signal, and the valid signal is used to indicate whether the query result data is valid;

[0032] Through the initialization circuit, write the query result data into the data bit of the data entry corresponding to the address entry where the data address of the query result data is located in the data entry queue, and write the valid signal into the valid signal bit of the data entry corresponding to the address entry where the data address of the query result data is located;

[0033] Through the data selector, when the data address in the address entry matches the data address of the backfill data, extracting the data corresponding to the data address in the address entry from the backfill data and storing the extracted data into the data bit of the data entry corresponding to the address entry, includes:

[0034] Through the data selector, when the data address in the address entry matches the data address of the backfill data and the valid signal in the valid signal bit of the data entry corresponding to the address entry indicates invalid, extracting the data corresponding to the data address in the address entry from the backfill data input to the data extraction device and storing the extracted data into the data bit of the data entry corresponding to the address entry.

[0035] In some embodiments, the data loading queue sub-circuit contains M of the data extraction devices, and the data loading queue sub-circuit further includes: a forwarding splitting unit, and a backfill splitting unit;

[0036] The method further includes:

[0037] Through the forwarding splitting unit, dispersedly writing the forwarding data of every M data addresses into the initialization circuits in the M data extraction devices;

[0038] Through the backfill splitting unit, dispersedly writing the backfill data of every M data addresses into the M data extraction devices.

[0039] In some embodiments, the data loading queue sub-circuit further contains a first pipeline register, and the data extraction device further contains a second pipeline register;

[0040] The step of, through the main pipeline circuit, dispersedly writing multiple data addresses into at least two of the data loading queue sub-circuits, includes:

[0041] Through the main pipeline circuit, writing the data address corresponding to the memory access instruction into the first pipeline register for pipelining processing;

[0042] The method further includes:

[0043] Through the backfill splitting unit, dispersedly writing the backfill data of consecutive M bytes into the second pipeline registers in the M data extraction devices for pipelining processing.

[0044] In some embodiments, the step of, through the output sub-circuit, obtaining and outputting the data queried by at least two of the data loading queue sub-circuits respectively, includes:

[0045] Through the output sub - circuit, in a polling manner, obtain the data queried by each of at least two of the data loading queue sub - circuits and output the data.

[0046] In some embodiments, the step of through the output sub - circuit, in a polling manner, obtain the data queried by each of at least two of the data loading queue sub - circuits and output the data includes:

[0047] Through the output sub - circuit, in a polling manner, obtain the data queried by each of at least two of the data loading queue sub - circuits, perform data shifting on the data queried by each of at least two of the data loading queue sub - circuits, and then output the data;

[0048] Wherein, the displacement amount of the data shifting is the low three bits of the address of the data queried by each of at least two of the data loading queue sub - circuits.

[0049] In some embodiments, the data buffer contains a missing status processing register; the method further includes:

[0050] Through the data buffer, when the data corresponding to the access request does not exist in the data buffer and the missing status processing register is not full, cache the access request into one of the entries in the missing status processing register;

[0051] Through the data buffer, send the access request in the missing status processing register to the lower - level cache or memory to query the data corresponding to the access request;

[0052] Through the data buffer, when receiving the data corresponding to the access request returned by the lower - level cache or memory, remove the access request from the missing status processing register.

[0053] In some embodiments, the method further includes:

[0054] Through the data buffer, when the data corresponding to the access request does not exist in the data buffer and the access request already exists in the missing status processing register, discard the access request.

[0055] In some embodiments, the method further includes:

[0056] Through the main pipeline circuit, when the data corresponding to the access request does not exist in the data buffer and the missing status processing register is full, trigger the re - transmission of the memory access instruction.

[0057] On the other hand, a processor is provided, and the processor includes at least one instruction processing component as described above.

[0058] On the other hand, a computer device is provided, which includes at least one processor, and the processor includes at least one instruction processing component as described above.

[0059] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0060] A data loading queue circuit is set in the processor. When the main pipeline circuit fails to query the data corresponding to the memory access instruction from the data cache, the data address corresponding to the memory access instruction is written into the data loading queue circuit. The data loading queue circuit receives the backfilled data, queries the corresponding data from the backfilled data based on the written data address, and outputs the queried data. That is to say, in the above solution, when the main pipeline circuit fails to query during the process of processing the memory access instruction, instead of directly resending the memory access instruction, the data loading queue circuit queries the corresponding data from the backfilled data of the data cache and outputs it, which can effectively reduce the number of resends of the memory access instruction, improve the utilization rate of the main pipeline, and further improve the efficiency of the processor in processing the memory access instruction. Description of the Drawings

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0062] Figure 1 is a schematic structural diagram of a processor;

[0063] Figure 2 is a schematic diagram of the working pipeline of an existing instruction processing component;

[0064] Figure 3 is a schematic diagram of the working pipeline of an instruction processing component provided by an exemplary embodiment of the present application;

[0065] Figure 4 is a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application;

[0066] Figure 5 is a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application;

[0067] Figure 6 is a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application;

[0068] Figure 7It is a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application;

[0069] Figure 8 It is a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application;

[0070] Figure 9 It is a schematic diagram of the instruction processing component structure provided by an exemplary embodiment of the present application;

[0071] Figure 10 It is a schematic diagram of the operation of the instruction processing component provided by an exemplary embodiment of the present application;

[0072] Figure 11 It is a schematic diagram of the pipeline structure involved in an exemplary embodiment of the present application;

[0073] Figure 12 It is a schematic diagram of the microarchitecture structure involved in an exemplary embodiment of the present application;

[0074] Figure 13 It is a schematic diagram of the operation of the refill split unit involved in an exemplary embodiment of the present application;

[0075] Figure 14 It is a schematic diagram of the operation of the right shifter involved in an exemplary embodiment of the present application;

[0076] Figure 15 It is a flowchart of the instruction processing method involved in an exemplary embodiment of the present application. Detailed implementation manners

[0077] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0078] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0079] The following first introduces some concepts involved in the present application:

[0080] 1) Main pipe: main pipeline, the main processing line; used to execute instructions and complete the computing tasks of the processor. In a processor, the processing of an instruction is divided into multiple stages, and different stages are executed by different circuits in the main pipeline.

[0081] 2) MSHR: miss status handling register, a register used to record the status information when cache misses occur in the processor, usually used in the control logic of the processor to take appropriate measures to handle cache misses when they occur.

[0082] 3) LDQ: load queue, a data structure inside the processor, used to record the status of the load instructions being executed and the data involved in these instructions.

[0083] 4) Dff: D-flip-flop, a basic sequential element in digital circuits, used to store the state of binary data.

[0084] 5) Mux: multiplexer, a digital circuit component used to select one of multiple input signals and pass it to a single output. Its function is similar to a switch, which can select one output from multiple inputs.

[0085] 6) RR: round robin, a round-robin arbiter; an arbiter design used to coordinate the access of multiple devices to shared resources. In this arbitration scheme, each device has the opportunity to obtain access to the shared resource within one round, ensuring fairness and uniformity. Its basic idea is to allocate arbitration signals to each device in sequence, so that each device has the opportunity to obtain the arbitration right within one round. If a device does not obtain the arbitration right in the current round, it will wait for the next round.

[0086] 7) LOD: leading one detect, a technique used to detect the first 1 at the most significant bit of data. In digital signal processing, communication systems, and hardware design, LOD is usually used to identify the highest non-zero bit of data to determine the range of data or perform other related operations.

[0087] Figure 1 The structural schematic diagram of a processor is shown as Figure 1 shown below:

[0088] The processor 100 consists of a controller 101, an arithmetic unit 102, and a memory 103. The controller 101 includes an instruction fetch unit 101a. The controller 101 is responsible for fetching instructions from the memory through the bus 110. The arithmetic unit 102 is used to perform operations according to the instructions fetched from the memory by the controller 101, such as performing addition and subtraction operations. The memory 103 is used to store the data used in the instructions. The instruction fetch unit 101a includes an instruction counter 101a1 and an instruction register 101a2. The instruction counter 101a1 is a register used to store the address of the currently executing instruction. When the computer executes a program, the instructions are executed one by one. The function of the instruction counter 101a1 is to record the address of the instruction to be executed currently. Whenever an instruction is executed, the address in the instruction counter 101a1 is updated to point to the next instruction to be executed, thereby realizing the sequential execution of the program. The instruction register 101a2 is a special register in the computer, used to store the currently executing instruction. The main function of the instruction fetch unit 101a is to read instructions from the memory and pass them to the instruction decoder for parsing and execution. When the computer runs a program, the instruction fetch unit 101a reads instructions from the memory through the bus 110, then stores the instructions in the instruction register 101a2, and then the instruction register 101a2 passes the operation code and operands of the instruction to the instruction decoder, and the instruction decoder performs corresponding operations according to the operation code.

[0089] Multi-level cache technology is widely used in modern processors. Since the direct access speed of the processor to the memory is very slow, usually up to several hundred clock cycles, but the computing speed of the processor is very fast. To match the computing speed of the processor, multi-level caches need to be added between the processor and the memory. The cache closer to the processor is faster, but the capacity is smaller. Due to the limited capacity of the cache, cache misses may occur during the process of the processor accessing the cache. If a cache miss occurs, it is necessary to search in the lower-level cache until the missing data is found.

[0090] When a cache miss occurs, the processor needs to perform a series of related processing for adaptation. The memory access instruction does not get the required data due to the cache miss and needs to be issued again. If a hit is found in the L2 cache after the L1 cache miss, the data can come back quickly, but if the data is missing in the L2 cache and subsequent cache levels, the required data needs to be searched in the memory, which will consume a lot of time. How the processor processes the scenario when a cache miss occurs is a key technical point in processor design.

[0091] Please refer to Figure 2 which shows a schematic diagram of the working pipeline of an existing instruction processing component. As Figure 2As shown, the existing solution is as shown in the above figure. The memory access instruction is issued from the issue queue. After being processed by the 4-stage main pipeline, if no exception occurs during the memory access process, the complete memory access data can be obtained at the last stage and written directly; but if an exception occurs, such as a cache miss being detected in the s3 stage, a miss replay is directly performed, that is, the instruction is re-put into the issue queue for execution.

[0092] However, in the above Figure 2 if a cache miss is detected in the third stage, the instruction is directly reissued, which will cause the instruction to be issued multiple times. Since the pipeline resources are relatively important, the multiple reissuance of an instruction will reduce the effective utilization rate of the pipeline;

[0093] If a lower-level cache miss occurs, the number of times the memory access instruction needs to be reissued increases significantly. During the execution of the memory access instruction, the data cache and the data tag memory need to be read, and the power consumption of reading the memory is very large, which greatly wastes power.

[0094] Please refer to Figure 3 , which shows a schematic diagram of the working pipeline of the instruction processing component provided by an exemplary embodiment of the present application. The instruction processing component includes: an instruction issuing circuit 301, a main pipeline circuit 302, a data buffer 303, and a data loading queue circuit 304.

[0095] One or more output ports of the instruction issuing circuit 301 are connected to one or more input ports of the main pipeline circuit 302. The memory access instruction is transmitted from the above output port of the instruction issuing circuit 301 to the above input port of the main pipeline circuit 302, so that the memory access instruction can be input into the main pipeline circuit 302.

[0096] The main pipeline circuit 302 is electrically connected to the data buffer 303, and the main pipeline circuit 302 is electrically connected to the data loading queue circuit 304.

[0097] One or more output ports of the main pipeline circuit 302 are connected to one or more input ports of the data buffer 303. The access request is transmitted from the above output port of the main pipeline circuit 302 to the above input port of the data buffer 303, so that the access request can be input into the data buffer 303.

[0098] One or more output ports of the data buffer 303 are connected to one or more input ports of the main pipeline circuit 302. The query result is transmitted from the above output port of the data buffer 303 to the above input port of the main pipeline circuit 302, so that the query result is input into the main pipeline circuit 302.

[0099] One or more output ports of the main pipeline circuit 302 are connected to one or more input ports of the data load queue circuit 304. The data address corresponding to the memory access instruction is transmitted from the above output port of the main pipeline circuit 302 to the above input port of the data load queue circuit 304, so that the data address corresponding to the memory access instruction is input into the data load queue circuit 304.

[0100] The main pipeline circuit 302 is configured to receive a memory access instruction and send an access request to the data buffer 303 based on the memory access instruction. The access request is used to request to query the data corresponding to the memory access instruction from the data buffer 303.

[0101] Among them, the above memory access instruction can also be called a memory access instruction; during the process of processing the memory access instruction by the main pipeline circuit 302, it is necessary to obtain and write out the data corresponding to the memory access instruction; during this process, the main pipeline circuit 302 can preferentially send an access request to the first-level cache (i.e., the above data buffer) to query the data corresponding to the memory access instruction from the data buffer 303.

[0102] The main pipeline circuit 302 is further configured to write the data address corresponding to the memory access instruction into the data load queue circuit 304 when the query for the data corresponding to the memory access instruction from the data buffer 303 fails.

[0103] Among them, the failure to query the data corresponding to the memory access instruction from the data buffer 303 may mean that all or part of the data corresponding to the memory access instruction does not exist in the data buffer 303; that is, when all the data corresponding to the memory access instruction does not exist in the data buffer 303, it is considered that the query for the data corresponding to the memory access instruction from the data buffer 303 fails, and when only a part of all the data corresponding to the memory access instruction exists in the data buffer 303, it is also considered that the query for the data corresponding to the memory access instruction from the data buffer 303 fails. In the embodiments of the present application, when the query for the data corresponding to the memory access instruction from the data buffer 303 fails, the main pipeline circuit 302 does not immediately trigger the retransmission of the memory access instruction, but writes the data address corresponding to the memory access instruction into the data load queue circuit 304, and the data load queue circuit 304 continues to obtain the corresponding data. In some embodiments, the above memory access instruction can be used to obtain the data corresponding to one or more data addresses. At this time, the main pipeline circuit 302 can write the above one or more data addresses into the data load queue circuit 304.

[0104] The data load queue circuit 304 is configured to receive the backfill data, query the corresponding data from the backfill data based on the written data address, and output the queried data; the backfill data is the data filled into the data buffer 303.

[0105] Among them, the above-mentioned backfilled data may refer to the data that is not queried from the data buffer and is written into the data buffer 303 after being queried from the lower-level buffer or memory of the data buffer 303 subsequently. For example, at a certain moment 1, the main pipeline circuit 302 queries the data corresponding to a data address A from the data buffer 303. The data buffer 303 does not store the data corresponding to the data address A at moment 1. Subsequently, at moment 2, it queries the data corresponding to the data address A from the lower-level buffer or memory and the query is successful. At this time, the data corresponding to the data address A can be written into the data buffer 303. The data corresponding to the data address A written into the data buffer 303 here can be referred to as backfilled data. Optionally, when writing the data corresponding to the data address A into the data buffer 303, the data block where the data corresponding to the data address A is located can be written into the data buffer 303. At this time, the data block where the data corresponding to the data address A is located can be referred to as the above-mentioned backfilled data. In the embodiments of the present application, in addition to being written into the data buffer 303, the backfilled data is also written into the data load queue circuit 304. For example, the line connecting the lower-level buffer or memory to the above-mentioned data buffer can also extend to the above-mentioned data load queue circuit 304. In this way, when there is backfilled data written into the data buffer 303, the backfilled data will also be delivered to the data load queue circuit 304.

[0106] In summary, in the solution shown in the embodiments of the present application, a data load queue circuit 304 is set in the processor. When the main pipeline circuit 302 fails to query the data corresponding to the memory access instruction from the data cache 303, it writes the data address corresponding to the memory access instruction into the data load queue circuit 304. The data load queue circuit 304 receives the backfilled data, queries the corresponding data from the backfilled data based on the written data address, and outputs the queried data. That is to say, in the above solution, when the main pipeline circuit 302 fails to query from the data buffer 303 during the process of processing the memory access instruction, it does not directly resend the memory access instruction, but the data load queue circuit 304 queries the corresponding data from the backfilled data of the data buffer 303 and outputs it, which can effectively reduce the resend times of the memory access instruction, improve the utilization rate of the main pipeline, and thus improve the efficiency of the processor in processing the memory access instruction.

[0107] Based on Figure 3 the instruction processing component shown, please refer to Figure 4 , which shows a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application. The data load queue circuit 304 may include: at least two data load queue sub-circuits 3041, and an output sub-circuit 305; at least two data load queue sub-circuits 3041 are electrically connected to the main pipeline circuit 302 respectively; at least two data load queue sub-circuits 3041 are electrically connected to the output sub-circuit 305 respectively.

[0108] One or more output ports of the main pipeline circuit 302 are respectively connected to one or more input ports of at least two data loading queue sub - circuits 3041. Multiple data addresses are transmitted from the above - mentioned output ports of the main pipeline 302 to the above - mentioned input ports of at least two data loading queue sub - circuits 3041, so that the multiple data addresses are split into two parts and are respectively input into at least two data loading queue sub - circuits 3041;

[0109] One or more output ports of at least two data loading sub - circuits 3041 are connected to one or more input ports of the output sub - circuit 305. The data queried by at least two data loading queue sub - circuits 3041 are transmitted from the above - mentioned output ports of at least two data loading queue sub - circuits 3041 to the above - mentioned input ports of the output sub - circuit, so that the queried data is input into the output sub - circuit 305.

[0110] The main pipeline circuit 302 is used to disperse and write multiple data addresses into at least two data loading queue sub - circuits 3041.

[0111] Exemplarily, the main pipeline circuit 302 splits the data addresses according to a preset algorithm or rule and writes them into at least two data loading queue sub - circuits 3041 according to the logic of parallel writing, ensuring uniform address distribution and improving the parallelism of data query in the data loading queue circuit 304.

[0112] At least two data loading queue sub - circuits 3041 are used to execute the steps of receiving back - filled data in parallel and querying corresponding data from the back - filled data based on the written data addresses.

[0113] The output sub - circuit 305 is used to obtain the data queried by at least two data loading queue sub - circuits 3041 respectively and output it.

[0114] In the embodiment of the present application, multiple data addresses can be dispersed into at least two data loading queue sub - circuits 3041, so that at least two data loading queue sub - circuits 3041 execute the step of querying data from the back - filled data in parallel, thereby improving the concurrency of data query and the efficiency of obtaining the data corresponding to the memory access instructions with query failures from the back - filled data.

[0115] Please refer to Figure 5 , which shows a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application, Figure 4 The data loading queue sub - circuit 3041 in [] includes: an address entry queue 3041a, an address comparison device 3041b, a data extraction device 3041c, and an output selection circuit 3041d.

[0116] One or more output ports of the address entry queue 3041a are connected to one or more input ports of the address comparison device 3041b. The data address in the address entry is transmitted from the above-mentioned output port of the address entry queue 3041a to the above-mentioned input port of the address comparison device 3041b, so that the data address in the address entry is input into the address comparison device 3041b.

[0117] One or more output ports of the address entry queue 3041a are connected to one or more input ports of the output selection circuit 3041d. The backfill field is transmitted from the above-mentioned output port of the address entry queue 3041a to the above-mentioned input port of the output selection circuit 3041d, so that the backfill field is input into the output selection circuit 3041d;

[0118] One or more output ports of the address comparison device 3041b are connected to one or more input ports of the data extraction device 3041c.

[0119] One or more input ports of the address comparison device 3041b are connected to one or more output ports of the main pipeline circuit 302. The data address of the backfill data is transmitted from the above-mentioned output port of the main pipeline circuit 302 to the above-mentioned input port of the address comparison device 3041b, so that the data address of the backfill data is input into the address comparison device 3041b.

[0120] The address entry queue 3041a contains multiple address entries, and each address entry is used to store a written data address.

[0121] The above-mentioned address entry contains information such as a number, a data address, and a backfill field.

[0122] The data extraction device 3041c contains a data entry queue 3041c1, a data selector 3041c2, and a data output device 3041c3. The data entry queue 3041c1 contains multiple data entries.

[0123] One or more input ports of the data entry queue 3041c1 are connected to one or more output ports of the data selector 3041c2. The data corresponding to the data address in the address entry is transmitted from the above-mentioned output port of the data selector 3041c2 to the above-mentioned input port of the data entry queue 3041c1, so that the above-mentioned data is input into the data entry queue 3041c1.

[0124] One or more output ports of the data entry queue 3041c1 are connected to one or more input ports of the data outputter 3041c3. The data entries in the data entry queue 3041c1 are transferred from the above-mentioned output ports of the data entry queue 3041c1 to the above-mentioned input ports of the data outputter 3041c3, enabling the above-mentioned data entries to be input into the data outputter 3041c3.

[0125] One or more input ports of the data selector 3041c2 are connected to one or more output ports of the main pipeline circuit 302. The backfill data is transferred from the above-mentioned output ports of the main pipeline circuit 302 to the above-mentioned input ports of the data selector 3041c2, enabling the backfill data to be input into the data selector 3041c2.

[0126] One or more input ports of the data outputter 3041c3 are connected to one or more outputs of the output selection circuit 3041d. The number of the address entry is transferred from the above-mentioned output ports of the output selection circuit 3041d to the above-mentioned input ports of the data outputter 3041c3, enabling the number of the address entry to be input into the data outputter 3041c3.

[0127] The address comparison device 3041b is used to compare the data address in the address entry with the data address of the backfill data, and when the data address in the address entry matches the data address of the backfill data, set the backfill field in the address entry to a specified value.

[0128] The above-mentioned address match means that when comparing the value of the data address in the address entry with the value of the data address of the backfill data, the values of the two addresses are the same.

[0129] The above-mentioned specified value is a specific numerical value or state for setting the backfill field during address matching. This specified value can be defined by developers during development and design, and can be a binary number, a status flag, or other content that needs to be set during matching. For example, the above-mentioned specified value can be 1 or 0.

[0130] The data selector 3041c2 is used to extract the data corresponding to the data address in the address entry from the backfill data when the data address in the address entry matches the data address of the backfill data, and store the extracted data into the data bit of the data entry corresponding to the address entry.

[0131] The output selection circuit 3041d is used to select the number of an address entry whose backfill field is the specified value and transfer the number of the address entry to the data output component.

[0132] The above-mentioned number is used to uniquely identify or recognize different address entries, and each address entry in the processor can be preset with a fixed number.

[0133] Alternatively, the numbers of the above address entries can also be generated by the processor. For example, in some embodiments, an incrementing number can be used as the number, and each time a new address entry is added, the number is automatically incremented; for example, 1, 2, 3, etc.; in some embodiments, a hash function can be used to generate the number, and the hash function can convert data into a hash value of a fixed length, and these hash values are used as the numbers; usually represented as a string or number of a fixed length. In some embodiments, the number can also be generated based on business rules, and the business rules include information such as specific prefixes, dates, locations, etc. For example, 20231127 represents November 27, 2023.

[0134] In the embodiment of the present application, the output selection circuit 3041d can sequentially search each address entry according to the specified value of the backfill field, find the address entry whose backfill field is the specified value, then query the number of the address entry, and pass the number to the data output component.

[0135] The data outputter 3041c3 is configured to transfer the data stored in the data entry corresponding to the number of the address entry passed by the output selection circuit 3041d in the data entry queue 3041c1 to the output sub-circuit 305.

[0136] In the solution shown in the embodiment of the present application, by setting an address comparison device to compare the data address of the backfill data with the data address that failed to be queried from the data buffer 303, the data selector 3041c2 is triggered to accurately extract the data that failed to be queried before from the backfill data, and the queried data is output through the output selection circuit 3041d and the data outputter 3041c3, thereby implementing a circuit structure for data extraction based on address comparison, ensuring that the data corresponding to the data address of the memory access instruction that failed to be queried before can be obtained from the backfill data through a hardware circuit, and ensuring the execution efficiency of data extraction.

[0137] Please refer to Figure 6 , which shows a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application. Based on Figure 5 Each data entry in the data entry queue 3041c1 in also includes a valid signal bit 3041c1a, and the data extraction device 3041c also includes an initialization circuit 3041c4.

[0138] One or more input ports of the valid signal bit 3041c1a are connected to one or more output ports of the initialization circuit 3041c4, and the valid signal is transmitted from the above output port of the initialization circuit 3041c4 to the above input port of the valid signal bit 3041c1a, so that the valid signal is input into the valid signal bit 3041c1a.

[0139] One or more input ports of the initialization circuit 3041c4 are connected to one or more output ports of the main pipeline circuit 302, and the forwarded data is transmitted from the above output ports of the main pipeline circuit 302 to the above input ports of the initialization circuit 3041c4, so that the forwarded data is input into the initialization circuit 3041c4.

[0140] The main pipeline circuit 302 is further configured to write the forwarded data into the initialization circuit 3041c4; the forwarded data is composed of the results queried from the data buffer; the forwarded data includes the query result data and a valid signal, and the valid signal is used to indicate whether the query result data is valid.

[0141] Optionally, the above valid signal bit can be 1 or 0. 1 indicates that the query result data is valid, and 0 indicates that the query result data is invalid; or, 0 indicates that the query result data is valid, and 1 indicates that the query result data is invalid.

[0142] The initialization circuit 3041c4 is configured to write the query result data into the data entry queue 3041c1, into the data bit of the data entry corresponding to the address entry where the data address of the query result data is located, and write the valid signal into the valid signal bit 3041c1a of the data entry corresponding to the address entry where the data address of the query result data is located.

[0143] The initialization circuit 3041c4 obtains data from the main pipeline circuit 302 and writes the valid signal into the valid signal bit 3041c1a of the data entry corresponding to the address entry where the data address of the query result data is located.

[0144] Exemplarily, when the data address in the address entry matches the data address of the backfilled data, it indicates that the data address in the address entry has been backfilled. The initialization circuit 3041c4 writes 0 into the valid signal bit of the data entry corresponding to the address entry, and the valid signal indicates invalid; when the data address in the address entry does not match the data address of the backfilled data, it indicates that the data address in the address entry has not been backfilled. The initialization circuit 3041c4 writes 1 into the valid signal bit of the data entry corresponding to the address entry, and the valid signal indicates valid.

[0145] The data selector 3041c2 is configured to extract the data corresponding to the data address in the address entry from the backfilled data of the input data extraction device and store the extracted data into the data bit of the data entry corresponding to the address entry when the data address in the address entry matches the data address of the backfilled data and the valid signal in the valid signal bit 3041c1a of the data entry corresponding to the address entry indicates invalid.

[0146] In the solution shown in the embodiment of the present application, the forwarded data is also written into the data entry queue 3041c1. When obtaining the data corresponding to the data address of the memory access instruction that failed to be queried previously from the backfilled data, for the partially obtained data, it is not necessary to repeatedly query from the backfilled data, thereby further improving the efficiency of obtaining the data corresponding to the data address of the memory access instruction that failed to be queried previously from the backfilled data.

[0147] Please refer to Figure 7 , which shows a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application. Based on Figure 6 In the data loading queue sub-circuit 3041, there are M data extraction devices 3041c. The data loading queue sub-circuit 3041 also includes: a forwarding splitting unit 3041e and a backfilling splitting unit 3041f.

[0148] One or more input ports of the forwarding splitting unit 3041e are connected to one or more output ports of the main pipeline circuit 302. The forwarded data is transmitted from the above output ports of the main pipeline circuit 302 to the above input ports of the forwarding splitting unit 3041e, so that the forwarded data is input into the forwarding splitting unit 3041e.

[0149] One or more output ports of the forwarding splitting unit 3041e are connected to one or more input ports of the initialization circuit 3041c4. The split forwarded data is transmitted from the above output ports of the forwarding splitting unit 3041e to the above input ports of the initialization circuit 3041c4, so that the split forwarded data is input into the initialization circuit 3041c4.

[0150] One or more input ports of the backfilling splitting unit 3041f are connected to one or more output ports of the main pipeline circuit 302. The backfilled data is transmitted from the above output ports of the main pipeline circuit 302 to the above input ports of the backfilling splitting unit 3041f, so that the backfilled data is input into the backfilling splitting unit 3041f.

[0151] One or more output ports of the backfilling splitting unit 3041f are connected to one or more input ports of the data selector 3041c2. The split backfilled data is transmitted from the above output ports of the backfilling splitting unit 3041f to the above input ports of the data selector 3041c2, so that the split backfilled data is input into the data selector 3041c2.

[0152] The forwarding splitting unit 3041e is configured to disperse and write the forwarded data of every M data addresses into the initialization circuit 3041c4 in the M data extraction devices.

[0153] Exemplarily, the data addresses are A1, A2, A3, A4, A5, A6, A7, A8, A9, and M is 3. The forward split unit 3041e dispersedly writes the data addresses into the initialization circuits of 3 data extraction devices, and the obtained groups may be Group 1: A1, A2, A3; Group 2: A4, A5, A6; Group 3: A7, A8, A9.

[0154] The backfill split unit 3041f is used to dispersedly write the backfill data of every M data addresses into M data extraction devices 3041c.

[0155] Exemplarily, the data addresses are B1, B2, B3, B4, B5, B6, B7, B8, B9, and M is 3. The backfill split unit 3041f dispersedly writes the data addresses into 3 data extraction devices, and the obtained groups may be Group 1: B1, B2, B3; Group 2: B4, B5, B6; Group 3: B7, B8, B9.

[0156] In the solution shown in the embodiments of the present application, M data extraction devices 3041c can be set in one data loading queue sub - circuit 3041. Moreover, the forward data and backfill data input into the data loading queue sub - circuit 3041 will be dispersed to M data extraction devices 3041c, so that the M data extraction devices 3041c can parallelly execute the step of querying data from the backfill data, thereby improving the concurrency of data query and the efficiency of obtaining the data corresponding to the memory access instructions with query failures from the backfill data.

[0157] Please refer to Figure 8 , which shows a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application. Based on Figure 7 the data loading queue sub - circuit 3041 further includes a first pipelining register 3041g, and the data extraction device 3041c further includes a second pipelining register 3041c5.

[0158] One or more input ports of the first pipelining register 3041g are connected to one or more output ports of the main pipeline circuit 302. The data address corresponding to the memory access instruction is transmitted from the above - mentioned output port of the main pipeline circuit 302 to the above - mentioned input port of the first pipelining register 3041g, so that the data address corresponding to the memory access instruction is input into the first pipelining register 3041g.

[0159] One or more output ports of the first beat register 3041g are connected to one or more input ports of the address comparison device 3041b. The data address corresponding to the memory access instruction after beat processing is transmitted from the above output port of the first beat register 3041g to the above input port of the address comparison device 3041b, so that the data address corresponding to the memory access instruction after beat processing is input into the address comparison device 3041b.

[0160] One or more input ports of the second beat register 3041c5 are connected to one or more output ports of the fill split unit 3041f. The split fill data is transmitted from the above output port of the fill split unit 3041f to the above input port of the second beat register 3041c5, so that the split fill data is input into the second beat register 3041c5.

[0161] One or more output ports of the second beat register 3041c5 are connected to one or more output ports of the data selector 3041c2. The fill data after beat processing is transmitted from the above output port of the second beat register 3041c5 to the above output port of the data selector 3041c2, so that the fill data after beat processing is input into the data selector 3041c2.

[0162] The main pipeline circuit 302 is used to write the data address corresponding to the memory access instruction into the first beat register 3041g for beat processing.

[0163] The fill split unit is used to disperse and write the fill data of continuous M bytes into the second beat register 3041c5 in M data extraction devices for beat processing.

[0164] In the solution shown in the embodiment of the present application, at the entrance of the data extraction device 3041c, beat delay processing can be performed on the fill data address and the fill data, avoiding the problem of too long processing cycle caused by too long processing flow in a single processing process, thereby avoiding affecting the main frequency design of the processor.

[0165] Please refer to Figure 9 , which shows a schematic structural diagram of an instruction processing component provided by an exemplary embodiment of the present application. Figure 4 In [reference], the output sub-circuit 305 is used to obtain and output the data queried by at least two data loading queue sub-circuits 3041 in a polling manner.

[0166] The input port of the output sub-circuit 305 is respectively connected to the output port of the data loading queue circuit 304; the data queried by each of at least two data loading queue sub-circuits 3041 is respectively transmitted from the above output port of the data loading queue circuit 304 to the above input port of the output sub-circuit 305, so that the queried data is input into the output sub-circuit 305.

[0167] The output sub-circuit 305 is used to obtain the data queried by each of at least two data loading queue sub-circuits 3041 in a polling manner, and perform data shifting on the data queried by each of at least two data loading queue sub-circuits 3041 and then output.

[0168] Wherein, the displacement amount of the data shift is the low three bits of the address of the data queried by each of at least two data loading queue sub-circuits 3041.

[0169] In some embodiments, the above output sub-circuit 305 may include a polling arbitration sub-circuit 3051 and a polling arbitration sub-circuit 3052.

[0170] The input port of the polling arbitration sub-circuit 3051 is respectively connected to the output port of the data loading queue sub-circuit 3041; the data queried by each of at least two data loading queue sub-circuits 3041 is respectively transmitted from the above output port of the data loading queue sub-circuit 3041 to the above input port of the polling arbitration sub-circuit 3051, so that the queried data is input into the polling arbitration sub-circuit 3051;

[0171] The input port of the polling arbitration sub-circuit 3052 is respectively connected to the output port of the data loading queue sub-circuit 3041; the data queried by each of at least two data loading queue sub-circuits 3041 is respectively transmitted from the above output port of the data loading queue sub-circuit 3041 to the above input port of the polling arbitration sub-circuit 3052, so that the queried data is input into the polling arbitration sub-circuit 3052.

[0172] Wherein, the above polling arbitration sub-circuit 3051 is used to obtain the address data output by each of at least two data loading queue sub-circuits 3041 in a polling manner, and obtain one address data from the address data output by at least two data loading queue sub-circuits 3041;

[0173] The above polling arbitration sub-circuit 3052 is used to obtain the address data output by each of at least two data loading queue sub-circuits 3041 in a polling manner, and obtain one address data from the address data output by at least two data loading queue sub-circuits 3041.

[0174] Please refer to Figure 10, which shows a schematic diagram of the operation of an instruction processing component provided by an exemplary embodiment of the present application. The data cache 303 contains a miss status processing register 3031.

[0175] Among them, one or more input ports of the miss status processing register 3031 are connected to one or more output ports of the main pipeline circuit 302. When the data corresponding to the access request does not exist in the data cache 303 and the miss status processing register 3031 is not full, the access request is transmitted from the above output port of the main pipeline circuit 302 to the miss status processing register 3031, so that the access request is input into the miss status processing register 3031.

[0176] The data cache 303 is used to cache the access request into an entry in the miss status processing register 3031 when the data corresponding to the access request does not exist in the data cache 303 and the miss status processing register 3031 is not full.

[0177] The above entry may contain information such as the address of the missing data, the access type (read or write), etc.

[0178] The data cache 303 is also used to send the access request in the miss status processing register 3031 to the lower-level cache or memory to query the data corresponding to the access request;

[0179] The above access request may contain a read or write request, a request address, etc.

[0180] The data cache 303 is also used to remove the access request from the miss status processing register when the data corresponding to the access request returned by the lower-level cache or memory is received.

[0181] In the embodiment of the present application, a miss status processing register 3031 is also set in the data cache 303. For a certain access request, when the access request fails to be queried in the data cache 303, it is not necessary to resend the corresponding memory access instruction, but it is temporarily stored in the miss status processing register 3031, and then the data cache 303 sends the access request to the lower-level cache or memory for data query, thereby reducing the number of resends of the memory access instruction and improving the processing efficiency of the main pipeline.

[0182] The data cache 303 is used to discard the access request when the data corresponding to the access request does not exist in the data cache 303 and the access request already exists in the miss status processing register 3031.

[0183] In the embodiment of the present application, for access requests that occur repeatedly and are not found in the data buffer 303, it is not necessary to temporarily store them in the miss status processing register 3031 every time. Instead, when there is an access request in the miss status processing register 3031 and a new identical access request arrives, the two access requests are merged (discarding one of them), thereby reducing the duplicate processing of access requests and only improving the efficiency of data requests.

[0184] The main pipeline circuit 302 is further configured to trigger the retransmission of the memory access instruction when the data corresponding to the access request does not exist in the data buffer 303 and the miss status processing register 3031 is full.

[0185] In the embodiment of the present application, for a certain access request, when the access request fails to be found in the data buffer 303 and the miss status processing register 3031 is full at this time, the retransmission of the memory access instruction can be triggered, thereby avoiding the situation where the subsequent access instruction cannot be correctly executed due to the overflow of the miss status processing register 3031.

[0186] Based on the above Figures 3 to 10 For any of the above - shown solutions, the present application can provide a load queue design that supports parallel backfill data detection. When a cache miss occurs and the corresponding miss status processing register is not full, the current request can enter the miss status processing register, the current memory access instruction does not need to be retransmitted, and the subsequent backfill data of the L1 cache is continuously detected, and parallel detection and capture are realized by means of dividing banks and dividing high - low queues, which greatly improves the performance of the processor's memory access instructions.

[0187] Based on Figures 3 to 10 For any of the above - shown solutions, the detailed description of the technical solution of this embodiment is as follows:

[0188] I. Product side

[0189] In the field of processors, the multi - level cache technology is widely used. The load queue design that supports parallel backfill data detection proposed in this embodiment improves the effective utilization rate of the pipeline, reduces the chip power consumption, and effectively improves the competitiveness of the product.

[0190] II. Technical side

[0191] Please refer to Figure 11 , which shows a schematic diagram of the pipeline structure involved in an exemplary embodiment of the present application. As Figure 11As shown, in the s1 stage of the pipeline, an access request to the L1 dcache is issued. The MSHR in the L1 dcache is used to handle miss access requests. If a cache miss occurs during the current request to access L1, it directly enters the MSHR. The L1 dcache selects a request from the MSHR and sends it to the lower-level L2 dcache.

[0192] The MSHR usually has multiple entries. When a miss instruction enters the MSHR, it occupies one entry. When the request of this instruction is sent to the lower-level cache and the returned data is received, its corresponding entry is released.

[0193] Another function of the MSHR is request merging. If the address corresponding to the existing entry in the MSHR matches the address of the current miss, the current miss instruction will not apply for an entry again, but share it with the matching entry.

[0194] Due to the existence of the MSHR, even if a memory access instruction misses the access to the L1 dcache, as long as the instruction can enter the MSHR, its request will be sent to the lower-level cache subsequently. So when a cache miss occurs and the MSHR is not full, the instruction flows backward normally. In the s4 stage, it is judged whether the dcache data is obtained. If it is obtained, it is directly written out; otherwise, it enters the LDQ to wait. At this time, the data bypassed from the store pipeline needs to be written into the LDQ, that is, forward_in in the figure.

[0195] Subsequently, the LDQ needs to continuously detect the refill data returned from the lower-level cache. On the one hand, the refill data needs to be sent into the L1 dcache to refill the missing cache line. On the other hand, it will be sent into the LDQ in parallel. If the refill address matches the address in the LDQ, the refill data will be merged into the corresponding entry. When the refill data in the LDQ is ready, it will be written out.

[0196] As Figure 12 shown, it shows a schematic diagram of the microarchitecture structure involved in an exemplary embodiment of the present application. As Figure 12As shown, the LDQ is divided into two queues, namely the L_LDQ (low LDQ) and the H_LDQ (high LDQ). Assuming the queue depth of the LDQ is N, the L_LDQ stores the entries numbered from 0 to N / 2 - 1, and the H_LDQ stores the entries numbered from N / 2 to N - 1. The refill data will be copied twice and then enter the high and low queues respectively. After a simple split in each queue, it directly enters the dff2 for clocking. Since the width of the refill data is relatively wide, up to 512 bits, this way of dividing into high and low queues can effectively reduce the congestion of the connections, and the way of clocking at the entrances of the high and low queues is also relatively friendly to the timing. Only the L_LDQ will be described below, and the H_LDQ is isomorphic to it.

[0197] The widest data type of the memory access instruction is double word, that is, 64 bits. The L_LDQ is divided into 8 banks, and each bank processes 1 byte (i.e., 8 bits) of data. The Refill data first enters the refill split unit for byte-by-byte splitting processing.

[0198] As Figure 13 shown, it shows a schematic diagram of the operation of the refill split unit involved in an exemplary embodiment of the present application. As Figure 13 shown, the schematic diagram of splitting by byte is as shown above. In the above figure, a small vertical bar is 1 byte, and the input is 512 bits, that is, 64 bytes of data volume. The splitting method is to extract the bytes at the same position in every 8 bytes together. For example, the green square is the 0th byte in the 8-byte data. After extracting them together, they are sent to bank0. This processing method processes the data of different bytes in different banks, and the processing between different banks is completely parallel and does not interfere with each other in wiring, which is highly friendly to layout and wiring in the physical implementation of the chip.

[0199] Forward_in is the data forwarded through the store pipeline, which consists of two parts. One is the forwarded data with a width of 64 bits, and the other is the data valid signal valid, which is 8 bits. valid[i] being high indicates that the i-th byte in the forwarded data is valid. Since the data forwarded through the store pipeline is updated, during the process of merging with the dcache refill data, the forwarded data has the highest priority. That is, if the i-th byte of the forwarded data is valid, the refill data cannot overwrite it. Forward_in needs to enter the forward split unit for splitting, and the splitting logic is the same as that of the refill split unit. After splitting, the forward valid signal and data are respectively written into the allocated entry.

[0200] Refill_addr_in is registered in dff1 and then sent to L_LDQ. The refill address is compared with the addr (address) of each entry in L_LDQ. If the addresses match, it means that the refill address matches the address of the current entry, and capture is required. The capture process is to select 8 bits from the 64-bit input data according to the address information, and this selection function is implemented by a mux. The capture enable signal enable = compare_sucess && (!forward), that is, when the addresses match and there is no data forwarding for the current entry, the refill data can be captured.

[0201] When it is detected that the refill address comparison is successful, the refilled field of the corresponding entry is set to 1. Subsequently, an entry with a refilled field of 1 is selected through the LOD circuit for writing. The LOD outputs the number of the first refilled entry found, and the data field is selected through mux1 using this number.

[0202] Finally, one refilled entry is selected from the high and low queues respectively, and round-robin arbitration is performed through RR to obtain the data that finally needs to be written.

[0203] As Figure 14 shown, it shows the working schematic diagram of the right shifter involved in an exemplary embodiment of the present application. As Figure 14 shown, the data output through round-robin arbitration also needs to be data-shifted, and the shift amount is the lower 3 bits of the address. Two scenarios of data shifting for different data types are as Figure 14As shown, for the case where the data type is half word, the data is located at addresses 4 and 5, and they need to be right-shifted by 4 bytes to concentrate the valid data in the lower bits. For the case where the data type is byte, the data is located at address 7, and it needs to be right-shifted by 7 bytes.

[0204] Please refer to Figure 15 , which shows the flowchart of the instruction processing method involved in an exemplary embodiment of the present application. As Figure 15 shown, this method can be executed by an instruction processing component, and the instruction processing component can be any one of the Figures 3 to 10 instruction processing components described above. As Figure 15 shown, this method may include the following steps:

[0205] Step 1501: Receive a memory access instruction through the main pipeline circuit, and send an access request to the data buffer based on the memory access instruction. The access request is used to request to query the data corresponding to the memory access instruction from the data buffer.

[0206] Step 1502: Through the main pipeline circuit, in the case of failure to query the data corresponding to the memory access instruction from the data buffer, write the data address corresponding to the memory access instruction into the data load queue circuit.

[0207] Step 1503: Through the data load queue circuit, receive the backfill data, query the corresponding data from the backfill data based on the written data address, and output the queried data; the backfill data is the data filled into the data buffer.

[0208] In some embodiments, the data load queue circuit includes: at least two data load queue sub-circuits, and an output sub-circuit; the at least two data load queue sub-circuits are respectively electrically connected to the main pipeline circuit; at least two of the data load queue sub-circuits are respectively electrically connected to the output sub-circuit;

[0209] The step of writing the data address corresponding to the memory access instruction into the data load queue circuit includes:

[0210] Through the main pipeline circuit, disperse and write multiple data addresses into at least two of the data load queue sub-circuits;

[0211] The step of receiving the backfill data through the data load queue circuit, querying the corresponding data from the backfill data based on the written data address, and outputting the queried data includes:

[0212] Through at least two of the data load queue sub-circuits, execute in parallel the steps of receiving the backfill data and querying the corresponding data from the backfill data based on the written data address;

[0213] Obtain the data queried by each of at least two of the data loading queue sub - circuits through the output sub - circuit and output it.

[0214] In some embodiments, the data loading queue sub - circuit includes: an address entry queue, an address comparison device, a data extraction device, and an output selection circuit;

[0215] The address entry queue contains a plurality of address entries, and each address entry is used to store a written data address;

[0216] The data extraction device includes a data entry queue, a data selector, and a data outputter, and the data entry queue contains a plurality of data entries;

[0217] The step of receiving back - filled data in parallel through at least two of the data loading queue sub - circuits and querying corresponding data from the back - filled data based on the written data address includes:

[0218] Through the address comparison device, compare the data address in the address entry with the data address of the back - filled data, and when the data address in the address entry matches the data address of the back - filled data, set the back - fill field in the address entry to a specified value;

[0219] Through the data selector, when the data address in the address entry matches the data address of the back - filled data, extract the data corresponding to the data address in the address entry from the back - filled data and store the extracted data in the data bit of the data entry corresponding to the address entry;

[0220] Through the output selection circuit, select the number of an address entry whose back - fill field is the specified value, and transfer the number of the address entry to the data output component;

[0221] Through the data outputter, according to the number of the address entry transferred by the output selection circuit, transfer the data stored in the data entry corresponding to the number of the address entry in the data entry queue to the output sub - circuit.

[0222] In some embodiments, each data entry in the data entry queue further includes a valid signal bit; the data extraction device further includes an initialization circuit;

[0223] The method further includes:

[0224] Through the main pipeline circuit, write the forwarded data into the initialization circuit; the forwarded data is composed of the results queried from the data buffer; the forwarded data includes query result data and a valid signal, and the valid signal is used to indicate whether the query result data is valid;

[0225] Through the initialization circuit, write the query result data into the data bit of the data entry corresponding to the address entry where the data address of the query result data is located in the data entry queue, and write the valid signal into the valid signal bit of the data entry corresponding to the address entry where the data address of the query result data is located;

[0226] The step of, through the data selector, when the data address in the address entry matches the data address of the backfill data, extract the data corresponding to the data address in the address entry from the backfill data and store the extracted data into the data bit of the data entry corresponding to the address entry, includes:

[0227] Through the data selector, when the data address in the address entry matches the data address of the backfill data and the valid signal in the valid signal bit of the data entry corresponding to the address entry indicates invalid, extract the data corresponding to the data address in the address entry from the backfill data input to the data extraction device and store the extracted data into the data bit of the data entry corresponding to the address entry.

[0228] In some embodiments, the data loading queue sub-circuit contains M of the data extraction devices, and the data loading queue sub-circuit further includes: a forwarding splitting unit and a backfill splitting unit;

[0229] The method further includes:

[0230] Through the forwarding splitting unit, disperse and write the forwarded data of every M data addresses into the initialization circuits in the M data extraction devices;

[0231] Through the backfill splitting unit, disperse and write the backfill data of every M data addresses into the M data extraction devices.

[0232] In some embodiments, the data loading queue sub-circuit further contains a first pipelined register, and the data extraction device further contains a second pipelined register;

[0233] The step of, through the main pipeline circuit, disperse and write multiple data addresses into at least two of the data loading queue sub-circuits, includes:

[0234] Through the main pipeline circuit, write the data address corresponding to the memory access instruction into the first register for pipelining processing;

[0235] The method further includes:

[0236] Through the backfill splitting unit, disperse and write the backfill data of consecutive M bytes into the second registers in the M data extraction devices for pipelining processing.

[0237] In some embodiments, the step of, through the output sub-circuit, obtaining and outputting the data queried by at least two of the data loading queue sub-circuits respectively includes:

[0238] Through the output sub-circuit, obtain and output the data queried by at least two of the data loading queue sub-circuits respectively in a polling manner.

[0239] In some embodiments, the step of, through the output sub-circuit, obtaining and outputting the data queried by at least two of the data loading queue sub-circuits respectively in a polling manner includes:

[0240] Through the output sub-circuit, obtain the data queried by at least two of the data loading queue sub-circuits respectively in a polling manner, perform data shifting on the data queried by at least two of the data loading queue sub-circuits respectively, and then output the data;

[0241] Wherein, the displacement amount of the data shifting is the lower three bits of the addresses of the data queried by at least two of the data loading queue sub-circuits respectively.

[0242] In some embodiments, the data buffer includes a miss status processing register; the method further includes:

[0243] Through the data buffer, when the data corresponding to the access request does not exist in the data buffer and the miss status processing register is not full, cache the access request into an entry in the miss status processing register;

[0244] Through the data buffer, send the access request in the miss status processing register to the lower-level cache or memory to query the data corresponding to the access request;

[0245] Through the data buffer, when receiving the data corresponding to the access request returned by the lower-level cache or memory, remove the access request from the miss status processing register.

[0246] In some embodiments, the method further includes:

[0247] Through the data buffer, when the data corresponding to the access request does not exist in the data buffer and the access request already exists in the missing status processing register, the access request is discarded.

[0248] In some embodiments, the method further includes:

[0249] Through the main pipeline circuit, when the data corresponding to the access request does not exist in the data buffer and the missing status processing register is full, a retransmission of the memory access instruction is triggered.

[0250] The solution shown in the above embodiments of the present application can be applied to a processor. Specifically, the present application further provides a processor, which includes at least one instruction processing component as shown in any of the above Figures 3 to 10 any one.

[0251] On the other hand, the present application further provides a computer device, which includes at least one processor, and the processor includes at least one instruction processing component as shown in any of the above Figures 3 to 10 any one.

[0252] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disc, etc.

[0253] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An instruction processing component, characterized in that, The instruction processing component includes: an instruction emission circuit, a main pipeline circuit, a data buffer, and a data load queue circuit; the data buffer is the first-level cache in the processor; The main pipeline circuit is electrically connected to the data buffer, and the main pipeline circuit is electrically connected to the data load queue circuit; The main pipeline circuit is configured to receive a memory access instruction, and send an access request to the data buffer based on the memory access instruction, where the access request is used to request to query the data corresponding to the memory access instruction from the data buffer; The main pipeline circuit is further configured to, when the query of the data corresponding to the memory access instruction from the data buffer fails, write the data address corresponding to the memory access instruction into the data load queue circuit; The data load queue circuit is configured to receive backfill data, query the corresponding data from the backfill data based on the written data address, and output the queried data; the backfill data is the data filled into the data buffer.

2. The instruction processing component according to claim 1, wherein The data load queue circuit includes: at least two data load queue sub-circuits, and an output sub-circuit; the at least two data load queue sub-circuits are respectively electrically connected to the main pipeline circuit; at least two of the data load queue sub-circuits are respectively electrically connected to the output sub-circuit; The main pipeline circuit is configured to disperse and write multiple data addresses into at least two of the data load queue sub-circuits; At least two of the data load queue sub-circuits are configured to execute in parallel the steps of receiving backfill data and querying the corresponding data from the backfill data based on the written data address; The output sub-circuit is configured to obtain the data queried by at least two of the data load queue sub-circuits respectively and output the data.

3. The instruction processing component according to claim 2, characterized in that The data load queue sub-circuit includes: an address entry queue, an address comparison device, a data extraction device, and an output selection circuit; The address entry queue contains multiple address entries, and each address entry is used to store a written data address; The data extraction device includes a data entry queue, a data selector, and a data output device, and the data entry queue contains multiple data entries; The address comparison device is configured to compare the data address in the address entry with the data address of the backfill data, and when the data address in the address entry matches the data address of the backfill data, set the backfill field in the address entry to a specified value; The data selector is configured to, when the data address in the address entry matches the data address of the backfill data, extract the data corresponding to the data address in the address entry from the backfill data, and store the extracted data into the data bit of the data entry corresponding to the address entry; The output selection circuit is configured to select the number of an address entry whose backfill field is the specified value, and transfer the number of the address entry to the data output component; The data outputter is configured to transfer the data stored in the data entry corresponding to the number of the address entry passed by the output selection circuit to the output sub-circuit according to the number of the address entry.

4. The instruction processing component according to claim 3, wherein Each data entry in the data entry queue further includes a valid signal bit; the data extraction device further includes an initialization circuit; The main pipeline circuit is further configured to write the forwarded data into the initialization circuit; the forwarded data is composed of the result queried from the data buffer; the forwarded data includes the query result data and a valid signal, and the valid signal is used to indicate whether the query result data is valid; The initialization circuit is configured to write the query result data into the data bit of the data entry corresponding to the address entry where the data address of the query result data is located in the data entry queue, and write the valid signal into the valid signal bit of the data entry corresponding to the address entry where the data address of the query result data is located; The data selector is configured to extract the data corresponding to the data address in the address entry from the backfill data input to the data extraction device and store the extracted data in the data bit of the data entry corresponding to the address entry when the data address in the address entry matches the data address of the backfill data and the valid signal in the valid signal bit of the data entry corresponding to the address entry indicates invalid.

5. The instruction processing component according to claim 4, wherein The data loading queue sub-circuit includes M of the data extraction devices, and the data loading queue sub-circuit further includes: a forward split unit, and a backfill split unit; The forward split unit is configured to disperse and write the forwarded data of every M data addresses into the initialization circuits in the M data extraction devices; The backfill split unit is configured to disperse and write the backfill data of every M data addresses into the M data extraction devices.

6. The instruction processing component according to claim 5, wherein The data loading queue sub-circuit further includes a first pipelined register, and the data extraction device further includes a second pipelined register; The main pipeline circuit is configured to write the data address corresponding to the memory access instruction into the first pipelined register for pipelining processing; The backfill split unit is configured to disperse and write the backfill data of consecutive M bytes into the second pipelined registers in the M data extraction devices for pipelining processing.

7. The instruction processing component according to claim 2, wherein The output sub-circuit is configured to obtain and output the data queried by at least two of the data loading queue sub-circuits in a polling manner.

8. The instruction processing component according to claim 7, wherein The output sub-circuit is configured to obtain the data queried by at least two of the data loading queue sub-circuits in a polling manner, and perform data shifting on the data queried by at least two of the data loading queue sub-circuits and then output the data; wherein the displacement amount of the data shift is the lower three bits of the addresses of the data queried by at least two of the data loading queue sub-circuits.

9. The instruction processing component according to any one of claims 1 to 8, characterized in that, The data buffer includes a miss status handling register; The data buffer is configured to cache the access request into an entry in the miss status handling register when the data corresponding to the access request does not exist in the data buffer and the miss status handling register is not full; The data buffer is further configured to send the access request in the miss status handling register to a lower-level cache or memory to query the data corresponding to the access request; The data buffer is further configured to remove the access request from the miss status handling register when receiving the data corresponding to the access request returned by the lower-level cache or memory; 10. The instruction processing component according to claim 9, wherein The data buffer is configured to discard the access request when the data corresponding to the access request does not exist in the data buffer and the access request already exists in the miss status handling register; 11. The instruction processing component according to claim 9, wherein The main pipeline circuit is further configured to trigger retransmission of the memory access instruction when the data corresponding to the access request does not exist in the data buffer and the miss status handling register is full; 12. An instruction processing method, characterized in that, The method is executed by an instruction processing component, and the instruction processing component is the instruction processing component according to any one of claims 1 to 11. The method includes: Receiving a memory access instruction through the main pipeline circuit, and sending an access request to the data buffer based on the memory access instruction, where the access request is used to request to query the data corresponding to the memory access instruction from the data buffer; Writing the data address corresponding to the memory access instruction into the data load queue circuit through the main pipeline circuit when failing to query the data corresponding to the memory access instruction from the data buffer; Receiving backfill data through the data load queue circuit, querying the corresponding data from the backfill data based on the written data address, and outputting the queried data; the backfill data is the data filled into the data buffer; 13. The method according to claim 12, wherein The data load queue circuit includes: at least two data load queue sub-circuits and an output sub-circuit; the at least two data load queue sub-circuits are electrically connected to the main pipeline circuit respectively; the at least two data load queue sub-circuits are electrically connected to the output sub-circuit respectively; The writing the data address corresponding to the memory access instruction into the data load queue circuit includes: Dispersedly writing a plurality of data addresses into at least two of the data load queue sub-circuits through the main pipeline circuit; The receiving backfill data through the data load queue circuit, querying the corresponding data from the backfill data based on the written data address, and outputting the queried data includes: Parallelly executing the steps of receiving backfill data and querying the corresponding data from the backfill data based on the written data address through at least two of the data load queue sub-circuits; Through the output sub-circuit, data queried by each of at least two of the data loading queue sub-circuits is obtained and output.

14. A processor, characterized in that, The processor includes at least one instruction processing component as described in any one of claims 1 to 11.

15. A computer device, characterized in that, The computer device includes at least one processor, and the processor includes at least one instruction processing component as described in any one of claims 1 to 11.

Citation Information

Cited By

  • Instruction processing device, system and method, processor and electronic equipment

    CN122431730A