Processor core, instruction processing method, electronic device, and storage medium

By designing multiple finger fetching decoding pipelines on the front end of the processor core and combining memory interleaving and multi-port technology, the problem of insufficient bandwidth at the front end of the processor core is solved, and higher processor performance and multi-threading performance are achieved.

CN119847608BActive Publication Date: 2025-07-22HYGON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411919034.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-07-22
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The existing processor core front-end has problems of latency and bandwidth deficiency during instruction decoding. Especially when using micro instruction cache, it is impossible to fully utilize the advantages of multiple pipelines, resulting in insufficiency of processor performance and power consumption.

Method used

Multiple finger-fetch decoding pipelines are designed. Each pipeline can work in instruction decoding mode and micro-instruction cache mode. The cache access is optimized through memory interleaving and multi-port technology to ensure that there is a backup pipeline to undertake prediction information when the pipeline is blocked in any mode, and improve front-end bandwidth and multi-threading performance.

Benefits of technology

It improves the front-end bandwidth and multi-threading performance of the processor core, improves the overall performance and power consumption efficiency of the processor, makes full use of the advantages of multiple pipelines, and enhances the utilization of branch prediction bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847608B_ABST
    Figure CN119847608B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a processor core including a front end, an instruction processing method, device, and medium. The front end of the processor core includes N fetch and decode pipelines, at least one instruction cache, and at least one micro-instruction cache, where N is an integer greater than or equal to 2. Each fetch and decode pipeline includes fetch selection logic, an instruction cache port, a micro-instruction cache port, a decoder, and a first micro-instruction queue. Prediction information enters the first micro-instruction queue through processing in the instruction decode pipeline or the micro-instruction pipeline among the N fetch and decode pipelines for distribution. The processor core has a larger average front-end bandwidth and improved processing performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a processor core, an instruction processing method, an electronic device, and a computer-readable storage medium. Background Art

[0002] The front-end of a processor core (CPU core front-end) generally includes branch prediction (BP, branch prediction), instruction fetch (IF, instruction fetch), instruction decoding (DE, decoder), and instruction dispatch (DI, dispatch) modules. In some high-performance processors, the instruction decoding module further includes a micro-instruction cache (u-op cache, i.e., micro-op cache). The micro-instruction cache stores the micro-instructions obtained by decoding instructions. When the addresses of these micro-instructions are accessed for the second time, the processor directly accesses the micro-instruction cache and no longer accesses the instruction cache or decodes the instructions. In some processor architectures, instruction decoding increases latency, reduces bandwidth, and increases power consumption. If a micro-instruction cache is adopted, in a large number of programs, these negative effects brought by decoding will be greatly reduced. The selection logic for accessing the instruction cache or the micro-instruction cache is between the branch prediction stage and the instruction fetch stage. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a processor core, including a front-end, where the front-end includes N instruction fetch and decoding pipelines, at least one instruction cache, and at least one micro-instruction cache, and N is an integer greater than or equal to 2; each of the instruction fetch and decoding pipelines includes instruction fetch selection logic, an instruction cache port, a micro-instruction cache port, a decoder, and a first micro-instruction queue; the instruction cache port is configured to read and write the instruction cache and provide the instruction obtained from the instruction cache to the decoder; the micro-instruction cache port is configured to read and write the micro-instruction cache and provide a first micro-instruction obtained from the micro-instruction cache to the first micro-instruction queue; the decoder is configured to decode the instruction obtained from the instruction cache port into a second micro-instruction and provide the decoded second micro-instruction to the first micro-instruction queue; the instruction fetch selection logic is configured to select the instruction cache port and the decoder to perform instruction fetch and decoding operations, or select the micro-instruction cache port to perform instruction fetch and decoding operations.

[0004] For example, in the processor core provided by at least one embodiment of the present disclosure, the front-end further includes: an instruction dispatch unit configured to dispatch the micro-instructions received from the first micro-instruction queue of one of the N instruction fetch and decoding pipelines for execution.

[0005] For example, the processor core provided by at least one embodiment of the present disclosure, wherein the instruction cache is configured to be coupled to the instruction cache ports of the N fetch-decode pipelines through memory interleaving technology or multi-port technology, so that the N fetch-decode pipelines can access the instruction cache; and / or the micro-instruction cache is configured to be coupled to the micro-instruction cache ports of the N fetch-decode pipelines through memory interleaving technology or multi-port technology, so that the N fetch-decode pipelines can access the micro-instruction cache.

[0006] For example, the processor core provided by at least one embodiment of the present disclosure, wherein the front end further includes: a branch prediction unit configured to generate a prediction result based on the received instruction address for use by the N fetch-decode pipelines.

[0007] For example, the processor core provided by at least one embodiment of the present disclosure, wherein the front end further includes: a boundary information determination unit configured to generate boundary information based on the prediction result and generate at least one information stream based on the prediction result; wherein the type of the boundary information includes the last byte of a jump branch instruction, the last byte of the prediction result, or the last byte of an intermediate instruction in the prediction result.

[0008] For example, the processor core provided by at least one embodiment of the present disclosure, wherein the front end further includes: a window selection unit configured to allocate the at least one information stream to the N fetch-decode pipelines for processing the at least one information stream, wherein the allocation strategy of the window selection unit includes any one of the number of prediction results, the number of bytes of the prediction result, the blocking degree of the N fetch-decode pipelines, and the window switching frequency in different modes.

[0009] For example, the processor core provided by at least one embodiment of the present disclosure, wherein the front end further includes: a reordering logic configured to reorder the micro-instructions obtained by the N fetch-decode pipelines processing the at least one information stream and then send them to the instruction distribution unit.

[0010] For example, the processor core provided by at least one embodiment of the present disclosure, wherein the front end further includes: a first arbitration logic and a second arbitration logic, wherein the prediction result includes prediction results for M threads, where M is an integer greater than or equal to 2; the first arbitration logic is configured to allocate the prediction results among the M threads to the N fetch-decode pipelines for processing; and the second arbitration logic is configured to sequentially select and send the micro-instructions corresponding to the M threads obtained after being processed by the N fetch-decode pipelines to the instruction distribution unit.

[0011] For example, the processor core provided by at least one embodiment of the present disclosure, wherein each of the instruction fetch and decode pipelines further includes a second micro-instruction queue, and the second micro-instruction queue is configured to receive micro-instructions from the decoder and / or the micro-instruction cache port in each of the instruction fetch and decode pipelines.

[0012] At least one embodiment of the present disclosure further provides an instruction processing method, including: in response to an instruction processing request, selecting a target instruction fetch and decode pipeline among N instruction fetch and decode pipelines included in the front end of a processor core to respond to the instruction processing request, where N is an integer greater than or equal to 2; in the target instruction fetch and decode pipeline, selecting to perform instruction fetch and decode operations through an instruction cache port and a decoder, or selecting to perform instruction fetch and decode operations through a micro-instruction cache port to respond to the instruction processing request, where, in response to selecting the micro-instruction cache port to perform instruction fetch and decode operations to respond to the instruction processing request, includes: obtaining a first micro-instruction from a micro-instruction cache through the micro-instruction cache port and providing the first micro-instruction to a first micro-instruction queue, or in response to selecting to perform instruction fetch and decode operations through the instruction cache port and the decoder to respond to the instruction processing request, includes: obtaining an instruction from an instruction cache through the instruction cache port and providing the instruction to the decoder, and the decoder decodes the instruction into a second micro-instruction and provides the second micro-instruction to the first micro-instruction queue.

[0013] For example, the instruction processing method provided by at least one embodiment of the present disclosure further includes: distributing the micro-instructions received from the first micro-instruction queue of one of the target instruction fetch and decode pipelines among the N instruction fetch and decode pipelines for execution.

[0014] For example, the instruction processing method provided by at least one embodiment of the present disclosure, wherein the instruction cache is coupled to the instruction cache port of the target instruction fetch and decode pipeline among the N instruction fetch and decode pipelines through a memory interleaving technique or a multi-port technique, so that the target instruction fetch and decode pipeline among the N instruction fetch and decode pipelines can access the instruction cache; and / or the micro-instruction cache is coupled to the micro-instruction cache port of the target instruction fetch and decode pipeline among the N instruction fetch and decode pipelines through a memory interleaving technique or a multi-port technique, so that the target instruction fetch and decode pipeline among the N instruction fetch and decode pipelines can access the micro-instruction cache.

[0015] For example, the instruction processing method provided by at least one embodiment of the present disclosure further includes: generating a prediction result based on the received instruction address to generate the instruction processing request for the target instruction fetch and decode pipeline among the N instruction fetch and decode pipelines.

[0016] For example, the instruction processing method provided by at least one embodiment of the present disclosure further includes: generating boundary information according to the prediction result corresponding to the instruction processing request, and generating at least one information stream according to the prediction result, where the type of the boundary information includes the last byte of a jump branch instruction, the last byte of the prediction result, or the last byte of an intermediate instruction in the prediction result.

[0017] For example, the instruction processing method provided by at least one embodiment of the present disclosure further includes: allocating the at least one information stream to a target fetch and decode pipeline among the N fetch and decode pipelines through window selection logic to process the at least one information stream, where the window selection logic includes any one of the number of prediction results, the number of bytes of the prediction result, the blocking degree of the target fetch and decode pipeline among the N fetch and decode pipelines, and the window switching frequency in different modes.

[0018] For example, the instruction processing method provided by at least one embodiment of the present disclosure further includes: reordering the microinstructions obtained by processing the at least one information stream by the target fetch and decode pipeline among the N fetch and decode pipelines for distribution.

[0019] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, the prediction result includes prediction results for M threads, where M is an integer greater than or equal to 2, and the method further includes: allocating the prediction result corresponding to the instruction processing request among the M threads to the target fetch and decode pipeline among the N fetch and decode pipelines for processing; and sequentially selecting the microinstructions corresponding to the M threads obtained by processing by the target fetch and decode pipeline among the N fetch and decode pipelines for distribution.

[0020] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, the responding to the selection by selecting the microinstruction cache port to perform a fetch and decode operation in response to the instruction processing request includes: obtaining a third microinstruction from the microinstruction cache through the microinstruction cache port and providing the third microinstruction to a second microinstruction queue; and the responding to the selection by the instruction cache port and the decoder to perform a fetch and decode operation in response to the instruction processing request includes: obtaining an instruction from the instruction cache through the instruction cache port and providing the instruction to the decoder, and the decoder decoding the instruction into a fourth microinstruction and providing the fourth microinstruction to the second microinstruction queue.

[0021] At least one embodiment of the present disclosure further provides an electronic device including a processor core as described in any of the above embodiments.

[0022] At least one embodiment of the present disclosure further provides an electronic device, including: a processor; and a memory including one or more computer program instructions; wherein, when the one or more computer program instructions are run by the processor, they execute the instruction processing method provided in any of the above embodiments.

[0023] At least one embodiment of the present disclosure provides a computer-readable storage medium that non-temporarily stores computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the instruction processing method provided in any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0025] Figure 1 Shows a schematic diagram of a pipeline of a processor core;

[0026] Figure 2 Shows a schematic diagram of the basic structure division of a processor;

[0027] Figure 3 Shows a schematic diagram of an example of the front-end composition of a CPU core;

[0028] Figure 4 Shows a schematic diagram of a branch prediction module;

[0029] Figure 5 Shows a schematic diagram of the operation of decoding instructions in a processor core;

[0030] Figure 6 Shows a schematic diagram of the structure of a processor core according to at least one embodiment of the present disclosure;

[0031] Figure 7 Shows a schematic flowchart of the instruction processing method according to at least one embodiment of the present disclosure;

[0032] Figure 8 Shows a specific flowchart of the instruction processing method according to at least one embodiment of the present disclosure;

[0033] Figure 9 Shows an example of the structure of a processor core according to at least one embodiment of the present disclosure;

[0034] Figure 10 Shows an example of the structure of a processor core according to at least another embodiment of the present disclosure;

[0035] Figure 11 Shows an example of the structure of a processor core according to at least yet another embodiment of the present disclosure;

[0036] Figure 12 A schematic block diagram of an electronic device according to at least one embodiment of the present disclosure is shown;

[0037] Figure 13 A schematic structural diagram of a computer-readable storage medium according to at least one embodiment of the present disclosure is shown; and

[0038] Figure 14 A schematic block diagram of an electronic device according to at least another embodiment of the present disclosure is shown. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0040] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, words such as "including" or "comprising" mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items. "Connection" or "coupling" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0041] The processor core (CPU core) of a single-core processor or a multi-core processor improves the instruction execution efficiency through pipelining technology. The pipelining technology divides the complete operation steps of the CPU core into multiple sub-steps and executes these sub-steps in the form of a pipeline to improve efficiency.

[0042] Figure 1 An instruction pipeline of an exemplary scalar central processing unit (CPU) is shown, and the instruction pipeline includes a five-stage pipeline. As Figure 1As shown, each instruction can be issued in each clock cycle and executed within a fixed time (e.g., 5 clock cycles). The execution of each instruction is divided into 5 steps: the instruction fetch (IF) stage 1001, the instruction decode (ID) stage 1002, the execution (EX) stage 1003, the memory access (MEM) stage 1004, and the write-back (WB) stage 1005. In the IF stage 1001, the specified instruction is fetched from the instruction cache. A part of the fetched specified instruction is used to specify the source registers available for executing the instruction. In the ID stage 1002, the instruction is decoded and control logic is generated to fetch the contents of the specified source registers. According to the control logic, arithmetic or logical operations are performed in the EX stage 1003 using the fetched contents. In the MEM stage 1004, the instruction reads / writes data in the memory cache. Finally, in the WB stage 1005, the value obtained by executing the instruction can be written back to a certain register.

[0043] Generally, the earlier pipeline steps in the CPU are divided into the front end. For example, taking a five-stage pipeline as an example, the modules for the instruction fetch stage and the instruction decode stage are classified as the CPU front end; correspondingly, the later pipeline steps in the CPU are divided into the back end. For example, taking a five-stage pipeline as an example, the modules for the execution stage, the memory access stage, and the write-back stage are classified as the CPU back end.

[0044] To support a high operating frequency, each pipeline stage may further contain multiple (sub) pipeline levels (clock cycles). Although each pipeline level performs limited operations, in this way, each clock can be made the shortest, and the performance of the CPU core is improved by increasing the operating frequency of the CPU. Each pipeline level can also further improve the performance of the processor core by accommodating more instructions (i.e., the superscalar technology). Superscalar refers to a method of parallelly executing multiple instructions in one cycle. A processor with increased instruction-level parallelism that can process multiple instructions in one cycle is called a superscalar processor. For example, a superscalar processor can further support out-of-order execution. Out-of-order execution refers to a technology in which the CPU allows multiple instructions to be sent to the corresponding circuit units for processing in an order different from the order specified by the program.

[0045] The processor core translates each architectural instruction (abbreviated as "instruction") in the microarchitecture into one or more microinstructions. Each microinstruction performs only limited operations, which can ensure that each pipeline stage is very short to increase the operating frequency of the processor core. For example, a memory read instruction (load) can be translated into an address generation microinstruction and a memory read microinstruction. The second microinstruction depends on the result of the first microinstruction. Therefore, the second microinstruction will not start execution until the first microinstruction has been executed. A microinstruction contains multiple microarchitecture-related fields used to transfer relevant information between pipeline stages.

[0046] Speculative Execution is another technique to improve processor performance. This technique executes the instructions following an instruction before it has completed execution. One technique for speculative execution is branch prediction. As mentioned above, the instruction fetch unit is responsible for providing the processor with the next instruction to be executed. During the instruction fetch stage, in addition to fetching multiple instructions, it is also necessary to determine the instruction fetch address for the next cycle. Therefore, at this stage, it is necessary to determine whether there is a conditional branch instruction, and if there is a branch, whether to jump (direction) and the target address. The instruction fetch unit includes a branch prediction unit (branch predictor) to perform branch prediction. The branch prediction unit (branch predictor) at the front end of the processor core predicts the jump direction of the conditional branch instruction, prefetches, and executes the instructions in that direction. Another technique for speculative execution is to execute a memory read instruction before the addresses of all the previous memory write instructions have been obtained.

[0047] Speculative execution further improves the instruction-level parallelism, thus significantly improving the performance of the processor core. When a speculative execution error occurs, such as a branch prediction error being detected, or a write instruction before a memory read instruction overwrites the same address, all the instructions in all the subsequent pipelines of the faulty instruction need to be flushed (or called "cleared"), and then the program jumps to the error point to be re-executed to ensure the correctness of the program execution.

[0048] Figure 2 A schematic diagram showing the basic structural division of a processor is shown. The processor 200 includes at least one CPU core (processor core) and at least one level of cache. For example, the CPU core includes a front end 201 and a back end 202. For example, the at least one level of cache includes a first-level cache (not shown in the figure) provided inside the CPU core and a second-level cache 203 outside the CPU core. Here, the second-level cache 203 is a separate structure.

[0049] Figure 3A schematic diagram showing an example of the front-end structure of a CPU core is presented. The front-end 201 of this CPU core includes an instruction fetch unit, a decoding unit, and an issue unit. The instruction fetch unit includes a branch predictor 301 and a selection logic 302; the decoding unit includes an instruction cache 303, an instruction decoder 304, a micro-instruction cache 305, and a micro-instruction queue 306; the issue unit 307 is connected to the micro-instruction queue 306. The CPU core has both an instruction cache and a micro-instruction cache, thus having micro-architecture optimization. The instruction address obtained by the instruction fetch unit (which can be the program counter address, abbreviated as the PC address) is predicted by the branch predictor 301 to obtain the instruction address to be executed next. At the same time, the instruction address passes through the selection logic 302 to determine whether the instruction corresponding to this instruction address needs to be decoded. If "yes", it follows Figure 3 the left path in Figure 3 and the instruction needs to be decoded; if "no", it follows

[0050] Figure 4 the right path in Figure 3 and the instruction does not need to be decoded, but instead accesses the micro-instruction cache to obtain the corresponding micro-instruction group data.

[0051] Figure 5 A schematic diagram showing the operation of decoding an instruction in a processor core is presented. In conjunction with Figure 3, according to the instruction address A, the instruction cache 303 is queried to obtain the undecoded instruction data (such as one or more instructions) corresponding to the instruction address A, and then the instruction data is decoded into multiple microinstructions through the instruction decoder 304 (these microinstructions can be called a microinstruction group, for example, including microinstruction 1, microinstruction 2, microinstruction 3, etc.). On the one hand, the obtained microinstruction group is sent to the microinstruction queue 306 to wait for the issue unit 307 to allocate and issue it to the corresponding execution units in the backend of the CPU core for execution. On the other hand, under certain conditions (such as the microinstruction group is a commonly used group of microinstructions), it is saved to the microinstruction cache 305 and waits for possible subsequent access.

[0052] In one embodiment, the processor core can update the microtag cache of the microinstructions in other functional modules 3012 while saving the microinstruction group to the microinstruction cache 305.

[0053] The inventors of the present disclosure have noticed that there is only one pipeline including branch prediction, instruction fetching, decoding, and microinstruction cache in the above-mentioned processor. With the development trend of increasing the instruction dispatch bandwidth (dispatching more instructions per unit time), memory access bandwidth, and the number of various execution units, the front end of the processor core gradually fails to meet the requirement of the backend of the processor core to execute more instructions. Therefore, how to increase the front-end bandwidth has become a problem to be solved by mainstream processors.

[0054] Simultaneous Multithreading (SMT) technology is an important technology to improve the overall performance of the CPU. It uses mechanisms such as multi-issue and out-of-order execution of high-performance CPU cores to simultaneously execute the instructions of multiple threads, so that a physical CPU core presents multiple virtual CPU cores to the software and operating system. When a modern multi-issue high-performance CPU core executes a single thread, most of the time, its internal multiple execution units and hardware resources cannot be fully utilized. When the thread stalls due to certain reasons (such as an L2 cache miss), the hardware execution units can only idle, which causes waste of hardware resources and reduces the performance power ratio. In the SMT mode, for example, the dual-thread mode (i.e., SMT2), when one thread stalls, other threads can still run, which improves the utilization rate of hardware resources, thereby improving the multi-thread throughput, overall performance, and performance power ratio of the CPU core. It should be noted that since it needs to share CPU core resources with other threads, the performance of a thread running under SMT is often lower than its performance in the single-thread mode.

[0055] In SMT mode, at least one pipeline stage of the processor core is time-division multiplexed. That is, if a thread occupies a pipeline stage of a pipeline at a certain moment, other threads cannot occupy the pipeline stage of the pipeline at this moment. However, if there are two pipelines, each thread can occupy one pipeline exclusively. In this way, the overall performance of SMT will be greatly improved.

[0056] The inventors of the present disclosure also noticed that for a multi-threaded processor core, it can only allow multiple threads to use partially expanded multiple pipelines in one of the modes: instruction decoding mode or microinstruction mode, and cannot fully and effectively utilize the advantages of multiple pipelines.

[0057] At least one embodiment of the present disclosure provides a processor core, an instruction processing method, an electronic device, and a storage medium. The processor core of the embodiment of the present disclosure includes a front end, which includes N instruction fetch decoding pipelines, at least one instruction cache and at least one microinstruction cache, where N is an integer greater than or equal to 2. Each of the above-mentioned instruction fetch decoding pipelines includes instruction fetch selection logic, an instruction cache port, a microinstruction cache port, a decoder, and a first microinstruction queue. The instruction cache port is configured to read and write instruction caches, and provides instructions obtained from the instruction cache to the decoder; the microinstruction cache port is configured to read and write microinstruction caches, and provides the first microinstruction obtained from the microinstruction cache to the first microinstruction queue; the decoder is configured to decode the instruction obtained from the instruction cache port into a second microinstruction and provide the decoded second microinstruction to the first microinstruction queue; the instruction fetch selection logic is configured to select the instruction cache port and the decoder to perform an instruction fetch decoding operation, or select the microinstruction cache port to perform an instruction fetch decoding operation. The first microinstruction and the second microinstruction here are respectively used to refer to the microinstructions described in the two processing methods.

[0058] At least one embodiment of the present disclosure further provides an instruction processing method corresponding to the above-mentioned processor core.

[0059] The processor core provided by at least one embodiment of the present disclosure can, on the one hand, be designed with multiple instruction fetch and decode pipelines, each of which can operate in instruction decoding mode and microinstruction cache mode. No matter one pipeline in either instruction decoding mode or microinstruction cache mode is blocked, another pipeline can be used to take over the prediction information, so that more microinstructions can be accumulated in the microinstruction queue for execution by the processor core backend, providing a larger processor front-end bandwidth and improving the performance of the processor. On the other hand, the processor core provided by at least one embodiment of the present disclosure can allow multiple threads to use multiple pipelines regardless of whether it is in instruction decoding mode or microinstruction cache mode, and each thread can generate a larger average front-end bandwidth, thereby greatly improving the multi-threaded performance.

[0060] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0061] Figure 6 The structural schematic diagram of the processor core of at least one embodiment of the present disclosure is shown.

[0062] As Figure 6 shown, the processor core 600 includes a front end 601, and the front end 601 includes N instruction fetch and decode pipelines (for example, N = 2, that is, the instruction fetch and decode pipeline 1 and the instruction fetch and decode pipeline 2 as shown in the figure), at least one instruction cache 6031 - 603n and at least one micro-instruction cache 6041 - 604n. It should be noted that the number of instruction fetch and decode pipelines can be determined based on the specific design performance of the front end of the processor core, and the present disclosure does not limit this.

[0063] The instruction fetch and decode pipeline 1 will be taken as an example for illustration below. The instruction fetch and decode pipeline 1 includes a selection logic 6021, an instruction cache port 6051, a micro-instruction cache port 6061, an instruction decoder 6071, and a first micro-instruction queue 6081. For example, the instruction cache port 6051, the instruction cache (such as the instruction cache 6031), and the instruction decoder 6071 constitute an instruction decode pipeline (for the instruction decode mode); for example, the micro-instruction cache port 6061 and the micro-instruction cache (such as the micro-instruction cache 6041) constitute a micro-instruction pipeline (for the micro-instruction cache mode). The internal composition structures of the remaining N - 1 instruction fetch and decode pipelines are the same as that of the instruction fetch and decode pipeline 1, and will not be elaborated here.

[0064] In the above embodiment, for example, there is only one instruction cache (such as the instruction cache 6031) and one micro-instruction cache (such as the micro-instruction cache 6041), and thus the N instruction fetch and decode pipelines share the instruction cache and the micro-instruction cache.

[0065] Exemplarily, unless otherwise specified, all the following descriptions are introduced based on the instruction fetch and decode pipeline 1.

[0066] For example, in a possible implementation manner, the above front end 601 further includes: an instruction distribution unit, and the instruction distribution unit is configured to distribute the micro-instructions received from the first micro-instruction queue of one of the N instruction fetch and decode pipelines for execution.

[0067] For example, the instruction distribution unit (not shown in the figure) is configured to be after the first micro-instruction queue 6081, receive the micro-instructions processed in the instruction fetch and decode pipeline 1 and distribute them to the subsequent execution units for execution.

[0068] For example, an instruction cache (such as instruction cache 6031) is coupled to instruction cache ports 6051 - 605n through Bank-Interleaving technology, so that the instruction cache ports 6051 - 605n can access the instruction cache 6031; a micro-instruction cache (such as micro-instruction cache 6041) is coupled to micro-instruction cache ports 6061 - 606n through Bank-Interleaving technology, so that the micro-instruction cache ports 6061 - 606n can access the instruction cache 6041.

[0069] The Bank-Interleaving technology divides memory modules into multiple independent channels and alternately allocates data blocks to different channels, enabling multiple channels to be operated on simultaneously when accessing memory, thereby reducing waiting time and increasing data transfer rate.

[0070] Also for example, an instruction cache (such as instruction cache 6031) is coupled to instruction cache ports 6051 - 605n through Multiport technology, so that the instruction cache ports 6051 - 605n can access the instruction cache 6031; a micro-instruction cache (such as micro-instruction cache 6041) is coupled to micro-instruction cache ports 6061 - 606n through Multiport technology, so that the micro-instruction cache ports 6061 - 606n can access the instruction cache 6041.

[0071] Each port among the multiple ports provided by the Multiport technology can perform read and write operations independently, which allows multiple requests to be processed simultaneously, thereby reducing waiting time and improving the overall performance of the system.

[0072] The above-mentioned instruction cache 6031 and micro-instruction cache 6041 are only exemplary, and the present disclosure is not limited thereto. For example, if two ports (such as instruction cache port 6051 and instruction cache port 6052) access simultaneously and encounter an address conflict, the access of one of them (such as instruction cache port 6052) can be delayed to wait until the access of the non-delayed one ends.

[0073] For example, in a possible implementation, the above-mentioned front end 601 further includes a branch prediction unit, which is configured to generate a prediction result based on the received instruction address for use in the above-mentioned N fetch-decode pipelines.

[0074] For example, a branch prediction unit (not shown in the figure) is configured before the instruction fetch and decode pipeline 1 to the instruction fetch and decode pipeline N. The branch prediction unit is used in the processor to optimize the program execution flow, especially the performance when processing conditional branch statements. For example, the branch prediction unit can generate a prediction result based on the instruction address of the instruction (such as a jump instruction) entering the instruction pipeline for the above-mentioned instruction fetch and decode pipeline 1 to the instruction fetch and decode pipeline N. For example, the branch prediction method of the branch prediction unit can include any one or several of static branch prediction, dynamic branch prediction, and indirect branch prediction.

[0075] For example, in a possible implementation, the above-mentioned front end 601 further includes a boundary information determination unit, which is configured to generate boundary information according to the above-mentioned prediction result and generate at least one information stream according to the above-mentioned prediction result. For example, according to the type of the branch prediction unit, the type of the boundary information includes the last byte of the jump branch instruction, the last byte of the above-mentioned prediction result, or the last byte of the intermediate instruction in the above-mentioned prediction result. For example, in the case where a type of branch prediction unit hits the branch target buffer (BTB), even if the branch instruction does not jump, the prediction information will be forced to end at the last byte of a certain branch instruction, rather than at the last byte of the cache basic block (or prediction window), which will result in "the last byte of the prediction result".

[0076] For example, the boundary information determination unit (not shown in the figure) is configured after the above-mentioned branch prediction unit and before the instruction fetch and decode pipeline. For example, the number of generated information streams is determined based on the number of boundary information, and the present disclosure does not limit this.

[0077] For example, the boundary information can be stored in a new cache structure or an existing cache structure. When the boundary information is the last byte of the intermediate instruction in the above-mentioned prediction result, it is decoded first and then stored in any of the above-mentioned cache structures.

[0078] For example, in a possible implementation, the above-mentioned front end 601 further includes a window selection unit, which is configured to allocate the above-mentioned at least one information stream to the above-mentioned N instruction fetch and decode pipelines to process the at least one information stream. For example, the allocation strategy of the window selection unit includes any one of the number of the above-mentioned prediction results, the number of bytes of the above-mentioned prediction result, the blocking degree of the above-mentioned N instruction fetch and decode pipelines, and the window switching frequency in different modes.

[0079] For example, a window selection unit (not shown in the figure) is configured after the above-mentioned boundary information determination unit and before the instruction fetch and decode pipeline. For example, at least one information stream (e.g., 3) is sent to the instruction fetch and decode pipelines 1 and 2 through the window selection unit. For example, the at least one information stream may be sent to one pipeline or different pipelines. For example, 2 information streams are sent to the instruction fetch and decode pipeline 1, and the remaining 1 is sent to the instruction fetch and decode pipeline 2. The specific number of allocated streams can be dynamically adjusted based on the execution status (e.g., the degree of blockage) of the instruction fetch and decode pipelines from 1 to n. The present disclosure does not limit this.

[0080] For example, in a possible implementation, the front end 601 further includes a reordering logic, which is configured to reorder the microinstructions obtained by processing the at least one information stream by the above-mentioned N instruction fetch and decode pipelines and then send them to the above-mentioned instruction distribution unit.

[0081] For example, the reordering logic (not shown in the figure) is configured after the instruction fetch and decode pipeline and before the instruction distribution unit. For example, since the instruction fetch and decode pipelines from 1 to n (e.g., the instruction fetch and decode pipelines 1 and 2) receive the prediction results corresponding to a single thread, the instruction data after being processed by the boundary information determination unit and distributed by the window selection unit is in an out-of-order state, that is, the information streams processed in the instruction fetch and decode pipelines 1 and 2 are in an out-of-order state. For example, based on the first-in-first-out principle (FIFO), it is necessary to reorder the microinstructions accumulated in the first microinstruction queues 6081 to 608n obtained by processing in the instruction fetch and decode pipeline 1 in the reordering logic to restore their order before entering the above-mentioned instruction fetch and decode pipeline, so that the microinstructions obtained by processing the instructions with a higher order in the information stream can be distributed in a timely manner, which helps to efficiently and correctly process the instruction sequence.

[0082] It should be noted that the above-mentioned boundary information unit, window selection unit, and reordering logic mentioned in the above embodiments are applicable to the single-thread mode.

[0083] The processor core provided by at least one embodiment of the present disclosure is applicable to the single-thread mode. Multiple instruction fetch and decode pipelines are designed. Each instruction fetch and decode pipeline can work in the instruction decoding mode and the microinstruction cache mode. No matter which one of the instruction decoding mode or the microinstruction cache mode has a pipeline blocked, another pipeline can be used to receive the prediction information, so that more microinstructions can be accumulated in the microinstruction queue for the execution of the backend of the processor core, providing a larger processor front-end bandwidth and improving the performance of the processor.

[0084] For example, in a possible implementation, the above-mentioned front end 601 further includes a first arbitration logic and a second arbitration logic. Among them, the above-mentioned prediction results include prediction results for M hardware threads (hereinafter referred to as "threads"), where M is an integer greater than or equal to 2 (for example, M = 2, 4, or 8, etc.); the first arbitration logic is configured to allocate the prediction results among the M threads to the above-mentioned N fetch-decode pipelines for processing; the second arbitration logic is configured to sequentially select and send the microinstructions corresponding to the M threads obtained after being processed by the above-mentioned N fetch-decode pipelines to the instruction distribution unit.

[0085] For example, the first arbitration logic (not shown in the figure) is configured after the branch prediction unit and before the fetch-decode pipeline. The second arbitration logic (not shown in the figure) is configured after the fetch-decode pipeline and before the instruction distribution unit. Since this embodiment is applicable to the multi-thread mode, it does not include a boundary information determination unit, a window selection unit, and a reordering logic.

[0086] For example, the above-mentioned prediction results include prediction results for two (M = 2) threads (for example, thread 1 and thread 2). The first arbitration logic allocates the prediction results of the two threads to N fetch-decode pipelines (for example, two, namely fetch-decode pipeline 1 and fetch-decode pipeline 2) for processing. For example, the prediction result of thread 1 can be allocated to fetch-decode pipeline 1 or fetch-decode pipeline 2, and the specific allocation rule is not limited in this disclosure. For example, the second arbitration logic sequentially selects and sends the microinstructions corresponding to the two threads obtained after being processed by fetch-decode pipeline 1 and fetch-decode pipeline 2 to the instruction distribution unit.

[0087] For example, in a possible implementation, each of the above-mentioned N fetch-decode pipelines further includes a second microinstruction queue, and the second microinstruction queue is configured to receive the microinstructions from the instruction decoder and / or the microinstruction cache port in each of the above-mentioned fetch-decode pipelines.

[0088] For example, in a possible implementation, the second microinstruction queue (not shown in the figure) is juxtaposed (and equivalent) with the first microinstruction queue (for example, the first microinstruction queue 6081), and is configured after the instruction decoder 6071 and the microinstruction cache port 6061 and before the above-mentioned second arbitration logic.

[0089] Exemplarily, the above prediction results include prediction results for four (M = 4) threads (e.g., thread 1 to thread 4), and the first arbitration logic distributes the prediction results of these four threads to N fetch-decode pipelines (e.g., two, namely fetch-decode pipeline 1 and fetch-decode pipeline 2). For example, the prediction results from thread 1 and thread 2 are distributed to fetch-decode pipeline 1, and the prediction results from thread 3 and thread 4 are distributed to fetch-decode pipeline 2 for processing. It should be noted that the above result of thread allocation is only exemplary, and the present disclosure places no restrictions thereon.

[0090] Taking the above allocation result as an example for illustration, for example, in fetch-decode pipeline 1, the microinstructions obtained by processing thread 1 through the instruction decode pipeline or the microinstruction pipeline can be sent to the first microinstruction queue 6081, and the microinstructions obtained by processing thread 2 through the instruction decode pipeline or the microinstruction pipeline can be sent to the second microinstruction queue.

[0091] For example, the microinstructions obtained by processing thread 1 through the instruction decode pipeline or the microinstruction pipeline can also be sent to the second microinstruction queue, and the microinstructions obtained by processing thread 2 through the instruction decode pipeline or the microinstruction pipeline can be sent to the first microinstruction queue 6081.

[0092] Still for example, part of the microinstructions obtained by processing thread 1 through the instruction decode pipeline or the microinstruction pipeline can be sent to the first microinstruction queue 6081, and the remaining part of the microinstructions obtained by processing thread 1 through the instruction decode pipeline or the microinstruction pipeline can be sent to the second microinstruction queue, and the same applies to thread 2.

[0093] It should be noted that for fetch-decode pipeline 2 that processes thread 3 and thread 4, its way of distributing microinstructions is similar to that of fetch-decode pipeline 1, and will not be elaborated here.

[0094] The above processor core provided by at least one embodiment of the present disclosure is applicable to the multi-threaded mode, and allows multiple threads to use multiple pipelines whether in the instruction decode mode or the microinstruction cache mode. Each thread can generate a larger average front-end bandwidth, thereby greatly improving the multi-threaded performance.

[0095] The processor core of the embodiment of the present disclosure further includes a backend in addition to the above front-end. For example, the backend includes various execution units (such as arithmetic logic unit (ALU), floating-point unit (FPU), load / store unit (LSU), etc.), reorder buffer (ROB), end unit, etc., which will not be elaborated here.

[0096] The processor core of the embodiment of the present disclosure can be based on X86 microarchitecture, ARM microarchitecture, RISC-V microarchitecture, MIPS microarchitecture, etc., and the present disclosure places no restrictions thereon.

[0097] At least one embodiment of the present disclosure further provides an instruction processing method, which includes: in response to an instruction processing request, selecting a target fetch and decode pipeline among N fetch and decode pipelines included in the front end of a processor core to respond to the instruction processing request, where N is an integer greater than or equal to 2; in the above-mentioned target fetch and decode pipeline, selecting to perform fetch and decode operations through an instruction cache port and a decoder, or selecting to perform fetch and decode operations through a micro-instruction cache port to respond to the instruction processing request. Responding to the selection of performing fetch and decode operations through the above-mentioned micro-instruction cache port to respond to the instruction processing request includes: obtaining a first micro-instruction from the micro-instruction cache through the above-mentioned micro-instruction cache port and providing the first micro-instruction to a first micro-instruction queue. Alternatively, responding to the selection of performing fetch and decode operations through the above-mentioned instruction cache port and decoder to respond to the instruction processing request includes: obtaining an instruction from the instruction cache through the above-mentioned instruction cache port and providing the instruction to the above-mentioned decoder, and the decoder decodes the instruction into a second micro-instruction and provides the second micro-instruction to the above-mentioned first micro-instruction queue.

[0098] Figure 7 A schematic flowchart of the instruction processing method according to at least one embodiment of the present disclosure is shown, as Figure 7 shown, the instruction processing method includes step S710 and step S720.

[0099] Step S710: In response to an instruction processing request, select a target fetch and decode pipeline among N fetch and decode pipelines included in the front end of a processor core to respond to the instruction processing request, where N is an integer greater than or equal to 2;

[0100] Step S720: In the target fetch and decode pipeline, select to perform fetch and decode operations through an instruction cache port and a decoder, or select to perform fetch and decode operations through a micro-instruction cache port to respond to the instruction processing request.

[0101] For example, the above instruction processing method corresponds to a processor core as Figure 6 shown.

[0102] For example, in a possible implementation manner, before responding to an instruction processing request, the above instruction processing method includes: generating a prediction result based on the received instruction address for the target fetch and decode pipeline among the N fetch and decode pipelines. For example, the prediction result is issued by a branch prediction unit, and the specific function of the branch prediction unit has been described above and will not be elaborated here.

[0103] For example, in the above instruction processing method, the prediction result includes a prediction result for a single thread or multiple threads (i.e., multiple hardware threads in a multi-threaded processor).

[0104] For example, in a possible implementation, boundary information is generated according to the prediction result, and at least one information stream is generated according to the prediction result. For example, the types of boundary information include the last byte of a jump branch instruction, the last byte of the prediction result, or the last byte of an intermediate instruction in the prediction result. For example, the number of information streams is based on the number of boundary information. For example, if the last byte of an intermediate instruction in the prediction result is regarded as boundary information, it is generated through decoding and then saved in a cache structure. For example, the cache structure can be a newly added cache or an existing cache can be reused, and the present disclosure does not limit this.

[0105] For example, in a possible implementation, at least one information stream is allocated to a target fetch-decode pipeline among N fetch-decode pipelines through window selection logic to process the at least one information stream. For example, the window selection logic includes any one of the number of prediction results, the number of bytes of the prediction result, the blocking degree of the target fetch-decode pipeline among N fetch-decode pipelines, and the window switching frequency in different modes, etc., and the present disclosure does not limit this. For example, the above-mentioned multiple information streams can be sent to one fetch-decode pipeline or to different fetch-decode pipelines.

[0106] For example, in response to selecting to pass through the micro-instruction cache port to perform a fetch-decode operation in response to an instruction processing request, it includes: obtaining a first micro-instruction from the micro-instruction cache through the micro-instruction cache port and providing the first micro-instruction to a first micro-instruction queue.

[0107] For another example, in response to selecting to pass through the instruction cache port and the decoder to perform a fetch-decode operation in response to an instruction processing request, it includes: obtaining an instruction from the instruction cache through the instruction cache port and providing the instruction to the decoder, and the decoder decodes the instruction into a second micro-instruction and provides the second micro-instruction to the first micro-instruction queue.

[0108] As described above, the micro-instruction cache port and the micro-instruction cache constitute a micro-instruction pipeline (i.e., for the micro-instruction cache mode), and the instruction cache port, the instruction decoder, and the instruction cache constitute an instruction decode pipeline (i.e., for the instruction decode mode). For example, the first micro-instruction refers to the micro-instruction obtained through processing by the micro-instruction pipeline; the second micro-instruction refers to the micro-instruction obtained through processing by the instruction decode pipeline.

[0109] For example, in one possible implementation, the instruction cache is coupled to the instruction cache port of the target instruction fetch and decode pipeline in the N instruction fetch and decode pipelines through memory interleaving technology or multi-port technology, so that the target instruction fetch and decode pipeline in the N instruction fetch and decode pipelines accesses the instruction cache; or, the microinstruction cache is coupled to the microinstruction cache port of the target instruction fetch and decode pipeline in the N instruction fetch and decode pipelines through memory interleaving technology or multi-port technology, so that the target instruction fetch and decode pipeline in the N instruction fetch and decode pipelines accesses the microinstruction cache. The present disclosure does not limit the cache technology that can achieve multi-port access, and the specific cache technology is based on the actual design.

[0110] For example, the microinstruction cache port of the target instruction fetch and decode pipeline in N instruction fetch and decode pipelines accesses the same microinstruction cache, and the instruction cache port of the target instruction fetch and decode pipeline in N instruction fetch and decode pipelines accesses the same instruction cache. It should be noted that if multiple ports (for example, two instruction cache ports access the instruction cache) access at the same time and encounter an address conflict, the access is performed in sequence.

[0111] For example, in a possible implementation, the instruction processing method further includes: reordering the microinstructions obtained by processing at least one information stream by the target instruction fetch and decode pipeline among the N instruction fetch and decode pipelines for distribution. For example, since there are two information streams in a certain target instruction fetch and decode pipeline, the running states of the two information streams in their respective instruction fetch and decode pipelines are determined by their respective pipelines, and it may happen that the information stream that enters first exits the pipeline later. Since the order of the information stream needs to be guaranteed at this stage, the microinstructions need to be reordered for distribution.

[0112] For example, in a possible implementation, the instruction processing method further includes: distributing the microinstructions received from the first microinstruction queue of one of the target instruction fetch and decode pipelines among the N instruction fetch and decode pipelines for execution.

[0113] The instruction processing method provided by at least one embodiment of the present disclosure is implemented by designing multiple target instruction fetch and decode pipelines, each of which can operate in instruction decoding mode and microinstruction cache mode. No matter if one pipeline in either instruction decoding mode or microinstruction cache mode is blocked, another pipeline can be used to take over the prediction information, so that more microinstructions can be accumulated in the microinstruction queue for execution by the processor core back end, thereby providing a larger processor front-end bandwidth and improving the performance of the processor.

[0114] For example, in a possible implementation, the prediction result includes prediction results for M threads, where M is an integer greater than or equal to 2. The instruction processing method further includes: allocating the prediction results in the M threads to a target fetch and decode pipeline among N fetch and decode pipelines for processing; and sequentially selecting the microinstructions corresponding to the M threads obtained by processing in the target fetch and decode pipeline among the N fetch and decode pipelines for distribution.

[0115] For example, if M < N (e.g., M = 3, N = 4), then each thread is sent to its respective fetch and decode pipeline (e.g., thread 1 corresponds to pipeline 1, and so on in order), where there are remaining fetch and decode pipelines (e.g., remaining pipeline 4).

[0116] For example, if M = N (e.g., M = N = 4), then each thread is sent to its respective fetch and decode pipeline (e.g., thread 1 corresponds to pipeline 1, and so on in order), where there are no remaining fetch and decode pipelines.

[0117] For example, if M > N (e.g., M = 5, N = 4), then in addition to the threads allocated to each pipeline (e.g., thread 1 corresponds to pipeline 1, and so on in order), the remaining threads (e.g., remaining thread 5) are overallocated to the existing pipelines for processing based on the execution status (e.g., blocking status) of all pipelines.

[0118] It should be noted that the above allocation order of each thread is only exemplary and can be specifically determined based on the actual design. The present disclosure places no restrictions on this.

[0119] For example, the microinstructions corresponding to the M threads (e.g., threads 1 to 4) obtained by processing in the target fetch and decode pipeline (e.g., fetch and decode pipelines 1 - 4) among the N fetch and decode pipelines are sequentially selected for distribution (e.g., fetch and decode pipeline 1 corresponds to thread 1, and so on).

[0120] For example, in a possible implementation, in response to selecting a microinstruction cache port to perform a fetch and decode operation in response to an instruction processing request, it further includes: obtaining a first microinstruction from the microinstruction cache through the microinstruction cache port and providing the first microinstruction to a second microinstruction queue; and in response to selecting an instruction cache port and a decoder to perform a fetch and decode operation in response to an instruction processing request, it further includes: obtaining an instruction from the instruction cache through the instruction cache port and providing the instruction to the decoder, and the decoder decodes the instruction into a second microinstruction and provides the second microinstruction to the second microinstruction queue.

[0121] For example, the first micro-instruction and the second micro-instruction can come from the same thread or different threads. For example, the instruction data from thread 1 and thread 2 is processed in the instruction fetch and decode pipeline 1. Exemplarily, the first micro-instruction and the second micro-instruction obtained by processing in the instruction fetch and decode pipeline 1 from thread 1 can be separately provided to the first micro-instruction queue or the second micro-instruction queue, or a part of the first micro-instruction and the second micro-instruction obtained by processing in the instruction fetch and decode pipeline 1 from thread 1 can be provided to the first micro-instruction queue, and the remaining part can be provided to the second micro-instruction queue, and the same applies to the micro-instruction mode of thread 2.

[0122] The instruction processing method provided by at least one embodiment of the present disclosure allows multiple threads to use multiple pipelines whether in the instruction decoding mode or the micro-instruction caching mode. Each thread can generate a larger average front-end bandwidth, thereby greatly improving the multi-thread performance.

[0123] Figure 8 The specific flowchart of the instruction processing method according to at least one embodiment of the present disclosure is shown, which is applicable to the single-thread mode. As Figure 8 shown, the instruction processing method includes steps S801 - S812.

[0124] For steps S801 - S804, S806 - S808, S810, and S812, they have been described above and will not be elaborated here.

[0125] Step S805: Whether to enter the instruction decoding mode.

[0126] For example, according to the cache hit or cache miss information given by the micro-instruction cache, it is determined whether to enable the micro-instruction caching mode. For example, when the micro-instruction cache gives cache hit information that several consecutive micro-instruction groups exist in the micro-instruction cache, the micro-instruction cache extraction mode is enabled, otherwise the prediction result is continued to be processed in the instruction cache mode.

[0127] Step S809: Whether the micro-instruction has become the head of the queue.

[0128] For example, if the micro-instruction of the currently processed information stream becomes the head of the micro-instruction queue (FIFO queue), it means that the micro-instructions of the previous information stream have been dispatched (issued). Thus, there are no longer micro-instructions of the previous information stream in the current micro-instruction queue. Therefore, the micro-instructions from different instruction fetch and decode pipelines can be reordered to restore the order before entering the N instruction fetch and decode pipelines.

[0129] Step S811: Whether it is possible to enter the instruction dispatch.

[0130] For example, in one possible implementation, instruction dispatch is determined based on one or more factors of resource availability, data dependency, scheduling policy, and branch prediction and prefetching.

[0131] For example, each execution unit (such as arithmetic logic unit ALU, floating point unit FPU, load / store unit LSU, etc.) has a certain throughput limit. If there is currently an idle execution unit, the corresponding instruction can be issued. Otherwise, the instruction needs to wait until the required resources become available.

[0132] For example, to ensure program correctness, instruction dispatch needs to take into account data dependencies. For example, if the result of one instruction is the input of another instruction, the latter cannot be executed before it. Reordering logic tracks these dependencies to ensure that instructions are dispatched only when all preconditions are met.

[0133] For example, complex scheduling algorithms are often used to decide which instructions should be dispatched first. Common scheduling algorithms include:

[0134] Oldest First: Select the instruction that arrived at the reservation station earliest and has no unresolved dependencies.

[0135] Least Delay: Prioritizes the dispatch of instructions with the shortest estimated execution time to reduce the possibility of pipeline stalls.

[0136] Maximize resource utilization: Try to balance the workload of each execution unit to avoid overloading some units while leaving others idle.

[0137] The instruction processing method provided by at least one embodiment of the present disclosure can be applied to a single-threaded processor core or a multi-threaded processor core. No matter which of the multiple instruction fetch and decode pipelines is blocked, there is a spare pipeline to take over the prediction information, so that more microinstructions can be accumulated in the microinstruction queue for back-end consumption, providing a larger front-end bandwidth, thereby further improving processor performance. Moreover, no matter whether the processor is in instruction decoding mode or microinstruction mode, the processor has two pipelines to consume the increased branch prediction bandwidth, so it can also make more full use of the increased branch prediction bandwidth.

[0138] Figure 9 A structural example of a processor core according to at least one embodiment of the present disclosure is shown, where the processor core is suitable for single-threaded mode.

[0139] like Figure 9As shown, the front end of the processor core includes two instruction fetch and decode pipelines (e.g., instruction fetch and decode pipeline 1 and instruction fetch and decode pipeline 2). It should be noted that in other examples, the number of instruction fetch and decode pipelines can be determined based on the specific design performance of the front end of the processor core and can include two or more instruction fetch and decode pipelines.

[0140] Specifically, the front end of the processor core includes a branch prediction unit 901, selection logics 9021 / 9022, instruction caches 903, instruction cache ports 9031 / 9032, instruction decoders 9041 / 9042, micro-instruction caches 905, micro-instruction cache ports 9051 / 9052, micro-instruction queues 9061 / 9063 (examples of "first micro-instruction queues"), instruction distribution units 907, boundary information determination units 908, window selection units 909, and reordering logics 910. These components or logics can be implemented by hardware, firmware, software, or any combination thereof, and their function descriptions are as described above and will not be elaborated here.

[0141] Taking one of the two instruction fetch and decode pipelines as an example (e.g., instruction fetch and decode pipeline 1) for illustration. The instruction fetch and decode pipeline 1 includes a selection logic 9021, an instruction cache 903, an instruction cache port 9031, an instruction decoder 9041, a micro-instruction cache 905, a micro-instruction cache port 9051, and a micro-instruction queue 9061. The instruction cache 903, the instruction cache port 9031, and the instruction decoder 9041 constitute the instruction decode pipeline 1; the micro-instruction cache 905 and the micro-instruction cache port 9051 constitute the micro-instruction pipeline 1. For the instruction fetch and decode pipeline 2, its composition is similar to that of the instruction fetch and decode pipeline 1 and will not be elaborated here.

[0142] The instruction cache ports 9031 and 9032 in the two instruction fetch and decode pipelines access the same instruction cache 903; the micro-instruction cache ports 9051 and 9052 in the two instruction fetch and decode pipelines access the same micro-instruction cache 905.

[0143] The instruction cache 903 is configured to be coupled to the instruction cache ports of the two instruction fetch and decode pipelines (e.g., instruction cache ports 9031 and 9032) through memory interleaving technology or multi-port technology to achieve non-blocking multi-port access, so that the two instruction fetch and decode pipelines can access the instruction cache; the micro-instruction cache 905 is configured to be coupled to the micro-instruction cache ports of the two instruction fetch and decode pipelines (e.g., micro-instruction cache ports 9051 and 9052) through memory interleaving technology or multi-port technology to achieve non-blocking multi-port access, so that the two instruction fetch and decode pipelines can access the micro-instruction cache.

[0144] For example, if the above instruction cache ports (e.g., instruction cache port 9031 and instruction cache port 9032) simultaneously access instruction cache 903 and encounter an address conflict, they are accessed in order, and the present disclosure does not restrict the access order. For the microinstruction cache, the processing method is consistent with the instruction cache, and will not be repeated here.

[0145] Figure 9 The processor core shown is suitable for single-threaded mode, and two instruction fetch and decode pipelines including an instruction decode pipeline and a microinstruction pipeline are designed. No matter which pipeline is blocked, there is a spare pipeline to take over the instructions to be processed, providing a larger front-end bandwidth, thereby further improving processor performance.

[0146] Figure 10 A structural example of a processor core according to at least another embodiment of the present disclosure is shown. The processor core is suitable for a dual-thread mode, that is, the processor core is an SMT2 processor core.

[0147] like Figure 10 As shown, the front end of the processor core also includes two instruction fetch and decode pipelines. Figure 9 The processor core of the illustrated embodiment does not include the boundary information confirmation unit 908 , the window selection unit 909 and the reordering logic 910 on the one hand, but includes the first arbitration logic 911 and the second arbitration logic 912 on the other hand.

[0148] The prediction results from the branch prediction unit include prediction results for two threads, and the first arbitration logic 911 is configured to distribute the prediction results in the two threads to the two instruction fetch decoding pipelines for processing.

[0149] The microinstruction queues 9061 / 9063 of the processor core receive microinstructions obtained by processing the prediction results of their respective threads respectively, and the second arbitration logic 912 is configured to select and send the microinstructions corresponding to the two threads obtained by processing the two instruction fetch and decode pipelines to the instruction distribution unit in sequence.

[0150] Figure 10 The processor core shown is suitable for dual-thread mode, which allows dual threads to use two instruction fetch and decode pipelines regardless of instruction decoding or micro-instruction mode, so that each thread can generate a larger average front-end bandwidth, thereby improving multi-thread performance.

[0151] Figure 11 A structural example of a processor core according to at least another embodiment of the present disclosure is shown. The processor core is suitable for a four-thread mode, that is, the processor core is an SMT4 processor core.

[0152] like Figure 11 As shown, the front end of the processor core also includes two instruction fetch and decode pipelines.Figure 10 The processor core of the illustrated embodiment adds micro-instruction queues 9062 / 9064 (examples of "second micro-instruction queues") in two instruction fetch and decode pipelines respectively.

[0153] The micro-instruction queues 9062 / 9064 are configured to receive micro-instructions from a decoder and / or a micro-instruction cache port in each corresponding instruction fetch and decode pipeline.

[0154] In Figure 11 the embodiment of, the prediction results include prediction results for four threads (e.g., thread 1 - thread 4), and there are two instruction fetch and decode pipelines (e.g., instruction fetch and decode pipeline 1 and instruction fetch and decode pipeline 2). After arbitration by the first arbitration logic 911, for example, the prediction results from thread 1 and thread 2 can be assigned to instruction fetch and decode pipeline 1 for processing, and the prediction results from thread 3 and thread 4 can be assigned to instruction fetch and decode pipeline 2 for processing.

[0155] For example, in instruction fetch and decode pipeline 1, all the micro-instructions obtained by processing thread 1 through the instruction decode pipeline or the micro-instruction pipeline can be sent to micro-instruction queue 9061 or micro-instruction queue 9062; for example, some of the micro-instructions obtained by processing thread 1 through the instruction decode pipeline or the micro-instruction pipeline can also be sent to micro-instruction queue 9061, and the remaining micro-instructions are sent to micro-instruction queue 9062. For the micro-instructions from thread 2 in instruction fetch and decode pipeline 1, their assignment method is the same as that of the above micro-instructions from thread 1, and will not be elaborated here. The second arbitration logic 912 is configured to sequentially select and send the micro-instructions corresponding to the four threads obtained by processing through the two instruction fetch and decode pipelines to the instruction distribution unit.

[0156] Figure 11 The illustrated processor core is applicable to the four-thread mode, allowing four threads to use two instruction fetch and decode pipelines both in the instruction decode mode and the micro-instruction mode, which can increase the number of instructions received by the front end of the processor core and thus improve the multi-thread performance.

[0157] At least one embodiment of the present disclosure also provides an electronic device. Figure 12 A schematic block diagram of the electronic device of at least one embodiment of the present disclosure is shown.

[0158] For example, as Figure 12As shown, the electronic device 1200 includes a processor 1210 and a memory 1220. The memory 1220 is used to store non-transitory computer-readable instructions (such as one or more computer program modules). The processor 1210 is used to run the computer program instructions, and when the computer program instructions are run by the processor 1210, they execute the instruction processing method provided in any embodiment of the present disclosure. The memory 1220 and the processor 1210 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0159] The processor 1210 can be a central processing unit (CPU), a tensor processing unit (TPU), a network processor (NP), or a graphics processing unit (GPU), etc., which have data processing capabilities and / or program execution capabilities. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. For example, the central processing unit (CPU) can be of the X86 or ARM architecture, etc. The processor 1210 can be a general-purpose processor or a special-purpose processor, and can control other components in the electronic device 1200 to perform desired functions.

[0160] For example, the memory 1220 can include any combination of one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules can be stored on the computer-readable storage media. The processor 1210 can run one or more computer program modules to implement various functions of the electronic device 1200. Various application programs and various data, as well as various data used and / or generated by the application programs, etc., can also be stored in the computer-readable storage media.

[0161] It should be noted that in the embodiments of the present disclosure, the specific functions and technical effects of the electronic device 1200 can refer to the description of the instruction processing method in the above text, and will not be elaborated here.

[0162] At least one embodiment of the present disclosure also provides a non-transitory storage medium. Figure 13 Schematic diagram of the computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 13As shown, the storage medium 1300 non - temporarily stores computer - executable instructions 1310, which can execute the instruction - processing method of any embodiment of the present disclosure when executed by a computer (including a processor).

[0163] For example, one or more computer instructions can be stored on the storage medium 1300. Some of the computer instructions stored on the storage medium 1300 can be, for example, instructions for implementing one or more steps in the above - mentioned instruction - processing method.

[0164] For example, the storage medium can include the storage component of a tablet computer, the hard disk of a personal computer, random access memory (RAM), read - only memory (ROM), erasable programmable read - only memory (EPROM), compact disc read - only memory (CD - ROM), flash memory, or any combination of the above - mentioned storage media, and can also be other applicable storage media. For example, the storage medium 1300 can include the memory 1220 in the aforementioned electronic device 1200.

[0165] The technical effects of the storage medium provided by the embodiments of the present disclosure can refer to the corresponding descriptions of the instruction - processing method in the above - mentioned embodiments, and will not be elaborated here.

[0166] At least one embodiment of the present disclosure also provides an electronic device, which includes the processor core of any of the above - mentioned embodiments.

[0167] Figure 14 A schematic block diagram of an electronic device according to at least another embodiment of the present disclosure is shown. Figure 14 The shown electronic device 1400 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0168] As Figure 14 shown, in some examples, the electronic device 1400 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 1401, which can include the processor core of any of the above - mentioned embodiments, and can perform various appropriate actions and processes according to the program stored in the read - only memory (ROM) 1402 or the program loaded from the storage device 1408 into the random access memory (RAM) 1403. In the RAM 1403, various programs and data required for the operation of the computer system are also stored. The processor 1301, the ROM 1402, and the RAM 1403 are connected to each other through a bus 1404. The input / output (I / O) interface 1305 is also connected to the bus 1404.

[0169] For example, the following components may be connected to the I / O interface 1405: an input device 1406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1407 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1408 including, for example, a magnetic tape, a hard disk, etc.; a communication device 1409 which may also include, for example, a network interface card such as a LAN card, a modem, etc. The communication device 1409 may allow the electronic device 1400 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The driver 1410 is also connected to the I / O interface 1405 as needed. A removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 1410 as needed so that a computer program read from it can be installed into the storage device 1408 as needed.

[0170] Although Figure 14 the electronic device 1400 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.

[0171] For example, the electronic device 1400 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 1409 may communicate with a network and other devices through wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.

[0172] For example, the electronic device 1400 may include any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, a server, etc., or may be any combination of a data processing device and hardware. The embodiments of the present disclosure are not limited thereto.

[0173] For the present disclosure, the following points need to be noted:

[0174] (1) In the drawings of the embodiments of the present disclosure, only the structures related to the embodiments of the present disclosure are involved, and other structures can refer to the general design.

[0175] (2) Without conflict, the features in the same embodiment and different embodiments of the present disclosure can be combined with each other.

[0176] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A processor core includes a front end, wherein, The front end includes N instruction fetch and decode pipelines, at least one instruction cache, and at least one micro-instruction cache, where N is an integer greater than or equal to 2. Each of the instruction fetch and decode pipelines includes an instruction fetch selection logic, an instruction cache port, a micro-instruction cache port, a decoder, and a first micro-instruction queue. The instruction cache port is configured to read and write the instruction cache and provide the instructions obtained from the instruction cache to the decoder. The micro-instruction cache port is configured to read and write the micro-instruction cache and provide the first micro-instructions obtained from the micro-instruction cache to the first micro-instruction queue. The decoder is configured to decode the instructions obtained from the instruction cache port into second micro-instructions and provide the decoded second micro-instructions to the first micro-instruction queue. The instruction fetch selection logic is configured to select the instruction cache port and the decoder for instruction fetch and decode operations, or select the micro-instruction cache port for instruction fetch and decode operations.

2. The processor core according to claim 1, wherein, The front end further includes: An instruction distribution unit configured to distribute the micro-instructions received from the first micro-instruction queue of one of the N instruction fetch and decode pipelines for execution.

3. The processor core according to claim 1, wherein The instruction cache is configured to be coupled to the instruction cache ports of the N instruction fetch and decode pipelines through memory interleaving technology or multi-port technology, so that the N instruction fetch and decode pipelines can access the instruction cache; and / or The micro-instruction cache is configured to be coupled to the micro-instruction cache ports of the N instruction fetch and decode pipelines through memory interleaving technology or multi-port technology, so that the N instruction fetch and decode pipelines can access the micro-instruction cache.

4. The processor core according to claim 2 or 3, wherein The front end further includes: A branch prediction unit configured to generate a prediction result based on the received instruction address for the N instruction fetch and decode pipelines.

5. The processor core according to claim 4, wherein, The front end further includes: A boundary information determination unit configured to generate boundary information according to the prediction result and generate at least one information stream according to the prediction result. Wherein, the type of the boundary information includes the last byte of a jump branch instruction, the last byte of the prediction result, or the last byte of an intermediate instruction in the prediction result.

6. The processor core according to claim 5, wherein, The front end further includes: A window selection unit configured to allocate the at least one information stream to the N instruction fetch and decode pipelines to process the at least one information stream. Wherein, the allocation strategy of the window selection unit includes any one of the number of prediction results, the number of bytes of the prediction result, the blocking degree of the N instruction fetch and decode pipelines, and the window switching frequency in different modes.

7. The processor core according to claim 5, wherein, The front end further includes: A reordering logic configured to reorder the micro-instructions obtained by the N instruction fetch and decode pipelines processing the at least one information stream and then send them to the instruction distribution unit.

8. The processor core according to claim 4, wherein The front end further includes: A first arbitration logic and a second arbitration logic, wherein, The prediction result includes prediction results for M threads, where M is an integer greater than or equal to 2; The first arbitration logic is configured to allocate the prediction results among the M threads to the N instruction fetch and decode pipelines for processing; and The second arbitration logic is configured to sequentially select and send the micro-instructions corresponding to the M threads processed by the N fetch-decode pipelines to the instruction distribution unit.

9. The processor core according to claim 8, wherein, Each of the fetch-decode pipelines further includes a second micro-instruction queue. The second micro-instruction queue is configured to receive the micro-instructions from the decoder and / or the micro-instruction cache port in each of the fetch-decode pipelines.

10. An instruction processing method, comprising: In response to an instruction processing request, selecting a target fetch-decode pipeline among the N fetch-decode pipelines included in the front end of a processor core to respond to the instruction processing request, where N is an integer greater than or equal to 2; In the target fetch-decode pipeline, selecting to perform fetch-decode operations through an instruction cache port and a decoder, or selecting to perform fetch-decode operations through a micro-instruction cache port to respond to the instruction processing request, where, in response to selecting to perform fetch-decode operations through the micro-instruction cache port to respond to the instruction processing request, includes: obtaining a first micro-instruction from a micro-instruction cache through the micro-instruction cache port and providing the first micro-instruction to a first micro-instruction queue, or In response to selecting to perform fetch-decode operations through the instruction cache port and the decoder to respond to the instruction processing request, includes: obtaining an instruction from an instruction cache through the instruction cache port and providing the instruction to the decoder, and the decoder decodes the instruction into a second micro-instruction and provides the second micro-instruction to the first micro-instruction queue.

11. The instruction processing method according to claim 10, further comprising: Distributing the micro-instructions received from the first micro-instruction queue of one of the target fetch-decode pipelines among the N fetch-decode pipelines for execution.

12. The instruction processing method according to claim 10, wherein, The instruction cache is coupled to the instruction cache port of the target fetch-decode pipeline among the N fetch-decode pipelines through a memory interleaving technique or a multi-port technique, so that the target fetch-decode pipeline among the N fetch-decode pipelines accesses the instruction cache; and / or The micro-instruction cache is coupled to the micro-instruction cache port of the target fetch-decode pipeline among the N fetch-decode pipelines through a memory interleaving technique or a multi-port technique, so that the target fetch-decode pipeline among the N fetch-decode pipelines accesses the micro-instruction cache.

13. The instruction processing method according to claim 11 or 12, further comprising: Generating a prediction result based on the received instruction address to generate the instruction processing request for the target fetch-decode pipeline among the N fetch-decode pipelines.

14. The instruction processing method according to claim 13, further comprising: Generating boundary information according to the prediction result corresponding to the instruction processing request, and generating at least one information stream according to the prediction result, where the type of the boundary information includes the last byte of a jump branch instruction, the last byte of the prediction result, or the last byte of an intermediate instruction in the prediction result.

15. The instruction processing method according to claim 14, further comprising: Allocating the at least one information stream to a target fetch and decode pipeline among the N fetch and decode pipelines through window selection logic to process the at least one information stream. Wherein, the window selection logic includes any one of the number of the prediction results, the number of bytes of the prediction results, the blocking degree of the target fetch and decode pipeline among the N fetch and decode pipelines, and the window switching frequency in different modes.

16. The instruction processing method according to claim 15, further comprising: Reordering the microinstructions obtained by processing the at least one information stream by the target fetch and decode pipeline among the N fetch and decode pipelines for distribution.

17. The instruction processing method according to claim 13, wherein, The prediction results include prediction results for M threads, where M is an integer greater than or equal to 2, and the method further comprises: Allocating the prediction results corresponding to the instruction processing request among the M threads to the target fetch and decode pipeline among the N fetch and decode pipelines for processing; and Sequentially selecting the microinstructions corresponding to the M threads obtained by processing by the target fetch and decode pipeline among the N fetch and decode pipelines for distribution.

18. The instruction processing method according to claim 17, wherein, The responding to the selection by selecting the microinstruction cache port to perform a fetch and decode operation to respond to the instruction processing request includes: Obtaining a first microinstruction from the microinstruction cache through the microinstruction cache port and providing the first microinstruction to a second microinstruction queue; and The responding to the selection by the instruction cache port and the decoder to perform a fetch and decode operation to respond to the instruction processing request includes: Obtaining an instruction from the instruction cache through the instruction cache port and providing the instruction to the decoder, and the decoder decoding the instruction into a second microinstruction and providing the second microinstruction to the second microinstruction queue.

19. An electronic device, comprising the processor core according to any one of claims 1-9.

20. An electronic device, comprising: A processor; And A memory, including one or more computer program instructions; Wherein, when the one or more computer program instructions are run by the processor, they execute the instruction processing method according to any one of claims 10-18.

21. A computer-readable storage medium that non-transitorily stores computer-readable instructions, wherein, When the computer-readable instructions are executed by the processor, the instruction processing method according to any one of claims 10-18 is implemented.

Citation Information

Patent Citations

  • Implementation method and device for processor front-end instruction reading queue

    CN118444984A

  • Apparatus and method for non-speculative resource deallocation

    US20210200552A1