Method and apparatus for accelerating critical path instructions

The PGO tool recognizes and recompiles the critical path instructions in the RISC-V processor, and combines the operation codes and key prompt information of custom key instructions to achieve accelerated execution of all types of critical path instructions, solving the limitations of identifying and accelerating critical path instructions in the existing technology and significantly improving system performance.

CN119597302BActive Publication Date: 2025-06-27INTEL CHINA RES CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510142159.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-27
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The prior art lacks a unified framework when identifying and accelerating critical path instructions in RISC-V processors, and is limited to identifying and accelerating loading instructions, and cannot effectively process other types of critical path instructions.

Method used

Critical path analysis is performed using profiling guidance optimization (PGO) tools, critical path instructions are identified, and recompiled into custom critical instructions with critical prompt information through compiler tools. The RISC-V processor realizes accelerated execution of custom key instructions based on the operation code and/or critical prompt information of custom key instructions.

Benefits of technology

It realizes free identification and acceleration of all types of critical path instructions, significantly improves system performance, and provides a general and flexible software and hardware collaborative design framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119597302B_ABST
    Figure CN119597302B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for accelerating critical path instructions. The method includes: performing critical path analysis on a program using a profile-guided optimization (PGO) tool to identify critical path instructions; recompiling the identified critical path instructions into custom critical instructions with criticality hint information using a compiler tool, wherein the opcode of the custom critical instruction is an opcode reserved by the RISC-V instruction set architecture for custom instructions, and wherein the opcode is used to indicate that the instruction is a specific type of critical instruction; and implementing accelerated execution of the custom critical instructions by an RISC-V processor based on the opcode and / or criticality hint information of the custom critical instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and more particularly, to methods and devices for accelerating critical path instructions in a fifth-generation Reduced Instruction Set Computer (RISC-V) processor. Background Art

[0002] Out-of-order execution is a key technology in modern processor architectures. The processor core executes instructions based on the data dependencies in the program. If there are no dependencies between instructions, they can be executed in parallel to hide latency. At this time, the instruction with the longest execution time (for example, a load operation that misses all caches) forms the critical path and ultimately determines the overall execution duration of the program. Any delay in executing these instructions will affect the performance of the entire system. On the contrary, the system can tolerate delays in non-critical path instructions as long as such delays do not result in the generation of new critical path instructions. Therefore, identifying critical path instructions in a program and designing specific microarchitecture algorithms to accelerate them are crucial for developing high-performance processors. Summary of the Invention

[0003] According to an embodiment of the present disclosure, there is provided a method for accelerating critical path instructions in a RISC-V processor, the method including: performing critical path analysis on a program using a profile guided optimization (PGO) tool to identify critical path instructions; recompiling the identified critical path instructions into custom critical instructions with criticality hint information using a compiler tool, wherein an opcode of the custom critical instruction is an opcode reserved by the RISC-V instruction set architecture for custom instructions, and wherein the opcode is used to indicate that the instruction is a specific type of critical instruction; and implementing accelerated execution of the custom critical instruction by the RISC-V processor based on the opcode and / or criticality hint information of the custom critical instruction.

[0004] According to an embodiment of the present disclosure, there is provided a device for accelerating critical path instructions in a RISC-V processor, the device including means for performing the steps of the above method.

[0005] According to an embodiment of the present disclosure, there is provided a computer-readable storage medium having executable code stored thereon, wherein the executable code, when executed by a processing circuit, causes the processing circuit to execute the method according to the above embodiment.

[0006] According to an embodiment of the present disclosure, there is provided a computer program product including executable code, wherein the instructions, when executed by a processing circuit, cause the processing circuit to execute the method according to the above embodiment. Brief Description of the Drawings

[0007] Embodiments of the present disclosure will be described by way of example and not limitation in conjunction with the figures in the accompanying drawings, where like reference numerals refer to like elements and in which:

[0008] Figure 1 A block diagram of an exemplary processor and / or SoC 100 is illustrated, which can have one or more cores and have an integrated memory controller.

[0009] Figure 2 A flowchart of a method for accelerating critical path instructions in a RISC-V processor according to an embodiment of the present disclosure is shown.

[0010] Figure 3 An exemplary schematic diagram of an event dependency graph according to an embodiment of the present invention is shown.

[0011] Figure 4 A schematic diagram of the existing instruction encoding of RISC-V is shown.

[0012] Figure 5 A schematic flowchart of the BPU predicting custom critical branch instructions according to an embodiment of the present disclosure is shown.

[0013] Figure 6 A schematic flowchart of the hardware cache prefetcher predicting load instructions according to an embodiment of the present disclosure is shown.

[0014] Figure 7 A block diagram showing components that can read instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and execute any one or more of the methods discussed herein according to some example embodiments is shown. Detailed Description

[0015] Features and exemplary embodiments of various aspects of the present application will be described in detail below. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to those skilled in the art that the present application may be practiced without some of these specific details. The following description of the embodiments is merely provided to better understand the present application by way of illustrating examples of the present application. The present application is in no way limited to any specific configuration set forth below, but covers any modification, replacement, and improvement of elements, components, and algorithms without departing from the spirit of the present application. Well-known structures and techniques are not shown in the drawings and the following description so as not to unnecessarily obscure the present application.

[0016] Additionally, various operations will be described as multiple discrete operations in a manner that is most helpful in understanding the illustrative embodiments; however, the order of description should not be construed as implying that these operations are necessarily order dependent. In particular, these operations do not have to be performed in the order presented.

[0017] The phrases "in an embodiment", "in one embodiment", and "in some embodiments" are used repeatedly herein. This phrase generally does not refer to the same embodiment; however, it may. Unless the context dictates otherwise, the terms "comprising", "having", and "including" are synonyms. The phrases "A or B" and "A / B" mean "(A), (B), or (A and B)".

[0018] Figure 1 A block diagram of an example processor and / or SoC 100 is illustrated, which may have one or more cores and have an integrated memory controller. The processor 100 illustrated by the solid-line block has a single core 102(A), a system agent unit circuit 110, and a set of one or more interface controller unit circuits 116, while the optionally added dashed-line block illustrates an alternative processor 100 as having multiple cores 102(A)-(N), a set of one or more integrated memory control unit circuits 114 in the system agent unit circuit 110, dedicated logic 108, and a set of one or more interface controller unit circuits 116.

[0019] Different implementations of the processor 100 may include: 1) a CPU, where the dedicated logic 108 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and the cores 102(A)-(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 102(A)-(N) are a large number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where the cores 102(A)-(N) are a large number of general-purpose in-order cores. Thus, the processor 100 may be a general-purpose processor, a coprocessor, or a special-purpose processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor may be implemented on one or more chips. The processor 100 may be part of one or more substrates and / or may be implemented on one or more substrates using any of a variety of process technologies, such as complementary metal oxide semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0020] The memory hierarchy includes one or more levels of cache unit circuits 104(A)-(N) within cores 102(A)-(N), a group of one or more shared cache unit circuits 106, and an external memory (not shown) coupled to the group of integrated memory controller unit circuits 114. The group of one or more shared cache unit circuits 106 may include one or more intermediate-level caches, such as a second-level (L2), third-level (L3), fourth-level (L4), or other levels of cache, such as a last-level cache (LLC), and / or combinations thereof. Although in some examples the interface network circuit 112 (e.g., a ring interconnect) provides an interface to the dedicated logic 108 (e.g., integrated graphics logic), the group of shared cache unit circuits 106, and the system agent unit circuit 110, alternative examples use any number of well-known techniques to provide an interface to these units. In some examples, coherence is maintained between one or more of the circuits in the shared cache unit circuits 106 and the cores 102(A)-(N). In some examples, the interface controller unit circuit 116 couples these cores to one or more other devices 118, such as one or more I / O devices, storage devices, one or more communication devices (e.g., wireless networks, wired networks, etc.), and so on.

[0021] In some examples, one or more of the cores 102(A)-(N) have multithreading capabilities. The system agent unit circuit 110 includes those components that coordinate and operate the cores 102(A)-(N). The system agent unit circuit 110 may include, for example, a power control unit (PCU) circuit and / or a display unit circuit (not shown). The PCU may be (or may include) the logic and components required to regulate the power states of the cores 102(A)-(N) and / or the dedicated logic 108 (e.g., integrated graphics logic). The display unit circuit is used to drive one or more externally connected displays.

[0022] The cores 102(A)-(N) may be homogeneous in terms of the instruction set architecture (ISA). Alternatively, the cores 102(A)-(N) may be heterogeneous in terms of the ISA; that is, a subset of the cores 102(A)-(N) may be capable of executing one ISA, while other cores may be capable of executing only a subset of that ISA or may be capable of executing another ISA. In one embodiment, the processor cores 102(A)-(N) may wholly or partially adopt the RISC-V instruction set architecture.

[0023] Existing solutions for identifying and accelerating critical path instructions mainly focus on predicting the criticality of load instructions and using prefetching techniques to accelerate the identified critical load instructions. However, existing solutions are limited to identifying and accelerating a single type of critical path instruction (i.e., only load instructions), and fail to effectively handle other types of critical path instructions. In addition, existing solutions lack a unified framework for identifying and accelerating critical path instructions.

[0024] In view of this, the present disclosure provides a software-hardware co-design framework capable of freely identifying and accelerating critical path instructions. Specifically, in the software-hardware co-design framework provided by the present disclosure, the software part can be used for: (1) using the PGO technology to identify critical path instructions in the critical path; (2) establishing an instruction extension for prompting the criticality of instructions by utilizing the characteristics of the RISC-V instruction set architecture; and (3) recompiling the identified critical path instructions using the criticality prompt to obtain different types of custom critical instructions. In the hardware part, different microarchitectures can adopt various strategies to process the instructions marked as critical (the recompiled custom critical instructions). Specifically, prediction components (e.g., branch prediction components (BPU), prefetchers, etc.) can accurately predict the custom critical instructions and prepare the required data in advance for the custom critical instructions, and buffer components (e.g., arithmetic logic units (ALU)) and computing components can ensure the availability of hardware resources for the custom critical instructions and preferentially allocate hardware resources to the custom critical instructions.

[0025] The software-hardware co-design framework of the present disclosure has high generality and flexibility. It can not only identify load instructions, but also identify all other types of critical instructions, and supports the implementation of various algorithms or strategies customized for accelerating these instructions. By comprehensively accelerating all types of critical instructions, compared with the previous solutions that only focused on load instructions, the framework of the present disclosure can significantly improve system performance.

[0026] Figure 2 A flowchart of a method for accelerating critical path instructions in a RISC-V processor according to an embodiment of the present disclosure is shown. Method 200 may include steps S202, S204, and S206. However, in some embodiments, method 200 may include more or fewer different steps, and the present disclosure places no limitation thereon.

[0027] In step S202, perform a critical path analysis on the program using a PGO tool to identify critical path instructions;

[0028] In step S204, the compiler tool recompiles the identified critical path instructions into custom critical instructions with criticality hint information, where the opcode of the custom critical instruction is the opcode reserved for custom instructions in the RISC-V instruction set architecture, and where the opcode is used to indicate that the instruction is a specific type of critical instruction; and

[0029] In step S206, the RISC-V processor accelerates the execution of the custom critical instruction based on the opcode and / or criticality hint information of the custom critical instruction.

[0030] In an embodiment of the present disclosure, method 200 may further include: identifying a critical path and critical events on the critical path based on an event dependency graph (DEG); and determining critical path instructions that generate the critical events based on the identified critical events. Figure 3 An exemplary schematic diagram of an event dependency graph according to an embodiment of the present disclosure is shown. The DEG is widely used to determine which CPU events are the main causes of cycle loss during the entire micro-execution process. As Figure 3 shown, for example, a D-cache miss may be a critical event that contributes 100 cycles to the micro-execution. And the events that form the DEG originate from the execution of instructions. Therefore, by searching for critical events in the DEG, we can identify which instructions are critical. These instructions may have one or more critical events in the DEG.

[0031] In an embodiment of the present disclosure, different opcodes reserved for custom instructions in the RISC-V instruction set architecture are used to indicate different types of custom critical instructions. In the existing instruction encoding of RISC-V, there are some opcodes reserved for custom instructions. Figure 4 A schematic diagram of the existing instruction encoding of RISC-V is shown. As Figure 4 shown, in the RISC-V instruction architecture, 4 custom instruction types, custom-0, custom-1, custom-2, and custom-3, are reserved, and their opcodes can use the opcodes corresponding to custom-0, custom-1, custom-2, and custom-3 in the table.

[0032] In one embodiment of the present disclosure, the opcode 0001011 of custom-0 can be used to compile a custom critical branch instruction; the opcode 0101011 of custom-1 can be used to compile a custom critical load instruction; the opcode 1011011 of custom-2 can be used to compile a custom critical multiply-add instruction; and the opcode 1111011 of custom-3 can be used to compile a custom critical multiply-subtract instruction. In another embodiment of the present disclosure, the opcodes of custom-0, custom-1, custom-2, and custom-3 can be respectively used to compile various other types of custom critical instructions. In yet another embodiment of the present disclosure, any other reserved opcode in the RISC-V instruction architecture can be used to compile various types of custom critical instructions, and the present disclosure does not make specific limitations thereto.

[0033] In an embodiment of the present disclosure, the criticality hint information of the custom critical instruction can be encoded in the Funct3 field of the custom critical instruction. For example, 3 bits in the Funct3 field are utilized. In an embodiment of the present disclosure, the criticality hint information of the custom critical instruction can be used to indicate whether the instruction is a regular instruction or a critical instruction. In an embodiment of the present disclosure, the criticality hint information of the custom critical instruction can also be used to hint at the stability of the critical instruction.

[0034] In one embodiment, the criticality hint information of the custom critical branch instruction is used to indicate that the instruction is one of the following: a regular instruction; a critical instruction with a stable jump; a critical instruction with a stable non-jump; or a critical instruction with an unstable jump. For example, in the Funct3 field of the custom critical branch instruction: 000 can indicate that the instruction is a regular instruction; 001 can indicate that the instruction is a critical instruction with a stable jump; 010 can indicate that the instruction is a critical instruction with a stable non-jump; 011 can indicate that the instruction is a critical instruction with an unstable jump.

[0035] In one embodiment, the criticality hint information of custom critical instructions such as custom load critical instructions, custom multiply-add critical instructions, or custom multiply-subtract critical instructions is used to indicate that the instruction is one of the following: a regular instruction; or a critical instruction. For example, in the Funct3 field of the above instructions: 000 can indicate that the instruction is a regular instruction; 001 can indicate that the instruction is a critical instruction.

[0036] It should be understood that the encoding format of the Funct3 field of the above custom critical instructions is only for illustrative purposes. The present disclosure does not specifically limit the encoding format of the Funct3 field. On the contrary, the present invention aims to cover all possible variations and modifications as long as they fall within the technical scope claimed by the present invention.

[0037] In an embodiment of the present disclosure, the accelerated execution of custom critical instructions implemented by a RISC-V processor based on the opcode and / or criticality hint information of the custom critical instructions may include: predicting the control flow and / or data flow of custom critical instructions by a prediction component of the RISC-V processor based on the opcode and / or criticality hint information of the custom critical instructions.

[0038] In an embodiment of the present disclosure, the prediction component of the RISC-V processor may include, for example, a branch prediction unit (BPU) for predicting the control flow of custom critical instructions and an instruction cache prefetcher, as well as a data cache prefetcher for predicting the data flow of custom critical instructions, and so on.

[0039] Figure 5 Shows a schematic flow diagram of the BPU predicting a custom critical branch instruction according to an embodiment of the present disclosure. As Figure 5 shown, the BPU component is configured to directly predict the branch direction of the critical branch instruction based on the hint information compiled in the critical instruction when it is determined based on the criticality hint information of the incoming custom critical instruction that the branch instruction is a critical instruction with a stable hint, without performing the conventional process of the BPU component for predicting the direction of the branch instruction. The BPU component is also configured to perform the conventional process for predicting the direction of the branch instruction on the branch instruction when it is determined based on the criticality hint information of the incoming custom critical instruction that the branch instruction is a critical instruction with an unstable hint or a conventional instruction. In one embodiment, in the BPU history table of the BPU component, compared with conventional branch instructions, critical instructions with unstable hints may be assigned a higher replacement age (or a lower replacement priority), and a higher replacement age means that the entry of this instruction will remain in the history table for a longer time and is not easily replaced by a new entry. For a branch instruction with an unstable hint, keeping its prediction information resident in the history table for a long time helps to improve the prediction accuracy.

[0040] In an embodiment of the present disclosure, the hardware cache prefetcher is configured to prefetch the corresponding data block or instruction only for critical instructions. In one embodiment, the cache prefetcher may include a training data filter for filtering instructions without criticality hint information, such as conventional load instructions. In another embodiment, a training data filter may be provided before the cache prefetcher. Figure 6 Shows a schematic flow diagram of the hardware cache prefetcher predicting a load instruction according to an embodiment of the present disclosure. As Figure 6As shown, before the instruction enters the cache prefetcher 620, it first enters the training data filter 610, which is used to filter any memory access requests of instructions without critical hints. Only the critical memory accesses from critical instructions are used to train the cache prefetcher 620 and trigger prefetch requests. It should be understood that although Figure 6 only load instructions are shown in

[0041] the cache prefetcher according to the embodiments of the present disclosure can be used for any type of critical instruction that can utilize the cache prefetcher.

[0042] In an embodiment of the present disclosure, the RISC-V processor implementing the accelerated execution of custom critical instructions based on the opcode and / or critical hint information of the custom critical instructions may include: the hardware buffer component and / or the hardware computing component of the RISC-V processor preparing the corresponding hardware resources for the custom critical instructions preferentially.

[0043] In one embodiment, the buffer component is configured to allocate resources preferentially for custom critical instructions with critical hints and reduce the possibility of their being replaced. In other words, when a critical instruction arrives at the buffer component (such as a rename register file, an instruction queue, etc.), if the critical instruction and a normal instruction request hardware resources simultaneously, the buffer component will satisfy the resource requirements of the critical instruction preferentially. In an embodiment of the present disclosure, the buffer component is configured to, when a replacement operation is required, be able to replace the entries inserted by non-critical instructions preferentially rather than the entries inserted by critical instructions.

[0044] The software-hardware co-design framework of the present disclosure has high generality and flexibility. It can not only identify load instructions, but also identify all other types of critical instructions, and support the implementation of various algorithms or strategies customized for accelerating these instructions. By comprehensively accelerating all types of critical instructions, compared with the previous solutions that only focused on load instructions, the framework of the present disclosure can significantly improve the system performance.

[0045] Figure 7 is a block diagram showing components capable of reading instructions or executable code from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more of the methods discussed herein. Specifically, Figure 7A schematic diagram of hardware resource 700 is shown. Hardware resource 700 includes one or more processors (or processor cores) 710, one or more memory / storage devices 720, and one or more communication resources 730. Among them, each of these processors, memory / storage devices, and communication resources can be communicatively coupled via bus 740 or other interface circuits. For embodiments that utilize node virtualization (e.g., network function virtualization (NFV)), a hypervisor 702 can be executed to provide an execution environment for one or more network slices / sub-slices so as to utilize hardware resource 700.

[0046] Processor 710 can include, for example, processor 712 and processor 714. Processor 710 can be, for example, a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP) such as a baseband processing unit, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a radio frequency integrated circuit (RFIC), another processor (including those discussed herein), or any suitable combination thereof.

[0047] Memory / storage device 720 can include main memory, disk storage devices, or any suitable combination thereof. Memory / storage device 720 can include, but is not limited to, any type of volatile, non-volatile, or semi-volatile memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid state memory, etc.

[0048] Communication resource 730 can include an interconnect or network interface controller, component, or other suitable device to communicate with one or more peripheral devices 704 or one or more databases 706 or other network elements via network 708. For example, communication resource 730 can include wired communication components (e.g., for coupling via USB, Ethernet, etc.), cellular communication components, near field communication (NFC) components, Bluetooth® (or Bluetooth® low energy) components, Wi-Fi® components, and other communication components.

[0049] Instruction 750 may include software, a program, an application, an applet, an application, or other executable code for causing at least any one of the processors in processor 710 to perform any one or more of the methods discussed herein. Instruction 750 may reside, in whole or in part, in at least one of processor 710 (e.g., in the cache of the processor), memory / storage device 720, or any suitable combination thereof. Additionally, any portion of Instruction 750 may be transferred from any combination of peripheral device 704 or database 706 to hardware resource 700. Accordingly, the memory of processor 710, memory / storage device 720, peripheral device 704, and database 706 are examples of computer-readable and machine-readable media.

[0050] Some examples may be implemented using an article of manufacture or at least one computer-readable medium or be implemented as an article of manufacture or at least one computer-readable medium. The computer-readable medium may include a non-transitory storage medium to store logic. In some examples, the non-transitory storage medium may include one or more types of computer-readable storage media capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and the like. In some examples, the logic may include various software elements, such as software components, programs, applications, computer programs, applications, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.

[0051] According to some examples, the computer-readable medium may include a non-transitory storage medium to store or maintain instructions that, when executed by a machine, computing device, or system, cause the machine, computing device, or system to perform the methods and / or operations according to the described examples. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. The instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a machine, computing device, or system to perform a particular function. The instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.

[0052] One or more aspects of at least one example can be implemented by representative instructions stored on at least one machine-readable medium that represent various logics within a processor, which, when read by a machine, computing device, or system, cause the machine, computing device, or system to fabricate the logic to perform the techniques described herein. Such representations, referred to as “IP cores,” can be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to be loaded into a fabrication machine that actually fabricates the logic or processor.

[0053] The appearances of the phrase “one example” or “an example” do not necessarily all refer to the same example or embodiment. Any aspect described herein can be combined with any other aspect or similar aspect described herein, whether or not those aspects are described with respect to the same figure or element. The partitioning, omission, or inclusion of the block functions depicted in the figures does not infer that the hardware components, circuits, software, and / or elements for implementing those functions are necessarily partitioned, omitted, or included in an embodiment.

[0054] Some examples can be described using the terms “coupled” and “connected” and their derivatives. These terms are not necessarily synonyms of each other. For example, a description using the terms “connected” and / or “coupled” can indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “coupled” can also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0055] Terms such as “first,” “second,” etc. do not denote any order, quantity, or importance herein, but rather are used to distinguish one element from another. The term “a” herein does not denote a limitation on quantity, but rather denotes the presence of at least one of the items being referred to. The term “assert” when referring to a signal herein refers to a state of the signal in which the signal is valid, and that state can be achieved by applying any logic level (whether logic 0 or logic 1) to the signal. The term “subsequently” or “afterward” can mean immediately following or following after some other event or events. Depending on alternative embodiments, other sequences of steps can also be performed. Additionally, depending on the particular application, additional steps can be added or removed. Any combination of variations can be used, and those of ordinary skill in the art benefiting from this disclosure will understand many variations, modifications, and alternative embodiments thereof.

[0056] Unless otherwise specifically stated, disjunctive language such as the phrase "at least one of X, Y or Z" is understood within the context to generally state that an item, term, etc. can be X, Y or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language generally is not intended nor should it imply that certain embodiments require the presence of each of at least one X, at least one Y or at least one Z. Additionally, unless otherwise specifically stated, conjunctive language such as the phrase "at least one of X, Y and Z" should also be understood to refer to X, Y, Z or any combination thereof, including "X, Y, and / or Z".

Claims

1. A method for accelerating critical path instructions in a RISC-V processor, comprising: Perform critical path analysis on the program using the profile guided optimization (PGO) tool to identify critical path instructions; Recompiling the identified critical path instruction into a custom critical instruction with critical hint information using a compiler tool, wherein an opcode of the custom critical instruction is an opcode reserved for custom instructions by the RISC-V instruction set architecture, and wherein the opcode is used to indicate that the instruction is a critical instruction of a specific type; and The RISC-V processor implements accelerated execution of the custom key instruction based on the opcode and / or critical prompt information of the custom key instruction.

2. The method according to claim 1, wherein: The different opcodes reserved by the RISC-V instruction set architecture for custom instructions are used to compile different types of custom key instructions.

3. The method according to claim 1, wherein: Using the PGO tool to perform critical path analysis on a program to identify critical path instructions includes: Identifying a critical path and key events on the critical path based on the event dependency graph; and A critical path instruction that generates the critical event is determined based on the identified critical event.

4. The method according to claim 1, wherein: The Funct3 field of the custom key instruction is used to indicate the key prompt information.

5. The method according to any one of claims 1 to 4, wherein: The custom critical instruction is a custom critical branch instruction.

6. The method according to claim 5, wherein: The key prompt information of the custom key branch instruction is used to indicate that the instruction is one of the following items: General instructions; Key instructions for stable jumps; Critical instructions that are stable and do not jump; or Jump to unstable critical instructions.

7. The method according to any one of claims 1 to 4, wherein: The custom critical instruction is any one of the following: a custom load critical instruction, a custom multiply-add critical instruction, or a custom multiply-subtract critical instruction.

8. According to the method of claim 7, the key prompt information of the custom key instruction is used to indicate that the instruction is one of the following items: a standing order; or Key instructions.

9. The method according to any one of claims 1 to 4, wherein: Implementing accelerated execution of the custom key instruction by the RISC-V processor based on the opcode and / or critical prompt information of the custom key instruction includes: Predicting, by a prediction component of the RISC-V processor, a control flow and / or a data flow of the custom key instruction based on an opcode and / or critical hint information of the custom key instruction; and The buffer component and / or computing component of the RISC-V processor preferentially prepares corresponding hardware resources for the custom key instructions.

10. An apparatus for accelerating critical path instructions in a RISC-V processor, comprising means for performing the steps of the method according to any one of claims 1-9.

11. A computer-readable storage medium having instructions stored thereon, wherein the instructions, when executed by a processing circuit, cause the processing circuit to perform the method according to any one of claims 1-9.

12. A computer program product comprising instructions, wherein the instructions, when executed by a processing circuit, cause the processing circuit to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-core processor for execution of strands of instructions grouped according to criticality

    CN107567614A

  • Function as a service (FAAS) system enhancements

    CN112955869A