Intra-core instruction concurrent out-of-order execution system, method and equipment and medium
By introducing decoding identification module and transmission cache module into the SIMT architecture of GPGPU, concurrent and out-of-order execution of instructions in the kernel is realized, solving the problems of low computing efficiency and resource waste caused by instruction sequence in the SIMT architecture, and improving the utilization rate and computing efficiency of computing resources.
Patent Information
- Application Number
- CN202510051274.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-13
AI Technical Summary
In the classic SIMT architecture of GPGPU, the execution of instructions follows a strict order, resulting in pipeline stagnation, waste of computing resources and reduction of computing power caused by long-term instructions.
By adding a decoding identification module and a transmission cache module to the traditional GPGPU-SIMT architecture pipeline, instructions are divided into dependent execution and out-of-order execution. Dependency instructions must be executed sequentially. Out-of-order instructions can be executed across sequence and divided according to the calculation execution unit to realize concurrent out-of-order execution of instructions in the kernel.
It solves the problem of pipeline stagnation caused by long-term instructions, improves the utilization rate and computing efficiency of computing resources, and avoids waste of computing resources and reduction of computing power.
Smart Images

Figure CN119960835A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of instruction scheduling, and more specifically relates to a system, method, device and medium for concurrent out-of-order execution of instructions within a core. Background Art
[0002] With the advent of the information age, the scale and diversity of data are growing at an exponential rate, further triggering a sharp increase in the demand for data processing capabilities. Therefore, the development of high-performance computing chips has become a global research hotspot, among which general-purpose graphics processing unit (GPGPU) products are particularly prominent.
[0003] Currently, the architectural design of GPGPU products often achieves its excellent data processing capabilities by integrating multiple computing cores. Each computing core is equipped with a complete fetch, decode, issue, execute, and write-back pipeline, which forms the basis of an efficient execution unit. In addition, each execution unit further integrates a large number of dedicated computing units, which jointly support the implementation of the GPGPU single instruction multiple thread (SIMT) architecture. The SIMT architecture enables GPGPU to process a large number of threads in parallel under the guidance of a single instruction, thereby achieving high-concurrency computing of large-scale data sets.
[0004] However, in the classic SIMT architecture of GPGPU, the execution of instructions follows a strict order, and within a given computing core, all threads execute the same instructions in parallel. This execution model will cause the entire execution pipeline to stagnate when facing the execution of long-cycle instructions, thereby affecting the overall computing efficiency. In addition, since at any given moment, all threads are restricted to a single instruction and can only perform one computing operation, and the execution unit inside the GPGPU actually integrates multiple computing modules, this computing mode will result in only some of the computing modules in the GPGPU executing calculations at the same time, while other computing modules are idle, which will further lead to a waste of computing resources and a reduction in computing power. Summary of the invention
[0005] In view of the above problems, the purpose of the present invention is to provide a system, method, device and medium for concurrent out-of-order execution of instructions within the core, which realizes concurrent out-of-order execution of instructions within the core through sophisticated instruction classification, dependency management, flexible instruction execution strategy and efficient instruction processing flow, thereby improving computing efficiency and utilization of computing resources.
[0006] In order to achieve the above object, the present invention is implemented through the following technical solutions: In a first aspect, an embodiment of the present application provides a system for concurrent out-of-order execution of instructions within a core, comprising: Thread warp scheduling module, instruction fetch module, instruction decoding module, decoding identification module, emission cache module, operand collection module, calculation execution module and result write-back module; The thread warp scheduling module is used to maintain the register group of the number of thread warp slots, and to switch the computing core tasks at the thread warp granularity by sending the count value of the program counter to the instruction fetch module; An instruction fetch module is used to use the count value of the program counter as the instruction fetch PC, and use the instruction fetch PC as the address to read instruction data from the L1 cache outside the computing core, and then store the instruction data into the instruction cache inside the core; The instruction decoding module is used to read instruction data from the instruction cache in the core in a first-in-first-out manner, decode the instruction data into register information that can be recognized by the computing core, and send it to the decoding identification module; A decoding identification module is used to determine the classification of instructions in the core according to register information, classify the instructions in the core into five instruction types, classify and identify them according to the instruction types, generate identification instructions and send them to the emission cache module, and send corresponding control signals to ensure that the instructions are executed concurrently and out of order; the five instruction types include: jump instructions, memory access instructions, barrier instructions, dependent instructions and out of order instructions; A plurality of cache groups are provided in the transmitting cache module for storing identification instructions from the decoding identification module; the transmitting cache module is used to divide the cache groups into five types, store the identification instructions in the corresponding type of cache groups according to the instruction type of the identification instructions, and synchronously send the identification instructions in the cache groups to the calculation execution module for synchronous execution according to the storage state in the cache groups; the types of cache groups include out-of-order instruction cache groups, dependent shaping calculation cache groups, dependent floating-point calculation cache groups, dependent special function cache groups and memory access instruction cache groups; The operand collection module is used to maintain multiple groups of register values in units of threads, and read specific data from the register address corresponding to the corresponding thread according to the instruction information issued by the emission cache module; The calculation execution module is provided with four types of calculation execution units, which are used to synchronously receive the identification instructions issued by the transmission cache module, and synchronously execute various calculation task loads according to the instruction type of the identification instruction, and generate result data after the calculation is completed; the calculation execution unit includes an integer calculation unit, a floating point calculation unit, a special function calculation unit and a memory access calculation unit; The result writing module is used to receive the result data generated by the calculation execution module, and write it back to the corresponding position of the operand collection module according to the corresponding thread identifier and register address.
[0007] In an optional embodiment, the register information includes: 3 source operand register addresses, 1 destination register address, the calculation execution unit category required for the instruction data, the calculation type to be executed, the immediate number required for the calculation, and the instruction fetch PC of the current instruction required for the calculation.
[0008] In an optional implementation, the classifying and identifying instructions according to instruction types, generating identification instructions and sending them to the transmit cache module, and sending corresponding control signals to ensure that instructions are executed out of order concurrently, include: When the instruction type is a jump instruction, a pause instruction is sent to the instruction fetch module and the instruction decoding module to pause the execution of the instruction fetch module and the instruction decoding module; a clear signal is sent to the instruction cache in the core to clear all instruction information in the instruction cache in the core; When the instruction type is a memory access instruction, a dependency judgment is performed on the memory access instruction; if the memory access instruction does not have a dependency relationship, the memory access instruction is sent first; if a dependency relationship exists, the dependent superior instruction of the memory access instruction is sent first, and when all the dependent superior instructions are executed, the memory access instruction is sent first; When the instruction type is a barrier instruction, a pause instruction is sent to the instruction fetch module and the instruction decoding module to suspend the execution of the instruction fetch module and the instruction decoding module until all instructions and barrier instructions in the transmit cache modules are executed and then the execution of the instruction fetch module and the instruction decoding module is resumed.
[0009] In an optional implementation, the classifying and identifying instructions according to instruction types, generating identification instructions and sending them to the transmit cache module, and sending corresponding control signals to ensure that instructions are executed out of order concurrently, further includes: If the instruction type of the current in-core instruction is not a jump instruction, a memory access instruction, or a barrier instruction, the in-core instruction is regarded as a computing instruction; A dependency judgment is performed on the computing instruction; if the computing instruction has a dependency relationship with the instruction in the transmission cache module, the computing instruction is marked as a dependent instruction and sent to the transmission cache module, and the corresponding dependent superior instruction is identified; if the computing instruction has no dependency relationship with the instruction in the transmission cache module, the computing instruction is marked as an out-of-order instruction and sent to the transmission cache module.
[0010] In an optional implementation, the transmission buffer module is specifically used to: When the identification instruction sent by the decoding identification module is a jump instruction, a barrier instruction, and a disordered instruction, the identification instruction is placed in the highest bit of the disordered instruction cache group for maintenance; When the identification instruction sent by the decoding identification module is a dependent instruction, the dependent superior instruction of the dependent instruction is extracted from the out-of-order instruction cache group, and according to the calculation execution unit of the dependent instruction, the dependent superior instruction and the dependent instruction are sequentially stored in the corresponding instruction cache group for maintenance; When the identification instruction sent by the decoding identification module is a memory access instruction, determine whether the memory access instruction is a dependent memory access instruction; if the memory access instruction is a dependent memory access instruction, extract the dependent superior instruction of the memory access instruction from the out-of-order instruction cache group, and store the dependent superior instruction and the memory access instruction in the memory access instruction cache group in sequence; if the memory access instruction is not a dependent memory access instruction, put the memory access instruction in the lowest bit of the memory access instruction cache group for maintenance.
[0011] In a second aspect, an embodiment of the present application further provides a method for concurrent out-of-order execution of instructions within a core, comprising: The register information sent by the instruction decoding module is obtained through the decoding identification module, the digital register address, the destination register address, and the type information of the computing execution unit required by the instruction are extracted, and the extracted information is converted into binary information; According to the computing execution unit category information required by the instruction, the corresponding computing execution module identification bit in the instruction is set, and according to the register address information involved in the instruction execution, the corresponding register identification bit in the instruction is set; The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission cache module, identifies dependent instructions and out-of-order instructions based on the comparison results, and sends the instructions to the cache in the corresponding type of cache group based on the identification results.
[0012] When there are dependent instructions in the dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group in the emission cache module, the instructions in the emission dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group are preferentially synchronized and sent to the corresponding calculation execution unit of the calculation execution module for calculation; if there is no dependent instruction in any cache group among the dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group, the out-of-order instruction is extracted from the out-of-order instruction cache group and sent to the corresponding calculation execution unit of the calculation execution module for calculation.
[0013] In an optional implementation, setting a corresponding computing execution module identification bit in the instruction according to the computing execution unit category information required by the instruction, and setting a corresponding register identification bit in the instruction according to the register address information involved in the execution of the instruction, includes: If the calculation execution unit required by the instruction is an integer calculation unit, the calculation execution module identification position is set to 00; If the calculation execution unit required by the instruction is a floating point calculation unit, the calculation execution module identification position is set to 01; If the calculation execution unit required by the instruction is a special function calculation unit, the calculation execution module identification position is set to 10; If the computing execution unit required by the instruction is a memory access computing unit, the computing execution module identification position is set to 11; According to the register address information involved in the instruction execution, the corresponding register identification position in the instruction is set to 1.
[0014] In an optional implementation, the decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmit cache module, identifies dependent instructions and out-of-order instructions according to the comparison result, and sends the instruction to the cache in the corresponding type of cache group according to the identification result, including: The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission buffer module; If the same position is 1, the current instruction is a dependent instruction, and the instruction maintained in the emission cache module is a dependent superior instruction; the dependent superior instruction and the dependent instruction are sent together to the cache in the corresponding cache group according to the calculation execution module identification bit information; If there is no situation where all the same positions are 1, the current instruction is an out-of-order instruction, and the out-of-order instruction is sent to the cache in the out-of-order instruction cache group.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for concurrent out-of-order execution of instructions within a core as described in any one of the above items are implemented.
[0016] In a fourth aspect, an embodiment of the present application further provides a storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for concurrent out-of-order execution of instructions within a core as described in any of the above items are implemented.
[0017] It can be seen from the above technical solutions that the present invention has the following advantages: The present application adds a decoding identification module and a transmit cache module to the traditional GPGPU-SIMT architecture pipeline, and divides instructions into dependent execution and out-of-order execution. Dependent instructions must be executed sequentially, and out-of-order instructions can be executed across orders, thereby solving the pipeline stall problem caused by long-cycle instructions. At the same time, the present application uses the decoding identification module and the transmit cache module to further divide instructions according to the computing execution units, and the instructions of different computing execution units are executed synchronously, further improving the computing resource utilization efficiency in the GPGPU.
[0018] This application uses a decoding identification module to classify the instructions in the core into five types (jump instructions, memory access instructions, barrier instructions, dependent instructions and out-of-order instructions), and classifies and identifies them according to the instruction type and sends them to the corresponding transmit cache group. This fine instruction classification helps to optimize the execution order of instructions and reduce the waiting time between instructions, thereby improving computing efficiency.
[0019] The present application decodes and classifies instructions through a decoding identification module, and caches and synchronously executes instructions through a transmission cache module, thereby realizing an efficient instruction processing flow. This flow can reduce instruction processing time and waiting time, and improve computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0021] Figure 1 A schematic diagram of the structure of the concurrent out-of-order execution system of core instructions provided in this application.
[0022] Figure 2 A flowchart of the method for concurrent out-of-order execution of instructions within a core provided in this application.
[0023] Figure 3 A schematic diagram of the execution principle of the method for concurrent out-of-order execution of instructions within a core provided in this application.
[0024] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0025] Various embodiments of the present disclosure will be described more fully in the detailed description of the specific architecture and functions of the concurrent out-of-order execution system of the core instructions below. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.
[0026] Hereinafter, the terms "include" or "may include" used in various embodiments of the present disclosure indicate the presence of disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "include", "have", and their cognates are intended only to indicate specific features, numbers, steps, operations, elements, components, or a combination of the foregoing, and should not be understood as first excluding the presence of one or more other features, numbers, steps, operations, elements, components, or a combination of the foregoing or the possibility of adding one or more features, numbers, steps, operations, elements, components, or a combination of the foregoing.
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] See also Figure 1 The figure shows a system structure diagram of a concurrent out-of-order execution system of instructions in a specific embodiment. The overall system is a function addition to the traditional sequential execution SIMT architecture, and does not change the external interface content of the computing core, thereby realizing the transparency of the overall module to the upper layer of the computing core, that is, the system does not require the adaptation of software and the upper layer hardware of the computing core, and can be directly applied to the traditional GPGPU hardware implementation. The system includes: a thread warp scheduling module, an instruction fetch module, an instruction decoding module, a decoding identification module, a launch cache module, an operand collection module, a calculation execution module and a result write-back module.
[0029] The thread warp scheduling module is used to maintain the register group of the number of thread warp slots, and to switch the computing core tasks at the thread warp granularity by sending the count value of the program counter to the instruction fetch module.
[0030] The thread warp scheduling module is located at the top level of the pipeline in the computing core. It is used to maintain the register group of the number of thread warp slots and is responsible for the task switching of the computing core at the thread warp granularity. In this system, only when a specific computing execution unit of the computing execution module in the computing core is idle and the currently executed thread warp no longer has the work task of the computing execution unit, will the thread warp be switched to ensure the efficiency of the hardware resource utilization of GPGPU. Otherwise, the current thread warp will be executed first.
[0031] The instruction fetch module is used to use the count value of the program counter as the instruction fetch PC, and use the instruction fetch PC as the address to read instruction data from the L1 cache outside the computing core, and then store the instruction data in the instruction cache within the core.
[0032] For example, the instruction fetch module is responsible for completing the extraction of execution instructions in the current thread bundle. Specifically, this module is responsible for reading instruction data from the L1 cache outside the computing core using the instruction fetch PC as the address according to the instruction fetch PC sent by the thread bundle scheduling module, and then storing the instruction data in the instruction cache inside the core. In this system, the instruction fetch work of this module is to ensure the stability of the external interface signal of the computing core, so the instruction fetch is executed sequentially, that is, the instruction fetch PC is continuously accumulated. And the instruction fetch work of this module will only stop when the subsequent decoding identification module and the transmission cache module send a pause signal, otherwise it will continue. The setting of this mode further avoids the stagnation of the computing core instruction pipeline.
[0033] The instruction decoding module is used to read instruction data from the instruction cache in the core in a first-in-first-out manner, decode the instruction data into register information that can be recognized by the computing core, and send it to the decoding identification module.
[0034] For example, the instruction decoding module is responsible for reading instruction data from the instruction cache in the core, and then decoding the instruction data into various register values that can be recognized by the computing core, including 3 source operand register addresses, 1 destination register address, the execution unit computing module category required for the instruction, the calculation type performed by the computing module, the immediate number required for the calculation, and the current instruction PC required for the calculation. In this system, the instruction decoding module reads instruction data from the instruction cache in the core in a FIFO (first in, first out) manner for decoding, and the process will only stop when receiving the pause signal sent by the subsequent decoding identification module and the transmission cache module, otherwise it will continue.
[0035] The decoding identification module is used to determine the classification of the instructions in the core according to the register information, classify the instructions in the core into five instruction types, classify and identify them according to the instruction types, generate identification instructions and send them to the emission cache module, and send corresponding control signals to ensure that the instructions are executed concurrently and out of order; the five instruction types include: jump instructions, memory access instructions, barrier instructions, dependent instructions and out of order instructions.
[0036] A plurality of cache groups are arranged in the transmitting cache module for storing identification instructions from the decoding identification module; the transmitting cache module is used to divide the cache groups into five types, store the identification instructions in the corresponding type of cache groups according to the instruction type of the identification instructions, and synchronously send the identification instructions in the cache groups to the calculation execution module for synchronous execution according to the storage status in the cache groups; the types of cache groups include out-of-order instruction cache groups, dependent integer calculation cache groups, dependent floating-point calculation cache groups, dependent special function cache groups and memory access instruction cache groups.
[0037] The operand collection module is used to maintain multiple groups of register values in units of threads, and read specific data from the register address corresponding to the corresponding thread according to the instruction information issued by the emission cache module.
[0038] For example, the operand collection module maintains multiple sets of register values in units of threads, and is responsible for reading specific data from the corresponding register address of the corresponding thread according to the instruction information issued by the emission cache module. In this system, since the emission cache module will synchronously emit multiple instructions of different computing execution units, the registers of the operand collection module have multi-terminal read and write capabilities, that is, multiple sets of registers of a certain thread can be read and written at the same time.
[0039] The calculation execution module is provided with four types of calculation execution units, which are used to synchronously receive the identification instructions issued by the transmission cache module, and synchronously execute various calculation task loads according to the instruction type of the identification instruction, and generate result data after the calculation is completed; the calculation execution unit includes an integer calculation unit, a floating-point calculation unit, a special function calculation unit and a memory access calculation unit.
[0040] For example, the calculation execution module is a specific execution module for GPGPU in-core calculations, which is divided into four calculation execution units, namely, the integer calculation unit, the floating-point calculation unit, the special function calculation unit, and the memory access calculation unit. Each calculation execution unit has a hardware unit with the number of threads, and can execute the calculation operations of the number of threads simultaneously. In this system, the calculation execution module is different from the traditional GPGPU-SIMT architecture, which only receives one calculation instruction to execute one calculation operation. Instead, it will synchronously receive multiple calculation tasks and synchronously execute multiple calculation task loads.
[0041] The result writing module is used to receive the result data generated by the calculation execution module, and write it back to the corresponding position of the operand collection module according to the corresponding thread identifier and register address.
[0042] For example, the result write-back module is responsible for receiving the result data calculated by the calculation execution module, and writing it back to the corresponding position of the operand collection module according to the corresponding thread identifier and register address. In this system, since the calculation execution module will synchronously execute a variety of different calculation tasks, the result write-back module has the ability to write back multiple signals, that is, it can synchronously write back the result data of integer, floating point, special function, and memory loading to the operand collection module.
[0043] In this embodiment, by integrating modules such as thread warp scheduling, sequential instruction fetch and decoding, instruction classification identification, intelligent transmit cache management, multi-terminal read and write operand collection, and diversified synchronous computing execution, concurrent out-of-order execution of instructions within the core is achieved, which significantly improves the computing performance and resource utilization of GPGPU, while maintaining compatibility with existing software and hardware, and can be directly applied to traditional GPGPU hardware without additional adaptation.
[0044] In an embodiment of the present invention, based on the decoding identification module, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0045] The decoding identification module is a new module added by the present invention to ensure the efficient execution of concurrent out-of-order instructions. This module is responsible for receiving multiple register values sent from the instruction decoding module, and then classifying and identifying the instructions according to the corresponding register values. Specifically, the decoding identification module divides instructions into five types and performs different operations respectively. The five instruction types are jump instructions, memory access instructions, barrier instructions, dependent instructions and out-of-order instructions.
[0046] The decoding identification module can classify and identify instructions according to the instruction type, generate identification instructions and send them to the transmit cache module, and send corresponding control signals to ensure that instructions are executed concurrently and out of order. The specific implementation method is as follows: 1. When the instruction type is a jump instruction, a pause instruction is sent to the instruction fetch module and the instruction decoding module to pause the execution of the instruction fetch module and the instruction decoding module; a clear signal is sent to the instruction cache in the core to clear all instruction information in the instruction cache in the core.
[0047] For jump instructions, since the execution result of the jump instruction is to change the thread warp PC value in the top-level thread warp scheduling module of the pipeline, the execution of this instruction needs to ensure that all previous instructions are executed. That is, when the decoding identification module recognizes the jump instruction, it will send a pause instruction to the instruction fetch module and the instruction decoding module. While pausing the execution of the instruction fetch module and the instruction decoding module, it will send a clear signal to the instruction cache in the core to clear all instruction information in the instruction cache in the core.
[0048] 2. When the instruction type is a memory access instruction, a dependency judgment is performed on the memory access instruction; if the memory access instruction has no dependency relationship, the memory access instruction is sent first; if a dependency relationship exists, the dependent superior instruction of the memory access instruction is sent first, and when all the dependent superior instructions are executed, the memory access instruction is sent first.
[0049] When the identified instruction is a memory access instruction, since the memory access instruction is the instruction with the longest execution cycle in the entire GPGPU computing core, it needs to be sent and executed first. That is, when the decoding identification module identifies the memory access instruction, it will first determine the dependency of the memory access instruction. If the memory access instruction does not have a dependency relationship, the memory access instruction will be sent first; if there is a dependency relationship, the dependent superior instruction of the memory access instruction will be sent first, and the memory access instruction will be sent first after all the dependent superior instructions have been executed.
[0050] 3. When the instruction type is a barrier instruction, a pause instruction is sent to the instruction fetch module and the instruction decoding module to suspend the execution of the instruction fetch module and the instruction decoding module until all instructions and barrier instructions in the transmit cache modules are executed and then the execution of the instruction fetch module and the instruction decoding module is resumed.
[0051] When the recognized instruction is a barrier instruction, since the barrier instruction itself is a thread synchronization instruction defined by GPGPU, it is necessary to wait for all existing instructions to be executed before executing. That is, when the decoding identification module recognizes the barrier instruction, it will send a pause instruction to the instruction fetch module and the instruction decoding module to suspend the execution of the instruction fetch module and the instruction decoding module until all instructions and barrier instructions in the transmit cache module are executed and then resume the execution of the instruction fetch module and the instruction decoding module.
[0052] 4. If the instruction type of the current in-core instruction is not a jump instruction, memory access instruction, or barrier instruction, the in-core instruction is considered a computation instruction. A dependency check will be performed on the instruction later.
[0053] Specifically, when the decoding identification module recognizes a calculation instruction, it first performs a dependency judgment on the calculation instruction; if the calculation instruction has a dependency relationship with the instruction in the transmission cache module, the calculation instruction is marked as a dependent instruction and sent to the transmission cache module, and the corresponding dependent superior instruction is identified; if the calculation instruction has no dependency relationship with the instruction in the transmission cache module, the calculation instruction is marked as an out-of-order instruction and sent to the transmission cache module.
[0054] In an embodiment of the present invention, based on the transmit buffer module, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0055] The transmit cache module is a new module added by the present invention to ensure efficient execution of concurrent out-of-order instructions. The transmit cache module maintains n cache groups (n represents an unknown number and can be set freely) inside, which are responsible for storing the identification instructions sent from the decoding identification module. The n cache groups are divided into 5 categories as a whole, namely, out-of-order instruction cache group, dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group.
[0056] The specific functions of the launch cache module are as follows: When the identification instruction sent by the decoding identification module is a jump instruction, a barrier instruction, or an out-of-order instruction, the transmitting cache module puts the identification instruction into the highest bit of the out-of-order instruction cache group for maintenance.
[0057] When the identification instruction sent by the decoding identification module is a memory access instruction, it is first necessary to determine whether the memory access instruction is a dependent memory access instruction; if the memory access instruction is a dependent memory access instruction, extract the dependent superior instruction of the memory access instruction from the out-of-order instruction cache group, and store the dependent superior instruction and the memory access instruction in the memory access instruction cache group in sequence; if the memory access instruction is not a dependent memory access instruction, place the memory access instruction in the lowest bit of the memory access instruction cache group for maintenance.
[0058] When the identification instruction sent by the decoding identification module is a dependent instruction, the dependent parent instruction of the dependent instruction is extracted from the out-of-order instruction cache group, and according to the calculation execution unit of the dependent instruction, the dependent parent instruction and the dependent instruction are sequentially stored in the corresponding instruction cache group for maintenance. For example, when the identification instruction sent by the decoding identification module is a dependent instruction, the emission cache module extracts the dependent parent instruction of the dependent instruction from the out-of-order instruction cache group, and according to the calculation execution unit corresponding to the dependent instruction, the dependent parent instruction and the dependent instruction are sequentially stored in the corresponding integer / floating point / special function instruction cache group for maintenance.
[0059] It can be seen that in this system, since the launch cache module divides instructions into dependent instructions and out-of-order instructions, the dependent instructions and out-of-order instructions can be launched and executed synchronously, thus breaking the sequentiality of instruction execution in the original SIMT architecture in the GPGPU; at the same time, because the launch cache module also classifies instructions according to the computing execution unit that executes the instructions, the execution instructions of different computing execution units can be launched synchronously, thus ensuring the synchronous execution of all computing execution units in the GPGPU core at the same time, thereby improving the utilization efficiency of GPGPU core computing resources.
[0060] like Figure 2 As shown, the following is an embodiment of a method for concurrent out-of-order execution of instructions within a core provided by an embodiment of the present disclosure. This method and the concurrent out-of-order execution system for instructions within a core provided by the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the method for concurrent out-of-order execution of instructions within a core, reference can be made to the embodiment of the concurrent out-of-order execution system for instructions within a core provided by the above embodiments.
[0061] A method for concurrent out-of-order execution of instructions within a core comprises the following steps: S101: Obtain register information sent by the instruction decoding module through the decoding identification module, extract the data register address, the destination register address, and the type information of the computing execution unit required by the instruction, and convert the extracted information into binary information.
[0062] S102: According to the computing execution unit category information required by the instruction, the corresponding computing execution module identification bit in the instruction is set, and according to the register address information involved in the instruction execution, the corresponding register identification bit in the instruction is set.
[0063] For example, specifically, the decoding identification module first sets the corresponding computing execution module identification bit in the instruction according to the category information of the computing execution unit required by the instruction: If the integer calculation unit calculation is required, the corresponding calculation execution module identification position is 00; if the floating-point calculation unit calculation is required, the corresponding calculation execution module identification position is 01; if the special function calculation unit calculation is required, the corresponding calculation execution module identification position is 10; if the memory access calculation unit calculation is required, the corresponding calculation execution module identification position is 11.
[0064] Then, according to the register address information involved in the instruction execution, the corresponding register identification bit in the instruction is set. For example, if the register address involved in the instruction execution is 0 / 1 / 3, the 0 / 1 / 3 position in the register identification bit is set to 1.
[0065] S103: The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission cache module, identifies dependent instructions and out-of-order instructions according to the comparison result, and sends the instructions to the cache in the corresponding type of cache group according to the identification result.
[0066] For example, the decoding identification module compares the register identification information of the current instruction with the instruction register identification information maintained in the emission cache module. If the same positions are all 1, the current instruction is a dependent instruction, and the instruction maintained in the emission cache module is a dependent superior instruction. Then the dependent superior instruction and the dependent instruction are sent to the corresponding cache group (dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group, memory access instruction cache group) according to the calculation module identification information; if the same positions do not have 1, the current instruction is an out-of-order instruction, and the out-of-order instruction is sent to the out-of-order instruction cache group for cache.
[0067] S104: When there are dependent instructions in the dependent integer calculation cache group, the dependent floating-point calculation cache group, the dependent special function cache group and the memory access instruction cache group in the emission cache module, the instructions in the emission dependent integer calculation cache group, the dependent floating-point calculation cache group, the dependent special function cache group and the memory access instruction cache group are preferentially and synchronously sent to the corresponding calculation execution unit of the calculation execution module for calculation; if there is no dependent instruction in any cache group among the dependent integer calculation cache group, the dependent floating-point calculation cache group, the dependent special function cache group and the memory access instruction cache group, then the out-of-order instruction is extracted from the out-of-order instruction cache group and sent to the corresponding calculation execution unit of the calculation execution module for calculation.
[0068] In this step, when the four dependent computing module cache groups (dependent integer computing cache group, dependent floating-point computing cache group, dependent special function cache group, and memory access instruction cache group) in the emission cache module all have dependent instructions, the instructions in the four dependent computing module cache groups are preferentially emitted synchronously. If there is a dependent computing cache module that does not have a dependent instruction, the out-of-order instructions of the corresponding computing module are selected from the out-of-order instruction cache group to ensure that four groups of instructions to be executed are successfully emitted at the same time.
[0069] The method for concurrent out-of-order execution of instructions within a core provided in this embodiment realizes efficient and flexible instruction execution through the fine processing of the decoding identification module and the intelligent scheduling of the transmission cache module.
[0070] This method firstly accurately extracts key information in the instruction, such as register address, destination register address and computing execution unit category, through the decoding identification module, and converts this information into binary form, laying a solid foundation for subsequent processing. On this basis, the corresponding computing execution module identification bit is set according to the computing execution unit category information required by the instruction, and the corresponding register identification bit is set according to the register address information involved in the instruction execution. This step ensures that the instruction can be accurately classified and identified.
[0071] Next, through the cooperation of the decoding identification module and the transmission cache module, the dependent instructions and out-of-order instructions are effectively distinguished. The dependent instructions are sent to the corresponding dependent cache group together with the parent instructions, while the out-of-order instructions are sent to the out-of-order instruction cache group. This mechanism not only ensures the sequential execution of instructions, but also makes full use of the idle time of computing resources, improving the overall computing efficiency.
[0072] In the instruction emission phase, this method gives priority to synchronously emit instructions in the four dependent computing module cache groups. When there is no dependent instruction in a dependent computing cache module, the out-of-order instruction of the corresponding computing module is selected from the out-of-order instruction cache group. This strategy ensures that four groups of instructions to be executed are successfully emitted at the same time, maximizing the computing power of the computing execution unit and avoiding the waste of computing resources.
[0073] In summary, this method of concurrent out-of-order execution of instructions within the core achieves efficient utilization of computing resources and continuous operation of the instruction pipeline through sophisticated instruction classification and intelligent scheduling strategies, which significantly improves the computing performance and flexibility of GPGPU.
[0074] Further, as a refinement and expansion of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, take 10 specific instructions as an example, refer to Figure 3 The execution principle diagram shown in FIG. 1 provides another method for concurrent out-of-order execution of instructions within the core. The specific steps are as follows: S201: The decoding identification module receives instruction 0, identifies that its instruction execution calculation module is an integer calculation, and sets the calculation module identification to 00; identifies that the required register address is 0 / 1 / 3, and pulls up the register identification bit 0 / 1 / 3; identifies the instruction register identification information in the transmit cache module, because the transmit cache module is empty at this time, there is no dependency relationship for instruction 0, so instruction 0 is stored in the out-of-order execution cache group.
[0075] S202: The decoding identification module receives instruction 1, identifies that its instruction execution calculation module is an integer calculation, and sets the calculation module identification to 00; identifies that the required register address is 5 / 6 / 7, and pulls up the register identification bit 5 / 6 / 7; identifies the instruction register identification information in the emission cache module, because at this time the emission cache module only has the 0 / 1 / 3 register identification information of instruction 0, and there is no register conflict with instruction 1, so there is no dependency relationship for instruction 1, and instruction 1 is stored in the out-of-order execution cache group.
[0076] S203: The decoding identification module receives instruction 2, identifies that its instruction execution calculation module is an integer calculation, and therefore sets the calculation module identification to 00; identifies that the required register address is 2 / 4 / 8, and correspondingly pulls up the register identification bit 2 / 4 / 8; identifies the instruction register identification information in the transmit cache module, because at this time the transmit cache module has the register identification information 0 / 1 / 3, 5 / 6 / 7 of instruction 0 and instruction 1, and there is no register conflict with instruction 2, so there is no dependency relationship for instruction 2, and instruction 2 is stored in the out-of-order execution cache group.
[0077] S204: The decoding identification module receives instruction 3, identifies that its instruction execution calculation module is a floating-point calculation, and therefore sets the calculation module identification to 01; identifies that the required register address is 9 / 10 / 11, and accordingly pulls up the register identification bit 9 / 10 / 11; identifies the instruction register identification information in the emission cache module, because at this time the emission cache module has the register identification information 0 / 1 / 3, 5 / 6 / 7, 2 / 4 / 8 of instructions 0, 1, 2, and there is no register conflict with instruction 3, so there is no dependency relationship for instruction 3, and instruction 3 is stored in the out-of-order execution cache group.
[0078] S205: The decoding identification module receives instruction 4, identifies that its instruction execution calculation module is a special function calculation, and therefore sets the calculation module identification to 10; identifies that the required register address is 12 / 13 / 14, and correspondingly pulls up the register identification bit 12 / 13 / 14; identifies the instruction register identification information in the transmission cache module, because at this time the transmission cache module has the register identification information 0 / 1 / 3, 5 / 6 / 7, 2 / 4 / 8, 9 / 10 / 11 of instructions 0, 1, 2, 3, and there is no register conflict with instruction 4, so there is no dependency relationship for instruction 4, and instruction 4 is stored in the out-of-order execution cache group.
[0079] S206: The decoding identification module receives instruction 5, identifies its instruction execution calculation module as a memory access module, and sets the calculation module identification to 11; identifies the required register address as 15 / 18 / 19, and correspondingly pulls up the register identification bit 15 / 18 / 19; identifies the instruction register identification information in the transmission cache module, because at this time the transmission cache module has the register identification information of 0 / 1 / 3, 5 / 6 / 7, 2 / 4 / 8, 9 / 10 / 11, 12 / 13 / 14 of instructions 0, 1, 2, 3, 4, and there is no register conflict with instruction 5, so there is no dependency relationship for instruction 5, but because instruction 5 is a memory access instruction, instruction 5 is stored in the memory access instruction cache group.
[0080] S207: The decoding identification module receives instruction 6, identifies that its instruction execution calculation module is an integer module, and sets the calculation module identification to 00; identifies that the required register address is 2 / 16 / 17, and correspondingly pulls up the register identification bit 2 / 16 / 17; identifies the instruction register identification information in the transmission cache module, because at this time the 2 / 4 / 8 register identification information of instruction 2 in the transmission cache module conflicts with register No. 2 of instruction 6, so instruction 5 is a dependent instruction and instruction 2 is a dependent superior instruction. Subsequently, since instruction 6 is an integer instruction, instruction 2 and instruction 6 are sequentially stored in the dependent integer calculation cache group.
[0081] S208: The decoding identification module receives instruction 7, identifies that its instruction execution calculation module is a floating-point module, and sets the calculation module identification to 01; identifies that the required register address is 9 / 21 / 22, and correspondingly pulls up the register identification bit 9 / 21 / 22; identifies the instruction register identification information in the transmission cache module, because at this time the 9 / 10 / 11 register identification information of instruction 3 in the transmission cache module conflicts with register No. 9 of instruction 7, so instruction 7 is a dependent instruction and instruction 3 is a dependent superior instruction. Subsequently, since instruction 7 is a floating-point instruction, instruction 3 and instruction 7 are sequentially stored in the dependent floating-point calculation cache group.
[0082] S209: The decoding identification module receives instruction 8, identifies that its instruction execution calculation module is a special function calculation module, and sets the calculation module identification to 10; identifies that the required register address is 13 / 20, and correspondingly pulls up the register identification bit 13 / 20; identifies the instruction register identification information in the transmission cache module, because at this time the 12 / 13 / 14 register identification information of instruction 4 in the transmission cache module conflicts with register No. 13 of instruction 8, so instruction 8 is a dependent instruction and instruction 4 is a dependent superior instruction. Subsequently, since instruction 8 is a special function calculation instruction, instruction 4 and instruction 8 are sequentially stored in the dependent special function cache group.
[0083] S210: The decoding identification module receives instruction 9, identifies its instruction execution calculation module as a memory access module, and sets the calculation module identification to 11; identifies the required register address as 19 / 23, and correspondingly pulls up the register identification bit 19 / 23; identifies the instruction register identification information in the transmission cache module, because at this time the 15 / 18 / 19 register identification information of instruction 5 in the transmission cache module conflicts with register No. 19 of instruction 9, so instruction 9 is a dependent instruction and instruction 5 is a dependent superior instruction. Subsequently, since instruction 9 is a memory access instruction, instruction 5 and instruction 9 are sequentially stored in the memory access instruction cache group.
[0084] Figure 4 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0085] The method for concurrent out-of-order execution of instructions in the core provided by the embodiment of the present application can be applied to electronic devices. Those skilled in the art will appreciate that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.
[0086] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, buttons, a camera, a display, and a SIM card interface, etc.
[0087] The processor may include one or more processing units, for example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated into one or more processors.
[0088] The processor can be the nerve center and command center of the electronic device. The controller can generate an operation control signal according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0089] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. The memory may store instructions or data that the processor has just used or is cyclically used. If the processor needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor, and thus improves system efficiency.
[0090] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to implement data storage functions. For example, files such as music and videos can be saved in the external memory card.
[0091] The internal memory can be used to store computer executable program codes, which include instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0092] The wireless communication function of an electronic device can be realized through an antenna, a wireless communication module, a modem processor, and a baseband processor.
[0093] The wireless communication module can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0094] Electronic devices can implement audio functions, etc. through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.
[0095] Electronic devices can achieve shooting functions through ISP, camera, video codec, GPU, display and application processor.
[0096] Electronic devices can achieve display functions through GPU, display screen and application processor.
[0097] The GPU is a microprocessor for image processing that connects the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs that execute program instructions to generate or change display information.
[0098] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0099] The electronic device described above realizes the following technical effects of the method for concurrent out-of-order execution of instructions within the core of the present application: First, the instruction information is accurately extracted through the decoding identification module, and the instructions are classified and identified according to the calculation execution unit category and register address information. Subsequently, the emission cache module is used to intelligently distinguish dependent instructions from out-of-order instructions, and cache them in the corresponding cache groups respectively. In the instruction emission stage, the instructions in the four dependent calculation module cache groups are preferentially emitted synchronously. When a dependent cache group is idle, out-of-order instructions are emitted to ensure full utilization of computing resources. The above-mentioned electronic device not only improves the computing efficiency of GPGPU by executing the concurrent out-of-order execution method of intra-core instructions, but also enhances the flexibility of instruction execution, effectively avoiding idleness and waste of computing resources. In general, the above-mentioned electronic device performs fine instruction classification and intelligent scheduling by executing the concurrent out-of-order execution method of intra-core instructions, realizes efficient configuration and utilization of computing resources, and provides strong support for improving GPGPU performance.
[0100] The storage medium provided in the present application stores a program product that can implement a method for concurrent out-of-order execution of instructions within a core.
[0101] The concurrent out-of-order execution methods of instructions within the core include: The register information sent by the instruction decoding module is obtained through the decoding identification module, the digital register address, the destination register address, and the type information of the computing execution unit required by the instruction are extracted, and the extracted information is converted into binary information; According to the computing execution unit category information required by the instruction, the corresponding computing execution module identification bit in the instruction is set, and according to the register address information involved in the instruction execution, the corresponding register identification bit in the instruction is set; The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission cache module, identifies dependent instructions and out-of-order instructions based on the comparison results, and sends the instructions to the cache in the corresponding type of cache group based on the identification results.
[0102] When there are dependent instructions in the dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group in the emission cache module, the instructions in the emission dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group are preferentially synchronized and sent to the corresponding calculation execution unit of the calculation execution module for calculation; if there is no dependent instruction in any cache group among the dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group, the out-of-order instruction is extracted from the out-of-order instruction cache group and sent to the corresponding calculation execution unit of the calculation execution module for calculation.
[0103] In some possible implementations, the method for concurrent out-of-order execution of instructions within a core of the present disclosure may be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above “Exemplary Method” section of this specification according to various exemplary implementations of the present disclosure.
[0104] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0105] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A concurrent out-of-order execution system for instructions within a core, characterized in that: include: Thread warp scheduling module, instruction fetch module, instruction decoding module, decoding identification module, emission cache module, operand collection module, calculation execution module and result write-back module; The thread warp scheduling module is used to maintain the register group of the number of thread warp slots, and to switch the computing core tasks at the thread warp granularity by sending the count value of the program counter to the instruction fetch module; An instruction fetch module is used to use the count value of the program counter as the instruction fetch PC, and use the instruction fetch PC as the address to read instruction data from the L1 cache outside the computing core, and then store the instruction data into the instruction cache inside the core; The instruction decoding module is used to read instruction data from the instruction cache in the core in a first-in-first-out manner, decode the instruction data into register information that can be recognized by the computing core, and send it to the decoding identification module; A decoding identification module is used to determine the classification of instructions in the core according to register information, classify the instructions in the core into five instruction types, classify and identify them according to the instruction types, generate identification instructions and send them to the emission cache module, and send corresponding control signals to ensure that the instructions are executed concurrently and out of order; the five instruction types include: jump instructions, memory access instructions, barrier instructions, dependent instructions and out of order instructions; A plurality of cache groups are provided in the transmitting cache module for storing identification instructions from the decoding identification module; the transmitting cache module is used to divide the cache groups into five types, store the identification instructions in the corresponding type of cache groups according to the instruction type of the identification instructions, and synchronously send the identification instructions in the cache groups to the calculation execution module for synchronous execution according to the storage state in the cache groups; the types of cache groups include out-of-order instruction cache groups, dependent shaping calculation cache groups, dependent floating-point calculation cache groups, dependent special function cache groups and memory access instruction cache groups; The operand collection module is used to maintain multiple groups of register values in units of threads, and read specific data from the register address corresponding to the corresponding thread according to the instruction information issued by the emission cache module; The calculation execution module is provided with four types of calculation execution units, which are used to synchronously receive the identification instructions issued by the transmission cache module, and synchronously execute various calculation task loads according to the instruction type of the identification instruction, and generate result data after the calculation is completed; the calculation execution unit includes an integer calculation unit, a floating point calculation unit, a special function calculation unit and a memory access calculation unit; The result writing module is used to receive the result data generated by the calculation execution module, and write it back to the corresponding position of the operand collection module according to the corresponding thread identifier and register address.
2. The system for concurrent out-of-order execution of instructions in a core according to claim 1, characterized in that: The register information includes: 3 source operand register addresses, 1 destination register address, the calculation execution unit category required by the instruction data, the calculation type to be executed, the immediate number required for the calculation, and the instruction fetch PC of the current instruction required for the calculation.
3. The concurrent out-of-order execution system for core instructions according to claim 1, characterized in that: The classifying and identifying the instruction type, generating an identification instruction and sending it to the transmit cache module, and sending a corresponding control signal to ensure that the instructions are executed concurrently and out of order, include: When the instruction type is a jump instruction, a pause instruction is sent to the instruction fetch module and the instruction decoding module to pause the execution of the instruction fetch module and the instruction decoding module; a clear signal is sent to the instruction cache in the core to clear all instruction information in the instruction cache in the core; When the instruction type is a memory access instruction, a dependency judgment is performed on the memory access instruction; if the memory access instruction does not have a dependency relationship, the memory access instruction is sent first; if a dependency relationship exists, the dependent superior instruction of the memory access instruction is sent first, and when all the dependent superior instructions are executed, the memory access instruction is sent first; When the instruction type is a barrier instruction, a pause instruction is sent to the instruction fetch module and the instruction decoding module to suspend the execution of the instruction fetch module and the instruction decoding module until all instructions and barrier instructions in the transmit cache modules are executed and then the execution of the instruction fetch module and the instruction decoding module is resumed.
4. The system for concurrent out-of-order execution of instructions in a core according to claim 3, characterized in that: The method of classifying and identifying the instruction type, generating an identification instruction and sending it to the transmit cache module, and sending a corresponding control signal to ensure that the instructions are executed concurrently and out of order, further includes: If the instruction type of the current in-core instruction is not a jump instruction, a memory access instruction, or a barrier instruction, the in-core instruction is regarded as a computing instruction; A dependency judgment is performed on the computing instruction; if the computing instruction has a dependency relationship with the instruction in the transmission cache module, the computing instruction is marked as a dependent instruction and sent to the transmission cache module, and the corresponding dependent superior instruction is identified; if the computing instruction has no dependency relationship with the instruction in the transmission cache module, the computing instruction is marked as an out-of-order instruction and sent to the transmission cache module.
5. The system for concurrent out-of-order execution of instructions in a core according to claim 4, characterized in that: The transmission buffer module is specifically used for: When the identification instruction sent by the decoding identification module is a jump instruction, a barrier instruction, or a disordered instruction, the identification instruction is placed in the highest bit of the disordered instruction cache group for maintenance; When the identification instruction sent by the decoding identification module is a dependent instruction, the dependent superior instruction of the dependent instruction is extracted from the out-of-order instruction cache group, and according to the calculation execution unit of the dependent instruction, the dependent superior instruction and the dependent instruction are sequentially stored in the corresponding instruction cache group for maintenance; When the identification instruction sent by the decoding identification module is a memory access instruction, determine whether the memory access instruction is a dependent memory access instruction; if the memory access instruction is a dependent memory access instruction, extract the dependent superior instruction of the memory access instruction from the out-of-order instruction cache group, and store the dependent superior instruction and the memory access instruction in the memory access instruction cache group in sequence; if the memory access instruction is not a dependent memory access instruction, put the memory access instruction in the lowest bit of the memory access instruction cache group for maintenance.
6. A method for concurrent out-of-order execution of instructions within a core, characterized in that: The method adopts the concurrent out-of-order execution system of core instructions as claimed in any one of claims 1 to 5; The method comprises: The register information sent by the instruction decoding module is obtained through the decoding identification module, the digital register address, the destination register address, and the type information of the computing execution unit required by the instruction are extracted, and the extracted information is converted into binary information; According to the computing execution unit category information required by the instruction, the corresponding computing execution module identification bit in the instruction is set, and according to the register address information involved in the execution of the instruction, the corresponding register identification bit in the instruction is set; The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission cache module, identifies dependent instructions and out-of-order instructions according to the comparison result, and sends the instruction to the cache in the corresponding type of cache group according to the identification result; When there are dependent instructions in the dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group in the emission cache module, the instructions in the emission dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group are preferentially synchronized and sent to the corresponding calculation execution unit of the calculation execution module for calculation; if there is no dependent instruction in any cache group among the dependent integer calculation cache group, dependent floating-point calculation cache group, dependent special function cache group and memory access instruction cache group, the out-of-order instruction is extracted from the out-of-order instruction cache group and sent to the corresponding calculation execution unit of the calculation execution module for calculation.
7. The method for concurrent out-of-order execution of instructions within a core according to claim 6, characterized in that: The step of setting a corresponding computing execution module identification bit in the instruction according to the computing execution unit category information required by the instruction, and setting a corresponding register identification bit in the instruction according to the register address information involved in the execution of the instruction, includes: If the calculation execution unit required by the instruction is an integer calculation unit, the calculation execution module identification position is set to 00; If the calculation execution unit required by the instruction is a floating point calculation unit, the calculation execution module identification position is set to 01; If the calculation execution unit required by the instruction is a special function calculation unit, the calculation execution module identification position is set to 10; If the computing execution unit required by the instruction is a memory access computing unit, the computing execution module identification position is set to 11; According to the register address information involved in the instruction execution, the corresponding register identification position in the instruction is set to 1.
8. The method for concurrent out-of-order execution of instructions within a core according to claim 7, characterized in that: The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission cache module, identifies dependent instructions and out-of-order instructions according to the comparison result, and sends the instructions to the cache in the corresponding type of cache group according to the identification result, including: The decoding identification module compares the register identification bit information of the current instruction with the instruction register identification bit information maintained in the transmission buffer module; If the same position is 1, the current instruction is a dependent instruction, and the instruction maintained in the transmit cache module is a dependent superior instruction; the dependent superior instruction and the dependent instruction are sent together to the cache in the corresponding cache group according to the calculation execution module identification bit information; If there is no situation where all the same positions are 1, the current instruction is an out-of-order instruction, and the out-of-order instruction is sent to the cache in the out-of-order instruction cache group.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for concurrent out-of-order execution of instructions within a core as described in any one of claims 6 to 8 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for concurrent out-of-order execution of instructions within a core as claimed in any one of claims 6 to 8 are implemented.
Citation Information
Patent Citations
Instruction classified multi-emitting method based on SPRAC V8 instruction set
CN105426160A
Instruction execution method and device, electronic equipment and storage medium
CN112559040A
Instruction processing method and device, equipment and storage medium
CN117931293A
Cited By
Instruction processing device and method and related equipment
CN120762763A