Optimization Method for Multi-Core and Multi-Pipeline Parallel Execution in Graphics Processors

By splitting and executing instruction requests in a multi-core graphics processor, the problem of execution efficiency reduction caused by multi-core simultaneous transmission of the same type of instructions is solved, and higher performance and parallelism are achieved.

CN115640052BActive Publication Date: 2025-06-24JINLING INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211300379.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-06-24
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

In multi-core graphics processors, instructions of the same type need to be queued for execution in order of arrival, resulting in a decrease in execution efficiency and the inability to achieve full parallel and out-of-order execution.

Method used

By receiving and cacheing instructions in the instruction buffer, selecting instructions that can be executed out of order, splitting the instruction request and reading data, independently processing instruction source operand data reading, realizing out of order execution within the pipeline between different cores.

Benefits of technology

The utilization rate of pipeline control execution logic and the utilization rate of special mathematical function calculation units are improved, achieving higher performance and parallelism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115640052B_ABST
    Figure CN115640052B_ABST
Patent Text Reader

Abstract

The present invention provides an optimization method for multi-core and multi-pipeline parallel execution in a graphics processor, specifically a design method for a multi-stream processing core unit in a processor to simultaneously issue multiple instructions into a pipeline for out-of-order execution. The simultaneous issue of multiple instructions by multiple cores includes issuing multiple independent instructions from single instruction multiple data threads (SIMD threads) executed by multiple cores in the GPU processing core part into the pipeline for out-of-order execution. The instructions are further segmented at the instruction granularity and executed at an executable request granularity smaller than the instruction, making full use of the performance of each data path to parallelly read data, so that the arithmetic logical unit (ALU) or execution unit inside the pipeline is always in a busy working state, thereby achieving out-of-order parallel execution of instructions and improving the execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an optimization method for multi-core and multi-pipeline parallel execution in a graphics processing unit. Background Art

[0002] In current graphics processing unit (GPU) designs, in order to achieve a more parallel effect, a multi-pipeline structure is usually adopted. In addition to integer and floating-point instruction calculation units, there are also various other instruction processing units for parallel execution. These instruction pipelines can be divided into multiple types such as texture sampling unit (Sample) instructions, data load and store unit (Load / Store) instructions, special mathematical function (SFU) instructions, and shared memory unit access (Shared Memory) instructions. After being emitted from the instruction emission unit, each pipeline independently reads relevant data and performs relevant operations.

[0003] However, in a multi-core design, usually multiple cores may simultaneously emit the same type of single instruction multiple data (SIMD) processing instructions and enter the instruction designated module for execution. The instruction execution pipeline also usually executes according to the instruction granularity. Under the limitation of data transmission bandwidth, the instruction needs to collect all source operand data before the entire instruction can start execution. Since the GPU itself adopts a multi-core structure design and in order to save the huge hardware overhead cost of the entire execution pipeline, there is only one set of such execution pipelines for a single core in graphics processing. Therefore, it is necessary to block the upstream instruction emission module and wait for the current instruction to complete execution before the blocked instructions such as Figure 1 shown.

[0004] At this time, if the data reading of the instruction source operands for the currently executed instruction has not been completed, or arbitration during the reading of the source operands causes a lag and other reasons, and it is not ready in time, it will cause the pipeline to wait all the time, and other cores cannot emit this type of instruction to execute either.

[0005] Therefore, the execution of this instruction unit will become a serial structure, as shown in Figure 2 the serial execution pipeline processing effect display, which describes the multi-stream processing core instruction execution process from emission - obtaining operands - to execution. The green area is the part for continuously sending requests for reading operand data, and the red area is the part for executing instructions. Here, when the execution enters the pipeline, there is no longer any blocking behavior for the subsequent execution part. Therefore, it can be assumed that it is completed when sent out, and there is no need to count cycles according to the length of the entire pipeline. The horizontal axis in the figure represents the execution cycle duration, and the vertical axis shows the execution of 4 stream processors (SP). In the figure, at different cycles on the horizontal axis, the instruction execution processes continuously and alternately emitted by 4 cores are shown. By Figure 2It can be seen that the operand fetch operation and the pipeline execution part need to wait for each other, and it is impossible to achieve fully parallel execution, let alone out-of-order execution.

[0006] The current method of polling reception and sequential execution mainly has the following problems:

[0007] First, after a single execution pipeline receives an instruction, it will become sequential execution, resulting in the serialization of the process of reading source operands and instruction execution, and the execution efficiency will decrease.

[0008] Second, for data without dependencies sent from multiple cores, sequential waiting for processing is required, and the waiting leads to a decrease in performance.

[0009] Third, currently, it is basically executed in the order in the reception queue, resulting in the execution of complete instructions, with a large execution granularity and unable to achieve the effect of fully out-of-order execution. Summary of the Invention

[0010] Object of the Invention: The object of the present invention is to solve the problem that when multiple cores simultaneously issue the same type of instructions and enter the corresponding pipelines for execution, they need to queue up and wait for execution in the order of arrival, and provide a design for out-of-order execution within the pipeline between different cores. The reading of source operand data of instructions from different cores should be independent of each other, and there is no problem of read / write data conflict. Therefore, the out-of-order parallel execution method can enable instructions simultaneously issued by multiple cores or instructions without dependency relationships between different threads (SIMD threads) issued by a single core to be executed out of order when the relevant resources are ready, improving the utilization rate of the pipeline control execution logic and the utilization rate of some special mathematical function calculation units, and achieving better performance.

[0011] The present invention specifically provides an optimization method for multi-core multi-pipeline parallel execution in a graphics processor. This method can convert the processing of single-instruction granularity execution into a parallel execution process of executable small-granularity requests. The feature of parallel processing lies in the design of smaller execution granularity of the execution pipeline inside the processor, including the following steps:

[0012] Step 1: The instruction buffer receives and caches instructions;

[0013] Step 2: Select instructions that can be executed out of order;

[0014] Step 3: Split the instruction request and read data;

[0015] Step 4: Receive the data returned by the read request and store it;

[0016] Step 5: Control the executable small-granularity requests;

[0017] Step 6: Extract instruction information;

[0018] Step 7, execute an executable small-grained request;

[0019] Step 8, output the result of the executed completed request;

[0020] Step 9, clear the executed completed instruction.

[0021] Step 1 includes: The parallel pipeline first needs to poll and receive the same external processing unit-related processing instructions (non-integer and non-floating-point processing instructions, generally a type of instruction processing pipeline shared by multiple streaming cores, such as the instruction pipeline for loading data) simultaneously sent from N streaming cores of the graphics processor (usually 2 or 4 streaming cores, each streaming core has an independent instruction emission system and can emit more than a dozen types of instructions such as floating-point instructions, integer instructions, special function instructions, texture sampling instructions, etc. Here, except for floating-point instructions, integer instructions, etc. which are usually in the streaming core, others can be outside the stream processor. And the method of the present invention is used in the instruction pipeline outside the stream processor), and write the received instructions into the instruction buffer dedicated to storing instruction information. If the streaming core does not emit instructions, skip the processing of the streaming core and this execution pipeline; when the operation of writing the instructions into the buffer is completed, obtain the instruction entry as an index to look up the specific instruction information. At the same time, put the instruction index number, the single instruction multiple data thread number being executed, the streaming core number, and the priority information into the instruction retrieval queue.

[0022] Step 2 includes: The instruction selector of the instruction pipeline selects the instructions corresponding to each independent streaming core from the instruction retrieval queue according to the priority information, streaming core number, and the single instruction multiple data thread number carried by the instructions. The single instruction multiple data threads being executed from different streaming cores can be preferentially selected to obtain the instruction index number.

[0023] Step 3 includes: According to the obtained instruction index information, read the instructions in the instruction pipeline and split them internally, split the instructions into executable small-grained requests, and send the streaming core number, the single instruction multiple data thread number being executed, the instruction index number, the instance start information and the instance quantity information, the source address information, the request index in the instruction, and the request executable flag to the data memories in each independent core to obtain the corresponding operand data.

[0024] Step 4 includes: The instruction request splitting unit in the instruction pipeline is mainly used to split the instruction for reading source operands that needs to be executed into executable granularities and apply for the location address of the available storage area in the data buffer. When the instruction request splitting unit splits the request, it applies to the data buffer for the location address of the return data to be stored. When the read data request returns, it receives and writes the data into the data buffer for storage according to the location address of the return data to be stored; after obtaining the operands, the order in which each core reads the return of the general data register will be disordered, and at this time, the data is written into the data buffer in a polling manner. At the same time, it checks the request executable flag bit. If the request executable flag carried in the returned executable small-granularity request is valid, it is sent to the request controller of the execution pipeline, and the request controller of the execution pipeline will control the execution of the small-granularity request according to the returned ready status information.

[0025] Step 5 includes: For the ready status of two or more executable small-granularity requests, the instruction pipeline enters the ready queue in the order of readiness. The executable small-granularity request control sequentially reads the requests with the ready status from the ready queue and sends them to the instruction information extraction module unit and the data buffer module to read relevant information (according to the instruction index, find the instruction entry, and then extract the corresponding instruction information required during execution according to the required information). The instruction information extraction module unit is mainly used to extract the valid information of the currently executing instruction from the instruction buffer. According to the instruction index, find the instruction entry, and then extract the corresponding instruction information required during execution according to the required information. The data buffer module is mainly used to manage the allocation and release of the data buffer area.

[0026] Step 6 includes: When executing the corresponding small-granularity request, the request controller will send the information to the instruction information extraction module unit to extract the execution status data. By returning the carried instruction index, it extracts the valid instruction request in the instruction buffer, and then extracts the valid instance mask and instruction opcode of the corresponding request according to the request instance start position and the number of instances, and then sends them together with the source data read by the request control to the request execution subunit for actual execution.

[0027] Step 7 includes: The request execution subunit is the execution main unit of the executable small-granularity request, which contains an arithmetic logic unit related to the specific functions of the pipeline required for execution calculation. For example, when performing a calculation as a special mathematical function, when calculating the reciprocal operation, the calculation ALU that can be calculated by the third-order look-up table method of the Taylor formula can be used. When the request execution subunit collects the opcode, valid instance mask, source data, and status information of the instruction execution, it starts to perform the calculation of the pipeline itself to obtain the calculation result.

[0028] Step 8 includes: Based on the status type of the instruction destination address information, the calculation result of the pipeline determines whether to write back to the general-purpose data register file or output it to the specified buffer storage area for storage.

[0029] Step 9 includes: When the execution request ends, check the current instruction to confirm whether all requests have been executed. If not, continue to loop and execute each small-granularity request; if completed, execute the corresponding instruction entry in the instruction retrieval queue, and move the subsequent instruction index entry down to fill and keep the continuous area available. At the same time, clear the corresponding instruction entry in the instruction buffer to receive the subsequent instructions emitted.

[0030] The advantages and remarkable effects of the present invention are as follows: (1) Multiple instructions emitted by multiple cores enter different pipelines for execution, and are completely split into compatible execution granularities inside the pipeline, and are respectively issued out of order to read the relevant instruction source data resources inside and outside the core. (2) When the instruction source data resources of any executable granularity are ready, and the last signal of the last request of the small-granularity request is ready, it means that the request can enter the request controller, and then enter the execution pipeline to actually execute relevant ALU or conversion processing operations. (3) During this process, different selected instructions will be split into different small-granularity requests for execution. At this time, the small-granularity requests of different instructions are out of order and interleaved. The retrieved instruction source data also enters the execution pipeline out of order for execution, keeping the execution pipeline busy all the time, improving the utilization rate of the execution pipeline, and also improving the parallelism between different cores, achieving the purpose of out-of-order execution. (4) It no longer depends on the emission granularity of the small-granularity requests of instructions and execution, which can increase the utilization rate of the data buffer, and at the same time can achieve the effect of Vertical to Horizontal. Brief Description of the Drawings

[0031] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0032] Figure 1 It is a schematic diagram of multiple-core emitted instructions entering the same pipeline designed for the present invention, and a polling strategy is adopted to alternately emit data into the pipeline for execution.

[0033] Figure 2 It is a schematic diagram of the traditional serial design for reading source operands and the execution effect of the pipeline. The green area is the number of cycles for the read operand request, and the red area is the number of cycles for the pipeline to execute instructions. The connected red and green parts are the complete life cycle of an instruction.

[0034] Figure 3 It is a schematic diagram of the out-of-order execution of the execution pipeline of the present invention.

[0035] Figure 4 This is the complete processing flowchart of the executable instructions of the present invention.

[0036] Figure 5 This is a schematic diagram of the selection of executable instructions of the present invention.

[0037] Figure 6 This is a schematic diagram of splitting the request of a single-channel executable instruction of the present invention.

[0038] Figure 7 This is a schematic diagram of splitting the request of a multi-channel executable instruction of the present invention.

[0039] Figure 8 This is a schematic diagram of the comparison of the execution effects between the present invention and the traditional design.

[0040] Figure 9 This is a schematic diagram of the instruction pipeline processing and execution process. Detailed implementation manners

[0041] The present invention proposes an out-of-order execution optimization method for multi-core and multi-pipeline parallel execution in a graphics processing unit (GPU), specifically a design method for multi-cores to simultaneously issue multiple instructions into the pipeline for out-of-order execution. Generally, each stream processing core issues texture sampling instructions into the texture processing unit (usually a dedicated texture processing unit in each graphics processing unit) for execution. The texture sampling instructions are mainly used to obtain image data at a specified position from a texture image given by an application program according to the coordinates of pixel points, perform sampling, linear interpolation, filtering processing, etc., and return the sampled result to the general data register for program texture mapping, etc. If processing this instruction according to a three-dimensional image, it is necessary to collect the three-dimensional coordinate address information of all instances included in the entire instruction before starting execution. In the present invention, its granularity is divided into smaller execution granularities for execution, and when the three-dimensional coordinate address information of the executable granularity is collected, execution can start. The multi-cores simultaneously issuing multiple instructions includes multiple instructions without dependency relationships emitted by SIMD threads of multiple cores in the GPU processing kernel part entering the pipeline for out-of-order execution. The instruction granularity is further divided, and execution is performed according to an executable request granularity smaller than the instruction granularity, making full use of the performance of each data path to parallelly read data, so that the ALU (Arithmetic Logical Unit) or execution unit inside the pipeline is always in a busy working state, thereby achieving out-of-order parallel instructions, improving the execution efficiency, and processing as Figure 3 shown.

[0042] First, inside the execution pipeline, an instruction buffer is used to receive requests for instruction emission and other related instruction information sent from multiple cores, including various essential information for instruction execution such as the source operand data address, instruction opcode, and single-instance valid processing mask in SIMD instructions. Each instruction item / entry in the instruction cache corresponds to an instruction. When looking up or using instruction information, it is indexed and searched through the internal instruction index. During subsequent execution, when instruction information needs to be used or data needs to be written out after execution completion, relevant instruction information can be obtained through this internal instruction index, including information such as the instruction destination data and destination address type, valid data mask, and single-instance valid processing mask.

[0043] Design an instruction selector, as Figure 5 shown. Select different instructions according to the threadid, priority, and source data dependencies, and disassemble them into requests for execution. The conditions for instruction selection are as follows: First, select according to priority. If it is a specified high-priority instruction, it needs to be selected and executed first. Second, select according to different cores because there are fewer data dependencies between different cores. Third, select according to different threadids. Finally, for different instructions within the same core and the same thread, they need to be strictly split and emitted according to the instruction last valid instance signal carried by the instruction itself to prevent errors caused by the subsequent instruction completely exceeding the completion of the previous instruction.

[0044] The selected instructions will be disassembled here according to the execution granularity of different pipelines and the requirements for the number of instruction source data channels, and disassembled into multiple request requests, as Figure 6 shown. The disassembled request requests are sent to the General Register File to read the instruction source data, along with the request index and the execute granularity last request execution termination signal for this execution granularity. For sample and load / store instructions, according to the different data or address channels, the source data from multiple channels needs to be read and converted from vertical format to horizontal format (Vertical to Horizontal). When the request data is read back, it is written into the specified data buffer memory for caching, as Figure 3 shown in the data buffer memory part.

[0045] When the data buffer memory receives the data returned from the general data register file read, along with the last data marker signal with an executable granularity request, it indicates that all the data for the executable granularity request has been prepared and can then be executed.

[0046] At this time, it will be sent to the request controller to notify it to prepare for execution. Then the request controller will obtain the instruction information from the instruction buffer through the instruction index and extract information such as the valid mask for all instances corresponding to the request, and send it to the request execution subunit for partial execution.

[0047] After the execution is completed, the result data and the information related to the destination address will be written to the destination address location of the resource specified by the destination data type.

[0048] During this process, the instructions extracted from the instruction buffer, after being disassembled, request data resources from different in-core resources, and the return time sequence is inconsistent, and it is also out-of-order interspersed between instructions. However, the execution resources within an instruction are guaranteed to be in order according to the order in which various resources are received. Therefore, the requests within an instruction must be executed in order.

[0049] Such as Figure 8 shown, the upper part is the traditional serial execution process, which is executed at the instruction granularity and requires the data to be prepared before starting the execution, resulting in an instruction serial structure in the pipeline. The lower part is the execution process of the present invention, where the green part is the single read operand request split out, the red area is the process of a single execution request, and the overlap of green and red is the simultaneous execution of these two operations. By adopting the above technical solution, the present invention can achieve the technical effects shown in the lower part of Figure 8 The processing granularity of the present invention becomes smaller, and there is information on caching and reading data at the same time. Therefore, the read operation data request part in the execution part is carried out simultaneously. To better show its parallel process, the green area of the read operand part and the red area of the execution part are divided into two lines for display. During this process, mainly the starting part of the read operand part will take time. When four cores read operands simultaneously, the part of reading operands will save a lot of time. At this time, each core reads and returns enough operands to saturate the execution pipeline, ensuring the efficiency of the execution pipeline. Compared with the upper half of the prior art in Figure 8 the entire life cycle of the execution of the same number of the same instructions is greatly reduced compared to the traditional method.

[0050] Since the traditional design method is to execute only after collecting all the data of a single instruction, therefore Figure 8The upper-middle continuous large green area is for reading and collecting the operand data sent out by the entire instruction. When the operands are ready, the entire instruction behavior starts to execute. Therefore, the resulting effect is that the continuous green area plus the red area represents the life cycle of the execution of an entire instruction. When the first green area plus a red area is executed, it indicates that an instruction has been executed. At this time, the next instruction will be switched to for execution. When multiple cores emit instructions, a polling strategy is adopted. Therefore, Figure 8 The black rectangular area marks the process in which the second instruction is emitted from core 1 and fully executed. Continuing in this way, the number of cycles finally used in the horizontal area is the duration used to execute the current continuous multiple instructions. Figure 8 It shows the process of alternately executing a total of 10 instructions.

[0051] It should be noted that the number of cycles for each execution of 10 instructions is not necessarily the same. It may be Figure 8 around 90 cycles as shown in

[0052] This is determined by the different randomly generated instruction channels. If it is a single channel, the execution may be at a granularity of 16 instances. For 32 parallel processed data in the SIMD32 state, it is read 2 times, waiting for 1 cycle for the data to return in the middle (waiting for the last read data request of this instruction to return and complete), and executed 2 times. It can be completed in a total of 5 times. However, if there are multiple channels and a large amount of data, the granularity may be 4 instances at this time. For 32 parallel processed data in the SIMD32 state, it is read 8 times, waiting for 1 cycle for the data to return in the middle, and executed 8 times. It takes a total of 17 times to complete. Therefore, for the traditional serial execution part, different instructions may have green areas of different lengths. Figure 8 As shown in the lower part of the area, in the design solution of this article, since the processing granularity becomes smaller and the information of each core's data is cached at the same time, the read operation data requests in the execution part are carried out simultaneously. To better show the process, the green area for reading operands and the red area for execution are respectively divided into small-granularity reading operands and two rows for displaying small-granularity requests. For each instruction emitted by each stream processor core, it must first execute the part of reading the source operands. At this time, the green area in the lower part of each stream processor core is first presented. Then, after performing a fetch operation for a period of time, the situation of the red execution part and the green data fetch part overlapping in cycles on the horizontal axis starts to appear. This is because after dividing into smaller execution granularities, as Figure 8The read operand Req_3 of the lower part streaming processing core 0 and the execution of part Req_0 occur simultaneously, achieving a truly parallel effect without waiting for the read operand to be completely finished (as shown in Table 1 specifically). During this process, the part of each streaming processing core that emits instructions mainly takes time in the small part of the previous read operand part. When four cores read operands simultaneously, the cycle part of reading operands will overlap a large number of cycles with the execution part of the execution pipeline, which is the effect of hiding latency and reduces a large amount of time. At this time, the present invention designs a solution with small execution granularity. Adding up the red areas on the horizontal axis should be close to a continuous red area, indicating that there are enough operands at this time to make the execution pipeline execute saturatedly. Therefore, the entire life cycle of the execution of the same number of the same instructions is greatly reduced compared to the traditional method.

[0053] Table 1

[0054]

[0055] Embodiment

[0056] The processing process of out-of-order execution applied in the present invention can be used for single-operand instructions and multi-operand instructions. Except for the small-granularity request splitting and different execution granularities, the overall processing processes are similar, as Figure 4 shown, including the following steps:

[0057] Step 1: The instruction buffer receives instructions;

[0058] Step 2: Select instructions that can be executed out of order;

[0059] Step 3: Split instruction requests and read data;

[0060] Step 4: Receive the data returned by the read request and store it;

[0061] Step 5: Control of executable small-granularity requests;

[0062] Step 6: Extract instruction information;

[0063] Step 7: Execute executable small-granularity requests;

[0064] Step 8: Output the results of the requests with completed execution;

[0065] Step 9: Clear the instructions with completed execution.

[0066] Step 1 includes: The execution pipeline first needs to poll and receive the instructions sent simultaneously from multiple cores, and write the received instructions into the instruction buffer. If a corresponding core does not send out the instruction, this core and the current execution pipeline path will be skipped. When the operation of writing the instruction into the buffer is completed, the instruction entry is obtained as an index to look up the specific information of the instruction, which is convenient for subsequent extraction of relevant information and execution. At the same time, the instruction index number, the single instruction multiple data thread number being executed, the stream processing core number, and the priority information are put into the instruction retrieval queue.

[0067] Step 2 includes: The instruction selector selects the oldest instruction corresponding to each independent stream processing core from the instruction retrieval queue according to the priority information, the stream processing core number, and the single instruction multiple data thread number carried by the instruction. Different single instruction multiple data thread numbers being executed can be processed in parallel, so the instruction index can be obtained.

[0068] Step 3 includes: According to the obtained instruction index information, the instruction is read and split internally. These instructions are split into small-granularity requests, and are sent to data memories such as the general data register file in each independent core respectively through the stream processing core number, the single instruction multiple data thread number being executed, the instruction index number, the instance start information and the instance quantity information, the source address information, the request index in the instruction, and the request executable flag to obtain the corresponding operand data.

[0069] Step 4 includes: The instruction request splitting unit also applies to the data buffer for the position address where the returned data needs to be stored at the same time. When the data is returned, it is received and written into the data buffer according to this address. After obtaining the operands, the order of the general data registers returned by each core will be disordered. At this time, it is written into the data buffer in a polling manner, and the request executable flag bit is also checked. At this time, if the request executable flag carried in the returned small-granularity request is valid, the request controller will control the execution of the returned ready-state request.

[0070] Step 5 includes: There may be multiple ready states at the same time. At this time, they enter the ready queue in the order of readiness. The executable small-granularity request control reads the ready-state requests from it in turn, and sends them to the instruction information extraction module and the data buffer module to read the relevant information.

[0071] Step 6 includes: When the corresponding small-granularity request is executed, the information is sent to the information extraction unit to extract the execution status data. Through the instruction index carried in the return, the valid instruction requests in the instruction buffer are extracted. Then, according to information such as the request instance start position and the instance quantity, the valid instance mask corresponding to the request, as well as key information such as the instruction opcode, are sent to the request execution subunit together with the source data read by the request control for actual execution.

[0072] Step 7 includes: The request execution subunit collects the operation code, valid instance mask, source data, and status information of the instruction execution, starts to execute the relevant ALU calculation, and obtains the calculation result.

[0073] Step 8 includes: The calculation result will be written back to the general data register file or output to a specified buffer for storage according to the instruction status.

[0074] Step 9 includes: When the execution request ends, check the current instruction to confirm whether all requests have been executed. If not, continue to loop and execute each executable small-grain request. If completed, execute the corresponding entry in the instruction retrieval queue, move the subsequent instruction index entry down, and fill the continuous area to keep it available. At the same time, clear the corresponding instruction entry in the instruction buffer for subsequent instruction reception.

[0075] In this embodiment, when executing an instruction with multiple operands, a smaller-grain request splitting method will be adopted for execution. At this time, in step 3, it is divided according to the grain size of the executable request. The multiple operands will complete the conversion from Vertical to Horizontal in the data reception buffer. For texture sampling processing instructions or data load and storage type instructions, there is no dependence between the instructions sent from multiple cores received, and for the data requests of multi-channel source data, the Vertical-to-Horizontal operation can be performed in the data buffer, and multiple requests will be issued. When the data request of the last channel comes back, there may be multiple executable grains for execution. At this time, just process these ready resources in order, as Figure 7 shown.

[0076] As Figure 9 shown, the flowchart of the execution process of the specific instructions in this embodiment after being split into each small-grain processing request. A single instruction is split into multiple small-grain requests. When each executable small-grain request is ready, it will directly enter the request controller for processing, and when the instruction information is ready, it will be sent to the execution pipeline for execution.

[0077] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the present invention for the multi-core multi-pipeline parallel execution optimization method in the graphics processor and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0078] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in a storage medium, including several instructions for causing a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MUU or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0079] The present invention provides an optimization method for parallel execution of multiple pipelines of multiple stream processing cores in a graphics processor. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. An optimization method for multi-core and multi-pipeline parallel execution in a graphics processing unit, characterized in that, Applied to the instruction execution pipeline in a graphics processing unit, where multiple stream processing cores in the graphics processing unit share the instruction execution pipeline, including the following steps: Step 1: The instruction buffer receives and caches instructions; Step 2: Select instructions that can be executed out of order. The selection of instructions that can be executed out of order includes: Through an instruction selector, instructions are sequentially selected in the order of the priority of the instructions, instructions of different cores, instructions of different thread identifiers threadid, and different instructions of the same core and the same thread; Step 3: Split the instruction request and read data according to the execution granularity of different pipelines and the requirements of the number of instruction source data channels. Among them, the read request for reading data carries the last data marker signal of the executable granularity request; Step 4: Receive the data returned by the read request and store it; Step 5: In response to the fact that the data returned by the read request carries the last data marker signal of the executable granularity request, control the executable small granularity request; Step 6: Extract instruction information; Step 7: Execute the request with an executable small granularity; Step 8: Output the result of the request with the execution completed; Step 9: Clear the instructions with the execution completed; Step 1 includes: The parallel pipeline first needs to perform polling reception processing on the same external processing instructions simultaneously sent from N stream processing cores of the graphics processing unit, and write the received instructions into the instruction buffer dedicated to storing instruction information. If the stream processing core does not send out instructions, the processing of the stream processing core and this execution pipeline is skipped; When the operation of writing the instruction into the buffer is completed, obtain the instruction entry as an index to look up the specific information of the instruction. At the same time, put the instruction index number, the single instruction multiple data thread number being executed, the stream processing core number, and the priority information into the instruction retrieval queue.

2. The method according to claim 1, wherein Step 2 includes: The instruction selector of the execution instruction pipeline selects instructions corresponding to each independent stream processing core from the instruction retrieval queue according to the priority information, stream processing core number, and the single instruction multiple data thread number carried by the instruction.

3. The method according to claim 2, wherein Step 3 includes: According to the obtained instruction index information in the instruction pipeline, read the instruction and split it internally. Split the instruction into executable small granularity requests, and send the stream processing core number, the single instruction multiple data thread number being executed, the instruction index number, the instance start information and the instance quantity information, the source address information, the request index in the instruction, and the request executable flag to the data memories in each independent core to obtain the corresponding operand data.

4. The method according to claim 3, characterized in that, Step 4 includes: when the instruction request splitting unit in the instruction pipeline splits a request, it applies to the data buffer for the location address of the return data to be stored. When the read data request returns, it receives and writes the data into the data buffer for storage according to the location address of the return data to be stored; after obtaining the operands, the order in which each core reads the return of the general-purpose data register will be disordered. At this time, it is written into the data buffer in a polling manner, and at the same time, the request executable flag bit is checked. If the request executable flag carried in the returned executable fine-grained request is valid, it is sent to the request controller of the execution pipeline, and the request controller of the execution pipeline will control the execution of the fine-grained request according to the returned ready state information.

5. The method according to claim 4, wherein Step 5 includes: for the ready states of two or more executable fine-grained requests, the execution pipeline enters the ready queue in the order of readiness. The executable fine-grained request control sequentially reads the requests with the ready state from the ready queue, finds the instruction entry according to the instruction index, and then extracts the corresponding instruction information required during execution according to the required information.

6. The method according to claim 5, wherein Step 6 includes: when executing the corresponding fine-grained request, the request controller sends the information to the information extraction module unit to extract the execution status data. By returning the carried instruction index, it extracts the valid instruction requests in the instruction buffer, and then extracts the valid instance mask and instruction opcode of the corresponding request according to the request instance start position and the number of instances, and then sends them together with the source data read by the request control to the request execution subunit for actual execution.

7. The method according to claim 6, wherein Step 7 includes: the request execution subunit collects the opcode, valid instance mask, source data, and status information of the instruction execution, and starts the calculation of the execution pipeline itself to obtain the calculation result.

8. The method according to claim 7, wherein Step 8 includes: the calculation result of the pipeline will determine whether to write back to the general-purpose data register file or output it to the specified buffer storage area for storage according to the instruction destination address information status type.

9. The method according to claim 8, wherein Step 9 includes: when the execution request ends, check the current instruction to confirm whether all requests have been executed. If not, continue to loop and execute each fine-grained request; if completed, execute the corresponding instruction entry in the instruction retrieval queue, move the subsequent instruction index entry down, fill and keep the continuous area available, and at the same time clear the corresponding instruction entry in the instruction buffer to receive the subsequent emitted instructions.

Citation Information

Patent Citations

  • Processing core having shared front end unit

    CN110045988A

  • Instruction execution method and device, equipment and storage medium

    CN114675890A

  • Methods of breaking down coarse-grained tasks for fine-grained task re-scheduling

    US20220075622A1