Instruction processing method for realizing multi-thread synchronous execution and related device

By performing write validity control and status update on the jump instruction sequence in a single-instruction multi-threaded computing system, the problem of thread execution path differences caused by jump instructions is solved, the synchronous execution of multiple threads is achieved, and the correctness and performance of the computing system are improved.

CN120631447AActive Publication Date: 2025-09-12HEFEI CORE POWER TECHNOLOGY CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510691662.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-12
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In a single-instruction multi-threaded computing system, jump instructions cause different threads to execute instruction flow paths differently, affecting the correctness, stability and performance of the computing system.

Method used

By adding write validity control to the jump instruction sequence, converting it into a state update instruction, and inserting a state reset instruction at the target position, a write restriction mark is added to the intermediate instruction sequence to ensure that the thread that meets the jump conditions does not execute the intermediate instructions ineffectively, and the thread that does not meet the conditions executes the intermediate instructions effectively.

Benefits of technology

It achieves the synchronous execution of multiple threads while maintaining the original semantics and running results, improving the correctness, stability and performance of the computing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631447A_ABST
    Figure CN120631447A_ABST
Patent Text Reader

Abstract

The invention provides an instruction processing method for realizing multi-thread synchronous execution and a related device, and the method comprises the steps: obtaining an original instruction stream, the original instruction stream comprises at least one jump instruction sequence, and the jump instruction sequence comprises a jump instruction and an intermediate instruction sequence; adding write-in validity control to the jump instruction sequence to obtain a target instruction stream for indicating that the intermediate instruction sequence is sequentially executed in an invalid execution mode in the first thread and the intermediate instruction sequence is sequentially executed in an effective execution mode in the second thread, the first thread refers to a thread meeting the jump condition indicated by the jump instruction, and the second thread refers to a thread not meeting the jump condition indicated by the jump instruction; and outputting the target instruction stream. Thus, under the condition that original semantics and operation results are kept unchanged, different threads can execute the instruction stream in the same step, synchronous execution of multiple threads is achieved, and correctness, stability and performance of a computing system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and specifically relates to an instruction processing method and related device for realizing multi-threaded synchronous execution. Background Art

[0002] In a Single Instruction Multiple Threads (SIMT) parallel computing system, a computing task is completed by multiple threads executing synchronously, requiring them to execute the same instruction stream in lockstep. However, when the instruction stream contains jump instructions that rely on register values, the register values ​​in different threads may not be identical. This leads to different execution paths for the instruction stream, violating the principle of synchronization and severely impacting the correctness, stability, and performance of the computing system. Summary of the Invention

[0003] The embodiments of the present application provide an instruction processing method and related devices for realizing multi-threaded synchronous execution, in order to solve the problem that multi-threads cannot be executed synchronously when there are jump instructions in the instruction stream, while retaining the original semantics and operation results, and improving the correctness, stability and performance of the computing system.

[0004] In a first aspect, an embodiment of the present application provides an instruction processing method for implementing multi-threaded synchronous execution, comprising:

[0005] Acquire an original instruction stream, wherein the original instruction stream includes at least one jump instruction sequence, wherein the jump instruction sequence includes a jump instruction and an intermediate instruction sequence, wherein the intermediate instruction sequence refers to an instruction sequence between the jump instruction and a target location pointed to by the jump instruction;

[0006] Adding a write validity control to the jump instruction sequence to obtain a target instruction stream, wherein the write validity control is used to instruct a first thread to sequentially execute the intermediate instruction sequence in an invalid execution manner and a second thread to sequentially execute the intermediate instruction sequence in a valid execution manner, wherein the first thread is a thread that satisfies a jump condition indicated by the jump instruction and the second thread is a thread that does not satisfy the jump condition indicated by the jump instruction;

[0007] The target instruction stream is output.

[0008] Furthermore, before adding write validity control to the jump instruction sequence to obtain the target instruction stream, the method also includes: inserting an initialization instruction for the mask register at the starting position of the original instruction stream, and the initialization instruction is used to set all control bits of the mask register to the first logical value.

[0009] Furthermore, the adding of write validity control to the jump instruction sequence to obtain a target instruction stream includes: converting the jump instruction in the jump instruction sequence into a state update instruction, the state update instruction being used to determine the logical value of the specified control bit in the mask register; inserting a state reset instruction at the target position pointed to by the jump instruction, the state recovery instruction being used to reset the specified control bit to the first logical value; adding a restricted write mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence, the restricted write mark being used to indicate that the intermediate instruction sequence is to be executed sequentially in a valid execution manner when all control bits in the mask register are the first logical value, and that the intermediate instruction sequence is to be executed sequentially in an invalid execution manner when any control bit in the mask register is not the first logical value; and determining the target instruction stream based on the state update instruction, the state reset instruction and the updated intermediate instruction sequence.

[0010] Furthermore, the jump instruction is a conditional jump instruction, and converting the jump instruction in the jump instruction sequence into a state update instruction includes: converting the jump instruction into a first state update instruction, and the first state update instruction is used to invert the condition register in the jump condition indicated by the jump instruction and write it into a specified control bit in the mask register.

[0011] Furthermore, the jump instruction is an unconditional jump instruction, and converting the jump instruction in the jump instruction sequence into a state update instruction includes: converting the jump instruction into a second state update instruction, and the second state update instruction is used to update the specified control bit in the mask register to a second logical value.

[0012] Furthermore, adding a write restriction mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence includes: identifying active register operation instructions in the intermediate instruction sequence; and merging the active register operation instructions to obtain an updated intermediate instruction sequence.

[0013] Furthermore, the active register operation instructions are merged to obtain an updated intermediate instruction sequence, including: writing the original execution result of the active register operation instruction into a temporary register; inserting a merge instruction, wherein the merge instruction is used to instruct that the value of the temporary register be written into the active register when all control bits of the mask register are the first logical value.

[0014] In a second aspect, an embodiment of the present application provides an instruction processing device for implementing multi-threaded synchronous execution, comprising:

[0015] an acquisition unit, configured to acquire an original instruction stream, wherein the original instruction stream includes at least one jump instruction sequence, wherein the jump instruction sequence includes a jump instruction and an intermediate instruction sequence, wherein the intermediate instruction sequence refers to an instruction sequence between the jump instruction and a target location pointed to by the jump instruction;

[0016] an instruction update unit, configured to add a write validity control to the jump instruction sequence to obtain a target instruction stream, wherein the write validity control is configured to instruct a first thread to sequentially execute the intermediate instruction sequence in an invalid execution manner and a second thread to sequentially execute the intermediate instruction sequence in a valid execution manner, wherein the first thread is a thread that satisfies a jump condition indicated by the jump instruction and the second thread is a thread that does not satisfy the jump condition indicated by the jump instruction;

[0017] An output unit is used to output the target instruction stream.

[0018] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program comprises instructions for executing the steps in the method described in the first aspect of the present application.

[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present application.

[0020] It can be seen that in an embodiment of the present application, after obtaining the original instruction stream, the target instruction stream is obtained by adding a write validity control to the jump instruction sequence in the original instruction stream, so that the intermediate instruction sequence is sequentially executed in an invalid execution manner in the first thread that meets the jump condition, and the intermediate instruction sequence is sequentially executed in a valid execution manner in the second thread that does not meet the jump condition, thereby enabling different threads to execute instruction streams in the same synchronization while keeping the original semantics and running results unchanged, realizing the synchronous execution of multiple threads, and improving the correctness, stability and performance of the computing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1This is an example diagram of a traditional compilation result provided by an embodiment of the present application;

[0023] Figure 2 This is a structural block diagram of a parallel computing system provided in an embodiment of the present application;

[0024] Figure 3 This is a structural block diagram of a compiler provided in an embodiment of the present application;

[0025] Figure 4 This is a flow chart of an instruction processing method for implementing multi-threaded synchronous execution provided by an embodiment of the present application;

[0026] Figure 5 This is an example diagram of a target instruction flow after optimizing a traditional compilation result, provided by an embodiment of the present application;

[0027] Figure 6 This is a structural block diagram of an instruction processing device for implementing multi-threaded synchronous execution provided by an embodiment of the present application;

[0028] Figure 7 This is a structural block diagram of another instruction processing device for implementing multi-threaded synchronous execution provided by an embodiment of the present application;

[0029] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0031] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0032] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0033] At present, with the development of technologies such as deep learning and edge computing, the requirements for the computing power of parallel computing systems are getting higher and higher. GPGPU (General-Purpose Computing on Graphics Processing Units) has opened up a new paradigm for high-performance computing by generalizing the parallel computing capabilities of graphics hardware, and has become the mainstream architecture for achieving large-scale parallel computing. The RPP (Reconfigurable Parallel Processing) parallel computing system solves the problem of low computing efficiency of traditional GPGPU by further optimizing the hardware system architecture, and supports programming in general programming languages, greatly improving versatility. However, in terms of the execution of compilation results, the SIMT-based RPP computing system requires multiple threads to execute the compilation results in unison. When there are jump instructions in the compilation results, the paths for executing the instruction stream by different threads are different, which exceeds the hardware capabilities of the RPP parallel computing system. For example, please refer to Figure 1 , Figure 1 This is an example diagram of a traditional compilation result provided by the embodiment of this application. Figure 1 As shown, instruction line 01 compares the value of register IGH with the constant 15 for a value greater than or equal to 15, writing the result to register IE. Instruction line 02 determines if register IE is true and jumps to L_BB0_2, i.e., line 11. However, in an RPP parallel computing system, the value of register IGH varies for each thread, potentially resulting in different values ​​corresponding to instruction line 01—some true and some false. Consequently, when executing instruction line 02, some threads will jump to line 11, while others will sequentially execute line 04. This results in different threads executing different instructions, violating the execution principles of SIMT-based RPP parallel computing systems and impacting the correctness, stability, and performance of the computing system.

[0034] To solve the above problems, an embodiment of the present application provides an instruction processing method and related devices for realizing multi-threaded synchronous execution.

[0035] See also Figure 2 , Figure 2This is a block diagram of a parallel computing system provided by an embodiment of the present application. Figure 2 As shown, the parallel computing system 20 includes at least a compiler 21 and a parallel processor 22. The compiler 21 and the parallel processor 22 are connected to realize data exchange. The parallel processor 22 can be specifically implemented as a reconfigurable parallel processor (RPP) with multiple threads that can synchronously execute the same instruction stream to complete the computing task.

[0036] See also Figure 3 , Figure 3 This is a structural block diagram of a compiler provided in an embodiment of the present application. Figure 3 As shown, the compiler 21 includes a compiler front-end 211, an optimizer 212, and a compiler back-end 213. The compiler back-end 213 includes an instruction generation module 214 and an instruction optimization module 215. Specifically, source code is input into the compiler front-end 211 in text form. The compiler front-end 211 is used to convert the source code into an intermediate code (Intermediate Representation, IR) sequence and then output it to the optimizer 212. The optimizer 212 is used to optimize the intermediate code sequence independently of the target machine and output the optimized intermediate code sequence to the compiler back-end 213. The compiler back-end 213 is used to process and optimize the intermediate code sequence through the instruction generation module 214 to obtain an original instruction stream, and optimize the original instruction stream through the instruction optimization module 215 to obtain a target instruction stream.

[0037] The following describes an instruction processing method for implementing multi-threaded synchronous execution provided by an embodiment of the present application.

[0038] See also Figure 4 , Figure 4 This is a flow chart of an instruction processing method for implementing multi-threaded synchronous execution provided by an embodiment of the present application, which is applied to Figure 2 In the compiler 21 shown, as Figure 4 As shown, the method includes:

[0039] S401, obtaining the original instruction stream.

[0040] The original instruction stream includes at least one jump instruction sequence, the jump instruction sequence includes a jump instruction and an intermediate instruction sequence, and the intermediate instruction sequence refers to an instruction sequence between the jump instruction and the target location pointed to by the jump instruction. Figure 1Taking the instruction flow shown as an example, the jump instruction sequence refers to the instruction sequence consisting of the instructions on line 02 to line 11, where the instruction on line 02 is the jump instruction and is a conditional jump instruction (JPC), the instruction on line 11 is the target location pointed to by the jump instruction, and the remaining instruction sets are the intermediate instruction sequences. Figure 1 The instruction stream shown also includes another jump instruction sequence, namely, the instruction sequence consisting of instructions on line 09 to line 16, wherein the instruction on line 09 is the jump instruction and is an unconditional jump instruction (JP), the instruction on line 16 is the target location pointed to by the jump instruction, and the remaining instruction sets are the intermediate instruction sequences.

[0041] S402: Add write validity control to the jump instruction sequence to obtain a target instruction stream.

[0042] The write validity control is used to instruct the first thread to sequentially execute the intermediate instruction sequence in an invalid execution mode, and the second thread to sequentially execute the intermediate instruction sequence in a valid execution mode. The first thread refers to the thread that meets the jump condition indicated by the jump instruction, and the second thread refers to the thread that does not meet the jump condition indicated by the jump instruction. In traditional assembly results, if the value of register IGH in the first thread is greater than 15 and the value of register IGH in the second thread is not greater than 15, the first thread will jump to line 11 to continue execution, while the second thread will sequentially execute instructions on line 04, violating the synchronization principle. In this example, by adding write validity control, when the processor executes this instruction stream, the first thread sequentially executes the intermediate instruction sequence in an invalid execution mode, and the second thread sequentially executes the intermediate instruction sequence in a valid execution mode. In other words, regardless of whether a jump occurs, the first and second threads continue to execute instructions on line 04 and maintain sequential execution. This enables different threads to execute the instruction stream in sync. Due to the presence of write validity control, the original semantics and running results are preserved while executing synchronously.

[0043] S403: Output the target instruction stream.

[0044] It can be seen that in an embodiment of the present application, after obtaining the original instruction stream, the target instruction stream is obtained by adding a write validity control to the jump instruction sequence in the original instruction stream, so that the intermediate instruction sequence is sequentially executed in an invalid execution manner in the first thread that meets the jump condition, and the intermediate instruction sequence is sequentially executed in a valid execution manner in the second thread that does not meet the jump condition, thereby enabling different threads to execute instruction streams in the same synchronization while keeping the original semantics and running results unchanged, realizing the synchronous execution of multiple threads, and improving the correctness, stability and performance of the computing system.

[0045] In one possible example, before adding write validity control to the jump instruction sequence to obtain the target instruction stream, the method further includes: inserting an initialization instruction for the mask register at the starting position of the original instruction stream, and the initialization instruction is used to set all control bits of the mask register to a first logical value.

[0046] The mask register is a special register used to control the data processing range or selectively perform operations. In this example, the mask register is register PRED, which is 16 bits long. The initialization instruction for the mask register is, for example, PRED=0×FFFF, which means that all control bits of the mask register are set to 1. Figure 1 Taking the instruction stream shown as an example, an initialization instruction for the mask register is inserted at the starting position of the original instruction stream, which is specifically implemented by adding the 00th line instruction: PRED=0×FFFF.

[0047] It can be seen that in this example, the initialization instruction for the mask register is inserted at the starting position of the original instruction stream, which provides a judgment basis for the subsequent multi-threaded synchronous execution and ensures that the semantics remain unchanged, thereby improving the correctness and stability of the computing system.

[0048] In one possible example, adding a write validity control to the jump instruction sequence to obtain a target instruction stream includes: converting the jump instruction in the jump instruction sequence into a state update instruction, the state update instruction being used to determine the logical value of a specified control bit in the mask register; inserting a state reset instruction at the target position pointed to by the jump instruction, the state recovery instruction being used to reset the specified control bit to the first logical value; adding a restricted write mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence, the restricted write mark being used to indicate that the intermediate instruction sequence is to be executed sequentially in a valid execution manner when all control bits in the mask register are the first logical value, and that the intermediate instruction sequence is to be executed sequentially in an invalid execution manner when any control bit in the mask register is not the first logical value; and determining the target instruction stream based on the state update instruction, the state reset instruction and the updated intermediate instruction sequence.

[0049] The state update instruction is used to modify the specified control bits of the mask register when a jump occurs, and to maintain the control bits of the mask register unchanged when no jump occurs. Then, combined with the write restriction flag, the specified control bits of the mask register in the first thread that jumps are modified, thereby sequentially executing the intermediate instruction sequence in an invalid execution manner. All control bits of the mask register in the second thread that does not jump remain unchanged at the first logical value, thereby sequentially executing the intermediate instruction sequence in a valid execution manner. This achieves synchronous execution of multiple threads while maintaining the original semantics and running results. Finally, by inserting a state reset instruction at the target location, the specified control bits of the mask register are reset to the first logical value, restoring the state before the jump, so that subsequent instructions can be executed normally.

[0050] Among them, different jump instructions in the original instruction stream correspond to different designated control bits. Figure 1 Taking the instruction flow shown as an example, the jump instruction on line 02 can correspond to bit 0 of the mask register, that is, its corresponding state update instruction is used to modify or maintain the logical value of bit 0 of the mask register, and the jump instruction on line 09 can correspond to bit 1 of the mask register, that is, its corresponding state update instruction is used to modify the logical value of bit 1 of the mask register.

[0051] The write restriction mark is, for example, a @PRED prefix, which indicates that the following instructions must be executed effectively when PRED=0×FFFF, that is, when all control bits of register PRED are 1. In this way, when a jump occurs, the instructions are executed in sequence in an invalid manner because the specified control bits of register PRED are modified. When no jump occurs, the instructions are executed in sequence in a valid manner because all control bits of register PRED are 1.

[0052] It can be seen that in this example, the jump instruction sequence is optimized by converting the jump instructions in the jump instruction sequence into state update instructions, inserting state reset instructions at the target position, and adding a restricted write mark to the intermediate instruction sequence, so as to obtain the target instruction stream, so that the parallel processor can realize the synchronous execution of multiple threads while maintaining the original semantics and running results unchanged when executing the target instruction stream, thereby improving the correctness, stability and performance of the computing system.

[0053] In one possible example, the jump instruction is a conditional jump instruction, and converting the jump instruction in the jump instruction sequence into a state update instruction includes: converting the jump instruction into a first state update instruction, and the first state update instruction is used to invert the condition register in the jump condition indicated by the jump instruction and write it into a specified control bit in the mask register.

[0054] Among them, for the conditional jump instruction, the jump instruction is converted into a first state update instruction, that is, the condition register in the jump condition is inverted and written into the specified control bit in the mask register, so that when the jump condition is met, the specified control bit is modified to the second logic value, that is, "0", and when the jump condition is not met, the logic value of the specified control bit is maintained at the first logic value, that is, "1". Figure 1 Taking the instruction flow shown as an example, the instruction on line 02 is the conditional jump instruction, where register IE is the conditional register. This jump instruction corresponds to bit 0 of mask register PRED. The first state update instruction is, for example, PRED.0 = NOT IE. Thus, when the value of register IE is true, i.e., a jump occurs, bit 0 of register PRED takes the value 0. When the value of register IE is false, i.e., no jump occurs, bit 0 of register PRED takes the value 1. Accordingly, a state reset instruction, for example, PRED.0 = 1, is inserted at the target location corresponding to the jump instruction, i.e., line 11. This resets bit 0 of register PRED to 1 to ensure normal execution of subsequent instructions.

[0055] It can be seen that in this example, by inverting the condition register in the conditional jump instruction and writing it into the specified control bit in the mask register, the conditional jump instruction is converted into a state update instruction, which can be combined with the restricted write mark to enable different threads to execute the intermediate instruction sequence sequentially in a valid or invalid execution manner, thereby achieving synchronous execution of multiple threads while maintaining the original semantics and running results unchanged, thereby improving the correctness, stability and performance of the computing system.

[0056] In one possible example, the jump instruction is an unconditional jump instruction, and converting the jump instruction in the jump instruction sequence into a state update instruction includes: converting the jump instruction into a second state update instruction, and the second state update instruction is used to update the specified control bit in the mask register to a second logical value.

[0057] Among them, for unconditional jump instructions, since the jump must occur, in order to ensure that the semantics and running results remain unchanged while achieving synchronous execution of multiple threads, this example converts the unconditional jump instruction into a second state update instruction, that is, updating the specified control bit of the mask register to the second logic value, that is, "0". Figure 1 Taking the instruction flow shown as an example, the unconditional jump instruction on line 09 corresponds to bit 1 of the mask register PRED. Therefore, the second state update instruction includes: PRED.1 = 0. Accordingly, a state reset instruction, such as PRED.1 = 1, is inserted at the target location corresponding to the jump instruction, i.e., line 16. This resets bit 1 of the PRED register to 1, ensuring normal execution of subsequent instructions.

[0058] It should be noted that in general assembly programs, unconditional jump instructions are usually located in the middle instruction sequence corresponding to conditional jump instructions to prevent the program from continuing to execute sequentially to the target location indicated by the jump instruction when the jump condition is not met. Therefore, in this example, after the unconditional jump instruction is updated to the second state update instruction, a write restriction mark will also be added to avoid changing the original semantics and running results. Figure 1 Taking the instruction flow shown as an example, the unconditional jump instruction on line 09 is located in the intermediate instruction sequence corresponding to the conditional jump instruction on line 02. The conditional jump instruction on line 02 is converted into a first state update instruction, such as PRED.0 = NOT IE, and a state reset instruction, such as PRED.0 = 1, is inserted on line 11. The unconditional jump instruction on line 09 is converted into a second state update instruction, such as PRED.1 = 0, and a state reset instruction, such as PRED.1 = 1, is inserted on line 16. A write-restriction flag is added to the second state update instruction, updating it to @PRED PRED.1 = 0. This ensures that the instruction will only modify bit 1 of register PRED to 0 when all control bits of register PRED are 1. If the write-restriction flag is not added to the second state update instruction, the first thread that jumps will modify bit 1 of register PRED to 0 when executing line 09, and the original execution result of jumping from line 02 to line 11 cannot be achieved.

[0059] The above is an application scenario in a conventional assembler without redundant information. In some assemblers with redundant information, the unconditional instruction may be located outside the intermediate instruction sequence corresponding to the conditional jump instruction. For such application scenarios, the compiler can convert the jump instruction into a third-state update instruction. The third-state instruction is used to update the specified control bit in the mask register to the second logic value under the premise that all control bits of the mask register are the first logic value, thereby avoiding violating the original semantics and running results.

[0060] It can be seen that in this example, for the unconditional jump instruction, by modifying the specified control bit in the mask register to the second logical value, it is possible to combine the restricted write mark to enable different threads to execute the intermediate instruction sequence sequentially in an invalid execution manner, thereby achieving synchronous execution of multiple threads while maintaining the original semantics and running results unchanged, thereby improving the correctness, stability and performance of the computing system.

[0061] In one possible example, adding a write restriction mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence includes: identifying active register operation instructions in the intermediate instruction sequence; and merging the active register operation instructions to obtain an updated intermediate instruction sequence.

[0062] The active register operation instruction is an instruction for updating the value of an active register, and the active register is a register that will still be used in subsequent instructions. Figure 1 Taking the instruction flow shown as an example, instructions from line 04 to line 08, and instructions from line 12 to line 14 are all active register operation instructions. Among them, after the unconditional jump instruction on line 09 is converted into a state update instruction, for example, PRED.1=0, the instruction can be regarded as a register operation instruction for updating the logical value of the first bit of the mask register PRED. Correspondingly, after the state reset instruction is inserted in its target position, line 16, for example, PRED.1=1, the instruction can be regarded as a register operation instruction for the logical value of the first bit of the mask register PRED. That is, after the instruction on line 09 is converted into a state update instruction, the register PRED indicated by it for updating will still be used in subsequent instructions and is an active register. Therefore, after the instruction on line 09 is converted into a state update instruction, it is also an active register operation instruction. By adding a write restriction mark through merging processing, the updated operation instruction is obtained, that is, @PRED PRED.1=0.

[0063] In this example, for other instructions in the intermediate instruction sequence that are not active register operation instructions, since they will not be used in subsequent instructions, there is no need to merge them, that is, there is no need to add a write restriction mark. Figure 1 Taking the instruction stream shown as an example, line 11 is the target position of the jump instruction in line 02. After the state reset instruction is inserted, for example, PRED.0=1, since there is no register operation instruction for the logical value of the 0th bit of register PRED in the subsequent instruction stream, the state reset instruction does not belong to the active register operation instruction and will not affect the execution of the subsequent instruction stream. Therefore, the instruction does not need to be merged to add a write restriction mark.

[0064] It can be seen that in this example, by merging the active register operation instructions in the intermediate instruction sequence, a restricted write mark is added to the active register operation instructions to retain the original semantics and operation results, thereby improving the correctness of the computing system.

[0065] In one possible example, the merging processing of the active register operation instructions to obtain an updated intermediate instruction sequence includes: writing the original execution result of the active register operation instruction into a temporary register; inserting a merge instruction, wherein the merge instruction is used to instruct to write the value of the temporary register into the active register when all control bits of the mask register are the first logical value.

[0066] In the RPP parallel computing system, each thread has a temporary register. By writing the original execution result of the active register operation instruction into the temporary register and then inserting the merge instruction, the active register operation instruction is merged, which is equivalent to adding a write restriction mark to the active register operation instruction. The temporary register is, for example, the register VAB, and the merge instruction is, for example, the MERGE instruction. Figure 1 Taking the instruction flow shown as an example, the active register operation instruction on line 04, "IEF = SHL.B32 ICD, 1," writes its execution result to register VAB, i.e., "VAB = SHL.B32 ICD, 1." Then, a merge instruction, "MERGE IEF = MERGE IEF, VAB, PRED," is inserted. This instruction returns VAB only if the value of PRED is 0xFFFF, meaning all control bits are 1. Otherwise, it returns IEF. This is equivalent to an active register operation instruction with the write restriction flag @PRED added, i.e., "@PRED IEF = SHL.B32 ICD, 1." Thus, when a jump occurs, the logic value of a control bit in register PRED is 0, and the merge instruction returns the original value, meaning the original execution result is not written, resulting in invalid execution. When no jump instruction occurs, all control bits in register PRED are 1, and the merge instruction returns the value of VAB, meaning the original execution result is returned, resulting in valid execution.

[0067] It can be seen that in this example, by writing the original execution result of the active register operation instruction into the temporary register, and then inserting the merge instruction to control the writing of the value of the temporary register into the active register when all the control bits of the mask register are the first logical value, it is achieved that the intermediate instruction sequence can be executed sequentially in a valid execution manner only when no jump occurs, and the intermediate instruction sequence is executed sequentially in an invalid execution manner when a jump occurs. While keeping the original semantics and running results unchanged, different threads can execute the instruction stream in the same synchronization, thereby realizing the synchronous execution of multiple threads and improving the correctness, stability and performance of the computing system.

[0068] For a possible example, see Figure 5 , Figure 5 This is an example diagram of a target instruction flow after optimizing the traditional compilation result provided by the embodiment of the present application. Figure 5As shown, the target instruction stream is optimized as follows compared to the original instruction stream: an initialization instruction for the mask register, PRED = 0×FFFF, is added to line 00. The conditional jump instruction on line 02 is converted to a first state update instruction, PRED.0 = NOT IE, and a corresponding state reset instruction, PRED.0 = 1, is inserted on line 11. Similarly, the unconditional jump instruction on line 09 is converted to a second state update instruction, PRED.1 = 0, and a corresponding state reset instruction, PRED.1 = 1, is inserted on line 16. A write restriction flag is added to the active register operation instructions in the instruction sequence between lines 02 and 11, i.e., the @PRED prefix is ​​added to the instructions from lines 04 to 09. Similarly, a write restriction flag is added to the active register operation instructions in the instruction sequence between lines 09 and 16, i.e., the @PRED prefix is ​​added to the instructions from lines 12 to 14. In this way, when the value of register IE is true, that is, a jump occurs, PRED.0=0, and the instructions from line 04 to line 09 are all invalid, which is equivalent to jumping to line 11 to continue execution; when the value of register IE is false, that is, no jump occurs, PRED.0=1, subsequent instructions are validly executed, and then executed sequentially to line 09, an unconditional jump is performed, that is, PRED.1=0, and the instructions from line 12 to line 14 are all invalid, which is equivalent to jumping to line 16 to continue execution.

[0069] It can be seen that the embodiment of the present application optimizes the original instruction stream, thereby enabling different threads to execute the instruction stream in the same synchronization while maintaining the original semantics and running results unchanged, thereby improving the correctness, stability and performance of the computing system.

[0070] In accordance with the above-mentioned embodiment, please refer to Figure 6 , Figure 6 This is a block diagram of a structure of an instruction processing device for implementing multi-threaded synchronous execution provided by an embodiment of the present application, wherein the device is applied to Figure 2In the compiler 21 shown, the instruction processing device 60 for implementing multi-threaded synchronous execution includes: an acquisition unit 601, used to acquire an original instruction stream, wherein the original instruction stream includes at least one jump instruction sequence, wherein the jump instruction sequence includes a jump instruction and an intermediate instruction sequence, wherein the intermediate instruction sequence refers to an instruction sequence between the jump instruction and the target position pointed to by the jump instruction; an instruction update unit 602, used to add a write validity control to the jump instruction sequence to obtain a target instruction stream, wherein the write validity control is used to indicate that the intermediate instruction sequence is sequentially executed in an invalid execution manner in a first thread, and that the intermediate instruction sequence is sequentially executed in a valid execution manner in a second thread, wherein the first thread refers to a thread that meets the jump condition indicated by the jump instruction, and the second thread refers to a thread that does not meet the jump condition indicated by the jump instruction; and an output unit 603, used to output the target instruction stream.

[0071] In one possible example, before adding write validity control to the jump instruction sequence to obtain the target instruction stream, the instruction processing device 60 for implementing multi-threaded synchronous execution is also used to: insert an initialization instruction for the mask register at the starting position of the original instruction stream, and the initialization instruction is used to set all control bits of the mask register to the first logical value.

[0072] In one possible example, in terms of adding write validity control to the jump instruction sequence to obtain a target instruction stream, the instruction update unit 602 is specifically used to: convert the jump instruction in the jump instruction sequence into a state update instruction, the state update instruction is used to determine the logical value of the specified control bit in the mask register; insert a state reset instruction at the target position pointed to by the jump instruction, the state recovery instruction is used to reset the specified control bit to the first logical value; add a restricted write mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence, the restricted write mark is used to indicate that the intermediate instruction sequence is executed sequentially in a valid execution manner when all control bits in the mask register are the first logical value, and that the intermediate instruction sequence is executed sequentially in an invalid execution manner when any control bit in the mask register is not the first logical value; determine the target instruction stream based on the state update instruction, the state reset instruction and the updated intermediate instruction sequence.

[0073] In one possible example, the jump instruction is a conditional jump instruction, and in terms of converting the jump instruction in the jump instruction sequence into a state update instruction, the instruction update unit 602 is specifically used to: convert the jump instruction into a first state update instruction, and the first state update instruction is used to invert the condition register in the jump condition indicated by the jump instruction and write it into a specified control bit in the mask register.

[0074] In one possible example, the jump instruction is an unconditional jump instruction. In terms of converting the jump instruction in the jump instruction sequence into a state update instruction, the instruction update unit 602 is specifically used to: convert the jump instruction into a second state update instruction, and the second state update instruction is used to update the specified control bit in the mask register to a second logical value.

[0075] In one possible example, in terms of adding a write restriction mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence, the instruction update unit 602 is specifically used to: identify active register operation instructions in the intermediate instruction sequence; and merge the active register operation instructions to obtain an updated intermediate instruction sequence.

[0076] In one possible example, in terms of merging the active register operation instructions to obtain an updated intermediate instruction sequence, the instruction update unit 602 is specifically used to: write the original execution result of the active register operation instruction into a temporary register; and insert a merge instruction, wherein the merge instruction is used to instruct to write the value of the temporary register into the active register when all control bits of the mask register are the first logical value.

[0077] It can be understood that since the method embodiment and the device embodiment are different presentation forms of the same technical concept, the content of the method embodiment part in this application should be synchronously adapted to the device embodiment part and will not be repeated here.

[0078] In the case of integrated units, such as Figure 7 As shown, Figure 7 This is a structural block diagram of another instruction processing device for implementing multi-threaded synchronous execution provided by an embodiment of the present application. Figure 7 In the embodiment, the instruction processing device 60 for implementing multi-threaded synchronous execution includes: a processing module 62 and a communication module 61. The processing module 62 is used to control and manage the actions of the instruction processing device for implementing multi-threaded synchronous execution, for example, executing the steps of the acquisition unit 601, the instruction update unit 602 and the output unit 603, and / or other processes for implementing the technology described herein. The communication module 61 is used to support the interaction between the instruction processing device for implementing multi-threaded synchronous execution and other devices. Figure 6 As shown, the instruction processing device for implementing multi-threaded synchronous execution may further include a storage module 63, and the storage module 63 is used to store program codes and data of the instruction processing device for implementing multi-threaded synchronous execution.

[0079] Among them, all relevant contents of each scenario involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here. The above instruction processing device 60 for realizing multi-threaded synchronous execution can execute the above Figure 4 The instruction processing method for realizing multi-thread synchronous execution is shown.

[0080] Based on the description of the above method embodiment and device embodiment, please refer to Figure 8 , Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 The electronic device shown includes a memory 801, a processor 802, a communication interface 803, and a bus 804. The memory 801, the processor 802, and the communication interface 803 are communicatively connected to each other via the bus 804. The electronic device may be a compiler, or a chip or functional module in the compiler.

[0081] The memory 801 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM).

[0082] The memory 801 can store programs. When the program stored in the memory 801 is executed by the processor 802, the processor 802 and the communication interface 803 are used to execute the various steps of the instruction processing method for realizing multi-threaded synchronous execution in the embodiment of the present application.

[0083] The processor 802 can adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the functions required to be performed by the units in the electronic device of the embodiment of the present application, or to execute the instruction processing method for implementing multi-threaded synchronous execution of the method embodiment of the present application.

[0084] The processor 802 may also be an integrated circuit chip with signal processing capabilities. During implementation, the various steps of the instruction processing method for implementing multi-threaded synchronous execution of the present application may be completed by hardware integrated logic circuits or software instructions in the processor 802. The aforementioned processor 802 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application may be directly embodied as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, or the like. The storage medium is located in the memory 801, and the processor 802 reads the information in the memory 801 and combines its hardware to complete the functions required to be performed by the units included in the electronic device of the embodiment of the present application, or executes the instruction processing method of the method embodiment of the present application to realize multi-threaded synchronous execution.

[0085] The communication interface 803 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the electronic device and other devices or a communication network. For example, data can be obtained through the communication interface 803.

[0086] The bus 804 may include a path for transmitting information between various components of the electronic device (eg, the memory 801 , the processor 802 , and the communication interface 803 ).

[0087] It should be noted that although Figure 8 The electronic device shown only shows the memory 801, the processor 802, and the communication interface 803. However, in the specific implementation process, those skilled in the art should understand that the electronic device also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the electronic device may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the electronic device may also include only the devices necessary to implement the embodiments of the present application, and does not necessarily include Figure 8 All devices shown in .

[0088] An embodiment of the present application also provides a computer storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, part or all of the steps of any method described in the above method embodiments are implemented.

[0089] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0090] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely schematic; for example, the division of the units is merely a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection of some interfaces, devices or units, which may be electrical, mechanical or other forms.

[0091] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0092] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may be physically included separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional units.

[0093] Although the present invention is disclosed above, it is not limited thereto. Any person skilled in the art may readily conceive of variations or substitutions, and may make various modifications and alterations without departing from the spirit and scope of the present invention. Combinations of the above-described functions and implementation steps, including software and hardware implementations, are all within the scope of protection of the present invention.

Claims

1. A method for processing instructions to achieve multi-threaded synchronous execution, characterized in that: include: Acquire an original instruction stream, wherein the original instruction stream includes at least one jump instruction sequence, wherein the jump instruction sequence includes a jump instruction and an intermediate instruction sequence, wherein the intermediate instruction sequence refers to an instruction sequence between the jump instruction and a target location pointed to by the jump instruction; Adding a write validity control to the jump instruction sequence to obtain a target instruction stream, wherein the write validity control is used to instruct a first thread to sequentially execute the intermediate instruction sequence in an invalid execution manner and a second thread to sequentially execute the intermediate instruction sequence in a valid execution manner, wherein the first thread is a thread that satisfies a jump condition indicated by the jump instruction and the second thread is a thread that does not satisfy the jump condition indicated by the jump instruction; The target instruction stream is output.

2. The method according to claim 1, characterized in that Before adding write validity control to the jump instruction sequence to obtain the target instruction stream, the method further includes: An initialization instruction for a mask register is inserted at a start position of the original instruction stream, where the initialization instruction is used to set all control bits of the mask register to a first logic value.

3. The method according to claim 2, characterized in that Adding write validity control to the jump instruction sequence to obtain a target instruction stream includes: converting a jump instruction in the jump instruction sequence into a state update instruction, wherein the state update instruction is used to determine a logic value of a specified control bit in the mask register; Inserting a state reset instruction at the target location pointed to by the jump instruction, wherein the state restoration instruction is used to reset the designated control bit to the first logic value; adding a write restriction flag to the intermediate instruction sequence to obtain an updated intermediate instruction sequence, wherein the write restriction flag is used to indicate that the intermediate instruction sequence is to be sequentially executed in a valid execution manner when all control bits in the mask register are the first logic value, and to indicate that the intermediate instruction sequence is to be sequentially executed in an invalid execution manner when any control bit in the mask register is not the first logic value; A target instruction stream is determined according to the state updating instruction, the state resetting instruction, and the updated intermediate instruction sequence.

4. The method according to claim 3, characterized in that The jump instruction is a conditional jump instruction, and converting the jump instruction in the jump instruction sequence into a state update instruction includes: The jump instruction is converted into a first state update instruction, where the first state update instruction is used to invert a condition register in a jump condition indicated by the jump instruction and then write the inverted condition register into a designated control bit in the mask register.

5. The method according to claim 3, characterized in that The jump instruction is an unconditional jump instruction, and converting the jump instruction in the jump instruction sequence into a state update instruction includes: The jump instruction is converted into a second state update instruction, where the second state update instruction is used to update a designated control bit in the mask register to a second logic value.

6. The method according to any one of claims 3 to 5, characterized in that: Adding a write restriction mark to the intermediate instruction sequence to obtain an updated intermediate instruction sequence includes: identifying active register operation instructions in the intermediate instruction sequence; The active register operation instructions are merged to obtain an updated intermediate instruction sequence.

7. The method according to claim 6, characterized in that The merging of the active register operation instructions to obtain an updated intermediate instruction sequence includes: Writing the original execution result of the active register operation instruction into a temporary register; A merge instruction is inserted, wherein the merge instruction is used to instruct to write the value of the temporary register into the active register when all control bits of the mask register are the first logic value.

8. An instruction processing device for realizing multi-threaded synchronous execution, characterized in that: include: an acquisition unit, configured to acquire an original instruction stream, wherein the original instruction stream includes at least one jump instruction sequence, wherein the jump instruction sequence includes a jump instruction and an intermediate instruction sequence, wherein the intermediate instruction sequence refers to an instruction sequence between the jump instruction and a target location pointed to by the jump instruction; an instruction update unit, configured to add a write validity control to the jump instruction sequence to obtain a target instruction stream, wherein the write validity control is configured to instruct a first thread to sequentially execute the intermediate instruction sequence in an invalid execution manner and a second thread to sequentially execute the intermediate instruction sequence in a valid execution manner, wherein the first thread is a thread that satisfies a jump condition indicated by the jump instruction and the second thread is a thread that does not satisfy the jump condition indicated by the jump instruction; An output unit is used to output the target instruction stream.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for executing the steps in the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Hardware multithreading control method for microprocessor and device thereof

    CN101957744A

  • Multithreaded processor and instruction execution and synchronization method thereof and computer program product

    CN101980147A

  • Loop control flow diversion

    CN102193777A

  • Controlling the execution of adjacent instructions that are dependent upon a same data condition

    CN103348318A

  • Systems, apparatuses, and methods for jumps using a mask register

    CN103718157A