Multi-threaded bundle debugging method, device and artificial intelligence chip
By storing debugging information and instruction addresses in the target register group and judging the expected thread bundle with thread bundle identification, the problem of parallel computing and debugging of multi-threaded bundles is solved, and efficient and accurate multi-threaded bundle debugging is achieved.
Patent Information
- Application Number
- CN202510399997.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing program debugging methods are mainly applicable to single-threaded programs and cannot meet the debugging requirements of multi-threaded bundle parallel computing.
By pre-stored the instruction information for triggering debugging and the instruction address of the debugging program in the target register group, the instruction information executed by each thread bundle is compared with the instruction information in the target register group, the expected thread bundle is judged in combination with the thread bundle identification, and the debug enable and stop information is sent to the processor, and the thread bundle is controlled to wait and continue to execute the original program.
It realizes precise debugging of specific thread bundles in multi-threaded beam parallel computing, avoids data pollution, improves debugging accuracy and efficiency, and does not require modification of the original program.
Smart Images

Figure CN119902968B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-threaded bundle debugging method, device and artificial intelligence chip. Background Art
[0002] As the size of programs continues to grow, software developers have an increasingly urgent need for program functionality. They expect hardware to provide precise debugging support, allowing developers to check the execution of instructions line by line or section by section, and then deeply analyze the execution characteristics of the program. In this way, developers can more efficiently optimize program performance and ensure efficient operation of the software.
[0003] However, in parallel computing, multiple warps participate in scheduling and arbitration at the same time, and the execution content of each warp is both independent and different from each other. Most existing program debugging methods are only applicable to single-threaded programs and cannot meet the debugging needs of multi-warp parallel computing.
[0004] Therefore, how to implement multi-threaded bundle debugging in parallel computing has become a technical problem that needs to be solved urgently in the industry. Summary of the invention
[0005] The present invention provides a multi-threaded bundle debugging method, device and artificial intelligence chip, which are used to solve the technical problem of how to implement multi-threaded bundle debugging in parallel computing.
[0006] The present invention provides a multi-thread warp debugging method, comprising:
[0007] Determine the expected warp to trigger debugging based on the comparison result of the instruction information of the original program executed by each warp and the instruction information stored in the target register group, and the warp identifier of each warp;
[0008] Sending debug trigger information to a processor so that the processor sends debug enable information to each thread warp, and controls an unexpected thread warp to wait for the expected thread warp to be debugged;
[0009] Based on the debugging enable information, controlling the expected thread warp to execute the debugging program based on the instruction address of the debugging program stored in the target register group;
[0010] Debug end information is sent to the processor, so that the processor sends debug stop information to each thread warp, and controls each thread warp to continue to execute the original program.
[0011] In some embodiments, the target register set is a memory-mapped input / output register; the target register set includes a mode enable register, a trigger information register, and a debugger instruction address register;
[0012] The mode enable register is used to define the trigger mode of debugging and the number of the trigger information registers; the trigger mode includes an instruction address trigger mode or an instruction type trigger mode;
[0013] The trigger information register is used to store the instruction information for triggering debugging; the instruction information includes an instruction address or an instruction type;
[0014] The debug program instruction address register is used to store the instruction address of the debug program.
[0015] In some embodiments, determining that an expected warp triggers debugging based on the comparison result of the instruction information of the original program executed by each warp and the instruction information stored in the target register file, and the warp identifier of each warp, includes:
[0016] Determining that each warp triggers debugging based on the comparison result of the instruction information of the original program executed by each warp and the instruction information stored in the target register file;
[0017] Determining that the expected warp triggers debugging based on the comparison result of the warp identifier of each warp and the warp identifier of the expected warp.
[0018] In some embodiments, the debug trigger information includes the warp identifier of the warp that triggers debugging; the processor is configured to determine that the warp is the expected warp based on the warp identifier of the warp.
[0019] In some embodiments, after controlling the expected warp to execute the debug program based on the debug program instruction address stored in the target register file based on the debug enable information, the method further includes:
[0020] Recording the executed instruction address of the original program; the processor is configured to control each warp to continue executing the original program based on the executed instruction address after debugging ends.
[0021] In some embodiments, the processor is configured to configure the target register file based on the trigger mode, instruction address, instruction type of the original program, and the instruction address of the debug program.
[0022] The present invention provides a multi-warp debugging device, including:
[0023] A comparison module, configured to determine that an expected warp triggers debugging based on the comparison result of the instruction information of the original program executed by each warp and the instruction information stored in the target register file, and the warp identifier of each warp;
[0024] A first sending module, configured to send debugging trigger information to a processor, so that the processor sends debugging enable information to each warp, and controls non-expected warps to wait for the expected warp to perform debugging;
[0025] An execution module, configured to, based on the debugging enable information, control the expected warp to execute the debugging program based on the instruction address of the debugging program stored in the target register group;
[0026] A second sending module, configured to send debugging end information to the processor, so that the processor sends debugging stop information to each warp, and controls each warp to continue executing the original program.
[0027] The present invention provides an artificial intelligence chip, including:
[0028] A target register group, configured to define a triggering mode of debugging, and store instruction information for triggering debugging and an instruction address of a debugging program;
[0029] A scheduling control unit, connected to the target register group, configured to execute the multi-warp debugging method as described above, and control an expected warp to perform debugging among multiple warps.
[0030] The present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the multi-warp debugging method as described above is implemented.
[0031] The present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-warp debugging method as described above is implemented.
[0032] For the multi-warp debugging method, device, and artificial intelligence chip provided by the present invention, since instruction information for triggering debugging and an instruction address of a debugging program are pre-stored in a target register group, during the process of each warp executing the instruction information of the original program, it is determined whether a warp triggers debugging by comparing the instruction information, and it is determined whether it is an expected warp by the warp identifier, so that debugging for a specific warp can be implemented during the multi-warp parallel computing process; through the processor for debugging enabling and debugging stopping, non-expected warps are controlled to wait for the expected warp to perform debugging, avoiding data pollution, and precise debugging for a specific warp is achieved; in addition, only the target register group needs to be configured, and the original program does not need to be modified, avoiding changes in the instruction address caused by replacing the original program instructions or inserting debugging instructions, improving the accuracy of implementing multi-warp debugging in parallel computing and improving the efficiency of multi-warp debugging. Description of the Drawings
[0033] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0034] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0035] Figure 1 It is a schematic flowchart of the multi-thread warp debugging method provided by the present invention.
[0036] Figure 2 It is a schematic diagram of the software program provided by the present invention.
[0037] Figure 3 It is a schematic diagram of the target register bank provided by the present invention.
[0038] Figure 4 It is a schematic structural diagram of the multi-thread warp debugging device provided by the present invention.
[0039] Figure 5 It is one of the schematic structural diagrams of the artificial intelligence chip provided by the present invention.
[0040] Figure 6 It is another schematic structural diagram of the artificial intelligence chip provided by the present invention.
[0041] Figure 7 It is a schematic working diagram of the scheduling control unit provided by the present invention.
[0042] Figure 8 It is a schematic flowchart of the execution of the debugging program provided by the present invention.
[0043] Figure 9 It is a schematic diagram of the jump of the warp state provided by the present invention.
[0044] Figure 10 It is a schematic diagram of the multi-thread warp debugging provided by the present invention.
[0045] Figure 11 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0046] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0047] It should be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units or modules does not necessarily have to be limited to those steps or units or modules clearly listed, but may include other steps or units or modules not clearly listed or inherent to these processes, methods, products or devices.
[0048] Setting breakpoints and single-stepping are the basis for precise debugging. The following methods are used for program debugging in the related art:
[0049] (1) Set breakpoints in the software program. When debugging the program, insert a breakpoint instruction at the place where problems may occur. When the program is executed by the hardware and reaches the breakpoint, it enters the debugging mode and executes the debugging program. Using this method requires modifying the content of the original program. The modification methods include replacing a certain instruction in the original program with a breakpoint instruction or inserting a breakpoint instruction at a certain position. There is generally a data dependency between the front and back instructions in the program. When replacing an instruction, the program content needs to be carefully considered and then replaced at a position that does not affect the program function. The usage requirements and limitations are relatively high. Inserting a breakpoint instruction will change the program sequence after the breakpoint, resulting in inaccurate debugging.
[0050] (2) Set hardware breakpoints. That is, use the debugging address register defined by the hardware to implement breakpoints. After the program executes to the defined debugging address, it enters the debugging mode and executes the debugging program. This method depends on the number of debugging address registers reserved in the hardware, is easily limited by the number, and can only be used for debugging of instruction addresses, lacking flexibility.
[0051] (3) Set a single-step execution register. After configuring this register, each time the program executes an instruction, it jumps to the debugging program for exception handling and execution. After using this method, the original program enters the debug mode every time it executes an instruction, and then executes the debugging program. This method may introduce redundant debugging and trigger at positions where debugging is not required.
[0052] The above method can only be applied to the execution debugging of a single thread and cannot meet the debugging requirements of multi-threaded warp parallel computing.
[0053] To solve the deficiencies of the related technology, Figure 1 is a schematic flow chart of the multi-threaded warp debugging method provided by the present invention. As Figure 1 shown, the method includes step 110, step 120, step 130, and step 140.
[0054] Step 110: Based on the comparison result between the instruction information of the original program executed by each warp and the instruction information stored in the target register group, and the warp ID of each warp, determine the expected warp to trigger debugging.
[0055] Specifically, the execution entity of the multi-threaded warp debugging method provided by the embodiments of the present invention is a multi-threaded warp debugging device. This device can be implemented by software, such as a multi-threaded warp debugging program running in an artificial intelligence chip, etc.; it can also be implemented by hardware, such as a warp scheduler (WarpScheduler) that executes the multi-threaded warp debugging method in an artificial intelligence chip, etc.
[0056] A warp is a set composed of a group of threads. These threads share the same instruction stream during execution. These threads will run the same instruction simultaneously in a single instruction multiple data (SIMD) manner, but process different data. For example, in some artificial intelligence chips, 1 warp usually contains 32 threads.
[0057] The warp ID is the unique number of the warp. The warp ID can be calculated through the thread index and is used to distinguish different warps. The expected warp is a specific warp that is expected to be debugged. The number of expected warps can be one or more. Correspondingly, the remaining warps other than the expected warps can be considered non-expected warps. The number of non-expected warps can be zero or more.
[0058] The original program refers to the initial version of the program, that is, the program code that has not been modified, optimized, or debugged. The debugging program refers to the program code used to detect, locate, and fix errors (bugs) in the original program.
[0059] Figure 2 is a schematic diagram of the software program provided by the present invention. As Figure 2As shown, PC represents the Program Counter, and PC0 to PCN represent the 0th to Nth instruction addresses. Inst represents Instruction, and Inst0 to InstN represent the 0th to Nth instructions. Each instruction has a unique corresponding instruction address.
[0060] Instruction information refers to the information related to instruction execution, usually including the instruction address and the instruction type. The instruction address is the value of the program counter, which is used to identify the position of the currently executing instruction in memory. It determines the execution order of the program and supports operations such as branching, jumping, and function calls. The instruction type refers to the functional classification of the instruction, such as arithmetic instructions, logical instructions, memory access instructions, control flow instructions, and special instructions, etc.
[0061] The target register bank refers to the storage unit in the artificial intelligence chip used to temporarily store data, instructions, or addresses and other information. Instruction information for triggering debugging can be stored in the target register bank. These instruction information can be pre-set according to debugging requirements. In the embodiments of the present invention, debugging specifically refers to breakpoint debugging, where the execution of the program is paused by setting breakpoints in the code. A breakpoint is a marker that indicates the debugger to pause when the program execution reaches this position, allowing developers to check the state of the program.
[0062] During the process of each warp executing the original program, the instruction information executed by each warp can be compared with the instruction information stored in the target register bank, for example, comparing the instruction address or the instruction type. If the comparison result is consistent, it can be determined that each warp triggers debugging. Then, the warp identifier of each warp that triggers debugging is compared with the warp identifier of the expected warp. If the comparison result is consistent, it can be determined that the expected warp triggers debugging.
[0063] Step 120: Send debugging trigger information to the processor, so that the processor sends debugging enable information to each warp, and controls the non-expected warps to wait for the expected warp to perform debugging.
[0064] Specifically, the processor is used to manage each warp. In the embodiments of the present invention, the processor can be a central processing unit (CPU) independent of the artificial intelligence chip for managing warps, or a processor set inside the artificial intelligence chip for managing warps, such as the Streaming Multiprocessor (SM) set in some Graphics Processing Units (GPUs). The Streaming Multiprocessor and the warp scheduler jointly implement the management of warps.
[0065] Debug trigger information can be sent to the processor, which is used to prompt the processor that the expected warp has run to the instruction address where debugging can be triggered or the type of instruction being processed can trigger debugging.
[0066] After receiving the debug trigger information, the processor can send debug enable information (debugmode on) to each warp. For the expected warp, the execution of the debug program can be controlled; for the unexpected warp, it needs to wait for the expected warp to perform debugging. This is to prevent the unexpected warp from continuing to execute the original program and updating the data in the debug mode, causing data contamination.
[0067] Step 130: Based on the debug enable information, control the expected warp to execute the debug program based on the instruction address of the debug program stored in the target register group.
[0068] Specifically, after receiving the debug enable information sent by the processor, the execution of the expected warp can be controlled to obtain the code of the debug program according to the instruction address of the debug program stored in the target register group, so as to execute the debug program.
[0069] Step 140: Send debug end information to the processor, so that the processor sends debug stop information to each warp and controls each warp to continue executing the original program.
[0070] Specifically, after the expected warp finishes executing the debug program, debug end information can be sent to the processor.
[0071] After receiving the debug end information, the processor can send debug stop information (debugmode off) to each warp. For each warp, the instruction information of the original program can be continuously obtained and the original program can be executed.
[0072] The multi-threaded warp debugging method provided by the embodiments of the present invention determines the expected warp to trigger debugging based on the comparison result between the instruction information of each warp executing the original program and the instruction information stored in the target register group, and the warp identifier of each warp; sends a debugging trigger message to the processor, so that the processor sends a debugging enable message to each warp, controls the non-expected warps to wait for the expected warp to perform debugging; based on the debugging enable message, controls the expected warp to execute the debugging program based on the instruction address of the debugging program stored in the target register group; sends a debugging end message to the processor, so that the processor sends a debugging stop message to each warp, controls each warp to continue executing the original program; since the instruction information for triggering debugging and the instruction address of the debugging program are pre-stored in the target register group, during the process of each warp executing the instruction information of the original program, it is judged whether the warp triggers debugging by comparing the instruction information, and it is judged whether it is the expected warp by the warp identifier, so that it is possible to perform debugging on a specific warp during the multi-threaded warp parallel computing process; through the processor to enable and stop debugging, controls the non-expected warps to wait for the expected warp to perform debugging, avoids data contamination, and realizes precise debugging for a specific warp; in addition, only the target register group needs to be configured, and the original program does not need to be modified, avoiding the change of the instruction address caused by replacing the original program instructions or inserting debugging instructions, improving the accuracy of multi-threaded warp debugging in parallel computing and improving the efficiency of multi-threaded warp debugging.
[0073] It should be noted that each embodiment of the present invention can be freely combined, the order can be swapped, or it can be executed independently, and does not need to rely on or depend on a fixed execution order.
[0074] In some embodiments, the target register group is a memory-mapped input / output register; the target register group includes a mode enable register, a trigger information register, and a debugging program instruction address register;
[0075] The mode enable register is used to define the trigger mode of debugging and the number of trigger information registers; the trigger mode includes an instruction address trigger mode or an instruction type trigger mode;
[0076] The trigger information register is used to store the instruction information for triggering debugging; the instruction information includes an instruction address or an instruction type;
[0077] The debugging program instruction address register is used to store the instruction address of the debugging program.
[0078] Specifically, the target register bank can be implemented using memory-mapped input / output registers (MMIO registers). Memory-mapped input / output registers refer to registers that map the control registers and data registers of a hardware device to the processor's address space through Memory-Mapped I / O (MMIO) technology. In this way, the operating system and programs can access the hardware device through conventional memory access instructions (such as read and write operations) without using dedicated input / output instructions.
[0079] Figure 3 is a schematic diagram of the target register bank provided by the present invention. As Figure 3 shown, the target register bank includes a mode enable register, a trigger information register, and a debug program instruction address register. Each register can include 32 bits.
[0080] The mode enable register can be numbered R0 and is used to define the trigger mode of debugging and the number of trigger information registers; the trigger mode includes an instruction address trigger mode or an instruction type trigger mode. For example, bits 1 and 2 in the register can represent bit[1:0], and when these two bits take the value 00, it represents the instruction address trigger mode; when these two bits take the value 01, it represents the instruction type trigger mode; other values are reserved for the time being. Bits 3 to 32 in the register can represent bit[31:2] and can be used to define the number of trigger information registers.
[0081] The number of trigger information registers can be set according to actual needs. There can be multiple trigger information registers (the number is N), and the numbers can be represented as R1 to RN. Generally, N is 8 or 16.
[0082] In the instruction address trigger mode, the trigger information register is used to store the instruction address that triggers debugging. When the original program executes to this instruction address, an interrupt is triggered and debugging is entered.
[0083] In the instruction type trigger mode, the trigger information register is used to store the instruction type that triggers debugging. For example, bits 32 to 17 in the register can represent bit[31:16] and are used to store the instruction type; bits 16 to 1 in the register can represent bit[15:0] and are used to store the trigger count M corresponding to this instruction type. When the original program executes to this instruction type, an interrupt is triggered and debugging is entered. After this debugging ends, the trigger count M is updated to M - 1. When the same instruction type is executed next time, an interrupt is triggered again and debugging is entered, and the trigger count is updated again. Multiple instruction types can be configured, and the trigger count for each instruction type can be customized.
[0084] The number of the debug program instruction address register can be RN + 1, which is used to store the instruction address of the debug program, specifically, it can be the initial instruction address of the debug program.
[0085] The multi-thread warp debugging method provided by the embodiments of the present invention can, by configuring the target register group, preset the trigger mode, the instruction information for triggering debugging, and the instruction address of the debug program, and through the comparison of the instruction information, implement debugging for a specific warp during the parallel calculation of multi-thread warps.
[0086] In some embodiments, based on the comparison result of the instruction information of each warp executing the original program and the instruction information stored in the target register group, and the warp identifier of each warp, determining that the expected warp triggers debugging includes:
[0087] Based on the comparison result of the instruction information of each warp executing the original program and the instruction information stored in the target register group, determining that each warp triggers debugging;
[0088] Based on the comparison result of the warp identifier of each warp and the warp identifier of the expected warp, determining that the expected warp triggers debugging.
[0089] Specifically, the multi-thread warp debugging device can specifically execute the comparison function of the instruction information, including the comparison of the instruction address and the instruction type. When performing the comparison of the instruction address, the multi-thread warp debugging device judges whether to jump according to the comparison result of the instruction address, and then executes the jump program. When performing the comparison of the instruction type, the multi-thread warp debugging device can also record the number of times the instruction type has been triggered, and judge whether to jump according to the comparison result of the instruction type and the number of times triggered, and then execute the jump program.
[0090] First, compare the instruction information of each warp executing the original program with the instruction information stored in the target register group. If the comparison result of the instruction information is consistent, it can be determined that each warp triggers debugging. Considering that the executions of each warp are fast or slow, the above comparisons are also carried out one after another, but each warp will be compared.
[0091] Further, compare the warp identifier of each warp that triggers debugging with the warp identifier of the expected warp. If the comparison result of the warp identifier is consistent, it can be determined that the expected warp triggers debugging.
[0092] The multi-thread warp debugging method provided by the embodiments of the present invention can, through the comparison of the instruction information and the warp identifier, implement debugging for a specific warp during the parallel calculation of multi-thread warps.
[0093] In some embodiments, the debug trigger information includes the warp identifier of the warp that triggers the debug; the processor is configured to determine that the warp is the expected warp based on the warp identifier of the warp.
[0094] Specifically, the debug trigger information sent by the multi-warp debugging device to the processor should at least include the warp identifier of the warp that triggers the debug.
[0095] After receiving the debug trigger information, the processor can also compare the warp identifier of the warp with the warp identifier of the expected warp. If the comparison result of the warp identifiers is consistent, it can be determined that the expected warp triggers the debug, and thus send debug enable information to each warp.
[0096] The multi-warp debugging method provided by the embodiments of the present invention determines that the expected warp triggers the debug through the processor, improving the accuracy of multi-warp debugging in parallel computing.
[0097] In some embodiments, after controlling the expected warp to execute the debug program based on the instruction address of the debug program stored in the target register group based on the debug enable information, the method further includes:
[0098] Recording the executed instruction address of the original program; the processor is configured to control each warp to continue executing the original program based on the executed instruction address after the debug ends.
[0099] Specifically, when the multi-warp debugging device enters the debug, it also records the executed instruction address of the original program. After the debug ends, it can control each warp to continue executing the instructions in the original program according to the executed instruction address.
[0100] The multi-warp debugging method provided by the embodiments of the present invention records the executed instruction address of the original program and controls each warp to continue executing the original program after the debug ends, which can ensure that the execution logic and state of the original program are not interfered by the debug process, improve the reliability of the debug, support complex debug operations, and reduce the impact of the debug on the program performance.
[0101] In some embodiments, the processor is configured to configure the target register group based on the trigger mode, instruction address, instruction type of the original program, and the instruction address of the debug program.
[0102] Specifically, before executing the original program through multiple warps, the trigger mode, instruction address, instruction type of the original program, and the instruction address of the debug program can be determined in advance.
[0103] Configure the mode enable register in the target register group according to the trigger mode; configure the trigger information register in the target register group according to the instruction address or instruction type; configure the debug program instruction address register in the target register group according to the instruction address of the debug program.
[0104] The multi-threaded warp debugging method provided by the embodiments of the present invention pre-configures the target register group, avoiding setting breakpoints in the original program and eliminating the need to set breakpoints through hardware. It can achieve debugging for specific warps during the parallel computing of multi-threaded warps, improving the reliability of debugging, supporting complex debugging operations, and reducing the impact of debugging on program performance.
[0105] The device provided by the embodiments of the present invention will be described below. The device described below can be correspondingly referred to the method described above.
[0106] Figure 4 is a schematic structural diagram of the multi-threaded warp debugging device provided by the present invention, as Figure 4 shown, the device includes:
[0107] A comparison module 410, configured to determine that an expected warp triggers debugging based on the comparison result of the instruction information of each warp executing the original program and the instruction information stored in the target register group, and the warp identifier of each warp;
[0108] A first sending module 420, configured to send debug trigger information to a processor, so that the processor sends debug enable information to each warp, and controls the non-expected warps to wait for the expected warp to perform debugging;
[0109] An execution module 430, configured to control the expected warp to execute a debug program based on the instruction address of the debug program stored in the target register group based on the debug enable information;
[0110] A second sending module 440, configured to send debug end information to the processor, so that the processor sends debug stop information to each warp, and controls each warp to continue executing the original program.
[0111] The multi-thread warp debugging device provided by an embodiment of the present invention determines an expected thread warp to trigger debugging based on the comparison result between the instruction information of the original program executed by each thread warp and the instruction information stored in the target register group, and the thread warp identifier of each thread warp; sends a debugging trigger message to a processor, so that the processor sends a debugging enable message to each thread warp, and controls non-expected thread warps to wait for the expected thread warp to perform debugging; based on the debugging enable message, controls the expected thread warp to execute a debugging program based on the instruction address of the debugging program stored in the target register group; sends a debugging end message to the processor, so that the processor sends a debugging stop message to each thread warp, and controls each thread warp to continue executing the original program; since the instruction information for triggering debugging and the instruction address of the debugging program are pre-stored in the target register group, during the process of each thread warp executing the instruction information of the original program, it is determined whether a thread warp triggers debugging by comparing the instruction information, and it is determined whether it is an expected thread warp by the thread warp identifier, which can achieve debugging for a specific thread warp during the multi-thread warp parallel computing process; through the processor for debugging enabling and debugging stopping, controls non-expected thread warps to wait for the expected thread warp to perform debugging, avoids data contamination, and achieves precise debugging for a specific thread warp; improves the accuracy of multi-thread warp debugging in parallel computing and improves the efficiency of multi-thread warp debugging.
[0112] Figure 5 is one of the schematic structural diagrams of the artificial intelligence chip provided by the present invention, as Figure 5 shown, the artificial intelligence chip 500 includes:
[0113] A target register group 510, which is used to store instruction information for triggering debugging.
[0114] A scheduling control unit 520, which is connected to the target register group 510, and is used to execute the multi-thread warp debugging method in the above embodiment, and control the expected thread warp to perform debugging among multiple thread warps.
[0115] Specifically, the artificial intelligence chip may be a Graphics Processing Unit (GPU), a General-purpose graphics processing units (GPGPU), a Domain Specific Architecture (DSA), etc.
[0116] In the artificial intelligence chip provided by the embodiment of the present invention, in a parallel computing scenario, due to the setting of the target register group and the scheduling control unit, instruction information for triggering debugging and the instruction address of the debugging program can be pre-stored in the target register group. During the process of each warp executing the instruction information of the original program, the scheduling control unit determines whether the warp triggers debugging by comparing the instruction information, and determines whether it is the expected warp by the warp identifier, so that debugging can be realized for a specific warp during the parallel computing of multiple warps; the accuracy of realizing multi-warp debugging in parallel computing is improved, and the efficiency of multi-warp debugging is improved.
[0117] Figure 6 is the second structural schematic diagram of the artificial intelligence chip provided by the present invention, as Figure 6 shown, the artificial intelligence chip 500 further includes an instruction cache unit 530, an instruction fetch unit 540, a decoding unit 550, and an execution unit 560 in addition to the target register group 510 and the scheduling control unit 520. The processor 600 can be a processor provided inside the artificial intelligence chip or a processor independent of the artificial intelligence chip. The external storage 610 is a storage device connected to the artificial intelligence chip. The target register group can specifically be an MMIO register group.
[0118] This artificial intelligence chip introduces an MMIO register group, and realizes the synchronization and debugging of multiple warps by configuring the MMIO register and the debugging program through a processor (Host side). Among them, the MMIO register group defines the setting and triggering mode of breakpoints, and supports triggering based on instruction address and instruction type. The debugging program and the Host realize the targeted debugging of a specific warp (Expect Warp), and support debugging one or more warps simultaneously. The scheduling control unit obtains the information in the MMIO register group and the program instruction information, triggers an interrupt and preserves the execution state of the original program.
[0119] Figure 7 is the working schematic diagram of the scheduling control unit provided by the present invention, as Figure 7 shown, the scheduling control unit can execute the comparison function of instruction information, including the comparison of instruction address and instruction type.
[0120] The scheduling control unit obtains the instruction information executed by the warp through the decoding unit. When the interrupt trigger comparator performs the comparison of the instruction type, the state machine records the number of times the instruction type has been triggered. According to the comparison result, it is judged whether to jump, resource inspection is performed, and then the jump program is executed.
[0121] In the above embodiments, the processor (Host side) can configure the MMIO register set; receive the debug requests of each warp, determine whether it is the expected warp (Expect Warp), if so, enable the debug mode, and if not, ignore it. When the expected warp completes the current breakpoint debugging, the debug mode is turned off. When the processor detects that the expected warp has completed all breakpoint debugging, when receiving a debug request from an unexpected warp (not Expect Warp), it directly notifies it to exit.
[0122] In the above embodiments, the functional pseudocode of the debug program is as follows:
[0123] send signal->host
[0124] loop:
[0125] if(warp_id==expect id)
[0126] wait debug mode on
[0127] save temp reg
[0128] dump reg / mem
[0129] …
[0130] send signal->host
[0131] wait exit
[0132] break
[0133] else(warp_id!=expect id)
[0134] wait exit
[0135] break
[0136] step return
[0137] Figure 8 is a schematic diagram of the execution process of the debug program provided by the present invention. As Figure 8 shown, the debug program is used to define the execution behavior of the hardware after entering the debug mode. The debug program and the Host side jointly implement the execution synchronization and execution behavior definition of multiple warps. The debug program has the following functions: sending requests to the Host side; comparing the warp identifiers of each warp with the warp identifier of the expected warp; exporting (Dump) registers and stored data; breakpoint recovery.
[0138] Based on the above embodiments, the overall process of the multi-threaded warp debugging method provided by the present invention is described as follows:
[0139] 1. The Host (processor) initializes and configures the MMIO registers, sets the interrupt trigger mode and trigger enable, and configures the trigger information. For example, configure trigger points such as instruction addresses or the number of trigger times corresponding to instruction types.
[0140] 2. The Host loads the original program into the instruction cache unit.
[0141] 3. Each warp starts running. The fetch unit sends an instruction fetch request to the instruction cache unit, and the instruction cache unit returns the instruction.
[0142] 4. The decoding unit parses the instruction content and transmits it to the scheduling control unit.
[0143] 5. The scheduling control unit manages and allocates according to the instruction information and sends it to the execution unit for processing. Since the executions of each warp are independent and different, when the first or several warps execute to the debug breakpoint, the scheduling control unit saves the instruction address of the original program execution, then jumps into the debug program and sends a signal to the Host. The Host and the scheduling control unit respectively judge whether it is the expected warp.
[0144] (1) Host: If so, send the debug mode on signal to all warps; if not, ignore it.
[0145] (2) Scheduling control unit: The scheduling control unit also compares whether the thread number identifier of this warp is the same as the warp identifier of the expected warp. If so, after waiting for the Host to enable it, execute the user-defined debug program, and wait for the Host to notify to exit the debug mode after execution. If not, directly wait for the Host to notify to exit the debug mode.
[0146] (3) When the Host receives the signal that the expected warp executes to the breakpoint, it enables the debug mode; when it receives the signal that the expected warp finishes executing the debug program, it disables the debug mode.
[0147] 6. The execution speed of the expected warp is uncertain. If the warp that first executes to the breakpoint is not the expected warp, it stops at the else branch of the debug program and waits to exit, while other warps continue to execute. When the expected warp executes to the breakpoint, the Host enables the debug mode. After the debug mode is enabled, some warps may not reach the breakpoint yet. At this time, it directly jumps to the debug program and then stops at the else branch.
[0148] In debug mode, other warps are stopped, and only the expected warp is working to prevent data contamination caused by the execution of unexpected warps.
[0149] 7. The debug trigger modes are divided into instruction address-based and instruction type-based. When triggered by instruction address, when the warp executes the original program to the corresponding instruction address, it is triggered and jumps to the debug program. When triggered by instruction type, the scheduling control unit obtains the triggered instruction type and the number of triggers, combines the current instruction information of the decoding unit to determine whether to trigger, and then jumps to the debug program. After the expected warp finishes executing the debug program once, all warps resume the execution state of the original program and continue to execute.
[0150] 8. When the expected warp finishes executing all debug breakpoints, when the Host receives an unexpected warp executing to a breakpoint, an exit signal is directly generated.
[0151] 9. All warps continue to execute the original program until the end.
[0152] Figure 9 is a schematic diagram of the state jump of the warp provided by the present invention. As Figure 9 shown, the states of the warp are idle (IDLE), running (RUN), and debug (DEBUG). In the running state, the warp executes the original program after startup; in the debug state, it executes the debug program state. These three states can jump to each other.
[0153] Figure 10 is a schematic diagram of multi-warp debugging provided by the present invention. As Figure 10 shown, there are 2 breakpoints in the original program, and there are 3 types of warps, namely Warp K, Warp M, and Warp N. Warp M is the expected warp (Expect Warp), and the number can be one or more. Warp K is an unexpected warp that executes faster than the expected warp, and the number can be zero or more. Warp N is an unexpected warp that executes slower than the expected warp, and the number can be zero or more.
[0154] When entering the debug mode for the first time, Warp K first executes to the 1st breakpoint and waits in the else branch of the debug program. After Warp M executes to the 1st breakpoint, the Host enables the debug mode, then Warp M executes the debug program, and at the same time forces Warp N to enter the debug state and wait. Note that at this time, Warp N has not executed to the 1st breakpoint.
[0155] When entering the debug mode for the second time, Warp N reaches the 1st breakpoint and enters the else branch of the debug program to wait. After Warp M reaches the 2nd breakpoint, the Host enables the debug mode, then Warp M executes the debug program, and at the same time forces Warp K to enter the debug state and wait. Note that Warp K has not reached the 2nd breakpoint at this time.
[0156] After Warp M completes the debugging of all breakpoints. Thereafter, when Warp K reaches the 2nd breakpoint, since it is an unexpected warp, the Host directly notifies to exit. When Warp N reaches the 2nd breakpoint, since it is an unexpected warp, the Host directly notifies to exit.
[0157] Figure 11 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 11 shown, the electronic device may include: a processor (Processor) 1110, a communication interface (Communications Interface) 1120, a memory (Memory) 1130, and a communication bus (Communications Bus) 1140. Among them, the processor 1110, the communication interface 1120, and the memory 1130 complete mutual communication through the communication bus 1140. The processor 1110 can call the logical commands in the memory 1130 to execute the methods described in the above embodiments, for example:
[0158] Based on the comparison result between the instruction information of the original program executed by each warp and the instruction information stored in the target register group, and the warp identifier of each warp, determine that the expected warp triggers debugging; send a debug trigger message to the processor, so that the processor sends a debug enable message to each warp, and control the unexpected warps to wait for the expected warp to perform debugging; based on the debug enable message, control the expected warp to execute the debug program based on the instruction address of the debug program stored in the target register group; send a debug end message to the processor, so that the processor sends a debug stop message to each warp, and control each warp to continue executing the original program.
[0159] In addition, when the logical commands in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several commands for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0160] The processor in the electronic device provided by the embodiments of the present invention can call the logical instructions in the memory to implement the above method. The specific implementation manner is the same as that of the foregoing method embodiment, and the same beneficial effects can be achieved, which will not be elaborated herein.
[0161] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the methods provided in the above-mentioned embodiments.
[0162] The specific implementation manner is the same as that of the foregoing method embodiment, and the same beneficial effects can be achieved, which will not be elaborated herein.
[0163] The embodiments of the present invention provide a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method as described above.
[0164] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0165] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-threaded bundle debugging method, characterized in that: include: Determine the expected warp triggering debugging based on the comparison result of the instruction information of the original program executed by each warp and the instruction information stored in the target register group, and the warp identifier of each warp; each warp executes the original program in parallel; Sending debug trigger information to a processor so that the processor sends debug enable information to each thread warp, and controls an unexpected thread warp to wait for the expected thread warp to be debugged; Based on the debugging enable information, controlling the expected thread warp to execute the debugging program based on the instruction address of the debugging program stored in the target register group; Debug end information is sent to the processor, so that the processor sends debug stop information to each thread warp, and controls each thread warp to continue to execute the original program.
2. The multi-thread warp debugging method according to claim 1, characterized in that: The target register group is a memory-mapped input and output register; the target register group includes a mode enable register, a trigger information register and a debugger instruction address register; The mode enable register is used to define the debugging trigger mode and the number of the trigger information registers; the trigger mode includes an instruction address trigger mode or an instruction type trigger mode; The trigger information register is used to store instruction information for triggering debugging; the instruction information includes an instruction address or an instruction type; The debugger instruction address register is used to store the instruction address of the debugger.
3. The multi-thread warp debugging method according to claim 1, characterized in that: The step of determining the expected warp to trigger debugging based on the comparison result between the instruction information of the original program executed by each warp and the instruction information stored in the target register group, and the warp identifier of each warp, includes: Determine that each thread warp triggers debugging based on a comparison result between instruction information of the original program executed by each thread warp and instruction information stored in the target register group; Based on the comparison result between the warp identifiers of each warp and the warp identifier of the expected warp, it is determined that the expected warp triggers debugging.
4. The multi-thread warp debugging method according to claim 1, characterized in that: The debugging triggering information includes a warp identifier of a warp that triggers debugging; and the processor is configured to determine that the warp is the expected warp based on the warp identifier of the warp.
5. The multi-thread warp debugging method according to claim 1, characterized in that: After controlling the expected thread warp to execute the debug program based on the instruction address of the debug program stored in the target register group based on the debug enable information, the method further includes: Recording the executed instruction addresses of the original program; The processor is configured to control each thread warp to continue executing the original program based on the executed instruction address after debugging ends.
6. The multi-thread warp debugging method according to claim 1, characterized in that: The processor is configured to configure the target register group based on a trigger mode, an instruction address, an instruction type of the original program and an instruction address of the debugger.
7. A multi-threaded bundle debugging device, characterized in that: include: A comparison module, configured to determine an expected warp to trigger debugging based on a comparison result between instruction information of each warp executing the original program and instruction information stored in the target register group, and warp identifiers of each warp; Each thread warp executes the original program in parallel; A first sending module, configured to send debugging trigger information to a processor, so that the processor sends debugging enable information to each thread warp, and controls an unexpected thread warp to wait for the expected thread warp to be debugged; an execution module, configured to control the expected thread warp to execute the debug program based on an instruction address of the debug program stored in the target register group based on the debug enable information; The second sending module is used to send debugging end information to the processor, so that the processor sends debugging stop information to each thread warp to control each thread warp to continue executing the original program.
8. An artificial intelligence chip, characterized in that: include: The target register group is used to define the debugging trigger mode and store the instruction information for triggering debugging and the instruction address of the debugger; A scheduling control unit is connected to the target register group and is used to execute the multi-thread warp debugging method according to any one of claims 1 to 6, and control the expected warp in multiple warps for debugging.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the multi-thread warp debugging method according to any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-thread warp debugging method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Real-time debugging method on basis of embedded real-time systems
CN108319555A