A SIMT architecture accelerator card debugging method, device and medium

By reading preset registers, obtaining the storage address information of the SIMT architecture accelerator card debugging thread and copying the debugging information from GDDR to the host DDR, the problem of high cost and low efficiency of the existing SIMT architecture accelerator card debugging method is solved, and efficient and intuitive debugging information display is achieved.

CN119557237BActive Publication Date: 2025-05-02YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510112596.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-02
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing SIMT architecture accelerated card debugging methods are costly and inefficient, and the debugging information is not intuitive.

Method used

By reading the preset register, the area storage address calculation information of the target debugging thread is obtained, and the address of the debugging information to be stored in the GDDR is accurately calculated based on this information, the debugging information in the GDDR is copied to the DDR of the host, and the debugging information is displayed through the console of the host.

Benefits of technology

It realizes orderly and efficient storage and seamless data transmission of debugging information, improves debugging efficiency and user experience, and simplifies the steps of obtaining debugging information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557237B_ABST
    Figure CN119557237B_ABST
Patent Text Reader

Abstract

The present application discloses a SIMT architecture accelerator card debugging method, device and medium, and relates to the field of computing technology. The method comprises: when the target debugging thread needs to be debugged, reading a preset register to obtain the regional storage address calculation information of the target debugging thread, and calculating the address of the area to be stored in the GDDR for the debug information to be output based on the regional storage address calculation information; after the execution of the accelerator card kernel function is completed, copying the debug information in the GDDR to the DDR of the host, and displaying the debug information through the console of the host. The present application solves the problems of the existing SIMT architecture accelerator card debugging method being high in cost, low in efficiency, and non-intuitive debug information through the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computing technology, and in particular to a SIMT architecture accelerator card debugging method, device and medium. Background Art

[0002] In the field of computing technology, especially when debugging accelerator cards using SIMT (single instruction multiple threads) architecture, there are currently a series of technical challenges. Traditional code debugging technologies, such as the combined tool chain of GDB+OpenOCD+JTAG, can meet the debugging needs of sequentially executed code to a certain extent, but for parallel executed code, especially the code running in parallel on SIMT architecture accelerator cards, its debugging effect is not ideal.

[0003] The SIMT architecture accelerator card contains a large number of threads, which can execute the same instructions at the same time but process different data. This parallel execution mode greatly improves computing efficiency, but it brings new problems during debugging. Since multiple threads may output debugging information at the same thread branch at the same time, it is easy to cause address conflicts, making it impossible to save and display debugging information correctly.

[0004] In order to solve the above problems, the debugging method of the traditional SIMT architecture accelerator card is to mount a UART interface for each thread and transmit the debugging information through the UART interface. However, this approach has many inconveniences in practical applications. First, when the number of threads is large, a large number of UART interfaces need to be mounted and connected to the PC side through cables for reception and display, which not only increases the hardware cost, but also makes the construction of the debugging environment more complicated and time-consuming. Secondly, since each thread outputs debugging information through an independent UART interface, it is difficult to intuitively see the execution order and relationship between the threads, which brings great difficulties to discovering and solving problems in parallel code. Summary of the invention

[0005] The embodiments of the present application provide a SIMT architecture accelerator card debugging method, device and medium to solve the following technical problems: the existing SIMT architecture accelerator card debugging method has high cost, low efficiency, and the debugging information is not intuitive.

[0006] In the first aspect, an embodiment of the present application provides a SIMT architecture accelerator card debugging method, which is applied to a SIMT architecture accelerator card debugging system. The SIMT architecture accelerator card debugging system includes a host and a SIMT architecture accelerator card, the host includes a CPU and a DDR, and the SIMT architecture accelerator card includes a GPU and a GDDR. The method includes: when it is necessary to debug a target debugging thread, reading a preset register to obtain regional storage address calculation information of the target debugging thread, and based on the regional storage address calculation information, calculating the address of the area to be stored in the GDDR for the debugging information to be output; after the execution of the accelerator card kernel function is completed, copying the debugging information in the GDDR to the DDR of the host, and displaying the debugging information through the console of the host.

[0007] In one implementation of the present application, the method also includes: determining the number of threads of the SIMT architecture accelerator card based on the hardware attribute information of the SIMT architecture accelerator card, and determining the accelerator card memory starting address in GDDR for storing debug information output by each thread; and determining the host memory starting address in DDR for storing debug information.

[0008] In one implementation of the present application, the method further includes: defining an upper limit on the number of bytes of debugging information that a thread in a SIMT architecture accelerator card can output at a single time; and defining an upper limit on the number of times that a thread in a SIMT architecture accelerator card can output debugging information when debugging.

[0009] In one implementation of the present application, the method also includes: configuring a preset register so that the preset register can be read when the accelerator card is debugged; wherein the preset register includes: a thread number register, a clock count register, and an instruction count register; initializing the preset register so that the initial value is cached in the preset register.

[0010] In one implementation of the present application, the regional address storage calculation information includes: the thread number of the target debugging thread and the debugging information count of the target debugging thread; based on the regional storage address calculation information, calculating the address of the area to be stored in the GDDR for the debugging information to be output, specifically including: inputting the regional storage address calculation information into a preset regional address calculation model to determine the address of the area to be stored in the GDDR for the debugging information to be output; wherein the regional address calculation model is represented by the following formula:

[0011]

[0012] in, is the address of the area to be stored. is the starting address of the accelerator card memory, The upper limit of the number of bytes of debugging information that a thread can output at a time. is the number of threads, Counts the number of debugging information for the target debugging thread. The thread number of the target debug thread.

[0013] In one implementation of the present application, copying the debugging information in the GDDR to the DDR of the host specifically includes: determining the memory area to be copied in the GDDR based on a preset memory area calculation model, and copying the storage data in the memory area to be copied to the DDR of the host; wherein the memory area calculation model is represented by the following formula:

[0014]

[0015] in, is the memory area to be copied, The upper limit of the number of bytes of debugging information that a thread can output at a time. is the number of threads, The upper limit of the number of times debugging information is output.

[0016] In one implementation of the present application, debugging information is displayed through the host console, specifically including: formatting the debugging information copied to the DDR based on the host's standard input and output library functions; and outputting the formatted debugging information to the host's console interface by calling the iostream function in the std standard library.

[0017] In one implementation of the present application, it also includes: when displaying debugging information through the host console, generating a plurality of selectable debugging information display order types; wherein the debugging information display order types include: displaying in ascending order by thread number, displaying in ascending order by clock count, and displaying in ascending order by instruction count; when the user selects any debugging information display order type, the debugging information copied to the DDR is sorted and then output to the host console interface.

[0018] In a second aspect, an embodiment of the present application also provides a SIMT architecture accelerator card debugging device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a SIMT architecture accelerator card debugging method such as any one of the above items.

[0019] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for debugging a SIMT architecture accelerator card, which stores computer executable instructions. When the computer executable instructions are executed, a SIMT architecture accelerator card debugging method such as any one of the above items is implemented.

[0020] The SIMT architecture accelerator card debugging method, device and medium provided in the embodiments of the present application have the following beneficial effects:

[0021] 1. The regional storage address calculation information of the target debugging thread is obtained by reading the preset register, and the address of the debugging information to be stored in GDDR is accurately calculated based on this information, avoiding the possible scattered information storage problem in the traditional debugging method, making the storage of debugging information more orderly and efficient, and greatly improving the speed and accuracy of information retrieval.

[0022] 2. After the accelerator card kernel function is executed, the debugging information in GDDR is automatically copied to the host DDR and displayed in real time through the console. This not only simplifies the steps of obtaining debugging information, but also realizes seamless data transmission from the accelerator card to the host, improving debugging efficiency and user experience.

[0023] 3. This application supports hardware attribute information based on SIMT architecture acceleration cards, dynamically determines the number of threads and memory starting addresses, and defines the upper limit of the number of bytes and the upper limit of the output count of debugging information output by a thread at a time. The flexible memory management mechanism ensures the effective storage of debugging information and avoids waste or shortage of memory resources.

[0024] 4. By configuring preset registers (such as thread number register, clock count register, instruction count register) and initializing their values, the relevant attributes of debugging information, such as thread number, clock count, instruction count, etc., can be accurately recorded. This information is crucial for locating debugging problems and tracing program execution paths, and helps improve debugging accuracy and efficiency.

[0025] 5. This application not only supports outputting debugging information to the console, but also provides a variety of selectable debugging information display order types, such as displaying in ascending order by thread number, ascending order by clock count, and ascending order by instruction count. The intelligent display method helps users understand the program execution process more intuitively and quickly locate potential problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0027] Figure 1 A flow chart of a SIMT architecture accelerator card debugging method provided in an embodiment of the present application;

[0028] Figure 2 A SIMT architecture accelerator card debugging system provided in an embodiment of the present application;

[0029] Figure 3A schematic diagram of a debugging information storage structure in a GDDR provided in an embodiment of the present application;

[0030] Figure 4 A schematic diagram of the internal structure of a SIMT architecture accelerator card debugging device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0032] The embodiments of the present application provide a SIMT architecture accelerator card debugging method, device and medium to solve the following technical problems: the existing SIMT architecture accelerator card debugging method has high cost, low efficiency, and the debugging information is not intuitive.

[0033] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0034] Figure 1 A flow chart of a SIMT architecture accelerator card debugging method provided in an embodiment of the present application. Figure 2 A SIMT architecture accelerator card debugging system is provided in an embodiment of the present application, such as Figure 2 As shown, the SIMT architecture accelerator card debugging system includes a host and a SIMT architecture accelerator card, the host includes a CPU and a DDR, and the SIMT architecture accelerator card includes a GPU and a GDDR.

[0035] like Figure 1 As shown, a SIMT architecture accelerator card debugging method provided in an embodiment of the present application specifically includes the following steps:

[0036] Step 101: When debugging a target debugging thread is required, a preset register is read to obtain regional storage address calculation information of the target debugging thread, and based on the regional storage address calculation information, an address of a region to be stored in GDDR for debugging information to be output is calculated.

[0037] In one embodiment of the present application, in order to implement the SIMT architecture accelerator card debugging method of the present application, it is first necessary to determine the number of threads of the SIMT architecture accelerator card based on the hardware attribute information of the SIMT architecture accelerator card, and determine the accelerator card memory start address in GDDR for storing the debug information output by each thread; and determine the host memory start address in DDR for storing the debug information. In the embodiment of the present application, THREAD_NUM represents the number of threads, and PRINT_ADDR_BASE represents the accelerator card memory start address of the debug information output by each thread.

[0038] In one embodiment of the present application, in order to implement the SIMT architecture accelerator card debugging method of the present application, it is also necessary to define the upper limit of the number of bytes of debugging information output by a thread in the SIMT architecture accelerator card at a time according to actual debugging needs, and define the upper limit of the number of times the thread in the SIMT architecture accelerator card outputs debugging information when debugging. In the embodiment of the present application, PRINT_BYTE_MAX represents the upper limit of the number of bytes of debugging information output by a thread at a time, and PRINT_TIME_MAX represents the upper limit of the number of times the thread outputs debugging information when debugging.

[0039] In one embodiment of the present application, in order to implement the SIMT architecture accelerator card debugging method of the present application, it is also necessary to configure a preset register so that the preset register can be read when the accelerator card is debugged, and initialize the preset register so that the initial value is cached in the preset register. In the embodiment of the present application, the preset register includes: a thread number register, a clock count register, and an instruction count register; wherein the thread number register is represented by CSR_THREAD_ID, the clock count register is represented by CSR_MCYCLE, the instruction count register is represented by CSR_MINSTRET, the thread number in the thread number register is represented by thread_id, the clock count in the clock count register is represented by csr_mcycle, and the instruction count in the instruction count register is represented by csr_minstret.

[0040] Preferably, the cache initial value thread_id of the thread number register is 0, the cache initial value csr_mcycle of the clock count register is 0, and the cache initial value csr_minstret of the instruction count register is 0.

[0041] In one embodiment of the present application, when the target debugging thread needs to be debugged, the values ​​of the thread number thread_id, the clock count csr_mcycle and the instruction count csr_minstret are preset and read in sequence to obtain the debugging information count of the target debugging thread. The debugging information count of the embodiment of the present application is represented by print_num.

[0042] Furthermore, after reading the preset register, the thread number thread_id, the clock count csr_mcycle, the instruction count csr_minstret and the debugging information to be output are connected into a string through a separator (such as a colon) to form the debugging information to be output.

[0043] Further, the thread number thread_id of the target debugging thread and the debugging information count print_num of the target debugging thread are determined as the regional storage address calculation information, and based on the regional storage address calculation information, the address of the area to be stored in the GDDR for the debugging information to be output is calculated.

[0044] Specifically, the regional storage address calculation information is input into a preset regional address calculation model to determine the address of the region to be stored in the GDDR where the debug information to be outputted is stored.

[0045] In one embodiment of the present application, the regional address calculation model is represented by the following formula:

[0046]

[0047] in, is the address of the area to be stored. is the starting address of the accelerator card memory, The upper limit of the number of bytes of debugging information that a thread can output at a time. is the number of threads, Counts the number of debugging information for the target debugging thread. The thread number of the target debug thread.

[0048] Further, after determining the address of the storage area of ​​the debugging information to be output in GDDR, the debugging information to be output composed of the above string is saved to the memory area of ​​GDDR with the starting address of print_addr, such as Figure 3 shown. Figure 3 A schematic diagram of a debugging information storage structure in GDDR provided in an embodiment of the present application.

[0049] Step 102: After the execution of the accelerator card kernel function is completed, the debugging information in the GDDR is copied to the DDR of the host, and the debugging information is displayed through the console of the host.

[0050] In one embodiment of the present application, after the execution of the accelerator card kernel function is completed, that is, after the debugging information of each thread is saved to the corresponding memory area in sequence according to the execution order of the kernel function, the debugging information in GDDR is copied to the DDR of the host.

[0051] Specifically, based on a preset memory area calculation model, a memory area to be copied in the GDDR is determined, and the stored data in the memory area to be copied is copied to the DDR of the host.

[0052] In one embodiment of the present application, the memory area calculation model is represented by the following formula:

[0053]

[0054] in, is the memory area to be copied, The upper limit of the number of bytes of debugging information that a thread can output at a time. is the number of threads, The upper limit of the number of times debugging information is output.

[0055] Furthermore, debugging information is displayed through the console of the host.

[0056] Specifically, based on the standard input and output library function of the host, the debugging information copied to the DDR is formatted; by calling the iostream function in the std standard library, the formatted debugging information is output to the console interface of the host.

[0057] In one embodiment of the present application, it also includes: when displaying debugging information through the host console, generating a plurality of selectable debugging information display order types; wherein the debugging information display order types include: displaying in ascending order by thread number, displaying in ascending order by clock count, and displaying in ascending order by instruction count; when the user selects any debugging information display order type, the debugging information copied to the DDR is sorted and then output to the console interface of the host.

[0058] Preferably, the thread numbers thread_id are displayed in ascending order, that is, the debugging information of thread number 0 is displayed in sequence according to the sequence in the DDR, and then the debugging information of thread number 1 is displayed, and so on.

[0059] Preferably, the clock counts csr_mcycle are displayed in ascending order, that is, the debugging information with a smaller clock count is displayed first, and if the clock counts are the same, the debugging information with a smaller thread number is displayed first.

[0060] Preferably, the instruction count csr_minstret is displayed in ascending order, that is, the debugging information with a smaller instruction count is displayed first, and if the instruction counts are the same, the debugging information with a smaller thread number is displayed first.

[0061] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a SIMT architecture acceleration card debugging device, whose structure is as follows: Figure 2 shown.

[0062] Figure 4 A schematic diagram of the internal structure of a SIMT architecture acceleration card debugging device provided in an embodiment of the present application. Figure 4 As shown, the device includes:

[0063] at least one processor 401;

[0064] and, a memory 402 communicatively coupled to the at least one processor;

[0065] The memory 402 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 401 to enable at least one processor 401 to:

[0066] When the target debugging thread needs to be debugged, the preset register is read to obtain the regional storage address calculation information of the target debugging thread, and based on the regional storage address calculation information, the address of the area to be stored in the GDDR of the debugging information to be output is calculated;

[0067] After the accelerator card kernel function is executed, the debugging information in the GDDR is copied to the host's DDR, and the debugging information is displayed through the host's console.

[0068] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for debugging a SIMT architecture accelerator card, storing computer executable instructions, wherein the computer executable instructions are set as:

[0069] When the target debugging thread needs to be debugged, the preset register is read to obtain the regional storage address calculation information of the target debugging thread, and based on the regional storage address calculation information, the address of the area to be stored in the GDDR of the debugging information to be output is calculated;

[0070] After the accelerator card kernel function is executed, the debugging information in the GDDR is copied to the host's DDR, and the debugging information is displayed through the host's console.

Claims

1. A SIMT architecture accelerator card debugging method, characterized in that: Applied to a SIMT architecture accelerator card debugging system, the SIMT architecture accelerator card debugging system includes a host and a SIMT architecture accelerator card, the host includes a CPU and a DDR, the SIMT architecture accelerator card includes a GPU and a GDDR, and the method includes: When the target debugging thread needs to be debugged, a preset register is read to obtain the regional storage address calculation information of the target debugging thread, and based on the regional storage address calculation information, an address of the region to be stored in the GDDR of the debugging information to be output is calculated; After the execution of the accelerator card kernel function is completed, the debugging information in the GDDR is copied to the DDR of the host, and the debugging information is displayed through the console of the host; The method also includes: Based on the hardware attribute information of the SIMT architecture accelerator card, determine the number of threads of the SIMT architecture accelerator card, and determine the accelerator card memory start address in the GDDR for storing the debugging information output by each thread; and Determine a host memory start address in the DDR for storing debug information; The method also includes: Define the upper limit of the number of bytes of debugging information output by a thread in the SIMT architecture accelerator card at a time; Defines an upper limit on the number of times debugging information is output when a thread in the SIMT architecture accelerator card is debugged.

2. A SIMT architecture accelerator card debugging method according to claim 1, characterized in that: The method further comprises: Configuring the preset register so that the preset register can be read when the accelerator card is debugged; wherein the preset register includes: a thread number register, a clock count register, and an instruction count register; The preset register is initialized so that an initial value is cached in the preset register.

3. A SIMT architecture accelerator card debugging method according to claim 2, characterized in that: The area address storage calculation information includes: the thread number of the target debugging thread and the debugging information count of the target debugging thread; Calculating the address of the area to be stored in the GDDR of the debug information to be output based on the area storage address calculation information specifically includes: Inputting the regional storage address calculation information into a preset regional address calculation model to determine the regional address to be stored in the GDDR for the debug information to be output; The regional address calculation model is represented by the following formula: in, is the address of the area to be stored. is the starting address of the accelerator card memory, The upper limit of the number of bytes of debugging information that a thread can output at a time. is the number of threads, Counts the number of debugging information for the target debugging thread. The thread number of the target debug thread.

4. A SIMT architecture accelerator card debugging method according to claim 3, characterized in that: Copying the debugging information in the GDDR to the DDR of the host specifically includes: Based on a preset memory area calculation model, determine a memory area to be copied in the GDDR, and copy the storage data in the memory area to be copied to the DDR of the host; The memory area calculation model is represented by the following formula: in, is the memory area to be copied, The upper limit of the number of bytes of debugging information that a thread can output at a time. is the number of threads, The upper limit of the number of times debugging information is output.

5. A SIMT architecture accelerator card debugging method according to claim 1, characterized in that: Displaying the debugging information through the console of the host specifically includes: Based on the standard input and output library functions of the host, the debugging information copied to the DDR is formatted; By calling the iostream function in the std standard library, the formatted debugging information is output to the console interface of the host.

6. A SIMT architecture accelerator card debugging method according to claim 5, characterized in that: The method further comprises: When displaying the debugging information through the console of the host, a plurality of selectable debugging information display order types are generated; wherein the debugging information display order types include: displaying in ascending order by thread number, displaying in ascending order by clock count, and displaying in ascending order by instruction count; When the user selects any debug information display sequence type, the debug information copied to the DDR is sorted and then output to the console interface of the host.

7. A SIMT architecture accelerator card debugging device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a SIMT architecture acceleration card debugging method as described in any one of claims 1-6.

8. A non-volatile computer storage medium for debugging a SIMT architecture accelerator card, storing computer executable instructions, characterized in that: When the computer executable instructions are executed, a SIMT architecture accelerator card debugging method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • GPU-based data-intensive task scheduling method and device and storage medium

    CN118964029A