GPU shader tracking and recording system
By introducing a shader tracking module into the shaders of the GPU chip, splicing and transmitting tokens for GPU performance analysis, the problem of lack of underlying information on the software side is solved, and more efficient and accurate performance analysis is achieved.
Patent Information
- Application Number
- CN202510704683.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In the prior art, the software side lacks underlying information when performing GPU performance analysis, resulting in low analysis accuracy and the data obtained by the shader tracking module are messy, which reduces the analysis efficiency.
A shader tracking module is introduced into the shaders of the GPU chip, which is used to obtain tracking data from multiple submodules and splice them into tokens, and is provided to the software side for analysis through the cache module.
It improves the accuracy and efficiency of GPU performance analysis, making the acquired data more comprehensive and detailed, and can more accurately identify performance bottlenecks.
Smart Images

Figure CN120235744A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of integrated circuit design, and particularly to a GPU shader tracing and recording system. Background Art
[0002] Currently, when performing GPU performance analysis on the software side, it is usually analyzed based on the directly obtained performance metrics of the GPU chip. The performance metrics may include CPU running duration, GPU running duration, storage utilization rate, etc. However, these performance metrics usually do not involve the underlying information of GPU operation, resulting in incomplete information for the software side to perform GPU performance analysis, and thus the accuracy of GPU performance analysis on the software side is relatively low.
[0003] In response to the above problems, the prior art has proposed GPU shader tracing. GPU shader tracing is a technical means for recording and analyzing shader-related events and operations in a GPU chip, which can obtain detailed information during the execution of the shader, such as the execution status of instructions, the execution duration of instructions, the statistical count of instruction types, etc., so as to help the software side deeply understand the operation of the shader, and thus perform debugging, performance optimization, and research on the graphics rendering process.
[0004] However, since the shader contains multiple sub-modules, and each sub-module can provide corresponding shader tracing data, the shader tracing module will obtain a large amount of shader tracing data, making it necessary for the software side to perform GPU performance analysis based on the numerous and messy shader tracing data, reducing the efficiency of GPU performance analysis. Therefore, how to improve the efficiency of GPU performance analysis has become an urgent problem to be solved. Summary of the Invention
[0005] In response to the above technical problems, the technical solution adopted by the present invention is as follows: A GPU shader tracing and recording system, the system includes: a software side, a GPU chip, and a device memory. Among them, the GPU chip includes a shader, the shader includes A sub-modules and a shader tracing module, A is an integer greater than zero, and the device memory includes a cache module.
[0006] The shader tracing module is used to respectively obtain corresponding shader tracing data from the A sub-modules.
[0007] The shader tracing module is further used to splice the obtained shader tracing data into tokens according to a preset method.
[0008] The shader tracing module is further used to send the tokens to the cache module.
[0009] The cache module is used to receive the access from the software side to provide tokens to support the software side in performing performance analysis of the GPU chip.
[0010] Compared with the prior art, the present invention has obvious beneficial effects. By means of the above technical solution, a GPU shader tracing and recording system provided by the present invention can achieve quite remarkable technological progressiveness and practicality, and has wide utilization value in the industry. It has at least the following beneficial effects: The present invention provides a GPU shader tracing and recording system, which includes: a software side, a GPU chip, and device memory. Among them, the GPU chip includes shaders, and the shaders include A sub-modules and a shader tracing module, where A is an integer greater than zero. The device memory includes a cache module. The shader tracing module is used to respectively obtain corresponding shader tracing data from the A sub-modules, and the shader tracing module is further used to splice the obtained shader tracing data into tokens in a preset manner. The shader tracing module is further used to send the tokens to the cache module, and the cache module is used to receive the access from the software side to provide tokens to support the software side in performing performance analysis of the GPU chip.
[0011] It can be seen that by obtaining the shader tracing data of each sub-module through the shader tracing module in the shader, splicing the shader tracing data of each sub-module into tokens, and then sending the tokens to the cache module to be provided to the software side, it is convenient for the software side to parse the tokens to perform performance analysis of the GPU chip, making the obtained shader tracing data more comprehensive, detailed and related, thereby improving the efficiency and accuracy of the software side to analyze the performance of the GPU chip according to the tokens to find out the performance bottleneck. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a schematic diagram of the architecture of a GPU shader tracing and recording system provided in Embodiment 1 of the present invention; Figure 2 It is a schematic diagram of the architecture of a token generation system based on GPU shader tracing data provided in Embodiment 2 of the present invention; Figure 3 It is a schematic diagram of the process when a computer program in a timestamp token generation system provided in Embodiment 3 of the present invention is processed and executed; Figure 4Schematic diagram of an architecture of a cache system based on shader trace data provided in the fourth embodiment of the present invention; Figure 5 Schematic diagram of an architecture of an address translation system based on shader trace data provided in the fifth embodiment of the present invention. Detailed implementation manners
[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] The first embodiment of the present invention provides a GPU shader trace recording system. Refer to Figure 1 , which is a schematic diagram of an architecture of a GPU shader trace recording system provided in the first embodiment of the present invention. The system includes: a software side, a GPU chip, and a device memory. Among them, the GPU chip includes a shader, the shader includes A sub-modules and a shader trace module, A is an integer greater than zero, and the device memory includes a cache module; The shader trace module is configured to respectively obtain corresponding shader trace data from the A sub-modules; The shader trace module is further configured to splice the obtained shader trace data into tokens in a preset manner; The shader trace module is further configured to send the tokens to the cache module; The cache module is configured to receive an access from the software side to provide tokens to support the software side to perform performance analysis of the GPU chip.
[0016] Among them, a shader is an important part of a graphics rendering pipeline. A shader can be used to implement various operations and effects on graphics. A shader can include multiple levels of shaders, such as a vertex shader, a pixel shader, a geometry shader, etc.
[0017] A sub-module may refer to a module in a shader that performs a specific function. A shader usually includes multiple sub-modules with different specific functions.
[0018] The device memory can adopt Double Data Rate Synchronous Dynamic Random Access Memory (DDR), High Bandwidth Memory (HBM), etc.
[0019] Shader trace data can be used to record shader-related events and operations in a GPU chip to obtain detailed information during the shader execution process, such as the execution status of instructions, the execution duration of instructions, the statistical count of instruction types, etc.
[0020] Tokens can be used to integrate related shader trace data together to obtain data with a higher degree of integration, so that the subsequent software side can quickly analyze based on the tokens, deeply understand the running situation of the shader, and thus perform debugging, performance optimization, and research on the graphics rendering process.
[0021] The cache module can be used to temporarily store tokens. Since the cache module is in the device memory, the cache module can support reading by the software side, that is, receive access from the software side, to provide tokens to support the software side to perform performance analysis of the GPU chip.
[0022] In a specific implementation manner, for any sub-module, the sub-module includes a number of target interfaces, each target interface corresponds to shader trace data of different trace data types, and all target interfaces are connected to the shader trace module; The obtaining the corresponding shader trace data from the A sub-modules respectively includes: For any sub-module, obtain the shader trace data of the corresponding trace data type from each target interface included in the sub-module.
[0023] Among them, different sub-modules have different functions, and the same sub-module can include multiple target interfaces. The shader trace data provided by different target interfaces is also different for the domestic trace data types. The trace data types can include instruction status, instruction execution duration, instruction type, etc.
[0024] Specifically, providing shader trace data to the shader trace module through each target interface of each sub-module improves the comprehensiveness and richness of the shader trace data. The provided shader trace data is also more detailed and more underlying, facilitating the software side to parse the tokens after the tokens are formed and analyze the GPU performance to find performance bottlenecks.
[0025] In a specific implementation manner, the shader trace data corresponds to the token type and the position information in the token of the corresponding token type; The splicing the obtained shader trace data into tokens according to a preset method includes: For any token type, splice a number of shader trace data belonging to the token type according to the position information of each shader trace data in the token of the corresponding token type to obtain the token corresponding to the token type.
[0026] Among them, the shader trace data corresponds to the token type. A single shader trace data corresponds to one token type, and a single token type can correspond to multiple different shader trace data.
[0027] Specifically, in this embodiment, for the convenience of transmission, storage, access, etc., the token adopts a fixed bit width form, and the fixed bit width can be 64bit. Since the bit width is fixed, for the tokens of a certain token type, the shader trace data corresponding to each position should also be fixed, which is convenient for the software side to parse the tokens. That is, the shader trace data corresponds to the token type and the position information in the tokens of the corresponding token type.
[0028] In a specific implementation manner, the token types at least include instruction type, instruction count type, performance count type, and timestamp type.
[0029] Among them, the token types can also include user data type, thread start type, thread end type, global start type, global end type, etc.
[0030] It should be noted that the implementer can construct new token types according to the actual requirements of the software side, and determine the shader trace data required for the new token types and the position information of the required shader trace data in the tokens of the new token types.
[0031] In a specific implementation manner, the shader trace module further includes a mask register, and the mask register includes a first mask sub-register; The first mask sub-register includes B first mask bits, where B is the number of token types, and each first mask bit corresponds to one token type respectively; For any first mask bit, if the first mask bit is the first preset value, then the tokens of the token type corresponding to the first mask bit are not generated.
[0032] Among them, the first mask bit can be binary, then the first mask bit can take the value of the first preset value or the second preset value. The first preset value can be 1, and the second preset value can be 0.
[0033] Specifically, since the shader trace data is continuously generated when the GPU executes the rendering task, the shader trace module needs to form a large number of tokens and then transmit the tokens to the cache module for storage. Considering the storage pressure of the cache module and the token generation pressure of the shader trace module, a mask register is set in the shader trace module. The mask register can be used to control the generation of only the tokens that the implementer is interested in, and the first mask sub-register can be used to control the generation of only the tokens of the token types that the implementer is interested in.
[0034] In a specific embodiment, the mask register includes a second mask sub-register; The second mask sub-register includes a plurality of second mask bits, where each second mask bit corresponds to an instruction type respectively; For any second mask bit, if the second mask bit is a first preset value, no token including the instruction type corresponding to the second mask bit is generated.
[0035] Wherein, the second mask bit can be binary, and the second mask bit can take a value of the first preset value or a second preset value.
[0036] Specifically, the second mask sub-register can be used to control the generation of only tokens of instruction types that the implementer is interested in.
[0037] In a specific embodiment, the mask register includes a third mask sub-register; The third mask sub-register includes a plurality of third mask bits, where each third mask bit corresponds to an instruction count type respectively; For any third mask bit, if the third mask bit is a first preset value, no token including the instruction count type corresponding to the third mask bit is generated.
[0038] Wherein, the third mask bit can be binary, and the third mask bit can take a value of the first preset value or a second preset value.
[0039] Specifically, the third mask sub-register can be used to control the generation of only tokens of instruction count types that the implementer is interested in.
[0040] In a specific embodiment, the token corresponding to the performance count type includes a high-order token and a low-order token, and the mask register includes a fourth mask sub-register; The fourth mask sub-register includes a plurality of fourth mask bits, where each fourth mask bit corresponds to a performance count type respectively; For any fourth mask bit, if the fourth mask bit is a first preset value, no low-order token corresponding to the performance count type is generated; If the fourth mask bit is a second preset value, no high-order token corresponding to the performance count type is generated.
[0041] Wherein, the fourth mask bit can be binary, and the fourth mask bit can take a value of the first preset value or a second preset value.
[0042] Specifically, since there are many performance count types, the token corresponding to the performance count type is divided into a high-order token and a low-order token, and the fourth mask sub-register can be used to control the generation of only the high-order token or the low-order token of the performance count type that the implementer is interested in.
[0043] In this embodiment, the shader tracing module in the shader is used to obtain the shader tracing data of each sub-module, splice the shader tracing data of each sub-module into a token, and then send the token to the cache module to be provided to the software side, facilitating the software side to parse the token for GPU chip performance analysis. This makes the obtained shader tracing data more comprehensive, detailed, and related, thereby improving the efficiency and accuracy of the software side to analyze the GPU chip performance based on the token to find performance bottlenecks.
[0044] Embodiment 2 of the present invention provides a token generation system based on GPU shader tracing data. Refer to Figure 2 , which is a schematic architecture diagram of a token generation system based on GPU shader tracing data provided by Embodiment 2 of the present invention. The system includes: a GPU chip, where the GPU chip includes a shader, and the shader includes A sub-modules and a shader tracing module, and A is an integer greater than zero; The shader tracing module is used as the master device to obtain the performance count data of A sub-modules as slave devices, and the sub-module includes a performance counter; The shader tracing module is also used to splice the obtained performance count data and the positioning data corresponding to the rendering task during the execution of the rendering task to obtain a token corresponding to the performance count type; The shader tracing module is also used to splice the first rendering command identifier and the first rendering sub-command identifier to obtain a token corresponding to the global start type at the start of the rendering task; The shader tracing module is also used to downsample the first rendering command identifier and the first rendering sub-command identifier in the token corresponding to the global start type to obtain a second rendering command identifier and a second rendering sub-command identifier when a thread starts to execute a rendering task, and splice the second rendering command identifier, the second rendering sub-command identifier, and the positioning data corresponding to the thread to obtain a token corresponding to the thread start type.
[0045] Among them, the performance counter can be used to collect and record the performance data of the sub-module in real time, and can quantify various performance-related metrics, such as CPU usage rate, memory occupancy rate, data transfer rate, bandwidth, etc.
[0046] In a sub-module, the performance counter in the sub-module records performance count data. Correspondingly, the shader tracing module is used as the master device to obtain the performance count data from the performance counters of A sub-modules as slave devices.
[0047] A rendering task can include a series of processes such as scene setup, material and texture processing, lighting calculation, and finally image generation. In the vertex processing stage, a vertex shader is required to process vertex data and pass information such as calculated vertex colors and texture coordinates to the pixel shader. The pixel shader combines the lighting model, material properties, texture data, etc. in the scene and determines the color of each pixel through complex mathematical calculations.
[0048] The positioning data can include one or more of queue identifier, workgroup identifier, processing unit identifier, array processor identifier, processing element unit identifier, thread identifier, block identifier, chain identifier, and event identifier. In this embodiment, the positioning data included in the token corresponding to the performance counting type can include block identifier, chain identifier, processing unit identifier, and event identifier.
[0049] In a rendering task, after the vertex data processing is completed, primitives are constructed based on the processing results and converted into pixels on the screen. For each pixel covered by a primitive, corresponding calculation processing is required. Then the rendering task will include multiple rendering commands in multiple stages. The first rendering command identifier can be used to identify the rendering command corresponding to the current processing object, and the first rendering sub-command identifier can be used to identify the rendering command of the sub-stage related to the current processing object. In the same rendering task, the first rendering sub-command identifier is incremented. Thus, the first rendering command identifier and the first rendering sub-command identifier effectively record the processing connection between multiple shaders, which is used to evaluate the balance of rendering task allocation, the size of the graphics corresponding to the rendering task, the rendering pressure of the lines or points generated by the rendering task on the shader, etc.
[0050] For a single thread, the thread executes corresponding processing according to the assigned rendering command. In order to enable the token corresponding to the thread to record the rendering command information of the first enabled shader, the first rendering command identifier and the first rendering sub-command identifier can be provided to the token corresponding to the thread. However, since both the first rendering command identifier and the first rendering sub-command identifier are 32-bit, if the token corresponding to the thread directly uses the first rendering command identifier and the first rendering sub-command identifier, it will be unable to record the positioning data of the thread, such as queue identifier, workgroup identifier, processing unit identifier, array processor identifier, processing element unit identifier, thread identifier, etc. Therefore, in this embodiment, downsampling is performed on the first rendering command identifier and the first rendering sub-command identifier to obtain the second rendering command identifier and the second rendering sub-command identifier.
[0051] Specifically, the second rendering command identifier can be the lower 16 bits of the first rendering command identifier, that is, by taking the modulo operation of the first rendering command identifier by 2 16 and the second rendering sub-command identifier can be the lower 8 bits of the first rendering sub-command identifier.
[0052] In the rendering task, some rendering operations may not result in the execution of the pixel shader due to factors such as clipping, depth testing, and stencil testing. Therefore, in order to ensure the accuracy of the recording of the second rendering command identifier, a 16-bit width is used instead of a smaller width. Similarly, the second rendering sub-command identifier may be continuous in the vertex shader but not continuous in the pixel shader. Therefore, in order to ensure the accuracy of the recording of the second rendering sub-command identifier, an 8-bit width is used instead of a smaller width.
[0053] In a specific implementation, the step of concatenating the performance count data obtained during the execution of the rendering task and the positioning data corresponding to the rendering task to obtain a token corresponding to the performance count type includes: Concatenating the performance count data obtained during the execution of the rendering task, the positioning data corresponding to the rendering task, and the performance count identifier to obtain a token corresponding to the performance count type; If the performance count identifier is a first preset value, then concatenating the performance count data obtained during the execution of the rendering task, the positioning data corresponding to the rendering task, and the performance count identifier to obtain a low-order token corresponding to the performance count type; If the performance count identifier is a second preset value, then concatenating the performance count data obtained during the execution of the rendering task, the positioning data corresponding to the rendering task, and the performance count identifier to obtain a high-order token corresponding to the performance count type.
[0054] Among them, the performance count identifier can be binary, so the performance count identifier can take the value of the first preset value or the second preset value. Since there are many performance count types, the tokens corresponding to the performance count types are divided into high-order tokens and low-order tokens.
[0055] In a specific implementation, the GPU chip further includes a temporary buffer; The step of obtaining the performance count data of A sub-modules serving as slave devices includes: Obtaining a token of the timestamp type, and determining a sampling time point according to the token of the timestamp type and a preset sampling period; If there is no unfinished signal at the sampling time point, then sending a performance count data sampling request to A sub-modules serving as slave devices at the sampling time point; After receiving the performance count data sampling request, the A sub-modules serving as slave devices respectively send the corresponding performance count data to the temporary buffer; When the performance count data corresponding to the A sub-modules serving as slave devices has not yet reached the temporary buffer, generating the unfinished signal.
[0056] Among them, the temporary buffer can refer to a peek buffer. During the shader processing, when intermediate viewing or debugging of performance counter data is required, the performance counter data will be sent to such a buffer, facilitating implementers to obtain information such as the intermediate state of the performance counter data without affecting the shader processing flow.
[0057] Specifically, a token of timestamp type can be used to represent an absolute timestamp. Based on the absolute timestamp and a preset sampling period, a number of sampling time points can be determined.
[0058] For each sampling time point, when there is no outstanding signal at that sampling time point, a performance counter data sampling request is sent to A slave sub-modules at the sampling time point.
[0059] When there is an outstanding signal at that sampling time point, the current sampling time point is skipped, and it is waited for the next sampling time point to determine whether there is an outstanding signal.
[0060] In one embodiment, the implementer can directly send a sampling signal to send a performance counter data sampling request to A slave sub-modules without waiting for the sampling time point and without considering the outstanding signal.
[0061] In a specific embodiment, the shader tracing module is further configured to generate a token corresponding to the global end type when receiving a rendering completion enable signal.
[0062] Among them, since tokens usually contain instruction types and positioning data, there can be associations between tokens based on the instruction types and positioning data. For tokens containing performance counter data, there may also be associations due to the performance counter data. When performing performance analysis on the software side, associated tokens can be processed simultaneously to improve the analysis efficiency and accuracy.
[0063] In a specific embodiment, the shader tracing module further includes a target array register; The target array register includes C array identification bits, where C is the number of array processors, and each array identification bit corresponds to an array processor respectively; For any array identification bit, if the array identification bit is a first preset value, tokens are not generated according to the instructions in the array processor corresponding to the array identification bit.
[0064] Among them, since a large number of tokens are formed during the execution of the rendering task, resulting in excessive storage pressure on the cache module, in this embodiment, a specific array processor is specified through the target array register, and only the tokens formed by the rendering instructions executed in the specific array processor are sent to the cache module. In this embodiment, by default, only the first array identification bit is set to the second preset value, and the other array identification bits are set to the first preset value.
[0065] In a specific implementation manner, the GPU chip further includes a device memory, and the device memory includes a cache module. The tokens have corresponding first priorities according to their corresponding token types. When there are at least two tokens that need to be stored in the cache module at the same time point, according to the first priorities corresponding to each token respectively, the token with the highest first priority is stored in the cache module, and the other tokens are discarded.
[0066] Among them, when there are at least two tokens that need to be stored in the cache module at the same time point, the token with the highest first priority is selected according to the pre-set priority and stored in the cache module, and the other tokens are discarded.
[0067] In an implementation manner, the implementer can adjust the first priorities corresponding to each token according to the actual situation, but it should be ensured that the first priorities corresponding to each token according to its corresponding token type are different. It should be noted that for the tokens of the timestamp type, by default, its first priority is the highest and cannot be changed.
[0068] In a specific implementation manner, the shader tracing module further includes a second priority register, and the second priority register includes B priority identification bits, where B is the number of token types, and each priority identification bit corresponds to a token type respectively. Correspondingly, when there are at least two tokens that need to be stored in the cache module at the same time point, according to the first priorities corresponding to each token respectively, the token with the highest first priority is stored in the cache module, and the other tokens are discarded, including: When there are at least two tokens that need to be stored in the cache module at the same time point, according to the first priorities corresponding to each token respectively, the token with the highest first priority is stored in the cache module. Among the other tokens, the tokens corresponding to the instruction type with the priority identification bit being the second preset value are stored later, and the tokens corresponding to the instruction type with the priority identification bit being the first preset value are discarded.
[0069] Among them, the implementer can set the second priority register to identify the token type. When the token corresponding to the instruction type whose priority identification bit is the second preset value needs to be discarded, it is changed to delayed storage, that is, stored in the next cycle, thereby ensuring that the token that the implementer is interested in will not be discarded due to the first priority.
[0070] In a specific implementation, when there are at least two tokens that need to be stored in the cache module at the same time point, according to the first priorities corresponding to the respective tokens, the token with the highest first priority is stored in the cache module, and among the other tokens, the token corresponding to the instruction type whose priority identification bit is the second preset value is stored in a delayed manner, and the token corresponding to the instruction type whose priority identification bit is the first preset value is discarded, including: When there are at least two tokens that need to be stored in the cache module at the same time point, if there are tokens of the same token type, the tokens of the same token type are merged into a single token; Then, according to the first priorities corresponding to each token, the token with the highest first priority is stored in the cache module. Among other tokens, the token corresponding to the instruction type whose priority identification bit is the second preset value is stored later, and the token corresponding to the instruction type whose priority identification bit is the first preset value is discarded.
[0071] Among them, tokens of the same token type can be merged into a single token, thereby reducing the number of tokens.
[0072] In this embodiment, tokens of different token types are generated by the shader tracking module in the shader, so that the tokens of different token types can accurately characterize the execution status of different aspects of a specific rendering task, making the acquired shader tracking data more comprehensive, detailed and correlated, and thread start type tokens are formed through downsampling operations, which reduces connecting lines, improves the efficiency of token generation, and also improves the efficiency and accuracy of the software side in analyzing the GPU chip performance based on tokens to find performance bottlenecks.
[0073] This embodiment 3 provides a timestamp token generation system. Figure 3 , is a flowchart of a computer program being processed and executed in a timestamp token generation system provided in Embodiment 3 of the present invention, the system comprising: a GPU chip, a processor, and a memory storing the computer program, wherein the GPU chip comprises a shader, and the shader comprises a shader tracking module, and when the computer program is executed by the processor, the following steps are implemented: S301, after the shader tracking module generates a token, performing cycle statistics to obtain a cycle statistics value; S302. When the shader tracing module generates a token other than the timestamp type, determine the relative timestamp of the current token according to the cycle statistic value after the last generated token. S303. When the cycle statistic value after the last generated token meets the first preset condition, generate a token of the timestamp type, and the token of the timestamp type includes an absolute timestamp.
[0074] Among them, the cycle may refer to the clock cycle of the GPU chip, and the cycle statistic value can be used to represent the number of cycles from the generation time point of the last generated token to the current time point. Then, the cycle statistic value can be regarded as a relative timestamp.
[0075] When the shader tracing module generates a token of other token types except the timestamp type, determine the relative timestamp of the current token according to the cycle statistic value after the last generated token.
[0076] In a specific implementation manner, the first preset condition is that the cycle statistic value after the last generated token is greater than a preset cycle threshold.
[0077] Among them, the reason for generating the token of the timestamp type may be that no token is generated within a certain time, that is, the cycle statistic value after the last generated token is greater than the preset cycle threshold.
[0078] In a specific implementation manner, the bit width corresponding to the relative timestamp is 4 bits.
[0079] Among them, since the relative timestamp only records the relative time, a lower bit width can be adopted, such as a 4-bit width.
[0080] In a specific implementation manner, when the computer program is executed by a processor, the following steps are also implemented: S304. When a token discard operation occurs, generate a token of the timestamp type.
[0081] Among them, the reason for generating the token of the timestamp type may be generated when a token discard operation occurs. The token discard operation may be due to multiple tokens being generated simultaneously, and according to the first priority, tokens with non-highest first priority are discarded. The token discard operation may also be a token discard operation caused by the need to replace tokens in the cache module.
[0082] In a specific implementation manner, when the computer program is executed by a processor, the following steps are also implemented: S305. When the shader tracing module is turned on, generate a token of the timestamp type.
[0083] Among them, the generation reason of the timestamp type token can be generated when the shader tracing module is turned on to record the time when the shader tracing module is turned on. The time point when the shader tracing module is turned on can refer to the time point when the shader tracing module receives the start tracing instruction.
[0084] In a specific implementation manner, when the computer program is executed by a processor, the following steps are further implemented: S306, when the shader tracing module is turned off, generate a timestamp type token.
[0085] Among them, the generation reason of the timestamp type token can be generated when the shader tracing module is turned off to record the time when the shader tracing module is turned off. The time point when the shader tracing module is turned off can refer to the time point when the shader tracing module receives the stop tracing instruction.
[0086] In a specific implementation manner, the generation reason is also included in the timestamp type token.
[0087] Among them, the timestamp type token includes the generation reason. The bit width corresponding to the generation reason in the timestamp type token is G, and G is a positive integer. In this embodiment, G is 4. Each bit corresponding to the generation reason in the timestamp type token corresponds to a generation reason. When a bit corresponding to the generation reason in the timestamp type token is the first preset value, it means that the generation reason corresponding to this bit takes effect. Generally, only one bit in each bit corresponding to the generation reason in the timestamp type token is the first preset value.
[0088] In a specific implementation manner, the bit width corresponding to the absolute timestamp is 56bit, and every time a cycle passes, the absolute timestamp is incremented by one; When the shader tracing module receives a reset signal or a refresh signal, reset the absolute timestamp; When the absolute timestamp exceeds the range represented by its bit width, reset the absolute timestamp.
[0089] Among them, the absolute timestamp can be used to record the real time when the shader tracing module performs tracing. Therefore, a relatively long bit width is required, such as 56bit.
[0090] When the shader tracing module receives a reset signal or a refresh signal, or when the absolute timestamp exceeds the range represented by its bit width, reset the absolute timestamp to record the real time again.
[0091] In this embodiment, the relative timestamp of the next token other than the timestamp type is determined according to the periodic statistical value after the token is generated, and a token of the timestamp type including the absolute timestamp is generated according to a preset condition, so that more detailed time information can be provided, GPU performance analysis can be avoided, and the efficiency of performance analysis is improved.
[0092] Embodiment 4 of the present invention provides a cache system based on shader tracing data. Refer to Figure 4 , which is a schematic architecture diagram of a cache system based on shader tracing data provided by Embodiment 4 of the present invention. The system includes: a software side, a GPU chip, and device memory. Among them, the GPU chip includes a shader, the shader includes a shader tracing module, the device memory includes a cache module, the cache module includes a first cache sub-module, a second cache sub-module, and a sub-module pointing register. The first cache sub-module corresponds to a first flag bit of a first cache register, and the second cache sub-module corresponds to a second flag bit of a second cache register; The shader tracing module is configured to send the formed token to the first cache sub-module when the sub-module pointing register is a first preset value; The shader tracing module is further configured to send the formed token to the second cache sub-module when the sub-module pointing register is a second preset value; The first cache sub-module is configured to store the token in the first cache sub-module when the token is received; The second cache sub-module is configured to store the token in the second cache sub-module when the token is received; When the first flag bit is a second preset value, the first cache sub-module generates first cache interrupt information and sends it to the software side; When the second flag bit is a second preset value, the second cache sub-module generates second cache interrupt information and sends it to the software side; The software side is configured to read the token from the first cache sub-module after receiving the first cache interrupt information; The software side is further configured to read the token from the second cache sub-module after receiving the second cache interrupt information.
[0093] Among them, the sub-module pointing register can be used to identify the storage location, and the storage location is one of the first cache sub-module and the second cache sub-module.
[0094] The first flag bit of the first cache register can be used to identify the empty / full state of the first cache sub-module, and the second flag bit of the second cache register can be used to identify the empty / full state of the second cache sub-module.
[0095] The first cache interruption information can be used to identify the full state of the first cache sub-module, and the second cache interruption information can be used to identify the full state of the second cache sub-module.
[0096] Specifically, the first cache sub-module generates the first cache interruption information and sends it to the driver on the software side, and the second cache sub-module generates the second cache interruption information and sends it to the driver on the software side.
[0097] In a specific implementation manner, the cache module further includes a first control register and a second control register; When the first control register is a first preset value, and the first flag bit or the second flag bit is a second preset value, the cache module stops receiving tokens.
[0098] Among them, the first control register and the second control register can be used to determine how to store subsequent tokens after the currently stored cache sub-module is in a full state.
[0099] Specifically, when the first control register is a first preset value, regardless of which of the first cache sub-module and the second cache sub-module is the currently stored cache sub-module and is in a full state, the cache module stops receiving tokens.
[0100] In a specific implementation manner, when the first control register is a second preset value, the second controller is a first preset value, and the first flag bit is a second preset value, after the cache module receives a token, it overwrites the first cache sub-module; When the first control register is a second preset value, the second controller is a first preset value, and the second flag bit is a second preset value, after the cache module receives a token, it overwrites the second cache sub-module.
[0101] Among them, when the first control register is a second preset value and the second controller is a first preset value, at this time, regardless of which of the first cache sub-module and the second cache sub-module is the currently stored cache sub-module and is in a full state, overwriting is performed in the currently stored cache sub-module, rather than switching the cache sub-module.
[0102] In a specific implementation manner, when the first control register is a second preset value, the second controller is a second preset value, and the first flag bit is a second preset value, the sub-module pointer register is set to a second preset value; When the first control register is a second preset value, the second controller is a second preset value, and the second flag bit is a second preset value, the sub-module pointer register is set to a first preset value.
[0103] Wherein, when the first control register is at a second preset value and the second controller is at a second preset value, at this time, regardless of which of the first cache sub-module and the second cache sub-module is the currently stored cache sub-module in a full state, if the other cache sub-module is not in a full state, then a cache sub-module switch is performed.
[0104] In a specific implementation manner, when the first control register is at a second preset value, the second controller is at a second preset value, and the first flag bit is at a second preset value, setting the sub-module pointing register to a second preset value includes: When the first control register is at a second preset value, the second controller is at a second preset value, and the first flag bit is at a second preset value, obtain the second flag bit; If the second flag bit is at a first preset value, then set the sub-module pointing register to a second preset value; If the second flag bit is at a second preset value, then wait until the second flag bit is converted to a first preset value, and then set the sub-module pointing register to a second preset value; When the first control register is at a second preset value, the second controller is at a second preset value, and the second flag bit is at a second preset value, setting the sub-module pointing register to a first preset value includes: When the first control register is at a second preset value, the second controller is at a second preset value, and the second flag bit is at a second preset value, obtain the first flag bit; If the first flag bit is at a first preset value, then set the sub-module pointing register to a first preset value; If the first flag bit is at a second preset value, then wait until the first flag bit is converted to a first preset value, and then set the sub-module pointing register to a first preset value.
[0105] Wherein, when the first control register is at a second preset value and the second controller is at a second preset value, at this time, regardless of which of the first cache sub-module and the second cache sub-module is the currently stored cache sub-module in a full state, if the other cache sub-module is also in a full state, then wait until the other cache sub-module changes to an empty state, and immediately perform a cache sub-module switch.
[0106] In a specific implementation manner, the first preset value is 0 and the second preset value is 1.
[0107] In a specific implementation manner, the software side is further used to access the first flag bit. When the first flag bit is at a second preset value, read a token from the first cache sub-module; The software side is also used to access the second flag bit. When the second flag bit is a second preset value, a token is read from the second cache sub-module.
[0108] Among them, the software side can obtain the empty / full status of the first cache sub-module and the second cache sub-module by directly accessing the first flag bit and the second flag bit. Furthermore, when the first cache sub-module or the second cache sub-module is in a full state, the full cache sub-module is read without waiting for the first cache interrupt information and the second cache interrupt information.
[0109] In a specific implementation manner, after the software side finishes reading a token from the first cache sub-module, the software side is also used to send a clearing signal to the first cache register to clear the first cache register; After the software side finishes reading a token from the second cache sub-module, the software side is also used to send a clearing signal to the second cache register to clear the second cache register.
[0110] Among them, after the software side finishes reading a certain cache sub-module, a clearing signal is sent to the cache sub-module so that the cache sub-module can receive and store tokens again subsequently.
[0111] In this embodiment, the cache module is divided into a first cache sub-module and a second cache sub-module, enabling the software side to read tokens and the shader tracing module to write tokens simultaneously, improving the efficiency of performance analysis. Moreover, the empty / full status of the first cache sub-module and the second cache sub-module is represented by the first flag bit of the first cache register and the second flag bit of the second cache register. When any cache sub-module is full, the software side is notified through cache interrupt information, improving the timeliness of the software side reading tokens and also improving the efficiency of performance analysis.
[0112] Embodiment 5 of the present invention provides an address translation system based on shader tracing data. Refer to Figure 5 , which is a schematic architecture diagram of an address translation system based on shader tracing data provided in Embodiment 4 of the present invention. The system includes: a GPU chip and a device memory. Among them, the GPU chip includes a shader and an address translation unit. The shader includes a shader tracing module. The device memory includes a cache module. The shader tracing module includes a warm-up unit; The shader tracing module is used to send the virtual address corresponding to the formed token to the warm-up unit; The warm-up unit is used to send the received virtual address to the address translation unit; The address translation unit is used to translate the received virtual address into a corresponding physical address and send the correspondence between the virtual address and the physical address to the warm-up unit; The preheating unit is further configured to store the correspondence between the received virtual address and the physical address; The shader tracing module is further configured to send the formed token to the cache module according to the correspondence between the virtual address and the physical address stored in the preheating unit.
[0113] Among them, the address translation unit can be used to support the conversion from a virtual address to a physical address. The preheating unit can be used to send the virtual address to be translated to the address translation unit in advance, and store the translation result of the address translation unit in its own storage space for subsequent direct provision when the translation result is needed, without waiting for the address translation unit to perform the translation again.
[0114] In a specific implementation manner, the sending the received virtual address to the address translation unit includes: Sending the received virtual address to the address translation unit in sequence.
[0115] Among them, after the preheating unit receives a virtual address, it can send the received virtual address to the address translation unit, that is, send the received virtual address to the address translation unit in sequence. Each time the address translation unit returns a translation result, the translation result is stored in the preheating unit.
[0116] In a specific implementation manner, the sending the received virtual address to the address translation unit includes: Sending the received virtual address to the address translation unit in batches.
[0117] Among them, after the preheating unit receives a number of virtual addresses, it can send the received virtual addresses to the address translation unit in batches, that is, send the received virtual addresses to the address translation unit in batches. The address translation unit returns a batch of translation results, and the batch of translation results is stored in the preheating unit.
[0118] In a specific implementation manner, the preheating unit includes D preheating storage pages, where D is a positive integer; The storing the correspondence between the received virtual address and the physical address includes: Storing the correspondence between the received virtual address and the physical address in the corresponding preheating storage page.
[0119] Among them, the preheating storage page can be used to store the correspondence between the virtual address and the physical address.
[0120] In a specific implementation manner, the cache module includes E buffer areas, and the address translation module includes F translation storage pages; D = min(E, F), where min() is a function to take the minimum value.
[0121] Among them, the number of preheating storage pages can be set to the minimum of the number of buffer areas in the cache module and the number of translation storage pages in the address translation module, so as to avoid waste of storage space.
[0122] In a specific implementation manner, the shader further includes A sub-modules, where A is an integer greater than zero; The shader tracing module is further configured to respectively obtain corresponding shader tracing data from the A sub-modules.
[0123] Among them, for any one of the sub-modules, the sub-module includes a plurality of target interfaces, and each target interface corresponds to shader tracing data of a different tracing data type. All the target interfaces are connected to the shader tracing module. The step of respectively obtaining corresponding shader tracing data from the A sub-modules may refer to, for any one of the sub-modules, respectively obtaining shader tracing data of the corresponding tracing data type from each target interface included in the sub-module.
[0124] In a specific implementation manner, the shader tracing module is further configured to splice the obtained shader tracing data into tokens in a preset manner.
[0125] Among them, the shader tracing data corresponds to the token type and the position information in the token of the corresponding token type. The step of splicing the obtained shader tracing data into tokens in a preset manner may refer to, for any one token type, splicing a plurality of shader tracing data belonging to the token type according to the position information of each shader tracing data in the token of the corresponding token type, so as to obtain the token corresponding to the token type.
[0126] In a specific implementation manner, the system further includes a software side; The cache module is configured to receive access from the software side to provide tokens to support the software side to perform performance analysis of the GPU chip.
[0127] In this embodiment, by adding a preheating unit in the shader tracing module, the preheating unit pre-sends the virtual address of the token to the address translation unit for address translation, and stores the address translation result in the preheating unit, so that when the shader tracing module sends the token to the cache module, the address translation result can be directly obtained from the preheating unit, improving the storage efficiency of storing the token in the cache module, avoiding the address translation unit being in an unhit state for a long time, and thus improving the efficiency of GPU performance analysis.
[0128] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention disclosed is defined by the appended claims.
Claims
1. A GPU shader tracing and recording system, characterized in that The system includes: a software side, a GPU chip, and device memory. Among them, the GPU chip includes shaders, and the shaders include A sub-modules and a shader tracing module, where A is an integer greater than zero, and the device memory includes a cache module; The shader tracing module is used to respectively obtain corresponding shader tracing data from the A sub-modules; The shader tracing module is further used to splice the obtained shader tracing data into tokens in a preset manner; The shader tracing module is further used to send the tokens to the cache module; The cache module is used to receive access from the software side to provide tokens to support the software side to perform performance analysis of the GPU chip.
2. The GPU shader trace recording system according to claim 1, wherein For any one of the sub-modules, the sub-module includes a number of target interfaces, and each target interface corresponds to shader tracing data of a different tracing data type, and all the target interfaces are connected to the shader tracing module; The step of respectively obtaining corresponding shader tracing data from the A sub-modules includes: For any one of the sub-modules, respectively obtain shader tracing data of the corresponding tracing data type from each target interface included in the sub-module.
3. The GPU shader trace recording system according to claim 1, characterized in that, The shader tracing data corresponds to a token type and position information in the token of the corresponding token type; The step of splicing the obtained shader tracing data into tokens in a preset manner includes: For any one token type, splice a number of shader tracing data belonging to the token type according to the position information of each shader tracing data in the token of the corresponding token type to obtain the token corresponding to the token type.
4. The GPU shader trace recording system according to claim 3, wherein The token types at least include an instruction type, an instruction count type, a performance count type, and a timestamp type.
5. The GPU shader trace recording system according to claim 4, wherein The shader tracing module further includes a mask register, and the mask register includes a first mask sub-register; The first mask sub-register includes B first mask bits, where B is the number of token types, and each first mask bit corresponds to a token type; For any one of the first mask bits, if the first mask bit is a first preset value, then a token of the token type corresponding to the first mask bit is not generated.
6. The GPU shader trace recording system according to claim 5, wherein, The mask register includes a second mask sub-register; The second mask sub-register includes a number of second mask bits, where each second mask bit corresponds to an instruction type; For any one of the second mask bits, if the second mask bit is a first preset value, then a token including the instruction type corresponding to the second mask bit is not generated.
7. The GPU shader trace recording system according to claim 5, wherein The mask register includes a third mask sub-register; The third mask sub-register includes a number of third mask bits, where each third mask bit corresponds to an instruction count type; For any one of the third mask bits, if the third mask bit is a first preset value, then a token including the instruction count type corresponding to the third mask bit is not generated.
8. The GPU shader trace recording system according to claim 5, characterized in that, The token corresponding to the performance count type includes a high-order token and a low-order token, and the mask register includes a fourth mask sub-register; The fourth mask sub-register includes a number of fourth mask bits, where each fourth mask bit corresponds to a performance count type; For any fourth mask bit, if the fourth mask bit is the first preset value, a low-order token corresponding to the performance count type is not generated; If the fourth mask bit is the second preset value, a high-order token corresponding to the performance count type is not generated.
Citation Information
Patent Citations
A client identity verification method and device
CN109862009A
Programmable ray tracing with hardware acceleration on graphics processor
CN110858410A
Storage method for marks of shader tracing module
CN117745515A
Method and apparatus for generation of programmable shader configuration information from state-based control information and program instructions
US20040012597A1
Graphics processing unit with unified vertex cache and shader register file
US20080074430A1