Data visualization method and device, electronic equipment and computer readable medium

By inserting callback tasks into the execution task queue of the graphics processor, determining and visualizing the start and end time of the GPU task, the problem of difficulty in accurately visualizing the execution of large-scale distributed tasks on the GPU in the prior art is solved, and accurate recording and visualization of the execution of GPU tasks is achieved, helping users quickly locate performance problems.

CN120029729APending Publication Date: 2025-05-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411877189.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately visualize the execution of large-scale distributed tasks on the GPU, especially in the performance bottlenecks and load uneven problems caused by the asynchronousness of the GPU, and it is difficult to accurately locate the problem points through existing visualization tools.

Method used

By inserting callback tasks into the execution task queue of the graphics processor, the start and end time of the task is determined and visually displayed, the time deviation caused by the asynchronousness of the CPU and GPU is avoided.

Benefits of technology

It realizes accurate recording and visualization of GPU task execution, helps users quickly locate performance problems and reduces the time cost of debugging and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029729A_ABST
    Figure CN120029729A_ABST
Patent Text Reader

Abstract

The invention provides a data visualization method and device, electronic equipment and a computer readable medium, and relates to the technical field of data visualization, in particular to the technical field of distributed task, graphics processor task visualization and the like. According to the specific implementation scheme, callback tasks are inserted in a first position and a second position of an execution task queue corresponding to the graphics processor; determining task starting time of the graphics processor task according to an execution result of the callback task inserted into the first position; according to an execution result of the callback task inserted into the second position, determining task end time of the graphics processor task; and carrying out visual display on the task starting time and the task ending time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data visualization, and particularly to technical fields such as distributed task and graphics processing unit (GPU) task visualization. Specifically, the present disclosure relates to a data visualization method, an apparatus, an electronic device, and a computer-readable storage medium. Background Art

[0002] Large-scale distributed tasks, such as large-scale distributed neural network training, due to complex computing tasks and involving the collaborative work of multiple GPUs (Graphics Processing Units), result in long task times and difficult debugging.

[0003] Moreover, since the GPU executes tasks asynchronously, performance bottlenecks or uneven loads are likely to occur during the training process. These problems can be analyzed and located through visualization tools. Summary of the Invention

[0004] The present disclosure provides a data visualization method, an apparatus, an electronic device, and a computer-readable storage medium.

[0005] According to a first aspect of the present disclosure, there is provided a data visualization method, the method comprising:

[0006] Inserting callback tasks at a first position and a second position in an execution task queue corresponding to a graphics processing unit;

[0007] Determining a task start time of a graphics processing unit task according to an execution result of the callback task inserted at the first position; and determining a task end time of the graphics processing unit task according to an execution result of the callback task inserted at the second position;

[0008] Visualizing and displaying the task start time and the task end time;

[0009] Wherein, the execution task queue is a queue formed by a plurality of execution tasks executed by the graphics processing unit in an execution order; the graphics processing unit task includes one or more of the execution tasks; the first position is in front of a first task and adjacent to the first task, the second position is behind a second task and adjacent to the second task; the first task is the execution task with the first execution order among the execution tasks corresponding to the graphics processing unit task, and the second task is the execution task with the last execution order among the execution tasks corresponding to the graphics processing unit task.

[0010] According to a second aspect of the present disclosure, there is provided a data visualization apparatus, the apparatus comprising:

[0011] A function module, used for inserting a callback task into a first position and a second position of an execution task queue corresponding to a graphics processor;

[0012] The data processing module is used to determine the task start time of the graphics processor task according to the execution result of the callback task inserted at the first position; and determine the task end time of the graphics processor task according to the execution result of the callback task inserted at the second position;

[0013] A visualization module, used to visualize the task start time and the task end time;

[0014] Among them, the execution task queue is a queue composed of multiple execution tasks executed on the graphics processor in an execution order; the graphics processor task includes one or more of the execution tasks; the first position is located before the first task and adjacent to the first task, and the second position is located after the second task and adjacent to the second task; the first task is the execution task with the first execution order in the execution tasks corresponding to the graphics processor task, and the second task is the execution task with the last execution order in the execution tasks corresponding to the graphics processor task.

[0015] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device comprising:

[0016] at least one processor; and

[0017] A memory in communication with the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the visualization method.

[0019] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned visualization method.

[0020] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above-mentioned visualization method when executed by a processor.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0023] Figure 1 is a flowchart of a visualization method provided by an embodiment of the present disclosure;

[0024] Figure 2 is a flowchart of some steps of another visualization method provided by an embodiment of the present disclosure;

[0025] Figure 3 is a flowchart of some steps of another visualization method provided by an embodiment of the present disclosure;

[0026] Figure 4 is a process schematic diagram of a specific embodiment of another visualization method provided by an embodiment of the present disclosure;

[0027] Figure 5 is a structural schematic diagram of a visualization device provided by an embodiment of the present disclosure;

[0028] Figure 6 It is a block diagram of an electronic device used to implement the visualization method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] In some related technologies, CPU (Central Processing Unit) visualization tools (such as Nsight) are used to analyze distributed tasks. These visualization tools mainly infer the execution status of GPU tasks by recording the task start and task end time on the CPU side.

[0031] However, due to the asynchrony of the GPU, the task time recorded on the CPU side cannot accurately reflect the execution time of the GPU task, especially when there is an uncertain delay between the GPU task and the CPU task. This leads to the insufficient visualization accuracy of the visualization tool on the CPU side, which in turn makes it impossible for users to accurately locate performance bottlenecks.

[0032] Moreover, existing visualization tools can only display the execution of the overall task, and cannot support the execution of fine-grained GPU tasks such as forward calculation, back propagation, and parameter optimization. When users face problems, they cannot intuitively locate the problem points through visualization tools, which makes the analysis and optimization processes lengthy and cumbersome.

[0033] At the same time, existing visualization work lacks support for multiple GPU devices, making it difficult to accurately analyze the communication delays and task timing between multiple GPU devices, and unable to provide users with effective multi-device debugging methods.

[0034] The data visualization method and device, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure are intended to solve at least one of the above-mentioned technical problems in the prior art.

[0035] The data visualization method provided in the embodiments of the present disclosure may be executed by an electronic device such as a terminal device or a server, and the terminal device may be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method may be implemented by a processor calling a computer-readable program instruction stored in a memory. Alternatively, the method may be executed by a server.

[0036] Figure 1 FIG. 1 is a flow chart of a data visualization method provided by an embodiment of the present disclosure. Figure 1 As shown in , the data visualization method provided by the embodiment of the present disclosure may include step S110, step S120, and step S130.

[0037] In step S110, a callback task is inserted into a first position and a second position of an execution task queue corresponding to the graphics processor;

[0038] In step S120, the task start time of the GPU task is determined according to the execution result of the callback task inserted at the first position; the task end time of the GPU task is determined according to the execution result of the callback task inserted at the second position;

[0039] In step S130, the task start time and the task end time are visualized;

[0040] Among them, the execution task queue is a queue composed of multiple execution tasks executed on the graphics processor in the order of execution; the graphics processor task includes one or more execution tasks; the first position is located before the first task and adjacent to the first task, and the second position is located after the second task and adjacent to the second task; the first task is the execution task with the first execution order in the execution tasks corresponding to the graphics processor task, and the second task is the execution task with the last execution order in the execution tasks corresponding to the graphics processor task.

[0041] For example, the data visualization method according to the embodiment of the present disclosure is used to visualize a GPU task executed on a GPU.

[0042] GPU is a processor specially designed to efficiently process large amounts of data in parallel and has hundreds or thousands of CUDA (Compute Unified Device Architecture) Cores.

[0043] CUDA is a general-purpose parallel computing architecture that enables GPUs to solve complex computing problems. CUDA Core is the most basic processing unit of the GPU. Specific instructions and tasks are processed on the CUDA Core. The essence of GPU parallel processing of data is that multiple CUDA Cores are running and processing data at the same time.

[0044] The GPU tasks executed on the GPU can be split into multiple execution tasks and assigned to multiple CUDA Cores for parallel execution. Each CUDA Core has a corresponding CUDA Stream. All CUDA operations are executed on the CUDA Stream, which is a queue of execution tasks executed on the CUDA Core in the order of execution. It should be emphasized that the execution tasks on the CUDAStream can only be executed in the order of execution and cannot be executed in parallel. Only when the execution tasks in the previous order are completed, the execution tasks in the later order can be executed.

[0045] The same GPU task is split into multiple execution tasks, which may be assigned to multiple CUDA Cores for execution. The execution tasks corresponding to different GPU tasks may be assigned to the same CUDA Core for execution.

[0046] Therefore, the CUDA Stream corresponding to the same CUDA Core can be an execution task queue composed of execution tasks corresponding to different GPU tasks in the execution order.

[0047] In some possible implementations, in step S110, the task queue corresponding to the graphics processor may be a CUDA Stream of a CUDA Core of the GPU. A GPU has multiple CUDA Cores, and thus multiple CUDA Streams.

[0048] That is, there may be multiple execution task queues corresponding to a GPU, and the first position and the second position of the execution task queue corresponding to the GPU may be the first position and the second position of the execution task queue corresponding to a CUDA Core. That is, the callback task is inserted into the first position and the second position of the execution task queue corresponding to a specific CUDA Core.

[0049] It may also be the first position and the second position of the execution task queue corresponding to multiple CUDA Cores, that is, the callback task is inserted into the first position and the second position of the execution task queue corresponding to different CUDA Cores respectively.

[0050] The first position is located before the first task and adjacent to the first task, and the second position is located after the second task and adjacent to the second task.

[0051] The first task may be the execution task with the first execution order among the execution tasks corresponding to the GPU task in the execution task queue, and the second task may be the execution task with the last execution order among the execution tasks corresponding to the same GPU task in the execution task queue.

[0052] The first task and the second task may be the same task or different tasks.

[0053] Specifically, multiple execution tasks corresponding to the same GPU task can be obtained in the execution task queue, and the execution order of the multiple execution tasks corresponding to the same GPU task can be determined according to the execution task queue. The first task is the first execution task, and the last task is the second execution task.

[0054] In some special cases, there is only one execution task corresponding to the same GPU task in the execution task queue, and the execution task is both the first task and the second task.

[0055] For example, a CUDA Stream corresponding to a GPU is ABCDE, where A, B, C, D, and E are execution tasks corresponding to different GPU tasks.

[0056] B, C, and E are execution tasks corresponding to the same GPU task. B is the execution task with the first execution order in the execution tasks corresponding to the GPU task in CUDA Stream, that is, the first task. B is the execution task with the last execution order in the execution tasks corresponding to the GPU task in CUDA Stream, that is, the second task. The first position is between A and B, and the second position is after E. The new execution task queue is AXBCDEX, and X is the callback task.

[0057] If only C is the execution task corresponding to a GPU task, C is both the first execution task in the execution order of the execution tasks corresponding to the GPU task in CUDA Stream, that is, the first task, and the last execution task in the execution order of the execution tasks corresponding to the GPU task in CUDA Stream, that is, the second task, then the first position is between B and C, and the second position is between C and D. The new execution task queue is AXBCXDE, and X is the callback task.

[0058] In some possible implementations, the callback task may be a task executed by a CPU corresponding to the GPU, such as obtaining data from the CPU.

[0059] In some possible implementations, in step S120, the time when the CPU executes the callback task inserted at the first position can be used as the task start time of the GPU task, and the time when the CPU executes the callback task inserted at the second position can be used as the task end time of the GPU task.

[0060] Since the callback task is inserted into the execution task queue, that is, CUDA Stream, the callback task will be executed only when all other execution tasks before the first task are completed, and since the callback task is executed on the CPU, it will not affect the execution of the execution task corresponding to the GPU task. In other words, when executing the callback task, the GPU will also execute the first task, and the execution time of the callback task is the execution time of the first task.

[0061] At the same time, the callback task inserted at the second position will be executed only after the second task is finished, and its execution time is the time when the second task is finished.

[0062] Since the first task is the first task in the execution order of the execution tasks corresponding to the GPU task in the execution task queue, and the second task is the last task in the execution order of the execution tasks corresponding to the GPU task in the execution task queue, that is to say, the start of the first task is the start of the GPU task, and the end of the second task is the end of the GPU task.

[0063] Therefore, the execution time of the callback task inserted at the first position is the task start time of the GPU task, and the execution time of the callback task inserted at the second position is the task end time of the GPU task.

[0064] That is to say, the task start time of the GPU task can be obtained by determining the execution time of the callback task inserted at the first position based on the execution result of the callback task inserted at the first position, and the task end time of the GPU task can be obtained by determining the execution time of the callback task inserted at the second position based on the execution result of the callback task inserted at the second position.

[0065] In some possible implementations, in step S130, the task start time and task end time of the GPU task can be visualized in the form of a bar chart or the like and displayed on an interactive device, so that the user can obtain the task start time and task end time of the GPU task according to the interactive device.

[0066] The user can analyze the GPU task according to the obtained task start time and task end time to locate the problem when an exception occurs in the GPU task.

[0067] In the data visualization method provided by the embodiment of the present disclosure, a callback task is inserted into the execution task queue, and the task start time and task end time of the GPU task are obtained through the execution result of the callback task. Since the callback task is inserted into the execution task queue of the GPU, it reflects the execution order of tasks on the GPU and does not depend on the scheduling and recommendation of the CPU. This avoids the deviation caused by the asynchrony of the CPU and GPU (after the CPU issues a task schedule, the GPU may delay the actual startup due to the execution load of the current task, resulting in the task start time recorded on the CPU side being earlier than the actual execution time of the GPU task), thereby achieving accurate recording of the execution status of the GPU tasks.

[0068] The visualization method provided by the embodiment of the present disclosure is introduced in detail below.

[0069] As described above, the execution task is a task obtained by splitting the GPU task. Therefore, the execution task can be a computing task or other tasks other than the computing task, such as a communication task.

[0070] In some possible implementations, the user is more concerned about the execution status of the computing task in the GPU task, and hopes to obtain the start time and end time of the computing task, without obtaining the execution status of other tasks.

[0071] Therefore, in some possible time modes, the first task may be the computing task that is first in the execution order among the computing tasks corresponding to the GPU task, and the second task may be the computing task that is last in the execution order among the computing tasks corresponding to the GPU task.

[0072] Since the first task and the second task are computing tasks, it can be deduced that at least one of the multiple execution tasks in the execution task queue is a computing task, and at least one of the execution tasks corresponding to the GPU task is a computing task.

[0073] Of course, if you want to obtain the start and end time of other execution tasks, you only need to insert the callback function before the execution task starts and insert the callback function after the execution task ends.

[0074] In some possible implementations, the insert callback task may be implemented through a CUDA Callback (Uniform Device Architecture Callback) function.

[0075] CUDA Callback is a callback mechanism used when executing tasks on the GPU. Through the CUDA callback function, a callback operation can be triggered when a specific event occurs on the GPU (such as the start or end of a task), and the callback task is executed on the CPU.

[0076] This mechanism allows the start and end times of tasks to be accurately recorded directly in the GPU execution queue, avoiding scheduling delays between the CPU and GPU, thereby providing accurate task timing information.

[0077] Specifically, a CUDA Callback is inserted before a task start function corresponding to the first task, and a CUDA Callback is inserted before a task end function corresponding to the second task.

[0078] The task start function corresponding to the first task may be the first computing node of the GPU task, and the task end function corresponding to the second task may be the last computing node of the GPU task.

[0079] As described above, the callback task may be a task executed by the CPU corresponding to the GPU.

[0080] In some possible implementations, the callback task may be a task for obtaining the time of the CPU corresponding to the GPU, that is, obtaining the current time of the CPU when executing the callback task. Therefore, the CPU time obtained by the callback task is the execution time of the callback task.

[0081] Therefore, as described above, the CPU time obtained by the callback task inserted in the first position can be determined as the task start time of the GPU task; and the CPU time obtained by the callback task inserted in the second position can be determined as the task end time of the GPU task.

[0082] The callback task directly obtains the CPU time, so it is more convenient to directly obtain the execution time of the callback task than to analyze the execution process of the callback task to obtain the execution time of the callback task.

[0083] The callback task may also include other tasks executed on the CPU, such as reading the CPU status, etc., which will not be described here.

[0084] In some possible implementations, since the GPU and the corresponding CPU may be used to execute large-scale distributed tasks, the CPU corresponding to the GPU may be a CPU located in a distributed cluster. These CPUs and GPUs located in the distributed cluster need to work as a system, which requires that the timestamps of the CPUs located in the distributed cluster have consistency in order to record the execution timing of GPU tasks on different GPUs.

[0085] Because the system clocks of different CPUs may differ slightly, the system clocks of the CPUs need to be synchronized to ensure the consistency of the timestamps of the CPUs in the distributed cluster, to avoid affecting the visualization of GPU tasks on different GPUs, and thus to avoid affecting the user's analysis of the task status of distributed tasks.

[0086] Figure 2 A schematic diagram of a process for synchronizing the system clock of the CPU of a distributed cluster is shown. Figure 2 As shown, synchronizing the system clocks of the CPUs of the distributed cluster may include step S210.

[0087] In step S210, the system clock of the CPU corresponding to the graphics processor is synchronized with the system clocks of other CPUs of the distributed cluster to which the CPU belongs based on the network time protocol.

[0088] Among them, NTP (Network Time Protocol) is a standard network time synchronization protocol. By synchronizing the time of each server with a standard time source (such as an NTP server), the time deviation of different machines can be minimized, usually achieving millisecond-level or even higher accuracy.

[0089] That is, in some possible implementations, in step S210, the synchronization of the system clock may be achieved by synchronizing all CPUs of the distributed cluster with a standard time source (such as an NTP server).

[0090] In some possible implementations, NTP allows configuration of an automatic synchronization frequency to ensure that the system clock time remains consistent even if a slight time drift occurs in the CPU of the distributed system during the execution of distributed tasks.

[0091] In order to meet the needs of long-term distributed task execution, the synchronization frequency of NTP can be appropriately increased to minimize the time difference between the system clocks of different CPUs.

[0092] In some possible implementations, even with the support of NTP, network jitter or slight drift of the system may still cause tiny time errors, especially when distributed tasks run for a long time, such errors will gradually accumulate.

[0093] Therefore, it is necessary to dynamically correct the time offset of the system clock of each CPU during the running of distributed tasks.

[0094] Figure 3 A schematic diagram of a process for dynamically correcting the system clock of each CPU is shown, as shown in FIG. Figure 3 As shown, dynamically correcting the system clock of each CPU may include step S310, step S320, and step S330.

[0095] In step S310, a time synchronization request signal sent by the CPU corresponding to the graphics processor to other CPUs and a response signal sent by other CPUs to the CPU corresponding to the graphics processor after receiving the time synchronization request signal are obtained;

[0096] In step S320, a time offset corresponding to the CPU of the graphics processor is calculated according to the time of the CPU of the graphics processor when the time synchronization request signal is sent and the time of the CPU of the graphics processor when the response signal is received;

[0097] In step S330, the task start time and the task end time are adjusted based on the time offset corresponding to the CPU corresponding to the GPU.

[0098] In some possible implementations, in step S310, the CPU of the distributed system will periodically send a time synchronization request signal to other CPUs of the distributed system and record the time of sending the time synchronization request signal. After receiving the time synchronization request signal, other CPUs will send a response signal to the CPU that sent the time synchronization request signal. The CPU receives the response signal and records the time of receiving the response signal.

[0099] The time synchronization request signal includes a timestamp of the current time when the time synchronization request signal is sent, and the response signal includes a timestamp of the current time when the CPU receives the time synchronization request signal.

[0100] In some possible implementations, in step S320, after receiving the response signal, the CPU calculates the delay in the round-trip communication time between the CPUs based on the time difference between the time of sending the time synchronization request signal and the time of receiving the response signal, and calculates the time offset between the CPUs based on the delay in the round-trip communication time and the time when the CPU receives the time synchronization request signal.

[0101] Specifically, the time when other CPUs receive the synchronization request signal can be calculated based on the timestamp of the current time when the time synchronization request signal is sent and the delay of the communication round-trip time, and the difference between the calculation result and the timestamp of the current time when the CPU receives the time synchronization request signal contained in the response signal can be used as the time offset.

[0102] In some possible implementations, multiple communication delays may be obtained by sending request signals at different times multiple times, and the delays calculated each time may be averaged and filtered to exclude abnormal data and obtain a more accurate time offset.

[0103] In some possible implementations, in step S330 , the acquired task start time and task end time may be adjusted based on the time offset between the CPU and other CPUs to eliminate the time offset between different CPUs.

[0104] In some possible implementations, the timestamp corresponding to the CPU may be obtained based on the acquired time offset, that is, when recording the timestamp of the GPU task, correction is performed according to the currently measured time offset, thereby achieving the unification of time within the distributed cluster.

[0105] The time offset is continuously updated as the task is executed to ensure time consistency throughout the execution of the distributed task.

[0106] The above correction process is triggered regularly, making the time correction more accurate. By working in conjunction with NTP, dynamic correction ensures that the time between CPUs remains consistent, and even if there is a slight system time drift during the execution of distributed tasks, the deviation can be effectively eliminated.

[0107] In some possible implementations, time data corresponding to GPU tasks (such as task start time and task end time) are collected from multiple CPUs and corresponding GPUs, and these data are uniformly stored and managed.

[0108] Through network transmission, the time data of each CPU and the corresponding GPU are centralized, and the data is deduplicated and corrected to deal with data inconsistency problems caused by network jitter. At the same time, users can conduct systematic analysis of the data of all GPU tasks to help users better understand and optimize the distributed task process.

[0109] In some possible implementations, task information of the GPU task may also be obtained while obtaining the task start time and the task end time.

[0110] The task information of the GPU task includes information such as the task ID (identification) of the GPU task and the task type of the GPU task.

[0111] In order to facilitate analysis and use by users, the task identifier of the GPU task, the task type of the GPU task, the task start time of the GPU task, and the task end time of the GPU task may be stored and managed.

[0112] In some specific implementations, the InterpreterCore object can be used to manage time data, and a custom time recording class can be added to track the task start time and task end time of the GPU task.

[0113] The data of each GPU task will be stored together with the task type and GPU task ID of the GPU task to facilitate subsequent analysis and use.

[0114] After the task is completed, the time data will be automatically collected and stored in a local file.

[0115] In some possible implementations, after the data is collected, the data is visualized and displayed based on a browser.

[0116] In some possible implementations, the task start time and the task end time are formatted into a data format supported by the browser, and based on the formatting result, the task start time and the task end time are visualized in the browser.

[0117] Specifically, the collected time data can be converted into Chrome Trace Event format for visualization through the chrome: / / tracing tool. During the conversion process, each task start time and task end time will be formatted to ensure the integrity and accuracy of the data.

[0118] Furthermore, different colors and layers can be used to distinguish GPU tasks and GPU task types of different GPUs, so that users can easily understand the dependencies and scheduling of tasks, and intuitively view the execution range, dependencies, parallelism and other information of GPU tasks, so as to discover potential performance bottlenecks.

[0119] In addition, this visualization method can also help users locate bottlenecks and communication delays in task execution, thereby optimizing system scheduling and resource utilization.

[0120] The data visualization method of the embodiment of the present disclosure is introduced below with a specific example. Figure 4 is a process diagram of a specific embodiment of a visualization method provided by an embodiment of the present disclosure, such as Figure 4As shown, first, the user enables the visualization tool in the Paddle framework and completes the initialization by activating the corresponding configuration option. When distributed tasks (such as distributed model training tasks) are executed, the visualization tool will automatically record the task start time and task end time of each GPU task (such as forward propagation tasks, backpropagation tasks, parameter optimization tasks) through CUDA CallBack to ensure accurate capture of the execution timing of each GPU task on the GPU. All timestamp data will be collected and stored by the Paddle framework without user intervention.

[0121] Next, the data is converted into Chrome Trace Event format, and users can view the execution status of each task by simply using the chrome: / / tracing tool. The data shows the execution time of each task, the scheduling between different GPU devices, communication delay, etc., helping users to intuitively understand the parallelism and dependencies of tasks, so as to effectively identify performance bottlenecks and optimize them.

[0122] The entire process is designed to minimize the difficulty of use for users. Users only need to perform simple switch operations to obtain a comprehensive and detailed visual analysis of the GPU task execution timing.

[0123] By recording the timestamp information of each fine-grained task such as forward calculation, back propagation, parameter optimization, etc. in detail, it helps users conduct in-depth analysis of each stage of the training task and accurately find performance bottlenecks.

[0124] Through accurate task time recording and visualization, users can quickly locate performance issues and significantly reduce the time cost of debugging and optimization. This support is particularly important in large-scale distributed training and helps improve development efficiency.

[0125] Whether it is single-machine training or multi-machine distributed training, it can provide consistent time recording and visualization support, which is suitable for training tasks of different scales and complexities. In a distributed environment, through the callback mechanism of time synchronization and communication tasks, it can accurately measure the communication delay across devices, helping users better understand and optimize cross-device task scheduling and data transmission.

[0126] Based on Figure 1 The same principle as shown in the method, Figure 5 A schematic diagram of the structure of a visualization device provided by an embodiment of the present disclosure is shown. Figure 5 As shown, the visualization device 50 may include:

[0127] Function module 510, used to insert a callback task into a first position and a second position of an execution task queue corresponding to a graphics processor;

[0128] The data processing module 520 is used to determine the task start time of the graphics processor task according to the execution result of the callback task inserted at the first position; and determine the task end time of the graphics processor task according to the execution result of the callback task inserted at the second position;

[0129] A visualization module 530 is used to visualize the task start time and task end time;

[0130] Among them, the execution task queue is a queue composed of multiple execution tasks executed on the graphics processor in the order of execution; the graphics processor task includes one or more execution tasks; the first position is located before the first task and adjacent to the first task, and the second position is located after the second task and adjacent to the second task; the first task is the execution task with the first execution order in the execution tasks corresponding to the graphics processor task, and the second task is the execution task with the last execution order in the execution tasks corresponding to the graphics processor task.

[0131] In the visualization device provided by the embodiment of the present disclosure, a callback task is inserted into the execution task queue, and the task start time and task end time of the GPU task are obtained through the execution result of the callback task. Since the callback task is inserted into the execution task queue of the GPU, it reflects the execution order of tasks on the GPU and does not depend on the scheduling and recommendation of the CPU. This avoids the deviation caused by the asynchrony of the CPU and GPU (after the CPU issues a task schedule, the GPU may delay the actual startup due to the execution load of the current task, resulting in the task start time recorded on the CPU side being earlier than the actual execution time of the GPU task), thereby achieving accurate recording of the execution status of the GPU tasks.

[0132] In some possible implementations, the callback task is used to obtain the time of the central processing unit corresponding to the graphics processor; the data processing module is used to determine the time of the central processing unit obtained by the callback task inserted at the first position as the task start time of the graphics processor task; and the time of the central processing unit obtained by the callback task inserted at the second position as the task end time of the graphics processor task.

[0133] In some possible implementations, the data visualization device further includes: a clock synchronization module for synchronizing the system clock of the central processing unit corresponding to the graphics processor with the system clocks of other central processing units of the distributed cluster to which the central processing unit belongs based on the network time protocol.

[0134] In some possible implementations, the data visualization device also includes a time offset module including: a signal sending and receiving unit, used to obtain a time synchronization request signal sent by a central processor corresponding to the graphics processor to other central processors, and a response signal sent by other central processors to the central processor corresponding to the graphics processor after receiving the time synchronization request signal; an offset calculation unit, used to calculate the time offset corresponding to the central processor corresponding to the graphics processor according to the time of the central processor corresponding to the graphics processor when sending the time synchronization request signal and the time of the central processor corresponding to the graphics processor when receiving the response signal; an offset adjustment unit, used to adjust the task start time and the task end time based on the time offset corresponding to the central processor corresponding to the graphics processor.

[0135] In some possible implementations, at least one of the multiple execution tasks in the execution task queue is a computing task, and at least one of the execution tasks corresponding to the graphics processor task is a computing task; the first task is the computing task that is first in the execution order among the computing tasks corresponding to the graphics processor task, and the second task is the computing task that is last in the execution order among the computing tasks corresponding to the graphics processor task.

[0136] In some possible implementations, the function module is used to: insert a computing unified device architecture callback function before a task start function corresponding to a first task, and insert a computing unified device architecture callback function before a task end function corresponding to a second task; the computing unified device architecture callback function is a function corresponding to the callback task.

[0137] In some possible implementations, the data visualization device also includes an information management module, which is used to: obtain task information of the graphics processor task, the task information including the task identifier of the graphics processor task and the task type of the graphics processor task; store and manage the task identifier of the graphics processor task, the task type of the graphics processor task, the task start time, and the task end time.

[0138] In some possible implementations, the visualization module is used to: format the task start time and the task end time into a data format supported by the browser, and visualize the task start time and the task end time in the browser based on the formatting result.

[0139] It can be understood that the above modules of the visualization device in the embodiment of the present disclosure have the function of realizing Figure 1The functions of the corresponding steps of the visualization method in the embodiment shown in . The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated with multiple modules. For the functional description of each module of the above visualization device, please refer to Figure 1 The corresponding description of the visualization method in the embodiment shown in will not be repeated here.

[0140] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0141] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0142] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0143] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the visualization method provided in the embodiment of the present disclosure.

[0144] Compared with the prior art, the electronic device inserts a callback task into the execution task queue and obtains the task start time and task end time of the GPU task through the execution result of the callback task. Since the callback task is inserted into the execution task queue of the GPU, it reflects the execution order of tasks on the GPU and does not depend on the scheduling and recommendation of the CPU. It avoids the deviation caused by the asynchrony of the CPU and GPU (after the CPU issues a task schedule, the GPU may delay the actual startup due to the execution load of the current task, resulting in the task start time recorded on the CPU side being earlier than the actual execution time of the GPU task), thereby achieving accurate recording of the execution status of the GPU tasks.

[0145] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the visualization method provided by the embodiment of the present disclosure.

[0146] Compared with the prior art, the readable storage medium inserts a callback task in the execution task queue and obtains the task start time and task end time of the GPU task through the execution result of the callback task. Since the callback task is inserted into the execution task queue of the GPU, it reflects the execution order of tasks on the GPU and does not depend on the scheduling and recommendation of the CPU. It avoids the deviation caused by the asynchrony of the CPU and GPU (after the CPU issues a task schedule, the GPU may delay the actual startup due to the execution load of the current task, resulting in the task start time recorded on the CPU side being earlier than the actual execution time of the GPU task), thereby achieving accurate recording of the execution status of the GPU tasks.

[0147] The computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the visualization method provided in the embodiment of the present disclosure.

[0148] Compared with the prior art, the computer program product inserts a callback task into the execution task queue and obtains the task start time and task end time of the GPU task through the execution result of the callback task. Since the callback task is inserted into the execution task queue of the GPU, it reflects the execution order of the tasks on the GPU and does not depend on the scheduling and recommendation of the CPU. Therefore, the deviation caused by the asynchrony of the CPU and GPU (after the CPU issues a task schedule, the GPU may delay the actual startup due to the execution load of the current task, resulting in the task start time recorded on the CPU side being earlier than the actual execution time of the GPU task) is avoided, thereby achieving accurate recording of the execution status of the GPU tasks.

[0149] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0150] like Figure 6As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0151] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0152] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as visualization methods. For example, in some embodiments, the visualization method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the visualization method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the visualization method in any other appropriate manner (e.g., by means of firmware).

[0153] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0154] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0155] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0157] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0158] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0159] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0160] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data visualization method, comprising: Inserting a callback task at a first position and a second position of an execution task queue corresponding to the graphics processor; Determine a task start time of the graphics processor task according to an execution result of the callback task inserted at the first position; Determining a task end time of the graphics processor task according to an execution result of the callback task inserted at the second position; Visually display the task start time and the task end time; Among them, the execution task queue is a queue composed of multiple execution tasks executed on the graphics processor in an execution order; the graphics processor task includes one or more of the execution tasks; the first position is located before the first task and adjacent to the first task, and the second position is located after the second task and adjacent to the second task; the first task is the execution task with the first execution order in the execution tasks corresponding to the graphics processor task, and the second task is the execution task with the last execution order in the execution tasks corresponding to the graphics processor task.

2. The method according to claim 1, wherein: The callback task is used to obtain the time of the central processing unit corresponding to the graphics processor; The step of determining the task start time of the graphics processor task according to the execution result of the callback task inserted at the first position includes: Determine the CPU time obtained by the callback task inserted in the first position as the task start time of the GPU task; The step of determining the task end time of the graphics processor task according to the execution result of the callback task inserted at the second position includes: The CPU time obtained by the callback task inserted in the second position is determined as the task end time of the GPU task.

3. The method according to claim 2, further comprising: The system clock of the central processing unit corresponding to the graphics processor is synchronized with the system clocks of other central processing units of the distributed cluster to which the central processing unit belongs based on the network time protocol.

4. The method according to claim 3, further comprising: Acquire a time synchronization request signal sent by the central processor corresponding to the graphics processor to other central processors and a response signal sent by other central processors to the central processor corresponding to the graphics processor after receiving the time synchronization request signal; Calculate the time offset corresponding to the central processor corresponding to the graphics processor according to the time of the central processor corresponding to the graphics processor when sending the time synchronization request signal and the time of the central processor corresponding to the graphics processor when receiving the response signal; The task start time and the task end time are adjusted based on the time offset corresponding to the central processing unit corresponding to the graphics processor.

5. The method according to claim 1, wherein: At least one of the multiple execution tasks in the execution task queue is a computing task, and at least one of the execution tasks corresponding to the graphics processor task is a computing task; the first task is the computing task with the first execution order among the computing tasks corresponding to the graphics processor task, and the second task is the computing task with the last execution order among the computing tasks corresponding to the graphics processor task.

6. The method according to claim 1, wherein: The inserting of the callback task into the first position and the second position of the execution task queue corresponding to the graphics processor includes: A computing unified device architecture callback function is inserted before the task start function corresponding to the first task, and a computing unified device architecture callback function is inserted before the task end function corresponding to the second task; the computing unified device architecture callback function is the function corresponding to the callback task.

7. The method according to claim 1, further comprising: Acquire task information of the graphics processor task, where the task information includes a task identifier of the graphics processor task and a task type of the graphics processor task; The task identifier of the graphics processor task, the task type of the graphics processor task, the task start time, and the task end time are stored and managed.

8. The method according to claim 1, wherein: The visual display of the task start time and the task end time includes: The task start time and the task end time are formatted into a data format supported by the browser, and based on the formatting result, the task start time and the task end time are visualized in the browser.

9. A data visualization device, comprising: A function module, used for inserting a callback task into a first position and a second position of an execution task queue corresponding to a graphics processor; A data processing module, used to determine a task start time of the graphics processor task according to an execution result of the callback task inserted at the first position; Determining a task end time of the graphics processor task according to an execution result of the callback task inserted at the second position; A visualization module, used to visualize the task start time and the task end time; Among them, the execution task queue is a queue composed of multiple execution tasks executed on the graphics processor in an execution order; the graphics processor task includes one or more of the execution tasks; the first position is located before the first task and adjacent to the first task, and the second position is located after the second task and adjacent to the second task; the first task is the execution task with the first execution order in the execution tasks corresponding to the graphics processor task, and the second task is the execution task with the last execution order in the execution tasks corresponding to the graphics processor task.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

12. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.