Method and system for realizing CUDA call tracking based on eBPF
By using eBPF technology to dynamically mount programs at CUDA call sites and asynchronously transmit metadata, the intrusiveness and efficiency issues of existing CUDA call tracing methods are resolved, non-intrusive, low-overhead fine-grained tracing is achieved, and the observability and debugging efficiency of GPU applications are improved.
Patent Information
- Application Number
- CN202511239512.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-09-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing CUDA call tracing methods have compatibility and performance impacts caused by intrusive instrumentation or library injection, are unable to obtain accurate underlying call context information, and have low efficiency in cross-user and kernel data communication, making it difficult to meet real-time and low-overhead requirements.
eBPF technology is used to dynamically mount eBPF programs at the entry and exit of API functions in the user-state CUDA runtime library, capture the metadata of CUDA API call events, and asynchronously transmit them to the user state through a ring buffer of type BPF_MAP_TYPE_RINGBUF, achieving non-intrusive, low-overhead fine-grained tracing.
It enables comprehensive, real-time tracking of CUDA applications, improves debugging efficiency and system observability, and reduces monitoring overhead. It is suitable for AI training, high-performance computing, and GPU virtualization scenarios.
Smart Images

Figure CN120723587A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of GPU call tracing based on eBPF, and in particular to a method and system for implementing CUDA call tracing based on eBPF. Background Art
[0002] With the rapid development of artificial intelligence, machine learning, and high-performance computing, GPU computing has become the core infrastructure supporting complex computing tasks.
[0003] In this field, NVIDIA's CUDA platform, the de facto standard for GPU general-purpose computing, is widely used in scenarios such as scientific computing, deep learning, and graphics rendering. Related technologies have established a monitoring system for CUDA API calls through user-mode instrumentation, library injection, and proprietary tools (such as Nsight Systems and Nsight Compute). Specifically, this technology system covers the entire process from application behavior analysis to GPU resource scheduling, including key steps such as API call capture, performance data collection, and execution timing recording. However, existing CUDA call tracing methods directly utilize intrusive instrumentation or library injection techniques without fully considering the compatibility and performance impact on production environments. This can lead to complex deployments, reduced system stability, or the inability to obtain accurate low-level call context information, thus compromising debugging efficiency and system observability. Furthermore, traditional tools lack efficient mechanisms for cross-user-mode and kernel-mode data communication, making it difficult to meet the requirements of real-time and low-overhead collaboration.
[0004] Therefore, the industry urgently needs a more general, low-overhead, non-intrusive solution that can provide fine-grained information to achieve comprehensive and real-time tracking of GPU CUDA operations. Summary of the Invention
[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0006] To this end, the first objective of this invention is to propose a method for implementing CUDA call tracing based on eBPF to address the challenges of existing GPU operation tracing techniques, particularly the shortcomings of traditional tools in terms of non-invasiveness, low overhead, and fine-grained information capture. This invention leverages the powerful programmability, event-driven nature, and non-invasiveness of eBPF technology at the Linux kernel level to achieve precise interception and data capture of user-space CUDA API calls, thereby providing comprehensive, real-time insights for debugging, performance analysis, and security auditing of CUDA applications.
[0007] The second object of the present invention is to propose a system for implementing CUDA call tracing based on eBPF.
[0008] A third object of the present invention is to provide a computer-readable storage medium.
[0009] To achieve the above objectives, a first embodiment of the present invention provides a method for implementing CUDA call tracing based on eBPF, comprising: S1, through the eBPF uprobe technology, dynamically mount the eBPF program at the API function entry and exit of the user-mode CUDA runtime library to achieve non-intrusive monitoring of CUDA API calls; S2, the eBPF program captures the metadata of the CUDA API call event in kernel mode and encapsulates the metadata into a predefined structure; S3, asynchronously transfer the structure data to userland via a ring buffer of type BPF_MAP_TYPE_RINGBUF; S4: The user-mode program calls the standard function library to load the eBPF program and reads the data in the ring buffer in real time to complete the analysis and visualization of the CUDA call event.
[0010] In one embodiment of the present invention, the data structure information of the metadata includes process ID, process name, event type, event parameters, return value and event occurrence timestamp.
[0011] In one embodiment of the present invention, the S1 further includes: S11, mount the eBPF program to the specified API function entry and exit in the user-mode CUDA runtime library through the `bpf_program__attach_uprobe_opts` function. The mount operation supports specifying the target library path through the `lib_path` parameter and the target function name through the `func_name` parameter. S12, during the mounting process, selectively monitor the CUDA API calls of the preset process by setting the `env.target_pid` parameter.
[0012] In one embodiment of the present invention, the S2 further includes: S21, get the time information of the event through the `bpf_ktime_get_ns()` function and encapsulate it into the `timestamp` field of the structure; S22: The event parameters and return values are parsed through the context parameters of the eBPF program and dynamically filled into the corresponding fields in the `event_instance` union according to different CUDA API types.
[0013] In one embodiment of the present invention, the S3 includes: S31, the maximum number of entries in the ring buffer is set to `256 * 1024` to ensure data throughput in a high-concurrency CUDA call scenario; S32, the asynchronous transmission uses the `bpf_ringbuf_submit` function to submit the encapsulated structure data to the ring buffer, and uses the `bpf_ringbuf_consume` function to perform non-blocking reading in user mode.
[0014] In one embodiment of the present invention, it further comprises: S5, filtering and aggregating the structure data, filtering out target CUDA call events according to preset filtering rules, and storing the filtered data in a user-state persistent log file.
[0015] To achieve the above objectives, a second embodiment of the present invention provides a system for implementing CUDA call tracing based on eBPF, including: The mount configuration module is used to dynamically mount eBPF programs at the entry and exit of the API function of the user-mode CUDA runtime library through the eBPF uprobe technology to achieve non-intrusive monitoring of CUDA API calls; A metadata collection module is used to capture metadata of CUDA API call events in kernel mode and encapsulate the metadata into a predefined structure; The data transmission module is used to asynchronously transfer structure data to user mode through the ring buffer of type BPF_MAP_TYPE_RINGBUF; The event analysis module is used to call the standard function library to load the eBPF program and read the data in the ring buffer in real time to complete the analysis and visualization of CUDA call events.
[0016] To achieve the above-mentioned purpose, the third embodiment of the present invention proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method described in the first embodiment is implemented.
[0017] The method, system, and storage medium of the embodiments of the present invention implement non-intrusive, low-overhead, fine-grained tracing of CUDA API calls, can efficiently capture call parameters, return values, and nanosecond timestamps, and improve the observability and debugging efficiency of GPU applications.
[0018] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 is a flowchart of a method for implementing CUDA call tracing based on eBPF according to an embodiment of the present invention; Figure 2 1 is an architectural diagram of a method for implementing CUDA call tracing based on eBPF according to an embodiment of the present invention; Figure 3 This is a structural diagram of a system for implementing CUDA call tracing based on eBPF according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0022] The following describes a method, system, and computer-readable storage medium for implementing CUDA call tracing based on eBPF according to embodiments of the present invention with reference to the accompanying drawings.
[0023] The technical terms of the present invention are introduced below: eBPF: A user-written program that can run in the Linux kernel. It is essentially a virtual machine running in kernel state, allowing users to load and run custom programs in the kernel. Users and the kernel can exchange data on demand and manage kernel behavior at any time.
[0024] GPU: Graphics Processing Unit, a processor specialized for image processing and parallel computing.
[0025] CUDA: Compute Unified Device Architecture, a parallel computing platform and programming model launched by NVIDIA that allows developers to take advantage of the powerful computing capabilities of NVIDIA GPUs.
[0026] Uprobe: A user-mode dynamic tracing technology that allows probe points to be inserted into user-space applications to monitor and track the program's running status and behavior in real time without modifying or recompiling the application's source code.
[0027] Example 1 Figure 1 FIG. 1 is a flow chart of a method for implementing CUDA call tracing based on eBPF according to an embodiment of the present invention. Figure 1 As shown, including: S1, through the uprobe technology of eBPF, dynamically mount the eBPF program at the entry and exit of the API function of the user-mode CUDA runtime library to achieve non-intrusive monitoring of CUDA API calls.
[0028] Specifically, this invention uses eBPF's uprobe technology to dynamically mount eBPF programs at the entry and exit points of user-space CUDA runtime library API functions. This is the core step in achieving non-intrusive monitoring of CUDA API calls. This technology is based on the eBPF (extended Berkeley Packet Filter) framework provided by the Linux kernel and combines it with the uprobe mechanism to achieve dynamic instrumentation and data collection of user-space functions.
[0029] In detail, this step first involves loading a pre-written eBPF program into kernel space by calling the interface provided by the libbpf library through a userspace program. The `bpf_program__attach_uprobe_opts()` function is then used to dynamically mount the eBPF program to the entry and exit locations of the specified API functions in the CUDA runtime library (such as libcuda.so or libnvidia-ml.so). Mounting requires providing the target function name (such as `cudaMalloc`), the library path (such as ` / usr / lib / libcuda.so.1`), and the optional process ID, enabling monitoring of pre-defined processes or global calls. The uprobe mechanism inserts probes at the entry and exit points of the target function. When the function is called, the kernel triggers the execution of the corresponding eBPF program.
[0030] Regarding parameter settings, the eBPF program retrieves the current process's PID and TGID (thread group ID) via `bpf_get_current_pid_tgid()` and the process name via `bpf_get_current_comm()`. Event types are identified by the `enum cuda_event_type` enumeration, such as `CUDA_EVENT_MALLOC` and `CUDA_EVENT_FREE`. Event data is transmitted via a ring buffer of type `BPF_MAP_TYPE_RINGBUF`, with a maximum entry size of `256 * 1024` bytes to ensure data loss at high throughput. The eBPF program encapsulates collected event information into a `struct event` structure, including fields such as process ID, process name, event type, and event details, and submits it to user space via `bpf_ringbuf_submit()`.
[0031] This step plays a key role in the entire technical solution. Its non-invasive nature ensures that monitoring CUDA applications in production environments will not affect their normal operation. At the same time, the efficient execution mechanism in kernel mode significantly reduces monitoring overhead. In addition, by mounting eBPF programs at the API entry and exit respectively, the call context, parameters, and return values can be fully captured, providing fine-grained data support for performance analysis, resource usage monitoring, and anomaly detection. This technical solution has broad application value in scenarios such as AI training, high-performance computing (HPC), and GPU virtualization. It is particularly suitable for the development of system-level monitoring tools that require real-time and low-intrusion tracking of CUDA call behavior.
[0032] Furthermore, S1 includes: S11, mount the eBPF program to the specified API function entry and exit in the user-mode CUDA runtime library through the `bpf_program__attach_uprobe_opts` function. The mount operation supports specifying the target library path through the `lib_path` parameter and the target function name through the `func_name` parameter.
[0033] Specifically, this step mounts the eBPF program to the specified API function entry and exit in the user-space CUDA runtime library by calling the `bpf_program__attach_uprobe_opts` function, which is one of the core links in implementing CUDA call tracing. At the technical implementation level, this step is based on the Linux kernel's uprobe mechanism, allowing dynamic insertion of probe points in user-space shared libraries (such as `libcudart.so`), thereby achieving non-intrusive monitoring of CUDA API calls without modifying the application or library files. Specifically, the `bpf_program__attach_uprobe_opts` function is provided by the libbpf library. Its function is to bind the loaded eBPF program to the specified user-space function address. When the target function is called, the kernel automatically triggers the execution of the eBPF program to capture the call context information.
[0034] In terms of parameter settings, `lib_path` is used to specify the path of the target CUDA runtime library (such as ` / usr / local / cuda / lib64 / libcudart.so`), and `func_name` is used to specify the function name to which the eBPF program needs to be mounted (such as `cudaMalloc`, `cudaMemcpy`, etc.). With these two parameters, the system can accurately locate specific CUDA API functions in user space and attach eBPF programs at their entry (`UPROBE`) and exit (`RETURN UPROBE`). In addition, the `env.target_pid` parameter can optionally be used to limit tracing to only preset processes, thereby improving system flexibility and security.
[0035] Technically, this step achieves function-level mounting accuracy, supporting dynamic instrumentation of any user-mode function without modifying library files or application source code. The mounting process incurs extremely low overhead, typically in the microsecond range, without significantly impacting application execution efficiency. Combined with the parsing of call parameters and return values by the eBPF program in subsequent steps, this mounting mechanism provides the foundation for fine-grained, low-latency CUDA call tracing.
[0036] This step is widely applicable to scenarios requiring performance analysis, debugging, resource monitoring, or security auditing of GPU applications. For example, during deep learning training, by mounting the `cudaMalloc` and `cudaFree` functions, GPU memory allocation and release behavior can be monitored in real time, assisting in identifying memory leaks or fragmentation issues. Furthermore, this mechanism can be used to analyze key performance indicators such as CUDA API call frequency and time distribution, providing data support for optimizing GPU resource scheduling.
[0037] In summary, this step achieves non-intrusive, low-overhead, and high-precision tracing of CUDA API calls through a precise function-level mounting mechanism. It is an important technical support point for the present invention in terms of system observability and has significant practical value and innovative significance.
[0038] S12, during the mounting process, the `env.target_pid` parameter is set to selectively monitor the CUDA API calls of the preset process, thereby avoiding indiscriminate tracking of all processes in the system.
[0039] Specifically, in some implementations, selectively monitoring CUDA API calls from a pre-defined process by setting the `env.target_pid` parameter during the mount process is a key step in the present invention's eBPF-based CUDA call tracing. The core technical principle of this step lies in leveraging eBPF's uprobe mechanism, combined with userspace code, to filter the target process's PID (Process ID). This allows for precise interception and data collection of CUDA API calls from the pre-defined process, avoiding indiscriminate tracking of all processes in the system, thereby reducing system overhead and improving targeted monitoring.
[0040] In specific implementations, the user-space program attaches the eBPF program to the specified function entry and exit points in the CUDA runtime library (e.g., libcuda.so) by calling the `bpf_program__attach_uprobe_opts` function. The `env.target_pid` parameter specifies the PID of the target process, which is entered by the user when the user-space program starts or dynamically obtained through the process management interface. When `env.target_pid` is set to a non-zero value, the eBPF program only monitors processes matching that PID; if it is set to 0, all processes are monitored. This allows the system to achieve fine-grained process-level control, ensuring that data is collected only for CUDA API calls made by the target process.
[0041] At the parameter level, `env.target_pid` takes a 32-bit integer value, conforming to the Linux PID standard (typically 1 to 4194303). Furthermore, the `lib_path` parameter in the `bpf_program__attach_uprobe_opts` mount function specifies the absolute path to the CUDA runtime library (e.g., ` / usr / lib / libcuda.so`), and the `func_name` parameter specifies the specific CUDA API function name (e.g., `cudaMalloc`, `cudaMemcpy`, etc.). Through precise configuration of these parameters, the eBPF program can insert probes at specific locations in userspace functions, enabling dynamic tracing of CUDA API calls.
[0042] In application scenarios, this step is particularly suitable for scenarios requiring performance analysis, resource usage monitoring, or security auditing of specific GPU computing tasks. For example, in a distributed deep learning training environment, users can monitor only the main process responsible for model training while ignoring CUDA calls from other auxiliary processes, thereby reducing data redundancy and system load. Furthermore, in containerized deployments or virtualized environments, this mechanism ensures that only CUDA calls within a specific container or virtual machine are tracked, improving system isolation and security.
[0043] The technical benefit of this step is that, by introducing the `env.target_pid` parameter, process-level filtering of CUDA API calls is implemented, significantly reducing system resource consumption and data processing complexity. At the same time, this mechanism maintains the non-invasive and low-overhead characteristics of eBPF tracing, ensuring that the performance of the monitored process is not significantly impacted. Furthermore, this design enhances the system's flexibility and configurability, providing reliable technical support for fine-grained monitoring in multi-tenant and production environments.
[0044] S2: The eBPF program captures metadata of CUDA API call events in kernel mode and encapsulates the metadata into a predefined structure.
[0045] Specifically, in this invention, the eBPF program captures metadata about CUDA API call events in kernel mode and encapsulates it into a predefined structure, a core step in implementing CUDA call tracing. This step leverages eBPF's uprobe mechanism to dynamically insert probe points at the entry and exit points of specified functions in the user-space CUDA runtime library (e.g., libcuda.so). This enables non-intrusive monitoring of CUDA API calls without modifying application or library files.
[0046] At the technical implementation level, the eBPF program uses `bpf_get_current_pid_tgid()` to obtain the PID and TGID (thread group ID) of the currently calling process, and then extracts the process ID (PID) using bit shifting. The process name (comm) is obtained using `bpf_get_current_comm()`, which copies the name of the current process into the `comm` field of a structure. The event type (`cuda_event_type`) is defined by an enumeration type, such as `CUDA_EVENT_MALLOC`, which represents a `cudaMalloc` call event. Event parameters and return values are collected at function entry and exit by the eBPF program, such as the memory allocation size `size` and the return value `ret_val`. Timestamps are obtained using the kernel's high-precision timestamp APIs (such as `bpf_ktime_get_ns()`), which provide nanosecond accuracy, ensuring accurate event timing.
[0047] Regarding parameter metrics, the sizes and types of key fields in the event structure are defined. For example, the pid field is a 32-bit integer, the comm field is TASK_COMM_LEN long (typically 16 bytes), and the details field is MAX_DETAILS_LEN long (configurable to 256 bytes or larger) to support the parameter complexity of different CUDA APIs. The max_entries parameter of the ring buffer (BPF_MAP_TYPE_RINGBUF) is set to 256 * 1024 bytes to ensure sufficient buffering capacity in highly concurrent call scenarios and to prevent data loss due to memory overflow.
[0048] In application scenarios, this step is suitable for performance analysis, resource usage monitoring, or security audits of GPU applications. For example, during deep learning training, by capturing the call parameters and return values of key APIs such as `cudaMalloc` and `cudaMemcpy`, memory allocation patterns and data transfer efficiency can be analyzed in real time, thereby optimizing GPU resource scheduling strategies. Furthermore, in multi-tenant environments, this mechanism can be used to audit GPU resource usage by different processes, improving system security and resource isolation.
[0049] The technical benefit of this step is that it enables fine-grained, low-latency, and non-intrusive tracing of GPU call behavior through structured encapsulation of CUDA API call metadata. Its innovation lies in leveraging eBPF's uprobe mechanism to mount probe points in user-space functions, combined with the kernel's efficient data collection and encapsulation capabilities. This provides a reliable data foundation for full-lifecycle monitoring of CUDA calls, significantly improving system observability and debugging efficiency.
[0050] Furthermore, S2 includes: S21, obtain the time information of the event through the `bpf_ktime_get_ns()` function and encapsulate it into the `timestamp` field of the structure.
[0051] Specifically, in this invention, the metadata includes nanosecond-level timestamps. Calling the `bpf_ktime_get_ns()` function to obtain the time information of the event occurrence and encapsulating it into the `timestamp` field of the structure is a key step in achieving high-precision timing tracking of CUDA calls. This step is technically implemented based on the eBPF (extended Berkeley Packet Filter) mechanism provided by the Linux kernel. Using kernel-mode eBPF programs, probes are inserted at the entry and exit points of CUDA API calls, enabling non-intrusive, fine-grained monitoring of GPU call behavior without modifying the application or runtime library.
[0052] From a technical implementation perspective, `bpf_ktime_get_ns()` is a standard function in eBPF programs for obtaining the current system timestamp. Its return value is a nanosecond-resolution timestamp (ns-resolution timestamp) since system boot, offering extremely high time precision, suitable for millisecond or even microsecond-level analysis of GPU call timing. In eBPF programs, when a CUDA API (such as `cudaMalloc`) is called, this function is triggered to capture the precise time of the event and write the timestamp to the `timestamp` field of a predefined shared data structure, `struct event`. This structure communicates efficiently with userspace via a ring buffer of type `BPF_MAP_TYPE_RINGBUF`, ensuring real-time and low-latency timestamp data.
[0053] At the parameter level, the `bpf_ktime_get_ns()` function requires no external parameters and returns a 64-bit unsigned integer (`u64`), representing a nanosecond timestamp. The `timestamp` field in the `struct event` structure should be defined as `u64 timestamp` to ensure complete storage of timestamps. Furthermore, the ring buffer is configured with a capacity of `256 * 1024` bytes to support event buffering in high-concurrency scenarios and avoid data loss due to user-space processing delays. Timestamp accuracy reaches nanoseconds (1e-9 seconds). Compared to the traditional user-space time acquisition method based on `clock_gettime()`, it has lower kernel-mode call overhead and avoids clock drift and cross-process clock synchronization issues.
[0054] At the application level, this step is widely applicable to scenarios requiring performance analysis, debugging, and resource management of CUDA API calls. For example, during deep learning training, by recording the precise timestamps of each `cudaMalloc`, `cudaMemcpy`, `cudaLaunchKernel`, and other API calls, the timing relationships between GPU memory allocation, data transfer, and kernel launches can be analyzed to identify performance bottlenecks and optimize resource scheduling strategies. In distributed GPU computing environments, this timestamp information can also be used to trace call chains across nodes, enabling end-to-end performance monitoring.
[0055] From a technical perspective, this step significantly improves the timing accuracy and system observability of CUDA call tracing by capturing nanosecond timestamps in kernel state. Its non-invasive design ensures minimal disruption to application runtime behavior. Combined with eBPF's efficient data transmission mechanism, it enables low-overhead, high-throughput event recording. This technical solution provides a reliable time base for performance analysis and debugging of GPU computing tasks while ensuring system stability, and is one of the core technologies supporting fine-grained CUDA call tracing.
[0056] S22: The event parameters and return values are parsed through the context parameters of the eBPF program and dynamically filled into the corresponding fields in the `event_instance` union according to different CUDA API types.
[0057] Specifically, in this invention, the step of "parsing the event parameters and return values using the eBPF program's context parameters and dynamically populating the corresponding fields in the `event_instance` union based on different CUDA API types" is a core component of CUDA call tracing. This step uses the eBPF program to mount uprobe probes at the entry and exit of user-space CUDA API functions, capturing key parameters and return values from the call context in real time and encapsulating them in a structured manner into a predefined union, providing basic data support for subsequent event analysis and visualization.
[0058] At the technical level, the eBPF program extracts function arguments from the current execution context using helper functions such as `bpf_get_argX()`. For example, for the `cudaMalloc` function, its entry parameter includes the memory size (`size_t size`), and the return value is the allocation result (e.g., `int ret_val`). The eBPF program records the call arguments at function entry and captures the return value at function return. It then allocates an event structure instance from a kernel-space ring buffer using `bpf_ringbuf_reserve()`. The program then populates the `event` structure with metadata such as the event type (e.g., `CUDA_EVENT_MALLOC`), the process ID (obtained through `bpf_get_current_pid_tgid()` and right-shifted by 32 bits to extract the PID), and the process name (obtained through `bpf_get_current_comm()`). The program also writes the specific argument values into the corresponding fields of the `event_instance` union. Finally, `bpf_ringbuf_submit()` submits the event to the ring buffer for asynchronous reading by userspace programs.
[0059] At the parameter level, the design of the `event_instance` union must strictly correspond to the CUDA API parameter structure. For example, `cudaMemcpy` must record fields such as the source address, destination address, and copy size. The `pid` field in the event structure is a 32-bit integer, the `comm` field is a character array of a fixed length (e.g., `TASK_COMM_LEN=16`), and the `type` field is the enumerated type `cuda_event_type`, whose value is mapped from the CUDA API type. The `max_entries` parameter of the `ringbuf` ring buffer is set to `256 * 1024` bytes to support high-throughput event transfer while avoiding memory overflows.
[0060] At the application level, this step is widely applicable to performance analysis, debugging, resource monitoring, and security auditing in GPU computing environments. For example, during deep learning training, by tracking the call parameters and return values of key APIs such as `cudaMalloc`, `cudaMemcpy`, and `cudaLaunchKernel`, memory leaks, data transfer bottlenecks, or kernel execution anomalies can be accurately identified. Furthermore, this mechanism can be used to build a profile of GPU resource usage, providing data support for resource scheduling in containerized environments.
[0061] The technical benefit of this step is that it enables fine-grained, non-intrusive tracing of CUDA API calls through structured encapsulation and dynamic field filling. Its innovation lies in the efficient capture and processing of user-mode call information in kernel mode without modifying the CUDA runtime library or application code, thereby improving the observability and analyzability of GPU behavior while maintaining system performance.
[0062] S3, asynchronously transfers the structure data to userland via a ring buffer of type BPF_MAP_TYPE_RINGBUF.
[0063] Specifically, in this paper, asynchronously transferring structured data to userland via a ring buffer of type `BPF_MAP_TYPE_RINGBUF` is a key data communication mechanism in the CUDA call tracing system. This step is based on the Linux kernel's extended Berkeley Packet Filter (eBPF) technology, leveraging its efficient, lock-free, asynchronous communication capabilities to achieve low-latency, low-overhead data transfer between kernel and userland.
[0064] BPF_MAP_TYPE_RINGBUF is an eBPF map type based on a ring buffer. Its design goal is to provide a high-performance data output mechanism for eBPF programs. In this method, a kernel-mode eBPF program, triggered at the entry or exit of a CUDA API call, reserves memory space in the ring buffer by calling the bpf_ringbuf_reserve() function to store the encapsulated struct event structure. This structure contains fields such as the process ID (pid), process name (comm), event type (type), and detailed event information (such as requested memory size and return value). Once the data is populated, the bpf_ringbuf_submit() function submits the data to the ring buffer for asynchronous reading by user-mode programs. User-mode programs create and poll this buffer using libbpf library APIs (such as bpf_ringbuf__new() and bpf_ringbuf__poll()), enabling real-time consumption and processing of event data.
[0065] The ring buffer size is defined by `__uint(max_entries, 256 * 1024)`, indicating a maximum storage capacity of 256KB of data. This parameter can be adjusted based on the application scenario to balance data throughput and memory usage. The third parameter in the `bpf_ringbuf_reserve()` function is the reservation flag; setting it to 0 disables any special reservation policy. The `pid` field in the event structure retrieves the thread group ID of the current process using `bpf_get_current_pid_tgid()`, extracting the process ID through a bit shift operation. The `comm` field retrieves the current process name using `bpf_get_current_comm()`, with a maximum length of `TASK_COMM_LEN` (typically 16 bytes). The `enum cuda_event_type` event type is used to distinguish different CUDA API calls. For example, `CUDA_EVENT_MALLOC` indicates a `cudaMalloc` call event.
[0066] This step is widely used in scenarios requiring non-intrusive, fine-grained tracing of CUDA API calls, such as GPU performance analysis, resource usage monitoring, debugging support, and security auditing. User-mode programs can run in system monitoring tools, debuggers, or custom analysis platforms, asynchronously reading event data from the ring buffer to enable real-time visualization of CUDA calls or further statistical analysis. Because eBPF programs do not require modification of the source code or binary files of the traced program, this solution is particularly suitable for monitoring GPU applications in production environments, offering excellent compatibility and deployment flexibility.
[0067] Asynchronous data transmission via `BPF_MAP_TYPE_RINGBUF` effectively reduces communication latency between the kernel and user-mode, improving overall system responsiveness and throughput. This mechanism avoids performance bottlenecks that can arise from traditional blocking communication methods while ensuring data orderliness and integrity. Furthermore, this step provides user-mode programs with structured, parseable event data, laying a solid foundation for subsequent analysis, visualization, and alerting mechanisms. This is one of the core technologies supporting the low-overhead, high-precision CUDA call tracing implementation of this invention.
[0068] Furthermore, S3 includes: S31, the maximum number of entries in the ring buffer is set to `256 * 1024` to ensure data throughput in a high-concurrency CUDA call scenario.
[0069] Specifically, in this paper, setting the maximum number of entries in the ring buffer (BPF_MAP_TYPE_RINGBUF) to 256 * 1024 is a key technical step in achieving high-concurrency CUDA call tracing. This step is based on the extended Berkeley Packet Filter (eBPF) mechanism provided by the Linux kernel. By establishing an efficient, low-latency data transmission channel between kernel space and user space, it ensures stable data throughput even in high-concurrency scenarios.
[0070] At the technical implementation level, `BPF_MAP_TYPE_RINGBUF` is a lock-free, efficient producer-consumer data structure suitable for asynchronous transmission of high-frequency events. After capturing CUDA API call events in kernel space, the eBPF program encapsulates the structured event data (such as process ID, call timestamp, API type, parameters, and return value) into a `struct event` type and submits the event to the ring buffer via the `bpf_ringbuf_reserve()` and `bpf_ringbuf_submit()` interfaces. Userspace programs asynchronously read the event data from this buffer through the libbpf library for parsing and display. Because the ring buffer uses pre-allocated memory, frequent memory allocation and deallocation operations are avoided, thereby reducing system overhead.
[0071] Regarding parameter metrics, `max_entries` is set to `256 * 1024` (i.e., 262,144 entries). This value has been verified through performance testing to effectively prevent data loss caused by buffer overflows when supporting high-concurrency CUDA calls (e.g., tens of thousands of API calls per second). Each event entry is typically hundreds of bytes in size, depending on the definition of fields in `structevent` (e.g., `MAX_DETAILS_LEN`). In actual deployments, this parameter should be dynamically adjusted based on the target system's memory resources and expected call frequency to balance performance and resource usage.
[0072] In application scenarios, this setting is particularly suitable for environments with extremely high CUDA call frequency requirements, such as large-scale parallel computing tasks, deep learning training frameworks, and high-performance computing clusters. For example, in distributed GPU training, multiple processes may simultaneously call APIs such as `cudaMalloc` and `cudaMemcpy`. Insufficient buffer capacity will result in event loss, affecting the integrity and accuracy of tracking. By setting a sufficiently large buffer capacity, the system can continuously and stably record all CUDA call events, providing a reliable data foundation for subsequent performance analysis, resource scheduling optimization, and anomaly detection.
[0073] The technical effect of this step is that by rationally configuring the capacity of the ring buffer, the system's data throughput and stability in high-concurrency scenarios are significantly improved, ensuring the complete capture and low-latency transmission of CUDA call events, thereby enhancing the real-time and reliability of the entire tracing system, and providing a solid guarantee for non-intrusive GPU behavior analysis based on eBPF.
[0074] S32, the asynchronous transmission uses the `bpf_ringbuf_submit` function to submit the encapsulated structure data to the ring buffer, and uses the `bpf_ringbuf_consume` function to perform non-blocking reading in user mode.
[0075] Specifically, the asynchronous transmission mechanism of the present invention uses the `bpf_ringbuf_submit` function to submit encapsulated structured data to the kernel's ring buffer. The `bpf_ringbuf_consume` function then performs non-blocking reading in userland, enabling efficient and low-latency data collection and processing. In some implementations, this mechanism is based on the eBPF (extended Berkeley Packet Filter) framework provided by the Linux kernel, utilizing the `BPF_MAP_TYPE_RINGBUF` mapping structure as the communication channel between the kernel and userland, ensuring reliable data transmission and orderly consumption in high-concurrency scenarios.
[0076] From a technical implementation perspective, the `bpf_ringbuf_submit` function is called by the kernel-mode eBPF program to submit the populated `struct event` structure to the ring buffer. This structure contains fields such as the process ID (`pid`), the process name (`comm`), the CUDA event type (`type`), and event details such as the memory allocation size and return status. Its size must be pre-allocated at compile time using `bpf_ringbuf_reserve` to ensure memory alignment and data integrity. Furthermore, `bpf_ringbuf_submit` writes the data to the producer side of the buffer atomically and updates the buffer's write pointer, ensuring data consistency when submitting events from multiple threads or at high frequency.
[0077] In user mode, the application reads event data from the buffer in a non-blocking manner by calling the `bpf_ringbuf_consume` function. This function is implemented based on the libbpf library and supports polling or event-driven reading modes. Users can choose different consumption strategies based on actual needs. Optionally, the user program can set a read timeout (such as 100ms) to avoid long waits when there is no data, thereby improving the overall system response efficiency. In addition, the user-mode program must parse and format the read structure data, such as converting the event type into a readable string, extracting the timestamp, and calculating the call duration, to support subsequent visualization or log analysis.
[0078] In terms of parameter indicators, the capacity of the ring buffer in the present invention is set to `256 * 1024` bytes. This parameter can be dynamically adjusted according to the system load and event frequency to balance memory usage and data throughput. The size of the event structure should be controlled within a reasonable range to avoid frequent buffer overflows or submission failures due to excessively large structures. At the same time, the event type is defined by `enum cuda_event_type` to support extensibility, such as the addition of `CUDA_EVENT_KERNEL_LAUNCH`, `CUDA_EVENT_MEM_COPY` and other event types to meet the tracking requirements of different CUDA APIs.
[0079] This step is widely applicable in practical applications where real-time monitoring of CUDA API calls is required, such as GPU performance tuning, resource usage analysis, and abnormal behavior detection. In large-scale distributed computing environments, the use of asynchronous transmission mechanisms can effectively reduce the impact of the tracking process on application performance while ensuring the real-time and integrity of data collection.
[0080] In terms of technical effectiveness, this asynchronous transmission mechanism significantly improves the system's data acquisition efficiency under highly concurrent CUDA calls, avoiding the performance bottlenecks caused by traditional blocking reads. Through efficient memory management of the ring buffer and a non-blocking read strategy, this invention achieves low-latency, low-overhead communication between the kernel and user space, providing reliable data transmission for fine-grained tracing of CUDA calls.
[0081] S4: The user-mode program calls the standard function library to load the eBPF program and reads the data in the ring buffer in real time to complete the analysis and visualization of the CUDA call event.
[0082] Specifically, in some implementations, user-mode programs load eBPF programs by calling standard libraries (such as libbpf) and read CUDA call event data transmitted in kernel space via the `BPF_MAP_TYPE_RINGBUF` ring buffer in real time, thereby completing the analysis and visualization of CUDA API calls. This step is a key part of data consumption and processing in the entire eBPF-based CUDA call tracing system. Its technical implementation is based on the collaborative work between the eBPF framework provided by the Linux kernel and user-mode programs.
[0083] From a technical implementation perspective, a user-mode program first loads a pre-compiled eBPF bytecode program through the libbpf library and attaches it to the entry and exit locations of target functions in the CUDA runtime library (such as `cudaMalloc` and `cudaMemcpy`). This process is typically implemented through the `bpf_program__attach_uprobe_opts()` function, where `lib_path` specifies the path to the CUDA runtime library (such as `libcuda.so`) and `func_name` is the name of the API function to be traced. When running in kernel mode, the eBPF program encapsulates captured CUDA call events into a predefined structure called `struct event` and submits the event data to the ring buffer using the `bpf_ringbuf_submit()` function.
[0084] At the parameter level, the ring buffer size is defined by `__uint(max_entries, 256 * 1024)`, indicating a maximum storage capacity of 256KB of data, supporting high-throughput event collection. The event structure `struct event` contains fields such as the process ID (`pid`), process name (`comm`), event type (`type`), and event details (`details`). Timestamps are accurate to nanoseconds, ensuring accurate event timing. User-mode programs read the buffer data in real time through `bpf_ringbuf_read()` or the event callback mechanism provided by `libbpf` (such as `handle_event()`), perform structured parsing, and output formatted data.
[0085] At the application level, this step is widely applicable to scenarios such as high-performance computing, deep learning training, GPU driver debugging, resource monitoring, and performance optimization. User-mode programs can be deployed as standalone monitoring tools or integrated into system management platforms to enable non-intrusive, low-overhead tracking of CUDA API calls without modifying application source code or recompiling, making them suitable for real-time performance analysis and anomaly detection in production environments.
[0086] Furthermore, this step significantly reduces the performance overhead brought by traditional instrumentation tools through an efficient kernel-user mode communication mechanism, while supporting flexible event filtering and display logic. For example, through timestamp sorting, call stack backtracing, call frequency statistics, etc., it provides developers with fine-grained insights into GPU call behavior, thereby improving system observability and debugging efficiency.
[0087] The method for implementing CUDA call tracing based on eBPF in an embodiment of the present invention implements non-intrusive, low-overhead, fine-grained tracing of CUDA API calls, thereby improving the observability and performance analysis efficiency of GPU applications.
[0088] Furthermore, it also includes: S5, filters and aggregates the structure data, filters out the target CUDA call events according to the preset filtering rules, and stores the filtered data in the user-state persistent log file for subsequent offline analysis and auditing.
[0089] Specifically, this step involves filtering and aggregating structure data and persistently storing the filtered target CUDA call events in a user-space log file. This is a key step in implementing data processing and storage in the CUDA call tracing system. Technically, this step first filters raw event data read from the kernel-space eBPF ring buffer (BPF_MAP_TYPE_RINGBUF) based on pre-defined filtering criteria (such as process ID, event type, and time range). The user-space program reads the event structure (struct event) from the ring buffer via the libbpf library and, based on the configured filtering criteria, uses conditional statements or regular expression matching to perform multi-dimensional filtering based on event type (enumcuda_event_type), process ID (pid), and timestamp (nanosecond precision) to extract event data related to the target CUDA call. The system then aggregates the filtered event data, for example, by categorizing call frequency by process ID, analyzing call patterns by event type, or analyzing call timing using timestamps to improve data readability and analysis efficiency. At the parameter indicator level, filtering rules can be configured as whitelist or blacklist, supporting regular expression matching (such as `cudaMemcpy`, `cudaLaunchKernel`, etc.), the time range can be set to a nanosecond time window (such as `start_time` and `end_time`), and the process ID can be specified as a single or multiple target processes. Aggregation processing can be based on a sliding window mechanism (such as statistics every 100ms) or event type classification to summarize data. In application scenarios, this step is widely used in performance analysis, resource scheduling optimization, security auditing and other scenarios. For example, in a multi-process parallel computing environment, by specifying the process ID and event type, the GPU memory allocation behavior of the preset process can be accurately tracked, assisting in identifying memory leaks or resource contention issues. In terms of technical effects, this step significantly reduces the data processing overhead of user space programs through efficient filtering and aggregation mechanisms, and improves the real-time and accuracy of log analysis. Furthermore, data is persistently stored in user-mode log files (e.g., using the `fwrite` or `pwrite` interfaces to write data to disk in binary or JSON format). This ensures long-term data availability and provides a reliable data source for subsequent offline analysis, visualization, and automated alerting systems. This step not only enhances the system's flexibility and scalability but also demonstrates the technical advantages of this invention in low-overhead, non-intrusive tracking.
[0090] The method according to an embodiment of the present invention non-invasively mounts an eBPF program within the user-space CUDA runtime library, enabling comprehensive and fine-grained capture of CUDA API call parameters, return values, and execution timing. Compared to traditional tools, this method offers extremely low overhead and high efficiency, and utilizes a high-performance ring buffer to enable asynchronous communication between kernel and user space, ensuring minimal impact on CUDA application performance.
[0091] Example 2 like Figure 2 The following is an architecture diagram of the method for implementing CUDA call tracing based on eBPF. The specific process is as follows: 1) Define the shared data structures required for data communication between kernel-space eBPF programs and user-space applications. The key information in these structures is as follows: they primarily include key fields such as process ID, process name, event type, and event details. These fields can be expanded based on business needs.
[0092] 2) Because eBPF programs cannot directly perform file I / O or network communication, their communication with userspace applications is primarily achieved through eBPF maps. This paper uses a ring queue of type `BPF_MAP_TYPE_RINGBUF` as the primary communication mechanism, allowing eBPF programs to submit events in kernel space and userspace applications to retrieve these events.
[0093] For each type of CUDA call, a corresponding function needs to be designed to read the key information in the calling process. Taking memory application as an example, the implemented function needs to include the following content: parse the process ID, process name, memory application size, success or failure information, and fill it into the data structure, and submit it to the ring buffer for user-mode retrieval and processing.
[0094] 3) The userspace application is the consumer of kernel data. It is responsible for interacting with the operating system kernel, loading the eBPF program, and reading, parsing, formatting, and presenting the captured CUDA event data from the ring buffer. The userspace application uses libbpf to load and attach the eBPF program to the CUDA function.
[0095] This function accepts a function name (e.g. "cudaMalloc") and the corresponding entry eBPF program. It then attaches these programs as uprobes to the specified library to implement the mounting of the eBPF program.
[0096] After reading the data from the ring queue, you can visualize it according to your business needs.
[0097] 4) Compile the above program code to obtain an executable file, run the executable file to complete the loading of the eBPF program, and when the CUDA API is subsequently called, the corresponding call information will be visualized on the client.
[0098] The core concept of this invention is to use eBPF's uprobe technology to dynamically mount eBPF programs at the entry and exit points of specified user-space functions without modifying the CUDA runtime library binary, recompiling the application, or even requiring the application source code. This non-invasive feature greatly simplifies deployment, reduces risks to the production environment, and ensures that the original performance and behavior of the tracked application are not interfered with. It avoids the compatibility issues and performance degradation that may be caused by traditional instrumentation or injection techniques. Using eBPF technology in kernel space, detailed metadata for each CUDA event is captured, including the calling process ID, process name, event occurrence timestamp (nanosecond precision), and parameters and return values associated with a specific API call. This data is then uniformly encapsulated and transferred from kernel state to user space using the efficient data structures provided by eBPF. The user-space program performs two tasks: first, it mounts the eBPF program on the corresponding CUDA API so that CUDA call events are intercepted and recorded; second, the eBPF application can easily read structured data stored in kernel mode, further analyze the data, and display it in a readable format, so that key information about CUDA events can be discovered in a timely manner.
[0099] The specific implementation process is as follows: Defines the shared data structures required for data communication between kernel-space eBPF programs and user-space applications. The key information in the defined data structures is as follows, including key fields such as process ID, process name, event type, and event details. These fields can be expanded based on business needs.
[0100] Because eBPF programs cannot directly perform file I / O or network communication, they communicate with userspace applications primarily through eBPF maps. This paper uses a ring queue of type `BPF_MAP_TYPE_RINGBUF` as the primary communication mechanism, allowing eBPF programs to submit events in kernel space and userspace applications to retrieve these events.
[0101] For each type of CUDA call, a corresponding function needs to be designed to read the key information in the calling process. Taking memory application as an example, the implemented function needs to include the following content: parse the process ID, process name, memory application size, success or failure information, and fill it into the data structure, and submit it to the ring buffer for user-mode retrieval and processing.
[0102] The userspace application is the consumer of kernel data. It is responsible for interacting with the operating system kernel, loading the eBPF program, and reading, parsing, formatting, and presenting the captured CUDA event data from the ring buffer. The userspace program uses libbpf to load and attach the eBPF program to the CUDA function.
[0103] This function accepts a function name (e.g. "cudaMalloc") and the corresponding entry eBPF program. It then attaches these programs as uprobes to the specified library to implement the mounting of the eBPF program.
[0104] After reading the data from the ring queue, you can visualize it according to your business needs.
[0105] Compile the above program code to obtain an executable file, run the executable file to complete the loading of the eBPF program, and when the CUDA API is subsequently called, the corresponding call information will be visualized on the client.
[0106] The method of the embodiment of the present invention designs an eBPF program to non-invasively and comprehensively obtain detailed information during the CUDA API call process, has the advantages of high efficiency and low overhead, and designs a reasonable and efficient data structure to realize data interaction and sharing between kernel space and user space. The kernel space part of the eBPF program implements data acquisition and parsing corresponding to the CUDA call, fills the detailed information into the corresponding fields of the structure and submits it to the ring buffer queue for the user space program to read. The user space program calls the standard function library to complete the mounting of the eBPF program at the corresponding location, and reads the data in the ring buffer in real time to complete the data analysis and display. The compiler generates an executable file, and runs the executable file to realize CUDA call tracing.
[0107] In summary, the beneficial effects of the present invention are: non-invasive, deep tracing can be achieved without modifying the application or CUDA runtime library, which greatly reduces deployment and maintenance costs and is suitable for production environments. Low overhead, eBPF programs run efficiently in kernel state and communicate asynchronously through a ring buffer, minimizing the impact on the performance of the traced application. Fine-grained, capable of capturing entry parameters, return values, and precise timestamps of CUDA API calls, providing unprecedented insights into GPU behavior. Modularity and scalability: The clear architectural design makes the system easy to maintain and can be easily expanded to support new CUDA APIs or more complex tracing requirements.
[0108] Example 3 In order to implement the above embodiment, Figure 3 As shown, this embodiment also provides a system 10 for implementing CUDA call tracing based on eBPF, including: The mounting configuration module 100 is used to dynamically mount the eBPF program at the API function entry and exit of the user-mode CUDA runtime library through the eBPF uprobe technology to achieve non-intrusive monitoring of CUDA API calls; The metadata collection module 200 is used to capture metadata of CUDA API call events in kernel mode and encapsulate the metadata into a predefined structure; The data transmission module 300 is used to asynchronously transmit the structure data to the user state through the ring buffer of type BPF_MAP_TYPE_RINGBUF; The event analysis module 400 is used to call the standard function library to load the eBPF program and read the data in the ring buffer in real time to complete the analysis and visualization of CUDA call events.
[0109] Furthermore, the data structure information of the metadata includes process ID, process name, event type, event parameters, return value and event occurrence timestamp.
[0110] Furthermore, the mount configuration module is also used to: The `bpf_program__attach_uprobe_opts` function is used to attach the eBPF program to the specified API function entry and exit points in the user-mode CUDA runtime library. The attach operation supports specifying the target library path through the `lib_path` parameter and the target function name through the `func_name` parameter. During the mounting process, you can set the `env.target_pid` parameter to selectively monitor the CUDA API calls of the preset process.
[0111] Furthermore, the metadata collection module is also used to: The metadata includes nanosecond timestamps. The time information of the event is obtained through the `bpf_ktime_get_ns()` function and encapsulated into the `timestamp` field of the structure. The event parameters and the return value are parsed through the context parameters of the eBPF program and dynamically filled into the corresponding fields in the `event_instance` union according to different CUDA API types.
[0112] Furthermore, the data transmission module is also used to: The maximum number of entries in the ring buffer is set to `256 * 1024` to ensure data throughput in high-concurrency CUDA call scenarios; The asynchronous transmission uses the `bpf_ringbuf_submit` function to submit the encapsulated structure data to the ring buffer, and uses the `bpf_ringbuf_consume` function to perform non-blocking reading in user mode.
[0113] Furthermore, it also includes: The data processing module is used to filter and aggregate structure data, filter out target CUDA call events according to preset filtering rules, and store the filtered data in a user-state persistent log file for subsequent offline analysis and auditing.
[0114] The eBPF-based CUDA call tracing system of the present invention addresses the challenges of existing GPU operation tracing techniques, particularly the shortcomings of traditional tools in terms of non-invasiveness, low overhead, and fine-grained information capture. This system leverages the powerful programmability, event-driven nature, and non-invasiveness of eBPF technology at the Linux kernel level to precisely intercept and capture user-space CUDA API calls, thereby providing comprehensive, real-time insights for debugging, performance analysis, and security auditing of CUDA applications.
[0115] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for implementing CUDA call tracing based on eBPF is implemented.
[0116] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0117] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. A method for implementing CUDA call tracing based on eBPF, characterized in that: include: S1, through the eBPF uprobe technology, dynamically mount the eBPF program at the API function entry and exit of the user-mode CUDA runtime library to achieve non-intrusive monitoring of CUDA API calls; S2, the eBPF program captures the metadata of the CUDA API call event in kernel mode and encapsulates the metadata into a predefined structure; S3, asynchronously transfer the structure data to userland via a ring buffer of type BPF_MAP_TYPE_RINGBUF; S4: The user-mode program calls the standard function library to load the eBPF program and reads the data in the ring buffer in real time to complete the analysis and visualization of the CUDA call event.
2. The method according to claim 1, wherein The data structure information of the metadata includes process ID, process name, event type, event parameters, return value and event occurrence timestamp.
3. The method according to claim 1, wherein Said S1 comprises: S11, mount the eBPF program to the specified API function entry and exit in the user-mode CUDA runtime library through the `bpf_program__attach_uprobe_opts` function. The mount operation supports specifying the target library path through the `lib_path` parameter and the target function name through the `func_name` parameter. S12, during the mounting process, selectively monitor the CUDA API calls of the preset process by setting the `env.target_pid` parameter.
4. The method according to claim 2, wherein The S2 includes: S21, get the time information of the event through the `bpf_ktime_get_ns()` function and encapsulate it into the `timestamp` field of the structure; S22: The event parameters and the return value are parsed through the context parameters of the eBPF program and dynamically filled into corresponding fields in the `event_instance` union according to different CUDA API types.
5. The method according to claim 1, wherein The S3 includes: S31, the maximum number of entries in the ring buffer is set to `256 * 1024` to ensure data throughput in a high-concurrency CUDA call scenario; S32, the asynchronous transmission uses the `bpf_ringbuf_submit` function to submit the encapsulated structure data to the ring buffer, and uses the `bpf_ringbuf_consume` function to perform non-blocking reading in user mode.
6. The method according to claim 1, wherein Also includes: S5, filtering and aggregating the structure data, filtering out target CUDA call events according to preset filtering rules, and storing the filtered data in a user-state persistent log file.
7. A system for implementing CUDA call tracing based on eBPF, characterized in that: include: The mount configuration module is used to dynamically mount eBPF programs at the entry and exit of the API function of the user-mode CUDA runtime library through the eBPF uprobe technology to achieve non-intrusive monitoring of CUDA API calls; A metadata collection module is used to capture metadata of CUDA API call events in kernel mode and encapsulate the metadata into a predefined structure; The data transmission module is used to asynchronously transfer structure data to user mode through the ring buffer of type BPF_MAP_TYPE_RINGBUF; The event analysis module is used to call the standard function library to load the eBPF program and read the data in the ring buffer in real time to complete the analysis and visualization of CUDA call events.
8. The system according to claim 7, wherein: The data structure information of the metadata includes process ID, process name, event type, event parameters, return value and event occurrence timestamp.
9. The system according to claim 7, wherein: The mount configuration module is also used to: The `bpf_program__attach_uprobe_opts` function is used to attach the eBPF program to the specified API function entry and exit points in the user-mode CUDA runtime library. The attach operation supports specifying the target library path through the `lib_path` parameter and the target function name through the `func_name` parameter. During the mounting process, you can set the `env.target_pid` parameter to selectively monitor the CUDA API calls of the preset process.
10. The system according to claim 8, wherein The metadata collection module is further used for: The metadata includes nanosecond timestamps. The time information of the event is obtained through the `bpf_ktime_get_ns()` function and encapsulated into the `timestamp` field of the structure. The event parameters and the return value are parsed through the context parameters of the eBPF program and dynamically filled into the corresponding fields in the `event_instance` union according to different CUDAAPI types.
11. The system according to claim 7, wherein: The data transmission module is further used for: The maximum number of entries in the ring buffer is set to `256 * 1024` to ensure data throughput in high-concurrency CUDA call scenarios; The asynchronous transmission uses the `bpf_ringbuf_submit` function to submit the encapsulated structure data to the ring buffer, and uses the `bpf_ringbuf_consume` function to perform non-blocking reading in user mode.
12. The system according to claim 7, wherein: Also includes: The data processing module is used to filter and aggregate structure data, filter out target CUDA call events according to preset filtering rules, and store the filtered data in a user-state persistent log file for subsequent offline analysis and auditing.
13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Full-behavior monitoring method and system in bypass non-intrusive application operation based on Kubernetes
CN115617610A
Process crash information collection method and device based on eBPF
CN116594796A
Linux kernel performance tracking tool based on eBPF
CN116662134A
Method and device for monitoring data security
CN117251852A
User mode process behavior monitoring method and device based on eBPF
CN118051400A
Cited By
Kernel audit message transmission method based on structural body
CN121116899A
Kernel auditing message transmission method based on structure
CN121116899B
Unexported callback function address acquisition method and system based on pointer chain tracking
CN121166494A
A method and system for obtaining the address of unexported callback functions based on pointer chain tracing
CN121166494B
GPU scheduling observation and control method and device based on eBPF and storage medium
CN121277712A