GPU (Graphic Processing Unit) and parallel IO (Input / Output) collaborative optimization method based on mode heterogeneous calculation

By monitoring the computing task mode in real time and generating a multi-dimensional resource situation matrix, the equipment and time dimensions of the heterogeneous computing environment are optimized, and a three-level adaptive transmission mechanism is built, which solves the problems of missing I/O access pattern matching and insufficient utilization of storage hierarchy in heterogeneous computing, improves the data prefetch hit rate and storage bandwidth utilization, and reduces latency and task contention.

CN120295803AActive Publication Date: 2025-07-11YUNHAI ZHICHUANG (JIANGSU) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510788400.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

When the traditional heterogeneous computing architecture accelerates the calculation, it ignores the coordinated optimization between parallel I/O and computing tasks, resulting in prominent bottleneck problems in storage walls and I/O, missing matching between heterogeneous computing mode and I/O access pattern, low data prefetch hit rate, difficult to achieve data dependency synchronization caused by dynamic task migration, multi-level storage hierarchical access characteristics are not fully utilized, and the peak utilization rate of heterogeneous storage bandwidth is low.

Method used

By monitoring the pattern characteristics of the computing task in real time, a multi-dimensional resource situation matrix is generated, collaborative optimization operations are performed, and a three-level adaptive transmission mechanism is built, including binding device dimensions and scheduling time dimensions, optimizing data transmission at the register-level, storage-level and network-level, and dynamically adjusting the usage strategy of computing resources to match the task type.

Benefits of technology

It improves the data prefetch hit rate, solves the data dependency synchronization problem in dynamic task migration, improves the peak utilization of heterogeneous storage bandwidth, reduces end-to-end latency, and enhances the I/O throughput of storage-intensive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295803A_ABST
    Figure CN120295803A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU (Graphics Processing Unit) and parallel IO (Input / Output) collaborative optimization method based on mode heterogeneous computing, and particularly relates to the field of heterogeneous computing optimizing.The method comprises the following steps: monitoring mode characteristics during execution of a target computing task in real time through an instruction analyzer, and synchronously collecting dynamic parameters of a heterogeneous computing environment to generate a multi-dimensional resource situation matrix; based on the mode and the resource situation matrix, collaborative optimization operation of the binding equipment dimension and the scheduling time dimension is executed; and constructing a three-level adaptive transmission mechanism and periodically implementing feedback optimization according to an execution state of the three-level adaptive transmission mechanism. According to the GPU and parallel IO collaborative optimization method based on mode heterogeneous computing, the data prefetching hit rate is increased by dynamically sensing the matching relation between a heterogeneous computing mode and an I / O access mode; end-to-end delay is reduced and task contention conflicts are reduced through a dynamic matching strategy of an equipment capability portrait and a task type; and the heterogeneous storage bandwidth peak value utilization rate and the I / O throughput are improved by differentially utilizing the hierarchical characteristics of the multi-stage storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heterogeneous computing optimization, and more specifically, to a method for co - optimizing GPU and parallel I / O based on pattern - heterogeneous computing. Background Art

[0002] With the explosive growth of the demand for high - performance computing and big - data processing, heterogeneous computing architectures have become a key technical direction for improving computing power and energy efficiency. Especially in scientific computing and industrial applications, traditional heterogeneous computing architectures use the CPU as the control core and the GPU as the coprocessor, and realize task allocation through a master - slave programming model, which plays a key role in accelerating compute - intensive tasks. However, traditional heterogeneous systems often only focus on compute acceleration and ignore the co - optimization between parallel I / O and compute tasks, resulting in increasingly prominent memory wall and I / O bottleneck problems.

[0003] Currently, existing technologies propose dynamic load - balancing and asynchronous data - transfer optimization methods. By introducing a hierarchical task - scheduling mechanism, using MPI distributed communication between nodes, realizing the cooperation between CPU multi - threading and GPU stream processors within nodes, and using double - buffering technology to overlap computing and data transfer, the task - allocation granularity is refined to the thread level, and the CPU / GPU task ratio is dynamically adjusted through runtime performance prediction.

[0004] However, in actual use, there are still some disadvantages, such as the lack of matching between heterogeneous computing modes and I / O access modes, resulting in a decrease in storage locality and a reduction in data prefetch hit rate; data dependencies caused by dynamic task migration are difficult to effectively synchronize through existing communication protocols, and task contention leads to a sharp increase in end - to - end latency; the access characteristics of multi - level storage hierarchies are not differentially utilized, and the peak utilization rate of heterogeneous storage bandwidth is low. Summary of the Invention

[0005] In order to overcome the above - mentioned defects of the prior art, the present invention provides a method for co - optimizing GPU and parallel I / O based on pattern - heterogeneous computing, through the following solutions to solve the problems raised in the above - mentioned background art.

[0006] To achieve the above object, the present invention provides the following technical solutions: A method for co - optimizing GPU and parallel I / O based on pattern - heterogeneous computing, comprising: S1: During the execution of the target computing task, the instruction analyzer monitors and obtains the pattern features of the target computing task in real - time; S2: Dynamically collect the parameters of the heterogeneous computing environment where the target computing task is located in real - time, and generate a multi - dimensional resource situation matrix; S3: Based on the pattern features and the multi - dimensional resource situation matrix, perform co - optimization operations, and the co - optimization operations include binding the device dimension and scheduling the time dimension; S4: Based on the execution status of the collaborative optimization operation, construct a three-level adaptive transmission mechanism: S5: Based on the execution status of the three-level adaptive transmission mechanism, periodically execute feedback optimization operations.

[0007] Preferably, in S1, the mode features include a calculation density coefficient, a data locality strength, and a communication dependency graph. The calculation density coefficient is a quantitative index reflecting the relative relationship between the computational intensity and the memory access intensity; the data locality strength is used to measure the degree of concentration of data accessed during the execution of a computational task; the communication dependency graph is used to describe the data transfer relationship and load situation between different computational tasks in a heterogeneous computing environment.

[0008] Preferably, the steps for obtaining the mode features in S1 by an instruction analyzer include: Use a sliding window weighted average algorithm to calculate the calculation density coefficient The calculation is specifically expressed as: , where, and respectively represent the number of floating-point operation instructions and memory access instructions counted within the th time window; represents the weight assigned to the th time window, represents the number of time windows; an exponential decay is used to control the influence degree of data in each time window, specifically expressed as: , where, represents a pre-configured decay coefficient; Based on the memory access address sequence obtained within a sampling period, calculate the dispersion between addresses to represent the spatial locality metric , specifically expressed as: , where, represents the total number of memory access operations within the sampling period, represents the index of the memory access operation, represents the target address of the th memory access operation; represents calculating the Hamming distance between two consecutive memory access addresses; represents the address bit width for calculating the spatial locality; Statistically analyze the distribution of the reuse distance of memory access addresses and calculate the entropy value to represent the temporal locality metric , specifically expressed as: , Wherein, is represented as a pre-set maximum reuse distance threshold, is represented as the index of the reuse distance, is represented as the proportion of the memory access operation with the reuse distance of in the sampling period, and the value is the ratio of the number of accesses with the reuse distance of to the total number of memory access operations in the sampling period.

[0009] Preferably, in the step S2, the multi-dimensional resource situation matrix includes a device dimension, a resource type dimension, and a time dimension. Among them, the resource type dimension is used to identify specific dynamic parameters.

[0010] Preferably, in the step S3, the implementation method of binding the device dimension in the execution of the collaborative optimization operation includes: Judging the task type of the target computing task according to the pattern feature; Generating a device capability profile according to the dynamic parameters of the corresponding heterogeneous computing environment in the multi-dimensional resource situation matrix. The device capability profile includes the real-time computing power score of the GPU, the I / O throughput capability score of the storage node, and the transmission quality score of the communication path between nodes; Based on the matching strategy between the task type and the device capability profile, dynamically execute the device dimension binding decision.

[0011] Preferably, in the step S3, the implementation method of scheduling the time dimension in the execution of the collaborative optimization operation includes: 1 - 3 clock cycles before the GPU kernel function starts, trigger the prefetch strategy through CUDA stream events to load the data of the next calculation window; Realize the strict timing matching between the execution of the computing kernel function and the data transmission through the hardware and software cooperation mechanism, and monitor the computing cycle and the DMA transmission cycle in real time; If the computing cycle is less than the DMA transmission cycle, enable GPU dynamic frequency scaling to increase the SM clock frequency to the theoretical peak; If the computing cycle is greater than the DMA transmission cycle, increase the DMA queue depth.

[0012] Preferably, in the step S4, the three-level adaptive transmission mechanism includes register-level transmission optimization, storage-level transmission optimization, and network-level transmission optimization. Among them, the register-level transmission optimization is for the task type of the target computing task being compute-intensive and the real-time computing power score of the GPU in the device capability profile being greater than the GPU real-time computing power score threshold. The storage-level transmission optimization is for the high I / O throughput demand type and the I / O throughput capability score of the target storage node being greater than the I / O throughput capability score threshold. The network-level transmission optimization is for the communication-intensive type and the transmission quality score involving the communication between nodes being greater than the transmission quality score threshold.

[0013] Preferably, in S4, the priority of the three-level adaptive transmission mechanism is register-level transmission optimization > memory-level transmission optimization > network-level transmission optimization. When the target computing task satisfies the triggering conditions of multiple levels simultaneously, the optimization strategy with the highest priority will be executed first.

[0014] Technical effects and advantages of the present invention: 1. By monitoring the mode of the computing task in real time and dynamically perceiving the matching relationship between the heterogeneous computing mode and the I / O access mode, the present invention solves the problem of the lack of matching between the heterogeneous computing mode and the I / O access mode, and improves the data prefetch hit rate. 2. Through the dynamic matching strategy of device capability profiling and task type, combined with the phase alignment mechanism of CUDA streams and DMA engines, the present invention effectively solves the data-dependent synchronization problem in dynamic task migration, reduces the end-to-end latency, and reduces task contention conflicts. 3. By differentially utilizing the characteristics of multi-level storage hierarchies, the present invention improves the peak utilization rate of heterogeneous storage bandwidth, solves the defect that the access characteristics of traditional multi-level storage hierarchies are not fully utilized, and improves the I / O throughput of storage-intensive tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 FIG. is a flowchart of the steps of a GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to an embodiment of the present application.

[0016] Figure 2 FIG. is a flowchart of the steps of binding device dimensions in a GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular forms "a", "an", "the", "above", "the foregoing", "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items. In the description of the embodiments of the present application, unless otherwise specified, "a plurality" means two or more.

[0019] As shown in the Figure 1 GPU and parallel IO collaborative optimization method based on pattern heterogeneous computing shown, including real-time monitoring of the pattern characteristics during the execution of the target computing task through an instruction analyzer, and synchronously collecting dynamic parameters of the heterogeneous computing environment to generate a multi-dimensional resource situation matrix; based on the pattern and resource situation matrix, performing collaborative optimization operations for the bound device dimension and the scheduling time dimension; constructing a three-level adaptive transmission mechanism and periodically implementing feedback optimization according to its execution status.

[0020] In a specific embodiment, the GPU and parallel IO collaborative optimization method based on pattern heterogeneous computing can be applied to a heterogeneous computing environment including one or more host CPUs, one or more graphics processing units (GPUs), and a parallel storage system. The specific steps of the GPU and parallel IO collaborative optimization method based on pattern heterogeneous computing are as follows: S1: During the execution of the target computing task, the instruction analyzer is used to real-time monitor and obtain the pattern characteristics of the target computing task; S2: Real-time collect the dynamic parameters of the heterogeneous computing environment where the target computing task is located to generate a multi-dimensional resource situation matrix; S3: Based on the pattern characteristics and the multi-dimensional resource situation matrix, perform collaborative optimization operations, and the collaborative optimization operations include binding the device dimension and the scheduling time dimension; S4: Based on the execution status of the collaborative optimization operation, construct a three-level adaptive transmission mechanism: S5: Based on the execution status of the three-level adaptive transmission mechanism, periodically perform feedback optimization operations.

[0021] Specifically, in S1, the pattern characteristics include the computational density coefficient, the data locality strength, and the communication dependency graph. The computational density coefficient is a quantitative index used to reflect the relative relationship between the computational intensity and the memory access intensity; the data locality strength is used to measure the degree of data concentration accessed by the computing task during execution, and in this embodiment, it includes spatial locality measurement and temporal locality measurement; the communication dependency graph is used to describe the data transfer relationship and load situation between different computing tasks in the heterogeneous computing environment.

[0022] It should be noted that the computational density coefficient uses an instruction analysis tool to real-time capture the instruction stream of the target computing task executed on the GPU. The instruction analysis tool includes, but is not limited to, dynamic binary instrumentation, using a hardware performance counter interface, or a specific hardware tracing unit, etc. The number of floating-point operation instructions and memory access instructions is counted within a preset time window. The floating-point operation instructions include, but are not limited to, arithmetic logic instructions such as single-precision, double-precision, and half-precision; the memory access instructions include, but are not limited to, load and store instructions for global memory and access instructions for shared memory.

[0023] In this embodiment, to smooth the instantaneous fluctuations and reflect the recent trends, the sliding window weighted average algorithm is used to calculate the computational density coefficient The calculation is specifically expressed as: , where and respectively represent the number of floating-point operation instructions and the number of memory access instructions counted within the th time window; represents the weight assigned to the th time window, represents the number of time windows; further, the exponential decay is used to control the influence degree of the data in each time window, which is specifically expressed as: , where represents the pre-configured decay coefficient, which is set to 0.1 in this embodiment.

[0024] It should be noted that the spatial locality metric in the data locality strength is used to measure the proximity of consecutive memory access operations in the address space; the temporal locality metric is used to measure the frequency of repeated access to the same data.

[0025] Based on the memory access address sequence obtained within a sampling period, this embodiment calculates the degree of dispersion between addresses to represent the spatial locality metric which is specifically expressed as: , where represents the total number of memory access operations within the sampling period, represents the index of the memory access operation, represents the target address of the th memory access operation; represents the calculation of the Hamming distance between two consecutive memory access addresses; represents the address bit width for calculating the spatial locality; further, the closer the value is to 1, the better the spatial locality; the Hamming distance is a metric for measuring the difference between two strings, defined as the number of different bits at each position in the two strings.

[0026] This embodiment calculates the entropy value to represent the temporal locality metric by statistically analyzing the distribution of the reuse distance of memory access addresses. The reuse distance is the number of accesses to other different addresses that occur between two consecutive accesses to an address which is specifically expressed as: , where is represented as a pre-set maximum reuse distance threshold, is represented as the index of the reuse distance, is represented as the proportion of the memory access operation with a reuse distance of in the sampling period, and the value is the ratio of the number of accesses with a reuse distance of to the total number of memory access operations in the sampling period; further, the smaller the value, the more concentrated the address reuse pattern and the higher the predictability, that is, the better the temporal locality.

[0027] The steps for obtaining the communication dependency graph in this embodiment are as follows: S101: Periodically record communication event data by intercepting, monitoring, and utilizing standard inter-process communication interfaces. The communication event data includes communication endpoints, data characteristics, and time information; among them, the communication endpoints are the logical identifiers of the source device and the target device, the data characteristics include the data volume transmitted each time, the frequency of communication occurrence, and the transmission direction; the time information includes the timestamps of the start and end of the communication. S102: Based on the captured number of communication events, construct a weighted directed graph within a preset time window to represent the communication topology and load intensity, that is, the communication dependency graph; each vertex of the weighted directed graph represents a computing entity participating in the communication. By adding a directed edge from vertex i to vertex j in the weighted directed graph, it represents the data transmission observed from entity i to entity j within the time window; the weight of the edge of the weighted directed graph is used to quantify the communication load between device i and j.

[0028] Specifically, in S2, collect the dynamic parameters of the heterogeneous computing environment where the target computing task is located. The dynamic parameters are key performance indicators that reflect the real-time working status and load levels of each computing, storage, and communication unit during system operation. Further, the dynamic parameters include computing resource parameters, storage resource parameters, and I / O resource parameters. Among them, the computing resource parameters include but are not limited to CPU utilization rate, SM occupancy rate, GPU register file pressure value, GPU shared memory bank conflict rate, and GPU warp scheduling efficiency, etc. The storage resource parameters include the access latency and bandwidth utilization rate of the GPU video memory, the host DRAM memory bandwidth utilization rate and page hit rate, and the read / write bandwidth and IO queue depth of the NVM, etc. The communication / network resource parameters include the effective bandwidth of the PCIe link, the RDMA end-to-end latency, the network interface bandwidth utilization rate, the number of network hops between nodes, the network switch port congestion rate, the TCP / UDP packet loss rate, and the network interface IO queue status, etc.

[0029] In this embodiment, the acquisition work is performed periodically during operation, and the pre-set acquisition period is set to 10 ms.

[0030] Furthermore, a multi-dimensional resource situation matrix is constructed based on the dynamic parameters of the heterogeneous computing environment collected. The multi-dimensional resource situation matrix is used to integrate, quantify, and track the real-time status of each device in the heterogeneous computing environment on different resource dimensions and its change trend over time. The multi-dimensional resource situation matrix includes a device dimension, a resource type dimension, and a time dimension. Among them, the device dimension is used to identify each physical device participating in the calculation, and is represented by a device number; the resource type dimension is used to identify specific dynamic parameters, and is distinguished by a predefined subclass code; the time dimension is used to track historical status, represented by a time window serial number and adopts a rolling update mechanism.

[0031] It should be noted that the generation process of the multi-dimensional resource situation matrix includes: collecting the original dynamic parameter values of each device according to a preset collection period; data cleaning and preprocessing, smoothing the outliers caused by instantaneous noise, and unifying the data units; applying a normalization formula to each preprocessed parameter value to calculate a score in the range of [0, 1]; filling the calculated normalized score value into the corresponding position in the multi-dimensional resource situation matrix, and rolling up the time window, discarding the oldest data, and completing the update of the multi-dimensional resource situation matrix.

[0032] Specifically, in S3, based on the pattern features obtained in S1 and the multi-dimensional resource situation matrix obtained in S2, collaborative optimization operations are respectively performed. The collaborative optimization operations include binding the device dimension and scheduling the time dimension; Furthermore, the implementation method of binding the device dimension in the collaborative optimization operation includes: judging the task type of the target computing task according to the pattern features. It should be noted that if the computing density coefficient is greater than the preset computing density coefficient threshold, it is determined that the target computing task is a computing-intensive task; if the data locality intensity is less than the preset data locality intensity threshold and the communication dependency graph shows that the number of data transmissions between the task and the storage node is greater than the transmission threshold, it is determined that the target computing task is a high-I / O throughput demand task; if there is a cross-node communication path in the communication dependency graph, it is determined that the target computing task is communication-intensive; according to the dynamic parameters of the corresponding heterogeneous computing environment in the multi-dimensional resource situation matrix, a device capability portrait is generated. The device capability portrait includes the real-time computing power score of the GPU, the I / O throughput capability score of the storage node, and the transmission quality score of the communication path between nodes; In this embodiment, the device capability portrait is generated by a weighted aggregation and dynamic calibration algorithm, where the real-time computing power score of the GPU is calculated , specifically expressed as: , where, , , , respectively represent the weight coefficients of the real-time computing power score of the GPU, Expressed as the SM occupancy rate, Expressed as the theoretical maximum SM occupancy rate, Expressed as the GPU register file pressure value, Expressed as the upper limit of the register file capacity, Expressed as the GPU warp scheduling efficiency, Expressed as the bandwidth utilization rate of the GPU video memory, Expressed as the theoretical peak bandwidth of the video memory; Further, when > 90%, automatically reduce the weight to avoid overloading the selected device; When < 30%, increase the weight to prioritize bandwidth-sensitive tasks; Calculate the I / O throughput capacity score of the compute storage node, specifically expressed as: , wherein, , , respectively represent the weight coefficients of the I / O throughput capacity score of the compute storage node, Expressed as the read bandwidth of the NVM, Expressed as the write bandwidth of the NVM, Expressed as the nominal bandwidth of the NVM device, Expressed as the IO queue depth, Expressed as the maximum supported queue depth, Expressed as the effective bandwidth of the PCIe link, Expressed as the theoretical bandwidth of the PCIe link; Further, when > 80% at that time, trigger weight decay to avoid queue overflow; When > 90%, then the weight is reduced to 0 to avoid link congestion; Calculate the transmission quality score of the communication path between compute nodes, specifically expressed as: , wherein, , , respectively represent the weight coefficients of the transmission quality score of the communication path between compute nodes, Expressed as the RDMA end-to-end delay, Expressed as the network switch port congestion rate, Expressed as the TCP packet loss rate; Based on the matching strategy of the task type and the device capability profile, dynamically execute the device dimension binding decision: It should be noted that the device dimension binding decision is the matching strategy corresponding to the target computing task, and the matching strategy includes, but is not limited to, the matching strategy for compute-intensive tasks, the matching strategy for tasks with high I / O throughput requirements, the matching strategy for communication-intensive tasks, etc.; In this embodiment, the matching strategy for compute-intensive tasks is to select the GPU device with the highest real-time computing power score of the GPU, and preferentially bind it to the computing unit with an SM occupancy rate lower than 80% and a register pressure value < 70% to avoid resource contention; the matching strategy for tasks with high I / O throughput requirements is to map the task to the storage node with the highest I / O throughput capacity score, and preferentially select the storage device with an NVMe queue depth < 80% and a PCIe bandwidth utilization rate < 75%; the matching strategy for communication-intensive tasks is to construct a direct-connected subnet based on the minimum communication overhead according to the topological relationship of the communication dependency graph, and preferentially select the communication path with the highest transmission quality score.

[0033] Furthermore, the implementation method of the scheduling time dimension in the collaborative optimization operation includes: generating a prefetch strategy for the target computing task based on the device dimension binding decision and the task type of the target computing task, and triggering the prefetch strategy through a CUDA stream event 1 - 3 clock cycles before the GPU kernel function starts, and loading the data of the next computing window into the L2 cache; Furthermore, when the data locality intensity of the target computing task and the access latency between the storage resource class parameters at different storage levels , calculate the prefetch window length , specifically expressed as: , where, is the task adjustment factor; in this embodiment, the adjustment factor for compute-intensive tasks is recorded as 0.8, the adjustment factor for communication-intensive tasks is 0.5, and the adjustment factor for I / O-intensive tasks is 0.3; Implement strict timing matching between the execution of the compute kernel function and data transmission through the hardware and software collaboration mechanism; use the PTP protocol of the InfiniBand NIC to calibrate the clock reference of the GPU and the NIC, with the error controlled within ±50 ns, and insert a synchronization barrier in the CUDA stream to force the kernel function to wait for the DMA transmission completion signal; and monitor the computing cycle and the DMA transmission cycle in real time. If the computing cycle is less than the DMA transmission cycle, enable GPU dynamic frequency scaling to increase the SM clock frequency to the theoretical peak; if the computing cycle is greater than the DMA transmission cycle, increase the DMA queue depth.

[0034] Specifically, in S4, dynamically associate the execution status of the collaborative optimization operation to construct a three-level adaptive transmission mechanism; It should be noted that the three-level adaptive transmission mechanism is an optimal transmission strategy dynamically selected based on the data transmission bottlenecks at multiple resource levels in a heterogeneous computing environment. The resource levels include registers, storage, and network. The three-level adaptive transmission mechanism includes register-level transmission optimization, storage-level transmission optimization, and network-level transmission optimization. Among them, register-level transmission optimization is triggered when the task type of the target computing task is computationally intensive and the real-time computing power score of the GPU in the device capability profile is greater than the GPU real-time computing power score threshold. Storage-level transmission optimization is for high I / O throughput demand types and the I / O throughput capacity score of the target storage node is greater than the I / O throughput capacity threshold. Network-level transmission optimization is for communication-intensive types and the transmission quality score involving inter-node transmission is greater than the transmission quality score threshold. When the target computing task meets the trigger conditions of multiple levels simultaneously, the optimization strategy with the highest priority will be executed first. The priority of the three-level adaptive transmission mechanism is register-level transmission optimization level > storage-level transmission optimization > network-level transmission optimization.

[0035] In this embodiment, register-level transmission optimization is triggered when the task type of the target computing task is computationally intensive and the real-time computing power score of the GPU is greater than the GPU real-time computing power score threshold. The GPU real-time computing power score threshold is set to 0.8 in this embodiment. The GPU thread blocks are logically divided into odd and even groups and mapped to different physical regions of the register file respectively. PTX instructions are used to explicitly control data prefetching to achieve the overlap of computing and data loading. Based on the prefetch window length of the computing core, 4 to 8 times of loop unrolling instruction sequences are inserted into the GPU kernel function, and synchronized primitives are used to coordinate the computing and data transmission pipelines between the odd and even Warp groups to maximize the instruction-level parallelism. At the same time, the register file pressure value is monitored in real time, and register resources are allocated on demand according to the register file pressure value. When the register file pressure value > 70%, the register overflow detection mechanism is triggered, and some variables with lower activity are temporarily stored in the shared memory to relieve the register pressure.

[0036] In this embodiment, the storage-level transmission optimization is triggered when the task type of the target computing task is high I / O throughput demand type and the I / O throughput capacity score of the target storage node is greater than the I / O throughput capacity score threshold. The I / O throughput capacity score threshold is set to 0.7 in this embodiment. Using the libcufile library based on the NVMe driver, a direct memory access channel is established from the SSD to the GPU video memory. The data transmission is performed in data blocks aligned with 16KB, completely bypassing the host CPU and memory, eliminating unnecessary copy overhead. According to the NVMe queue depth and PCIe bandwidth in the real-time STS metrics, the block size of the data transmission is dynamically adjusted. In particular, when the IO queue depth > 80% of the maximum supported queue depth, it switches to the random access optimization block mode to adapt to the high-concurrency random read / write scenario. The GPU atomic operation is used to accurately track the completion status of each DMA transmission block, ensuring that the compute kernel function is only started and executed after all the required data is completely transferred to the GPU video memory, ensuring data consistency.

[0037] In this embodiment, the network-level transmission optimization is triggered when the task type of the target computing task is communication-intensive and the transmission quality score involving inter-node transmission is greater than the transmission quality score threshold. The transmission quality score threshold is set to 0.6 in this embodiment. When the topological hop count of the communication path ≥ 3, the UCX protocol is preferentially selected, and the hierarchical routing algorithm is combined to dynamically calculate and compress the communication path, reducing network latency and congestion. When the communication occurs within a node, the zero-copy mode of MPI_Isend / IRecv is preferentially enabled, and the GPURDMA technology is used to directly access the peer GPU video memory, avoiding data transfer between the host memories. The congestion rate of the network switch port on the critical path is monitored in real time. If the network switch port congestion rate > 30%, the TCP transmission window size is dynamically adjusted to ensure the quality of service for critical data transmission. For the critical communication path with high reliability requirements where the transmission quality score > 0.9, a dual-path redundant transmission mechanism is enabled. The receiving end verifies the integrity of the data packet through methods such as hash check and selects the first valid data packet that arrives, improving the communication robustness.

[0038] Specifically, in S5, the execution status of the three-level adaptive transmission mechanism is dynamically associated, and the feedback optimization operation is executed at a preset cycle. The feedback optimization operation is the closed-loop optimization of S1 - S4. In this embodiment, the condition for triggering the feedback optimization operation is to map the performance metrics of the three-level adaptive transmission mechanism to the mode features of S1, and dynamically adjust the feature extraction rules of the computational density coefficient and the data locality strength; according to the over-limit situation of the PCIe bandwidth utilization rate optimized by the storage-level transmission, trigger the real-time update of the multi-dimensional resource situation matrix in S2, including the retraining of the NVMe queue depth threshold and the GPU computing power score; when the cross-node delay fluctuation of the network-level transmission optimization exceeds 20%, S3 re-optimizes the device dimension binding decision and the prefetch window length; if the Warp double-buffering strategy of the register-level transmission optimization fails to achieve the expected optimization effect for 5 consecutive times, automatically trigger the degradation evaluation of the three-level adaptive transmission mechanism in S4 and adjust the trigger condition threshold.

[0039] Secondly, in the attached drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments of the present disclosure are involved, and other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other; Finally, the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for collaborative optimization of GPU and parallel IO based on pattern heterogeneous computing, characterized in that, Including: S1: During the execution of the target computing task, the instruction analyzer monitors and obtains the pattern features of the target computing task in real time; S2: Dynamically collect the dynamic parameters of the heterogeneous computing environment where the target computing task is located, and generate a multi-dimensional resource situation matrix; S3: Based on the pattern features and the multi-dimensional resource situation matrix, perform collaborative optimization operations, and the collaborative optimization operations include binding the device dimension and scheduling the time dimension; S4: Based on the execution status of the collaborative optimization operation, construct a three-level adaptive transmission mechanism: S5: Based on the execution status of the three-level adaptive transmission mechanism, periodically perform feedback optimization operations.

2. The GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to claim 1, wherein: In S1, the pattern features include the computing density coefficient, the data locality intensity, and the communication dependency graph. The computing density coefficient is a quantitative index used to reflect the relative relationship between the computing intensity and the memory access intensity; The data locality intensity is used to measure the degree of concentration of data access during the execution of the computing task; The communication dependency graph is used to describe the data transfer relationship and load situation between different computing tasks in the heterogeneous computing environment.

3. The GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to claim 2, wherein: In S1, the steps for obtaining the pattern features by the instruction analyzer include: The sliding window weighted average algorithm is used to calculate the density coefficient The calculation is specifically expressed as: , Among them, and respectively represent the number of floating-point operation instructions and memory access instructions counted within the th time window; represents the weight assigned to the th time window, represents the number of time windows; The exponential decay is used to control the influence degree of data in each time window, which is specifically expressed as: , Among them, is expressed as a pre-configured attenuation coefficient; Based on the memory access address sequence obtained within a sampling period, calculate the degree of dispersion between addresses to represent the spatial locality metric, specifically expressed as: , Among them, represents the total number of memory access operations within a sampling period, represents the index of the memory access operation, represents the target address of the represents calculating the Hamming distance between two consecutive memory access addresses; represents the address bit width for calculating spatial locality Statistically analyze the distribution of the reuse distance of memory access addresses and calculate the entropy value to represent the temporal locality metric , specifically expressed as: , wherein, is expressed as a pre-set maximum reuse distance threshold, is expressed as the index of the reuse distance, is expressed as the proportion of the memory access operation with the reuse distance of in the sampling period, and the value is the ratio of the number of accesses with the reuse distance of to the total number of memory access operations in the sampling period.

4. The GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to claim 1, characterized in that: In S2, the multi-dimensional resource situation matrix includes the device dimension, the resource type dimension, and the time dimension. Among them, the resource type dimension is used to identify specific dynamic parameters.

5. The GPU and parallel I / O co-optimization method based on pattern heterogeneous computing according to claim 1, wherein: In S3, the implementation method of binding the device dimension in the collaborative optimization operation includes: Judge the task type of the target computing task according to the pattern features; Generate a device capability profile according to the dynamic parameters of the corresponding heterogeneous computing environment in the multi-dimensional resource situation matrix. The device capability profile includes the real-time computing power score of the GPU, the I / O throughput capacity score of the storage node, and the transmission quality score of the communication path between nodes; Based on the matching strategy between the task type and the device capability profile, dynamically execute the device dimension binding decision.

6. The GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to claim 1, wherein: In S3, the implementation method of scheduling the time dimension in the collaborative optimization operation includes: 1-3 clock cycles before the GPU kernel function is started, trigger the prefetch strategy through the CUDA stream event to load the data of the next computing window; Realize the strict timing matching between the execution of the computing kernel function and the data transmission through the hardware and software cooperation mechanism, and monitor the computing cycle and the DMA transmission cycle in real time; If the computing cycle is less than the DMA transmission cycle, enable GPU dynamic frequency scaling to increase the SM clock frequency to the theoretical peak; If the computing cycle is greater than the DMA transmission cycle, increase the DMA queue depth.

7. The GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to claim 1, characterized in that: In S4, the three-level adaptive transmission mechanism includes register-level transmission optimization, memory-level transmission optimization, and network-level transmission optimization. Among them, register-level transmission optimization is for target computing tasks with a task type of compute-intensive and the real-time computing power score of the GPU in the device capability profile being greater than the GPU real-time computing power score threshold. Memory-level transmission optimization is for high I / O throughput demand types and the I / O throughput capacity score of the target storage node being greater than the I / O throughput capacity score threshold. Network-level transmission optimization is for communication-intensive types and the transmission quality score between involved nodes being greater than the transmission quality score threshold.

8. The GPU and parallel I / O collaborative optimization method based on pattern heterogeneous computing according to claim 7, wherein: In S4, the priority of the three-level adaptive transmission mechanism is register-level transmission optimization level > memory-level transmission optimization > network-level transmission optimization. When the target computing task meets the trigger conditions of multiple levels simultaneously, the optimization strategy with the highest priority will be executed first.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster multi-mode heterogeneous value task autonomous collaborative allocation method and system

    CN114675674A

  • Dynamic computing power scheduling method for distributed heterogeneous nodes

    CN120066808A

  • Efficient parallelization and deployment method of multi-objective service function chain based on CPU + DPU platform

    US11936758B1

  • Workload measures based on access locality

    US20230325257A1

Cited By

  • Generative AI heterogeneous computing resource dynamic scheduling method and system of PC terminal

    CN120803747A

  • GPU (Graphics Processing Unit) program optimization method for parallel environment

    CN121029422A

  • Dynamic resource demand characterization method and system for task life cycle

    CN121233304A

  • A method and system for dynamic resource requirement characterization of a task life cycle

    CN121233304B

  • I / O intensive task hardware acceleration optimization method fusing software and hardware collaboration

    CN121957861A