Deep learning inference engine tensor optimization method and system for GPU acceleration
By performing asymmetric quantization and blocking processing of the original model, combined with parallel computing of a single kernel function, the GPU memory bandwidth bottleneck problem is solved, the performance and efficiency of the deep learning inference engine are improved, and memory usage and computing latency are reduced.
Patent Information
- Application Number
- CN202510492105.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
GPU memory bandwidth and processing efficiency have become bottlenecks in deep learning inference, resulting in too long data transmission time and affecting data processing efficiency.
By asymmetric quantization of the weight tensor of the original model, the second weight tensor is generated, the model is updated, and the data tensor is loaded into shared memory in chunks, and the input sub-blocks are processed in parallel with a single kernel function, and dynamic output adjustment is performed in combination with GPU utilization and request queue length.
It improves the performance and efficiency of the deep learning inference engine, reduces memory footprint and computing latency, and optimizes memory access and parallel computing.
Smart Images

Figure CN120371566A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a tensor optimization method and system for a deep learning inference engine accelerated by GPU. Background Art
[0002] With the wide application of deep learning in many fields such as image recognition, natural language processing, and speech recognition, the demand for deep learning inference has shown an explosive growth. For example, in the field of intelligent security, it is necessary to analyze surveillance videos in real time to identify abnormal behaviors and targets; in autonomous driving, camera and sensor data need to be processed quickly to make accurate driving decisions. These application scenarios require the inference engine to be able to process a large amount of data in a short time and give accurate results.
[0003] In practical applications, the memory bandwidth and processing efficiency of the Graphics Processing Unit (GPU) are limited, and a large amount of data transmission is involved in deep learning inference, such as transferring input data from the host memory to the GPU memory and transferring the calculation results from the GPU memory back to the host memory. When the data volume is large, the memory bandwidth may become a bottleneck, resulting in too long data transmission time and affecting data processing efficiency. Summary of the Invention
[0004] This application provides a tensor optimization method and system for a deep learning inference engine accelerated by GPU, which can at least to some extent solve the problem of low GPU processing efficiency in the face of massive data.
[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0006] According to one aspect of this application, a tensor optimization method for a deep learning inference engine accelerated by GPU is provided, including: obtaining model information of the original model and data to be processed; the model information includes a first weight tensor representing model weights, and the data to be processed includes a data tensor; performing asymmetric quantization on the first weight tensor of the original model to generate a second weight tensor, updating the original model based on the second weight tensor to obtain an updated model; dividing the data tensor into input sub-blocks according to a preset dimension, and loading the input sub-blocks and the operation modules in the updated model into the shared memory; calling the input sub-blocks from the shared memory, and based on a single kernel function composed of the operation modules, processing different batches of input sub-blocks in parallel to generate a processing result; performing dynamic output adjustment based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed to output the processing result.
[0007] In this application, based on the foregoing solution, the method for asymmetrically quantizing the first weight tensor of the original model to generate a second weight tensor and updating the original model based on the second weight tensor to obtain an updated model includes: determining a quantization range based on the maximum value and the minimum value in the first weight tensor of the original model; asymmetrically quantizing the first weight tensor of the original model based on the quantization range to generate a second weight tensor; and updating the original model based on the second weight tensor to obtain an updated model.
[0008] In this application, based on the foregoing solution, the method for updating the original model based on the second weight tensor to obtain an updated model includes: replacing the first weight tensor in the original model with the second weight tensor and updating the format of the original model to generate an updated model.
[0009] In this application, based on the foregoing solution, the method for partitioning the data tensor into input sub-blocks according to a preset dimension and loading the input sub-blocks and the operation modules in the updated model into the shared memory includes: Partitioning the data tensor in the height dimension and the width dimension to generate input sub-blocks; dynamically adjusting the batches and channels of the input sub-blocks in the batch dimension and the channel dimension to generate adjusted input sub-blocks; determining the attribute parameters of each input sub-block in the data tensor based on the access frequency, data volume, and survival period of the input sub-blocks; and sequentially loading the input sub-blocks and the operation modules in the updated model into the shared memory based on the attribute parameters.
[0010] In this application, based on the foregoing solution, the method for calling the input sub-blocks from the shared memory and parallelly processing the input sub-blocks of different batches based on the single kernel function composed of the operation modules to generate a processing result includes: constructing a single kernel function based on the convolution kernel, normalization module, and activation function of the updated model; and calling the input sub-blocks from the shared memory and parallelly processing the input sub-blocks of different batches through the single kernel function to generate a processing result.
[0011] In this application, based on the foregoing solution, the method for dynamically adjusting the output based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed to output the processing result includes: obtaining the GPU utilization rate and the length of the request queue corresponding to the data to be processed; dynamically adjusting the output based on the length of the request queue and the GPU utilization rate to determine the current batch data volume of the GPU; and outputting the processing result based on the current batch data volume.
[0012] In this application, based on the foregoing solution, after obtaining the GPU utilization rate and the length of the request queue corresponding to the data to be processed, the following steps are further included: analyzing the length of the request queue and the GPU utilization rate within a preset time period to generate an analysis result; and displaying the analysis result on a control terminal in the form of a chart.
[0013] According to one aspect of the present application, there is provided a tensor optimization system for a deep learning inference engine accelerated by GPU, including: An acquisition unit for acquiring model information of an original model and data to be processed; the model information includes a first weight tensor representing model weights, and the data to be processed includes a data tensor; A quantization unit for asymmetrically quantizing the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model; A loading unit for dividing the data tensor into input sub-blocks according to a preset dimension, and loading the input sub-blocks and the operation modules in the updated model into a shared memory; A processing unit for calling the input sub-blocks from the shared memory, and parallelly processing different batches of input sub-blocks based on a single kernel function composed of the operation modules to generate a processing result; An output unit for dynamically adjusting the output based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed, so as to output the processing result.
[0014] According to one aspect of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for optimizing tensors of a deep learning inference engine accelerated by GPU as described in the above embodiments.
[0015] According to one aspect of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for optimizing tensors of a deep learning inference engine accelerated by GPU as described in the above embodiments.
[0016] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for optimizing tensors of a deep learning inference engine accelerated by GPU provided in the above various optional implementation manners.
[0017] In the technical solution of this application, model information of the original model and data to be processed are obtained; asymmetric quantization is performed on the first weight tensor of the original model to generate a second weight tensor, and the original model is updated based on the second weight tensor to obtain an updated model; the data tensor is divided into blocks according to a preset dimension to generate input sub-blocks, and the input sub-blocks and the operation modules in the updated model are loaded into the shared memory; the input sub-blocks are called from the shared memory, and different batches of input sub-blocks are processed in parallel based on a single kernel function composed of the operation modules to generate processing results; dynamic output adjustment is performed based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed, so as to output the processing results. By performing asymmetric quantization on the model information, dividing and processing the data tensor to optimize memory access, and at the same time implementing parallel processing through a single kernel function, the performance and efficiency of the deep learning inference engine are improved, the memory occupancy and calculation latency are reduced, which provides strong support for GPU-accelerated deep learning applications.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings
[0019] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0020] Figure 1 Schematically shows a flowchart of a tensor optimization method for a GPU-accelerated deep learning inference engine in an embodiment of this application.
[0021] Figure 2 Schematically shows a flowchart of loading input sub-blocks and convolution kernels in an updated model into the shared memory in an embodiment of this application.
[0022] Figure 3 Schematically shows a schematic diagram of a tensor optimization system for a GPU-accelerated deep learning inference engine in an embodiment of this application.
[0023] Figure 4 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of this application. Detailed Description of the Embodiments
[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0025] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.
[0026] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0028] The implementation details of the technical solutions of this application are elaborated in detail below: Figure 1 A flowchart of a tensor optimization method for a GPU-accelerated deep learning inference engine according to an embodiment of this application is shown. Referring to Figure 1 as shown, the tensor optimization method for the GPU-accelerated deep learning inference engine includes at least steps S110 to S150, which are introduced in detail as follows: In step S110, model information of the original model and data to be processed are obtained; the model information includes a first weight tensor representing model weights, and the data to be processed includes a data tensor.
[0029] In one embodiment of the present application, according to the storage path or identifier of the original model file, the model file is located and read from the storage device using the file system interface, and the file content is parsed according to the format specification of the model file by the data parsing module. For the first weight tensor representing the model weight, a specific area for storing weight data is identified, and according to the dimension, data type and other information of the tensor, the binary data is converted into a numerical form that can be processed by the computer.
[0030] Optionally, sufficient memory space is allocated for the first weight tensor, and the parsed first weight tensor is filled into the corresponding memory location. At the same time, other relevant information of the model is extracted, such as the number of layers of the model, the type of each layer (convolutional layer, fully connected layer, etc.), and the connection relationship between the layers. This information will be used for subsequent reasoning calculations and data flow control.
[0031] In one embodiment of the present application, a data request or data transmission signal is received from an external data source, and a corresponding interface or driver is called to establish a connection with the data source according to the type and communication protocol of the data source. Then, the data to be processed is read from the data source according to a predetermined data format and encoding method.
[0032] Transfer the data to be processed from the data source to a temporary buffer in memory. During the transmission process, the data to be processed can be initially verified and processed, such as checking the integrity of the data and performing format conversion. If the data to be processed is stored or transmitted in a compressed format, the data can also be decompressed. After that, the data to be processed in the form of data tensors is further organized and sorted, and stored in a memory layout suitable for calculation, so as to perform efficient calculation operations with the model weight tensor later.
[0033] In step S120, asymmetric quantization is performed on the first weight tensor of the original model to generate a second weight tensor, and the original model is updated based on the second weight tensor to obtain an updated model.
[0034] In this embodiment, the first weight tensor in the original model is asymmetric quantized, and each element in the first weight tensor is quantized according to a preset quantization parameter, and converted from a floating point number to a lower precision integer representation, thereby generating a quantized second weight tensor. Subsequently, the first weight tensor in the original model is replaced with the second weight tensor to update the structure and data of the model. This process is intended to reduce the storage space and computing requirements of the model while maintaining the accuracy and performance of the model as much as possible, thereby obtaining an updated model that can run more efficiently during the inference phase.
[0035] In one embodiment of the present application, asymmetric quantization is performed on the first weight tensor of the original model to generate a second weight tensor, and the original model is updated based on the second weight tensor to obtain an updated model, including: Determine the quantization range based on the maximum value and the minimum value in the first weight tensor of the original model; Perform asymmetric quantization on the first weight tensor of the original model based on the quantization range to generate a second weight tensor; Update the original model based on the second weight tensor to obtain an updated model.
[0036] In one embodiment of the present application, in the step of determining the quantization range, all elements in the first weight tensor of the original model are traversed to find the maximum value and the minimum value among them. After determining the maximum value and the minimum value, the quantization range is determined according to them, and this quantization range will be used as the benchmark for subsequent quantization operations to ensure the accuracy and effectiveness of the quantization process.
[0037] After that, each element in the first weight tensor is quantized according to the quantization range. Based on the scaling factor and offset of the extreme values, the original floating-point weights are mapped to a finite set of discrete values, so as to convert each weight value in the first weight tensor into its corresponding quantized value. These operations are performed one by one for each element in the first weight tensor, and the elements in the quantized second weight tensor are correspondingly generated.
[0038] Exemplarily, based on the quantization range composed of the maximum value and the minimum value in the first weight tensor, the element values in the second weight tensor are determined through the following process as:
[0039] wherein, represents the rounding operation; W represents the element value in the first weight tensor, represents the mean value of the element values in the first weight tensor, respectively represent the maximum value and the minimum value of the elements in the first weight tensor, b represents the number of quantization bits.
[0040] Through the above process, by performing asymmetric quantization on the first weight tensor of the original model, the storage space required for the weight tensor can be significantly reduced, while maintaining or approaching the accuracy of the original model. This is because asymmetric quantization can more flexibly adapt to the weight distribution, and can compress data more effectively than symmetric quantization, retain the weight distribution characteristics, reduce the quantization error, and retain the model accuracy. After updating the model, using the quantized weight tensor for inference calculation can speed up the calculation speed and reduce the memory occupancy.
[0041] In an embodiment of the present application, the process of updating the original model based on the second weight tensor in step S120 to obtain an updated model includes: replacing the first weight tensor in the original model with the second weight tensor and updating the format of the original model to generate an updated model.
[0042] In this embodiment, after generating the second weight tensor, the quantized second weight tensor is used to replace the first weight tensor in the original model. At the same time, the original model based on deep learning is converted into a GPU-specific format model through the Open Neural Network Exchange (ONNX) format to achieve cross-frame model conversion, which is used as the updated model for subsequent inference or training tasks using the quantized second weight tensor.
[0043] The above process supports dynamic shape adjustment and mixed-precision calculation through the optimized format, so as to reduce storage requirements and computational overhead while maintaining the model performance, and improve the running efficiency of the model. After updating the model, using the quantized weight tensor for inference calculation can speed up the calculation speed and reduce memory occupancy.
[0044] In step S130, the data tensor is divided into input sub-blocks according to a preset dimension, and the input sub-blocks and the operation modules in the updated model are loaded into the shared memory.
[0045] In this embodiment, the data tensor is divided into blocks according to a preset dimension to generate multiple input sub-blocks of appropriate sizes for subsequent parallel processing. Then, these input sub-blocks and the operation modules in the updated model are loaded into the shared memory together. The shared memory, as a cache area, can significantly improve the data access speed, enabling the convolution operation to obtain the required input data and weight parameters in a shorter time. This process aims to optimize memory usage and data access efficiency and provide strong support for subsequent parallel computing tasks.
[0046] It should be noted that the operation modules in the updated model in this embodiment include convolution kernels, normalization modules, and activation functions.
[0047] As Figure 2 shown, in an embodiment of the present application, dividing the data tensor into input sub-blocks according to a preset dimension and loading the input sub-blocks and the operation modules in the updated model into the shared memory includes: S210, dividing the data tensor in the height dimension and the width dimension to generate input sub-blocks; S220, dynamically adjusting the batches and channels of the input sub-blocks in the batch dimension and the channel dimension to generate the adjusted input sub-blocks; S230. Determine the attribute parameters of each input sub - block in the data tensor based on the access frequency, data volume, and survival period of the input sub - blocks. S240. Based on the attribute parameters, load the input sub - blocks and the operation modules in the update model into the shared memory in sequence.
[0048] In this embodiment, during the process of dividing the data tensor in the height dimension and width dimension to generate input sub - blocks, first read the storage information of the data tensor in the memory to clarify the specific values of its height and width. Then, according to the preset chunking strategy, cut it in the height dimension at a certain step size according to a fixed size or a size dynamically determined according to the hardware characteristics.
[0049] Specifically, calculate the continuous data address corresponding to the data tensor in the memory. Through pointer movement and data copying operations, divide the data tensor into multiple parts in the height direction. Similarly, in the width dimension, perform similar operations, calculate the data address along the width direction and divide the data. Each divided part constitutes an input sub - block, and these input sub - blocks will be re - organized and stored in the memory for subsequent processing.
[0050] Optionally, by precisely managing memory allocation and release, the efficiency and correctness of the chunking operation can be ensured, and problems such as memory leaks and data conflicts can be avoided.
[0051] After that, dynamically adjust the batches and channels of the input sub - blocks in the batch dimension and channel dimension to generate the adjusted input sub - blocks. First, obtain the current batch and channel information of the input sub - blocks by reading relevant metadata or variables. Then, according to the dynamic adjustment strategy, such as factors like the real - time requirements of the model and the availability of hardware resources, determine the new batch size and number of channels through logical judgment. During the adjustment process, the rearrangement and copying of the input sub - blocks may merge or split the data of multiple sub - blocks and re - organize their storage order in the memory. During the adjustment process in the channel dimension, operations such as screening, expanding, or compressing the input sub - blocks of each channel can be performed.
[0052] Optionally, by efficiently managing the memory, using the cache mechanism to improve the data access speed, and at the same time ensuring the consistency and integrity of the data of the adjusted input sub - blocks to meet the requirements of subsequent calculations.
[0053] After that, based on the access frequency, data volume, and survival period of the input sub-blocks, determining the attribute parameters of the input sub-blocks corresponding to the data tensor, and sequentially loading the input sub-blocks and the operation modules in the updated model into the shared memory according to the attribute parameters is the key to optimizing memory usage and computing efficiency. When performing this step, first, collect the access frequency information of the input sub-blocks by counting the memory access instructions; calculate the data volume of the input sub-blocks according to the data type and storage layout, and determine the survival period of the input sub-blocks by tracking the creation, use, and destruction processes of the data. Calculate the attribute parameters of each input sub-block based on this comprehensive information, and this parameter reflects the importance and urgency of the input sub-blocks in subsequent calculations.
[0054] Specifically, in an embodiment of the present application, based on the access frequency, data volume, and survival period of the data tensor, the attribute parameters of each input sub-block in the data tensor are determined as follows:
[0055] Among them, f represents the access frequency, represents the exponential weight of the access frequency of the input sub-block, s represents the data volume of the input sub-block, T represents the survival period of the input sub-block, represents the survival period normalization coefficient, represents the natural constant e is the exponential function with base represents the idle time when the input sub-block has not been accessed, represents the decay coefficient.
[0056] The above process determines the data loading order and memory allocation strategy based on the attribute parameters to maximize the utilization rate of the shared memory, reduce the memory access latency, and ensure the timely availability of data during the calculation process.
[0057] During the loading process, sort the input sub-blocks according to the attribute parameters, preferentially load the input sub-blocks with high attribute parameters into the shared memory and retain them for a long time; for the input sub-blocks with lower attribute parameters, release them preferentially.
[0058] In practical applications, the shared memory has the characteristic of high-speed access. By reasonably managing the space allocation of the shared memory, ensuring that there are no conflicts between different data, and using the locality principle to reduce global memory access, the shared memory resources can be efficiently utilized to improve the efficiency and performance of subsequent convolution operations.
[0059] In the above process, by dividing the data tensor into blocks according to a preset dimension and loading it into the shared memory, the memory access pattern of the GPU can be optimized. The shared memory has the characteristics of low latency and high bandwidth, and is suitable for storing frequently accessed data to achieve high-speed access. By dynamically adjusting the batches and channels of the input sub-blocks, the memory access efficiency can be further improved to ensure that the GPU can process data efficiently.
[0060] In step S140, the input sub-blocks are called from the shared memory, and based on the single kernel function composed of the operation modules, the input sub-blocks of different batches are processed in parallel to generate processing results.
[0061] In this embodiment, the pre-loaded input sub-blocks are called sequentially or in parallel from the shared memory. Subsequently, using the single kernel function constructed based on the operation modules, convolution operations can be performed on the input sub-blocks of different batches in parallel on multiple processing units. Utilizing the multi-core architecture and parallel computing capabilities effectively improves the data processing speed. Finally, each processing unit generates the processing results of its corresponding input sub-blocks, and these results can then be integrated for subsequent calculations or analyses.
[0062] In an embodiment of the present application, calling the input sub-blocks from the shared memory and processing the input sub-blocks of different batches in parallel based on the single kernel function composed of the operation modules to generate processing results includes: Constructing a single kernel function based on the convolution kernel, normalization module, and activation function of the update model; Calling the input sub-blocks from the shared memory and processing the input sub-blocks of different batches in parallel through the single kernel function to generate processing results.
[0063] In this embodiment, the parameter configuration of the convolution kernel, the related settings of the normalization module, and the type and parameters of the activation function in the update model are extracted, and these data are integrated into a unified single kernel function through the compilation and optimization of the underlying code. So that the convolution operation can efficiently utilize hardware resources, and the data normalization processing involved in the normalization module will be embedded inside the kernel function to reduce data transmission overhead. At the same time, as the key to non-linear transformation, the activation function is designed in a form that can seamlessly connect with the convolution and normalization operations to maximize the calculation efficiency and reduce the execution latency.
[0064] Specifically, constructing a single kernel function based on the convolution kernel, normalization module, and activation function of the update model is:
[0065] Wherein, represents the activation function, and represents a normalization parameter, represents a convolution kernel, represents convolution kernel parameters, represents the parameters of an activation function, is a constant to prevent division by zero.
[0066] During the stage of performing parallel processing, shared memory serves as a data storage area for fast access, which can significantly improve the data reading speed and is a key resource in parallel computing. Call the pre-divided input sub-blocks from the shared memory, allocate these input sub-blocks to multiple processing units, and each processing unit runs the single kernel function constructed previously.
[0067] Optionally, in this embodiment, a three-level parallel processing architecture is constructed, including thread level, block level, and stream level. Specifically, at the thread level, each thread in a Compute Unified Device Architecture (CUDA) processes a pixel point of an output feature map; at the block level, each thread block processes multiple sub-blocks of the same channel, such as a 16×16 pixel area; at the stream level, multiple CUDA streams process input sub-blocks of different batches in parallel.
[0068] This parallel processing method allows input sub-blocks of different batches to be processed simultaneously, greatly improving the data processing throughput. Each instance of the operation module works independently but shares the same logical structure, ensuring the consistency and accuracy of the processing results. In this way, the system can make full use of the parallel computing power of modern multi-core processors, quickly generate and summarize the processing results of each batch of input sub-blocks, and provide efficient data support for subsequent model training or inference tasks.
[0069] In the above process, a single kernel function is constructed based on the convolution kernel, normalization module, and activation function in the operation module to achieve parallel processing of input sub-blocks of different batches. This parallel processing method can make full use of the multi-core parallel computing power of the GPU and significantly improve the inference speed. At the same time, the design of the single kernel function simplifies the complexity of the GPU program and reduces the development and maintenance costs.
[0070] In step S150, based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed, dynamic output adjustment is performed to output the processing result.
[0071] In this embodiment, dynamic output adjustment is performed in real time according to the length of the request queue corresponding to the data to be processed and the current GPU utilization rate. By monitoring these two key metrics in real time and based on their real-time status, it is flexibly determined how to process and output the results of the convolution operation. For example, when the request queue is long or the GPU utilization rate is high, the system may adopt a batch processing or caching strategy to optimize resource utilization and ensure the timely output of the processing results; otherwise, the processing results may be directly output to improve the overall processing efficiency.
[0072] In an embodiment of the present application, based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed, dynamic output adjustment is performed to output the processing results, including: Obtain the GPU utilization rate and the length of the request queue corresponding to the data to be processed; Based on the length of the request queue and the GPU utilization rate, perform dynamic output adjustment to determine the current batch data volume of the GPU; Based on the current batch data volume, output the processing results.
[0073] In this embodiment, a circular queue or a linked list is used to monitor the request information, and the number of requests in the queue is periodically checked to determine the length of the request queue corresponding to the data to be processed. At the same time, the GPU utilization rate is obtained by using the monitoring tool or interface of the GPU to determine the current workload of the GPU, including the number of tasks being executed, the proportion of computing resources occupied, etc., and then the GPU utilization rate is determined.
[0074] In an embodiment of the present application, if the request queue length is long and the GPU utilization rate is low, it means that the GPU has sufficient computing resources to process more data. At this time, the current batch data volume can be appropriately increased to improve the processing efficiency and make full use of the computing power of the GPU. On the contrary, if the request queue length is short and the GPU utilization rate is high, it means that the GPU is already running close to full load, and the computer will reduce the current batch data volume to avoid performance degradation caused by GPU overload.
[0075] The adjusted batch data volume can not only meet the processing requirements but also will not cause excessive pressure on the system. After determining the current batch data volume, based on the current batch data volume, the processing results corresponding to a reasonable data volume are output. This dynamically adjusts the input data volume processed each time according to the real-time task load, flexibly allocates computing resources according to real-time requirements, and the elastic adjustment strategy ensures that the hardware resources are always in a highly efficient operating state, taking into account both the response speed and the processing efficiency, and finally achieving the goal of meeting both low latency and high throughput in a high-concurrency scenario.
[0076] In the above process, dynamic output adjustment is performed according to the GPU utilization rate and the length of the request queue corresponding to the data to be processed, which can ensure the reasonable allocation and utilization of GPU resources. When the request queue is long or the GPU utilization rate is high, the data volume of the current batch can be appropriately increased to improve the throughput; otherwise, the data volume of the current batch is reduced to avoid resource waste. This dynamic adjustment mechanism can ensure that the GPU operates efficiently under different load conditions.
[0077] In an embodiment of the present application, after obtaining the GPU utilization rate and the length of the request queue corresponding to the data to be processed, it further includes: Analyze the length of the request queue and the GPU utilization rate within a preset time period to generate an analysis result; Display the analysis result on the control terminal in the form of a chart.
[0078] In this embodiment, data collection of the length of the request queue and the GPU utilization rate is performed according to a preset time period. For the length of the request queue, by interacting with the request management module, the number of requests to be processed in the current queue is obtained and recorded. For the GPU utilization rate, the interface provided by the GPU monitoring tool is called to obtain the utilization rate data of the GPU at different time points, and these data are usually expressed in the form of percentages.
[0079] After the data is collected, the data is stored in a temporary buffer or a database for subsequent analysis and processing. In the analysis stage, a data analysis model is used to analyze the change trends of the length of the request queue and the GPU utilization rate. For example, calculate the average value, maximum value, minimum value, and fluctuation range of the length of the request queue, and analyze its change law within a preset time period. For the GPU utilization rate, count its peak value, trough value, and average utilization rate, and judge whether the load condition of the GPU is stable. At the same time, the correlation between the length of the request queue and the GPU utilization rate can also be analyzed, for example, whether there is a situation where the increase in the length of the request queue leads to an increase in the GPU utilization rate. Through these analyses, an analysis result including various statistical indicators, trend charts, and correlation analysis results is generated.
[0080] Display the analysis result on the control terminal in the form of a chart. Exemplarily, for the change trends of the length of the request queue and the GPU utilization rate over time, a line chart is selected to intuitively display the change of the data; for the comparison of the length of the request queue and the GPU utilization rate in different time periods, a bar chart is selected to present the differences between the data. Users can intuitively view the analysis result on the control terminal, understand the situation of the length of the request queue and the GPU utilization rate, and thus make corresponding decisions and adjustments.
[0081] The above process can help developers understand the performance bottlenecks and optimization space of the system by regularly analyzing the request queue length and GPU utilization within a preset duration and generating analysis results. Displaying the analysis results in the form of charts on the control terminal can intuitively reflect the changing trend of system performance and provide strong support for further optimization.
[0082] In the technical solution of this application, the model information of the original model and the data to be processed are obtained; the first weight tensor of the original model is asymmetrically quantized to generate a second weight tensor, and the original model is updated based on the second weight tensor to obtain an updated model; the data tensor is partitioned into input sub-blocks according to a preset dimension, and the input sub-blocks and the operation modules in the updated model are loaded into the shared memory; the input sub-blocks are called from the shared memory, and different batches of input sub-blocks are processed in parallel based on a single kernel function composed of the operation modules to generate processing results; dynamic output adjustment is performed based on the GPU utilization and the request queue length corresponding to the data to be processed to output the processing results. By asymmetrically quantizing the model information, partitioning the data tensor to optimize memory access, and at the same time implementing parallel processing through a single kernel function, the performance and efficiency of the deep learning inference engine are improved, the memory occupancy and calculation latency are reduced, and strong support is provided for GPU-accelerated deep learning applications.
[0083] The following introduces the device embodiments of this application, which can be used to execute the tensor optimization method for the GPU-accelerated deep learning inference engine in the above embodiments of this application. It can be understood that the device can be a computer program (including program code) running in a computer device, for example, the device is an application software; the device can be used to execute the corresponding steps in the method provided in the embodiments of this application. For details not disclosed in the device embodiments of this application, please refer to the embodiments of the tensor optimization method for the GPU-accelerated deep learning inference engine in the above of this application.
[0084] Figure 3 The block diagram of a tensor optimization system for a GPU-accelerated deep learning inference engine according to an embodiment of this application is shown.
[0085] Refer to Figure 3 As shown, a tensor optimization system for a GPU-accelerated deep learning inference engine according to an embodiment of this application includes: An acquisition unit 310, configured to acquire the model information of the original model and the data to be processed; the model information includes a first weight tensor representing the model weights, and the data to be processed includes a data tensor; A quantization unit 320, configured to perform asymmetric quantization on a first weight tensor of the original model to generate a second weight tensor, and update the original model based on the second weight tensor to obtain an updated model; A loading unit 330, configured to partition the data tensor into input sub-blocks according to a preset dimension, and load the input sub-blocks and an operation module in the updated model into a shared memory; A processing unit 340, configured to call the input sub-blocks from the shared memory, and based on a single kernel function composed of the operation modules, process the input sub-blocks of different batches in parallel to generate a processing result; An output unit 350, configured to perform dynamic output adjustment based on the GPU utilization rate and the length of a request queue corresponding to data to be processed, so as to output the processing result.
[0086] In this application, based on the foregoing solution, the performing asymmetric quantization on the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model includes: determining a quantization range based on a maximum value and a minimum value in the first weight tensor of the original model; performing asymmetric quantization on the first weight tensor of the original model based on the quantization range to generate a second weight tensor; and updating the original model based on the second weight tensor to obtain an updated model.
[0087] In this application, based on the foregoing solution, the updating the original model based on the second weight tensor to obtain an updated model includes: replacing the first weight tensor in the original model with the second weight tensor, and updating the format of the original model to generate an updated model.
[0088] In this application, based on the foregoing solution, the partitioning the data tensor into input sub-blocks according to a preset dimension, and loading the input sub-blocks and an operation module in the updated model into a shared memory includes: Partitioning the data tensor in the height dimension and the width dimension to generate input sub-blocks; dynamically adjusting the batches and channels of the input sub-blocks in the batch dimension and the channel dimension to generate adjusted input sub-blocks; determining attribute parameters of each input sub-block in the data tensor based on the access frequency, data volume, and survival period of the input sub-blocks; and sequentially loading the input sub-blocks and the operation module in the updated model into the shared memory based on the attribute parameters.
[0089] In this application, based on the foregoing solution, calling the input sub-block from the shared memory and parallelly processing input sub-blocks of different batches based on a single kernel function composed of the operation modules to generate a processing result includes: constructing a single kernel function based on the convolution kernel, normalization module, and activation function of the updated model; calling the input sub-block from the shared memory, and parallelly processing input sub-blocks of different batches through the single kernel function to generate a processing result.
[0090] In this application, based on the foregoing solution, performing dynamic output adjustment based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed to output the processing result includes: obtaining the GPU utilization rate and the length of the request queue corresponding to the data to be processed; performing dynamic output adjustment based on the length of the request queue and the GPU utilization rate to determine the current batch data volume of the GPU; and outputting the processing result based on the current batch data volume.
[0091] In this application, based on the foregoing solution, after obtaining the GPU utilization rate and the length of the request queue corresponding to the data to be processed, it further includes: analyzing the length of the request queue and the GPU utilization rate within a preset time period to generate an analysis result; and displaying the analysis result in a graphical manner on the control terminal.
[0092] In the technical solution of this application, obtaining the model information of the original model and the data to be processed; performing asymmetric quantization on the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model; dividing the data tensor into blocks according to a preset dimension to generate input sub-blocks, and loading the input sub-blocks and the operation modules in the updated model into the shared memory; calling the input sub-block from the shared memory, and parallelly processing input sub-blocks of different batches based on a single kernel function composed of the operation modules to generate a processing result; performing dynamic output adjustment based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed to output the processing result. By performing asymmetric quantization on the model information, dividing the data tensor into blocks to optimize memory access, and simultaneously implementing parallel processing through a single kernel function, the performance and efficiency of the deep learning inference engine are improved, the memory occupancy and calculation latency are reduced, which provides strong support for GPU-accelerated deep learning applications.
[0093] Figure 4 The structure diagram of the computer system of the electronic device suitable for implementing the embodiments of this application is shown.
[0094] It should be noted that the computer system of the electronic device in this embodiment is only an example, and should not bring any limitations to the functions and usage scopes of the embodiments of this application.
[0095] In this embodiment, the computer system includes a central processing unit 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory 402 or the program loaded from the storage section 408 into the random access memory 403. For example, it can execute the tensor optimization method of the deep learning inference engine for GPU acceleration described in the above embodiment. In the random access memory 403, various programs and data required for system operation are also stored. The central processing unit 401, the read-only memory 402, and the random access memory 403 are connected to each other via a bus 404. The input / output interface 405 is also connected to the bus 404.
[0096] The following components are connected to the input / output interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.
[0097] Specifically, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit 401, various functions defined in the system of the present application are executed.
[0098] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on a computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the units themselves.
[0101] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various alternative implementation manners.
[0102] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the tensor optimization method for the GPU-accelerated deep learning inference engine described in the above embodiments.
[0103] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0104] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of the present application.
[0105] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0106] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A tensor optimization method for a GPU-accelerated deep learning inference engine, characterized in that, Including: Obtaining model information of the original model and data to be processed; the model information includes a first weight tensor representing model weights, and the data to be processed includes a data tensor; Asymmetrically quantizing the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model; Partitioning the data tensor into input sub-blocks according to a preset dimension, and loading the input sub-blocks and operation modules in the updated model into shared memory; Calling the input sub-blocks from the shared memory, and based on a single kernel function composed of the operation modules, processing different batches of input sub-blocks in parallel to generate processing results; Performing dynamic output adjustment based on GPU utilization and the length of the request queue corresponding to the data to be processed, so as to output the processing results.
2. The tensor optimization method for a GPU-accelerated deep learning inference engine according to claim 1, wherein The asymmetrically quantizing the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model includes: Determining a quantization range based on the maximum value and minimum value in the first weight tensor of the original model; Asymmetrically quantizing the first weight tensor of the original model based on the quantization range to generate a second weight tensor; Updating the original model based on the second weight tensor to obtain an updated model.
3. The tensor optimization method for the GPU-accelerated deep learning inference engine according to claim 2, wherein The updating the original model based on the second weight tensor to obtain an updated model includes: Replacing the first weight tensor in the original model with the second weight tensor, and updating the format of the original model to generate an updated model.
4. The tensor optimization method for a GPU-accelerated deep learning inference engine according to claim 1, characterized in that The partitioning the data tensor into input sub-blocks according to a preset dimension, and loading the input sub-blocks and operation modules in the updated model into shared memory includes: Partitioning the data tensor in the height dimension and width dimension to generate input sub-blocks; Dynamically adjusting the batches and channels of the input sub-blocks in the batch dimension and channel dimension to generate adjusted input sub-blocks; Determining attribute parameters of each input sub-block in the data tensor based on the access frequency, data volume and survival period of the input sub-blocks; Loading the input sub-blocks and operation modules in the updated model into the shared memory in sequence based on the attribute parameters, and the operation modules include convolution kernels, normalization modules and activation functions.
5. The tensor optimization method for a GPU-accelerated deep learning inference engine according to claim 4, characterized in that The calling the input sub-blocks from the shared memory, and based on a single kernel function composed of the operation modules, processing different batches of input sub-blocks in parallel to generate processing results includes: Constructing a single kernel function based on the convolution kernels, normalization modules and activation functions in the updated model; Calling the input sub-blocks from the shared memory, and processing different batches of input sub-blocks in parallel through the single kernel function to generate processing results.
6. The tensor optimization method for a GPU-accelerated deep learning inference engine according to claim 1, wherein The performing dynamic output adjustment based on GPU utilization and the length of the request queue corresponding to the data to be processed, so as to output the processing results includes: Obtaining the GPU utilization and the length of the request queue corresponding to the data to be processed; Performing dynamic output adjustment based on the request queue length and the GPU utilization to determine the current batch data volume of the GPU; Output the processing result based on the current batch data volume.
7. The tensor optimization method for a GPU-accelerated deep learning inference engine according to claim 6, wherein After obtaining the GPU utilization rate and the length of the request queue corresponding to the data to be processed, it further includes: Analyze the length of the request queue and the GPU utilization rate within a preset duration to generate an analysis result; Display the analysis result on the control terminal in the form of a chart.
8. A tensor optimization system for a GPU-accelerated deep learning inference engine, characterized in that, It includes: An acquisition unit for acquiring the model information of the original model and the data to be processed; the model information includes a first weight tensor representing model weights, and the data to be processed includes a data tensor; A quantization unit for performing asymmetric quantization on the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model; A loading unit for dividing the data tensor into input sub-blocks according to a preset dimension and loading the input sub-blocks and the operation modules in the updated model into the shared memory; A processing unit for calling the input sub-blocks from the shared memory and parallelly processing the input sub-blocks of different batches based on a single kernel function composed of the operation modules to generate a processing result; An output unit for performing dynamic output adjustment based on the GPU utilization rate and the length of the request queue corresponding to the data to be processed to output the processing result.
9. The tensor optimization system for GPU-accelerated deep learning inference engine according to claim 8, characterized in that, The performing asymmetric quantization on the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model includes: Determine the quantization range based on the maximum value and the minimum value in the first weight tensor of the original model; Perform asymmetric quantization on the first weight tensor of the original model based on the quantization range to generate a second weight tensor; Update the original model based on the second weight tensor to obtain an updated model.
10. The tensor optimization system for GPU-accelerated deep learning inference engine according to claim 9, wherein, The updating the original model based on the second weight tensor to obtain an updated model includes: Replace the first weight tensor in the original model with the second weight tensor and update the format of the original model to generate an updated model.
Citation Information
Patent Citations
Processing unit, acceleration unit, related device and method
CN114443145A
Deep learning compiler optimization method special for CNN accelerator
CN114995822A
Method for accelerating deep learning model
CN115688905A
Efficient image recognition system based on embedded edge device
CN119540734A
Tensor-based optimization method for memory management of a deep-learning GPU and system thereof
US20210142178A1
Cited By
Data processing method and device, electronic equipment and storage medium
CN120610829A
Artificial intelligence model dynamic reasoning method and system based on consumption-level heterogeneous chip
CN121008922A