Utilizing GPU blending for neural network processing
Patent Information
- Application Number
- PCT/GR2025/000006
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure GR2025000006_01102026_PF_FP_ABST
Abstract
Description
UTILIZING GPU BLENDING FOR NEURAL NETWORK PROCESSINGTECHNICAL FIELD
[0001] The present disclosed subject matter relates to neural network processing. More particularly, the present disclosed subject matter relates to utilizing a graphics processing unit 5 (GPU) blending techniques for quantization and data handling in neural networks.BACKGROUND
[0002] Digital displays, critical to contemporary devices, utilize technologies ranging from Liquid-Crystal Displays (LCDs) to advanced Organic Light-Emitting Diode (OLED) displays, each supporting high-definition visuals that demand robust computational power. At the core of these 10 technologies is the pixel, the fundamental unit of digital imaging. Modern high-definition displays, comprising millions of pixels, are frequently updated to enable smooth video playback and responsive interfaces, necessitating sophisticated processing technologies for efficient data management and transmission.
[0003] GPU incorporates at least one blending engine that are components within graphics processing units, specifically engineered to manage the complex task of blending pixel data during the rendering process. These units are vital in combining pixel colors from various sources, using diverse blending modes to produce the final visual output on the display. Such modes are especially useful for creating effects like transparency and translucency, contributing to more realistic and visually appealing graphics in applications such as video games and digital c> media.
[0004] Despite their inherently two-dimensional nature, modern displays often simulate three- dimensional environments to enhance visual perception. This is achieved through advanced processing techniques like texture mapping, which not only adds depth to two-dimensional surfaces but also significantly increases the computational load. These operations, from texture 5 mapping to complex pixel transformations such as rotations and scaling, highlight the need for high-level computational efficiency and optimized memory usage to maintain system performance, particularly in portable devices like mobile phones and laptops.
[0005] Commercially available GPUs are widely used in advancing Deep Neural Networks (DNNs) and deep learning models due to their ability to manage computations necessary for deep learning. Originally designed for graphics processing, GPU excels in matrix and vector operations that are fundamental to training DNNs. Their adoption for DNNs significantly accelerated in the early 2010s when it became evident that GPUs could reduce neural network training times compared to traditional CPUs. This capability has spurred major advancements in fields like image recognition, autonomous driving, and natural language processing, making GPUs indispensable for modern artificial intelligence (AI) development.
[0006] Figures 1 A to 1C illustrate a process of quantization and reconstruction process in deep neural networks (DNNs).
[0007] This Quantization and reconstruction process is particularly beneficial for devices with limited processing capabilities or those lacking FP32 support, as it is more resource-efficient and faster to compute. Specifically, these figures detail the fundamental concept of quantization in deep neural networks (DNNs), demonstrating how FP32 weights are converted into 8-bit integers (INT8). The quantization process includes mapping floating-point numbers to a smaller integer range, adjusting scales and zero points, and minimizing quantization errors, ensuring the network maintains high performance despite reduced data precision.
[0008] Figure 1A provides an exemplary general overview of quantization 110 of converting 32-bit floating-point (FP32) representations into 8-bit integer (INT8) formats, highlighting a practical example where weights ranging from -4.7 to 5.64 are efficiently quantified.
[0009] In the example, FP32-matrix 111 includes FP32 values, which are typical weights used in neural network computations. These weights vary from small negative numbers, such as -0.9, to larger positive values up to 5.64, each conventionally requiring 32 bits of memory.
[0010] Quantization transformation 113 is a process in which each FP32 value undergoes conversion into an INT8 format. This includes scaling a range of FP32 values down to a compact range that can be represented within 8 bits. This scaling is necessary to adapt weights represented by FP32 values to fit within the limited scope of INT8 values, from 0 to 255, and to reduce computational load significantly.
[0011] INT8-matrix 111 consists of quantized values resulting from the quantization transformation 113. The values in INT8-matrix 111 are integers ranging between 0 and 255. The reduction from requiring 32 bits per value to just 8 bits substantially decreases memory demand, computational effort, power consumption, and enhances efficiency in real-time Al tasks, making 5 it ideal for deployment on resource-constrained devices.
[0012] Figure 1B provides an exemplary graphical representation 120 of the scale mapping process used during the quantization of neural network weights from 32 -bit floating-point FP32 formats to 8-bit integer INT8 formats.
[0013] The scale mapping begins with an integer-axis (Int-axis) 125, marked with points 126 ranging from 0 to 255. These points delineate the entire possible range for INT8 values, establishing the foundational scale for the quantization.
[0014] Similarly, a FP32-axis 121 is marked with points 127 ranging from the lowest value 122 through zero point 123 to the highest value 124. These points outline the entire possible range for FP32 values, setting the baseline scale for the quantization.*5
[0015] The mapping of points from the FP32-axis 121 to the Int-axis 125 is executed in a nonlinear manner. Points 127 ranging from the lowest value 122 through zero point 123 to the highest value 124 on the FP32-axis 121 are mapped to the points 126 ranging from 'O' to '255' on the Int-axis 125. This non-linear mapping demonstrates the scaling mechanics applied during quantization, ensuring that each FP32 value is appropriately transformed to fit within the limited 0 scope of INT8 values despite the broad range of input values.
[0016] It should be noted that each FP32 weight is carefully mapped to an appropriate INT8 value based on a predefined scale and a zero point. This zero point is significant as it represents the FP32 value that equates to zero in the transformed INT8 range, ensuring that this central benchmark is preserved across data transformations.5
[0017] Figure 1C presents a detailed illustration that encapsulates the quantization process 130, starting from the initial representation of original weights to their eventual reconstruction after quantization. This involves mapping neural network weights from 32-bit floating-point to 2-bit signed integer formats and subsequently reconstructing them back to 32-bit floating-point values.
[0018] Weights 131 is a matrix of 32-bit floating weights such as 2.09, 0.98, -1.48, and others, demonstrating an exemplary initial data state prior to quantization.[00191 Transformation of the weight's matrix 131 into the quantized weights 2-bit signed integers matrix 133 using transformer 132, which converts the original weights 131 into a quantized format, displaying exemplary 2-bit signed integer values such as 1, -2, 0, and -1. This step marks the transition from high-precision floating-point to low-precision integer representation, displaying the initial phase of the quantization process,
[0020] Following the weights quantization, reconstructor 134 reconstructs the weights values into the reconstructed weights matrix 135, which has 32-bit floating values. Reconstructor 134 includes a process of recalculating the approximate original weights from the quantized integers, resulting in exemplary values such as 2.14, 1.07, and 0. This step approximates the original data post-quantization and assesses the accuracy of the quantization process.
[0021] It should be appreciated that the quantization process 130, which includes an analysis of binary to decimal conversions and a detailed examination of quantization errors, is essential for maintaining data integrity during quantization. By mapping binary quantified data back to their original decimal forms, this analysis ensures that digital systems can accurately interpret and process the data, thereby preserving its functional integrity. Additionally, quantifying the discrepancies between the original floating-point data and the quantized integers highlights precision loss, allowing for adjustments to minimize data loss. This rigorous approach guarantees the reliability and performance of neural networks after quantization, despite the inherent data compression and simplification involved in the process.
[0022] It should be noted that the quantization process 130 summarizes the pathway from the initial representation of neural network weights, through their quantization, and final reconstruction, underlining the fundamental stages of data transformation essential for neural network optimization and efficiency.
[0023] it would therefore be advantageous to provide a solution that would overcome the challenges noted above, specifically reducing the computational demands and memory usage in Al processing within GPUs, enhancing data fidelity and precision in the quantization of neuralnetwork data, and optimizing hardware to improve performance and energy efficiency for advanced blending and quantization processes in AI applications.SUMMARY(0024] A summary of several example embodiments of the disclosure follows. This summary is provided for the convenience of the reader to provide a basic understanding of such embodiments and does not wholly define the breadth of the disclosure. This summary is not an extensive overview of all contemplated embodiments and is intended to either identify key or critical elements of all embodiments nor to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more embodiments in a simplified form as a prelude to the more detailed description that is presented later. For convenience, the term "some embodiments" or "certain embodiments” may be used herein to refer to a single embodiment or multiple embodiments of the disclosure.
[0025] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0026] In one general aspect, a method may include initializing data for neural network processing having: configuring at least one blending engine to set up tensor values necessary for quantization; aligning data preparation with efficient neural network operational standards; and preprocessing input data to match the input requirements of specific artificial intelligence (Al) models. The method may also include invoking quantization within the at least one blending engine having: converting high-precision floating-point inputs into 8-bit integers using the GPU's at least one blending engine; employing complex mathematical operations for effective data conversion; and encoding techniques that optimize bit utilization for enhanced quantization. The method may furthermore include applying activation functions and dynamic re-quantization having: implementing nonlinear activation functions to enhance model response to data intricacies followed by re-quantization to adjust data precision for subsequent processingphases; and including dynamic adjustment of quantization scales. The method may in addition include. The method may moreover include executing optimized Al kernels having: running Al- specific computational tasks; and leveraging hardware optimizations within the blending engine for efficient kernel operations. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0027] Implementations may include one or more of the following features. The method where initializing data further includes configuring the at least one blending engine to pre- process input data to match the input requirements of specific AI models before quantization. The method may further include having: employing specific algorithms within the at least one blending engine for multiplication and addition operations that are optimized for data types typically processed in Al applications. The method where applying activation functions includes using a range of functions selected based on the a type of neural network layer being processed and dynamically adjusting the quantization scale to optimize computational efficiency. The method where executing optimized Al kernels includes using tensor cores specifically designed for accelerating deep learning tasks, with configurations being adaptable based on the complexity of the neural network model being executed. The method where executing optimized Al kernels further may include leveraging hardware optimizations that include adjusting control bits and utilizing multiplexers and shift registers within the at least one blending engine for efficient kernel operations. The method may include: managing efficient system communications by coordinating command and data flow across the GPU components, ensuring synchronized operations and optimal data transfer speeds between internal and external system interfaces, and using high-bandwidth interconnects like PCIe for robust inter-system communication. The method where the step of managing efficient system communications includes using a designated control unit to prioritize data packets based on real-time processing needs and adjust the communication protocols dynamically to prevent bottlenecks. Implementations of the described techniques may include hardware, a method or process, or a computer tangible medium.
[0028] In one general aspect, a non-transitory computer-readable medium may include one or more instructions that, when executed by one or more processors of a device, cause the device to: initialize data for neural network processing having: configure configuring at least one blending engine to set up tensor values necessary for quantization aligning data preparation 5 with efficient neural network operational standards; and preprocessing input data to match the input requirements of specific artificial intelligence (AI) models. The non-transitory computer- readable medium may also include invoke quantization within the at least one blending engine having: converting high-precision floating-point inputs into 8-bit integers using the GPU's at least one blending engine employing complex mathematical operations for effective data io conversion; and including encoding techniques that optimize bit utilization for enhanced quantization. The medium may furthermore include apply activation functions and dynamic re-quantization having: implementing nonlinear activation functions to enhance model response to data intricacies followed by re-quantization to adjust data precision for subsequent processing phases; and include including dynamic adjustment of quantization scales. The is medium may in addition include execute optimized Al kernels having: running Al-specific computational tasks; and leverage leveraging hardware optimizations within the blending engine for efficient kernel operations. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.io
[0029] In one general aspect, a system may include a processing circuitry. The system may also include a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:. The system may furthermore initialize data for neural network processing.. The system may in addition configure at least one blending engine to set up tensor values necessary for quantization. The system may moreover aligning data >5 preparation with efficient neural network operational standards. The system may also include preprocessing input data to match the input requirements of specific artificial intelligence (Al) models. The system may furthermore invoke quantization within the at least one blending engine having:. The system may in addition include converting high-precision floating-point inputs into 8-bit integers using the GPU's at least one blending engine. The system maymoreover include employing complex mathematical operations for effective data conversion. The system may also include encoding techniques that optimize bit utilization for enhanced quantization. The system may furthermore apply activation functions and dynamic requantization. The system may in addition include implementing nonlinear activation functions 5 to enhance model response to data intricacies followed by re-quantization to adjust data precision for subsequent processing phases. The system may moreover include dynamic adjustment of quantization scales. The system may also execute optimized Al kernels.. The system may furthermore include running Al-specific computational tasks. The system may in addition include leveraging hardware optimizations within the blending engine for efficient io kernel operations. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0030] Implementations may include one or more of the following features. The system where initializing data further includes configuring the at least one blending engine to pre-process 15 input data to match the input requirements of specific Al models before quantization. The system where applying activation functions includes using a range of functions selected based on a type of neural network layer being processed and dynamically adjusting the quantization scale to optimize computational efficiency. The system where executing optimized AI kernels includes using tensor cores specifically designed for accelerating deep learning tasks, with zo configurations being adaptable based on the complexity of the neural network model being executed. The system where the memory contains further instructions which when executed by the processing circuitry further configure the system to: manage efficient system communications by coordinating command and data flow across the GPU components, ensuring synchronized operations and optimal data transfer speeds between internal and external 2$ system interfaces, and using high-bandwidth interconnects like PCIe for robust inter-system communication.; and. Implementations of the described techniques may include hardware, a method or process, or a computer tangible medium.
[0031] Implementations may include one or more of the following features. The system where the step of managing efficient system communications includes using a designated control unitto prioritize data packets based on real-time processing needs and adjust, the communication protocols dynamically to prevent bottlenecks. Implementations of the described techniques may include hardware, a method or process, or a computer tangible medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The subject matter disclosed herein is particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other objects, features, and advantages of the disclosure will be apparent from the following detailed description taken in conjunction with the accompanying drawings.In the drawings:K>
[0033] Figures 1A to 1C illustrate a process of quantization process and reconstruction in deep neural networks (DNNs);
[0034] Figure 2 shows a block diagram of a computing system, in accordance with some disclosed embodiments; and
[0035] Figure 3 is a flowchart of a method for enhanced neural network processing, in accordance with some disclosed embodiments.DETAILED DESCRIPTION
[0036] The embodiments disclosed herein are only examples of the many possible advantageous uses and implementations of the innovative teachings presented herein. In general, statements made in the specification of the present application do not necessarily limit 0 any of the various claimed embodiments. Moreover, same statements may apply to some inventive features but not to others, in general, unless otherwise indicated, singular elements may be plural and vice versa with no loss of generality. In the drawings, numerals refer to like parts through several views.
[0037] One technical problem addressed by the disclosed subject matter is the reduction of 5 computational demands and memory usage in artificial intelligence (Al) processing within the graphics processing unit (GPU).
[0038] Another technical problem is enhancing the fidelity and precision of data during the quantization of neural network data.
[0039] Yet another problem is optimizing hardware performance and energy efficiency for advanced blending and quantization processes in Al applications.
[0040] The objective of the present disclosure revolves around the utilization of GPUs and the introduction of novel GPU blending techniques aimed at enhancing the functionality and efficiency of neural network processing in artificial intelligence (AI) applications.
[0041] One technical solution is the optimization of the quantization process. In some embodiments, a GPU blending unit may be used to perform quantization of tensor values, converting high-precision floating-point inputs into lower-precision discrete outputs such as 8- bit integers. In these embodiments, the optimization can be achieved by integrating quantization processes directly within the GPU’s at least one blending engine, leveraging the hardware's inherent capabilities to manage complex mathematical operations efficiently.
[0042] Another technical solution is applying advanced blending techniques. In some embodiments, the application of sophisticated blending techniques extends beyond traditional graphics applications to support AI processing, which includes managing transparency and other complex blending modes for rendering realistic visual effects. In these embodiments, the advanced blending techniques may utilize mathematical operations like multiplication, addition, or shifting directly within the blending engine to facilitate these advanced processes.
[0043] Yet another technical solution involves hardware optimization. In some embodiments, hardware modifications to the GPU, such as adding control bits to blending engines, adding a multiplexer and a shift register, or any combination thereof, are implemented. Consequently, these optimizations enhance the execution of neural network layers directly within the GPU hardware. In these embodiments, these hardware modifications enhance the GPU's performance and energy efficiency, making it more suitable for computationally intensive tasks without significant overhauls.
[0044] One technical effect of utilizing the disclosed subject matter is enhanced computational efficiency. Commercially available GPUs have been optimized primarily for tasks related to graphics rendering. The innovative utilization of GPU hardware in the present disclosure, particularly the blending engine, for Al-related quantization represents a meaningful change in thinking. By processing data directly within the GPU, this approach substantially alleviates thecomputational load on ancillary components, such as CPUs or external memory systems. This advancement constitutes a considerable improvement over prior methodologies, where the potential of GPUs for such operations was not fully exploited.
[0045] Another technical effect is the enhancement of data fidelity and efficiency. Commercially available GPU technologies have inadequately addressed the challenges of maintaining efficiency and data fidelity during quantization operations. The novel mapping and reconstruction techniques delineated in the present disclosure are designed to preserve the integrity of data during high-speed processing. This approach markedly diminishes the loss of data fidelity, a prevalent issue in the quantization processes within neural network computations.
[0046] Yet another technical effect is the reduction of energy consumption. The implementation of these solutions within the extant GPU architecture markedly enhances energy efficiency, an imperative consideration for mobile and portable devices. Traditionally, commercially available GPUs have been associated with substantial power consumption, particularly when managing complex Ai tasks. By refining the optimization of the blending engine to process these tasks more efficiently, the proposed solutions substantially decrease the overall energy demand.
[0047] It should be appreciated that the solutions presented in this disclosure leverage existing GPU architectures to address new challenges posed by the demands of modern Al applications. These enhancements not only improve the performance of GPUs beyond their traditional roles but also introduce a more efficient, powerful, and versatile toolset for developers and engineers, pushing the boundaries of what can be achieved with standard hardware components.
[0048] Figure 2 shows a block diagram of a computing system 200, in accordance with some disclosed embodiments. This computing system (System) 200 is configured to perform quantization and reconstruction processes as depicted in Figures 1A to 1C, and the method depicted in Figure 3. In some embodiments, System 200 may be implemented using various platforms such as a computer, server, cloud computing server, graphics processing unit (GPU), and any combination thereof, or the like. In some embodiments, System 200 includes processingunit 210, memory 220, and input / output (I / O) 230, interconnected using DMA protocols or other commercially available high-speed physical interconnection protocols.(0049] in some embodiments, Processing Unit 210 may include at least one Shader Unit 211, at least one Tensor Core 212, at least one Rasterizer 213, at least one Texture Mapping Unit (TMU) 214, at least one Blending Engine 215, and a Control 216. Shader Unit 211 may be responsible for rendering graphics through vertex, geometry, and pixel shaders for generating real-time visual elements of a scene. In some embodiments, Tensor Core 212 is tailored for deep learning and Al-specific computations, enhancing tasks like neural network training and inference. Rasterizer 213 transforms vector graphics into raster images for converting 3D models into 2D screen representations. In some embodiments, TMU 214 may be configured to manage texture-related tasks, including filtering and blending, to enhance visual realism. Control 216 orchestrates the flow of commands within Processing Unit 210, ensuring efficient routing between the processing unit, memory, and input / output interfaces.
[0050] The Blending Engine 215 enhances the performance of Processing Unit 210 for artificial intelligence (Al) applications by utilizing advanced blending techniques for sophisticated Al tasks beyond traditional graphics. These include operations like multiplication and addition within the blending engine 215 to support deeper neural network functionality. Hardware enhancements such as adding control bits, multiplexers, and shift registers further.support these processes, improving execution of neural network layers, reducing the load on other components, and increasing energy efficiency, focusing primarily on the optimization of deep learning tasks.
[0051] Additionally, or alternatively, Processing Unit 210 can be a Central Processing Unit (CPU), a microprocessor, an electronic circuit, or an Integrated Circuit (IC). It can also be implemented as firmware written for or ported to a specific processor such as a Digital Signal Processor (DSP) or as hardware or configurable hardware such as a Field Programmable Gate Array (FPGA) or an Application Specific integrated Circuit (ASIC).
[0052] Processing Unit 210 may be utilized to perform computations required by System 200 or any of its subcomponents. In some embodiments, System 200 may utilize I / O 230 as an interface to transmit and / or receive Information and instructions between System 200 and other nodes, such as GPUs or the like where multiple devices each have a plurality of nodes(cores). In some embodiments, I / O 230 utilizes physical connections, such as PCI Express (PCie). PCie (not shown) is a standard interface that connects GPU or similar systems to one another and to their motherboard, allowing communication with the CPU or other processing units and other peripherals. Additionally, or alternatively, I / O 230 may use another high-bandwidth interconnect, such as NVLink or Infinity Fabric, to allow direct data exchange between nodes of a neural network.
[0053] It should be appreciated that In a neural network implemented within a multi system graphics processing unit (GPU), I / O 230 serves as a critical interface for transmitting and receiving information and instructions between System 200 and similar system nodes. Each, containing multiple processing units or cores, collaborates to manage complex computational tasks like training deep neural networks. This setup allows for efficient data and weight synchronization across GPU (systems), enhancing the network’s ability to learn from and adapting to complex data patterns, thereby optimizing performance and accuracy in tasks such as pattern recognition and data classification.
[0054] Memory 220 may include both volatile and non-volatile memories, utilizing technologies such as semiconductor, magnetic, optical, or flash, individually or in combination.
[0003] For example, Memory 220 can include devices like a Flash disk, Random Access Memory (RAM), memory chips, optical storage devices (CDs, DVDs, laser disks), magnetic storage devices (hard disks, SAN, NAS}, or semiconductor storage devices like Flash devices or memory sticks. In some exemplary embodiments. Memory 220 may store program code that activates Processing Unit 210 to execute tasks, such as depicted in Figures 1A to 1C and 3. These components may be implemented as various forms of software or firmware, arranged in executable files or libraries, and programmed in any language suitable for the computing environment. Additionally, or alternatively, Memory 220 may include Video RAM (VRAM) 221 dedicated to high-speed storage of textures, frames, and weights matrices as floating or integer data. It may also include caches 222 to optimize data retrieval times from the main memory.
[0055] Figure 3 is a flowchart of Method 300 configured for enhanced neural network processing, in accordance with some disclosed embodiments. The method 300 leverages at least one blending engine 215 of System 200 (of Fig. 2), such as a graphics processing unit (GPU), tostreamline neural network processing, focusing on optimizing the quantization process. In some embodiments, method 300 incorporates advanced hardware adjustments and software integrations within System 200 to enhance the overall execution of Al kernels, improve data handling efficiency, and ensure robust communication across system components.
[0056] At S301, tensor values may be initialized. In some embodiments, at least one blending engine 215 (of Fig. 2) is configured to set up tensor values for the quantization process. This configuration, by preparing data in the format required for neural network processing, aligns with strategies for efficient data acquisition and preprocessing, in the embodiments, the blending engine's hardware capabilities with software requirements, enhancing the functionality of system 2.00 (of Fig. 2) by ensuring data is appropriately formatted from the start, thereby facilitating subsequent neural network operations and overall system performance.
[0057] At S302, the quantization process may be invoked to convert high-precision floating-point inputs into lower-precision discrete outputs such as 8-bit integers by utilizing at least one blending engine 215 (of Fig. 2). In some embodiments, this step employs complex operations like multiplication and addition directly within at least one blending engine 215 (of Fig. 2) for efficient data conversion and compression, reflecting the focus on optimizing mathematical operations and data handling through convolutional processes on system 200 (of Fig. 2).
[0058] At S303, the application of activation functions and re-quantization may be activated. In some embodiments, this step involves applying nonlinear activation functions to introduce necessary nonlinearity into the neural model, thereby enhancing its ability to capture complex patterns in the data input. Post-activation, data values req_value are re-quantized using the fallowing mathematical formula:ACTIVATION_MIN ≤ req_value + OUTPUT_ZERO_POINT ≤ ACTIVATION_MAX .[00591 Here, ACTIVATION_MIN is typically set to OUTPUT_ZERO_POINT in the case of ReLU activation functions, reducing the range of output values and ensuring they meet the precision requirements of subsequent neural network layers. This critical step ensures data fidelity by adjusting data precision, aligning with high -efficiency computational targets. Additionally, quantization parameters (multiplier and shift) are dynamically retrieved per channel from texture base addresses / buffers, allowing one or more blending engines 215 (ofFig. 2) to manage this re-quantization without additional core involvement, hence optimizing both performance and energy usage.(0060] At S304, kernel execution and optimization may be implemented. In some embodiments, Al kernels are executed, leveraging optimized hardware configurations, including tensor cores tailored for A! computations. This process involves adjusting hardware settings, such as adding control bits to at least one blending engine 215 (of Fig. 2) and activating multiplexers and shift registers to enhance neural network layer execution. These enhancements focus on maximizing performance and energy efficiency, underlining innovative hardware optimization.
[0061] At S305, efficient communication may be activated. In some embodiments, Control 216 manages the routing and synchronization of commands and data within System 200 (of Fig. 2), coordinating computational tasks between processing unit 210, memory 220, and the I / O 230 interfaces (of Fig 2). Additionally, or alternatively, I / O 230 facilitates high-speed data and command exchanges between GPUs or with external peripherals using high-bandwidth interconnects like PCIe, This step ensures robust inter-system communication and synchronization, for maintaining consistent data handling and optimizing system performance as outlined in the present disclosure.
[0062] The various embodiments disclosed herein can be implemented as hardware, firmware, software, or any combination thereof. Moreover, the software is preferably implemented as an application program tangibly embodied on a program storage unit or computer-readable medium consisting of parts, or of certain devices and / or a combination of devices. The application program may be uploaded to, and executed by, a machine comprising any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units (" CPUs"), memory, and input / output interfaces. The computer platform may also include an operating system and microinstruction code. The various processes and functions described herein may be either part of the microinstruction code or part of the application program or any combination thereof, which may be executed by a CPU, whether or not such a computer or processor is explicitly shown. In addition, various other peripheral units may be connected to the computer platform, such as an additional datastorage unit and a printing unit. Furthermore, a non-transitory computer-readable medium is any computer-readable medium except for a transitory propagating signal.
[0063] All examples and conditional language recited herein are intended for pedagogical purposes to aid the reader in understanding the principles of the disclosed embodiment and the 5 concepts contributed by the inventor to further the art and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the disclosed embodiments, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known w equivalents as well as equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure.
[0064] It should be understood that any reference to an element herein using a designation such as "first," "second," and so forth does not generally limit the quantity or order of those elements. Rather, these designations are generally used herein as a convenient method of s distinguishing between two or more elements or instances of an element. Thus, a reference to the first and second elements does not mean that only two elements may be employed there or that the first element must precede the second element in some manner. Also, unless stated otherwise, a set of elements comprises one or more elements.
[0065] As used herein, the phrase "at least one of" followed by a listing of items means that any of the listed items can be utilized individually, or any combination of two or more of the listed items can be utilized. For example, if a system is described as including "at least one of A, B, and C," the system can include A alone; B alone; C alone; 2A; 2B; 2C; 3A; A and B in combination; B and C in combination; A and C in combination; A, B, and C in combination; 2A and C in combination; A, 3B, and 2C in combination; and the like.
Claims
CLAIMSWhat is claimed is:
1. A method for processing and optimizing neural network data on a graphics processing unit (GPU), comprising:initializing data for neural network processing comprising:configuring at least one blending engine to set up tensor values necessary for quantization;aligning data preparation with efficient neural network operational standards; andpreprocessing input data to match input requirements of specific artificial intelligence (Al) models;invoking quantization within the at least, one blending engine comprising:converting high-precision floating-point inputs into 8-bit integers using the GPU's at least one blending engine;employing complex mathematical operations for effective data conversion; andincluding encoding techniques that optimize bit utilization for enhanced quantization;applying activation functions and dynamic re-quantization comprising:implementing nonlinear activation functions to enhance model response to data intricacies followed by re-quantization to adjust data precision for subsequent processing phases; andincluding dynamic adjustment of quantization scales;andexecuting optimized AI kernels comprising:running Al-specific computational tasks; andleveraging hardware optimizations within the blending engine for efficient kernel operations.
2. The method of claim 1, wherein initializing data further includes configuring the at least one blending engine to pre-process input data to match the input requirements of specific Al models before quantization.
3. The method of claim 1, further comprising; employing specific algorithms within the at least one blending engine for multiplication and addition operations that are optimized for data types typically processed in Al applications.
4. The method of claim 1, wherein applying activation functions includes using a range of functions selected based on a type of neural network layer being processed and dynamically adjusting the quantization scale to optimize computational efficiency.
5. The method of claim 1, wherein executing optimized Al kernels includes using tensor cores specifically designed for accelerating deep learning tasks, with configurations being adaptable based on the complexity of the neural network model being executed.
6. The method of claim 1, wherein executing optimized Al kernels further comprises leveraging hardware optimizations that include adjusting control bits and utilizing multiplexers and shift registers within the at least one blending engine for efficient kernel operations.
7. The method of claim 1, further comprising: managing efficient system communications by coordinating command and data flow across GPU components, ensuring synchronized operations and optimal data transfer speeds between internal and external system interfaces, and using high-bandwidth interconnects like PCie for robust inter-system communication.
8. The method of claim 7, wherein managing efficient system communications includes using a designated control unit to prioritize data packets based on real-time processing needs and adjust communication protocols dynamically to prevent bottlenecks.PCT / GR2025 / 0000069. A non-transitory computer-readable medium storing a set of instructions for processing and optimizing neural network data on a graphics processing unit (GPU), the set of instructions comprising:5 one or more instructions that, when executed by one or more processors of a device, cause the device to:initialize data for neural network processing comprising:configuring at least one blending engine to set up tensor values necessary for quantizationaligning data preparation with efficient neural network operational standards; andpreprocessing input data to match input requirements of specific artificial intelligence (Al) models;invoke quantization within the at least one blending engine comprising:converting high-precision floating-point inputs into 8-bit integers using the GPU's at least one blending engineemploying complex mathematical operations for effective data conversion; andincluding encoding techniques that optimize bit utilization for enhanced quantization;apply activation functions and dynamic re-quantization comprising:implementing nonlinear activation functions to enhance model response to data intricacies followed by re-quantization to adjust data precision for subsequent processing phases; andincluding dynamic adjustment of quantization scales; execute optimized Al kernels comprising:running Al-specific computational tasks; andleveraging hardware optimizations within the blending engine for efficient kernel operations.
10. A system for processing and optimizing neural network data on a graphics processing unit (GPU) comprising:a processing circuitry;s a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:initialize data for neural network processing comprising:configuring at least one blending engine to set up tensor values necessary for quantization;aligning data preparation with efficient neural network operational standards; andpreprocessing input data to match input requirements of specific artificial intelligence (Al) models;invoke quantization within the at least one blending engine comprising:converting high-precision floating-point inputs into 8-bit integers using the GPU's at least one blending engine;employing complex mathematical operations for effective data conversion; andincluding encoding techniques that optimize bit utilization for enhanced quantization;apply activation functions and dynamic re-quantization comprising:implementing nonlinear activation functions to enhance model response to data intricacies followed by re-quantization to adjust data precision for subsequent processing phases; andincluding dynamic adjustment of quantization scales;execute optimized Al kernels comprising:running Al-specific computational tasks; andleveraging hardware optimizations within the blending engine for efficient kernel operations..
11. The system of claim 10, wherein initializing data further includes configuring the at least one blending engine to pre-process input data to match the input requirements of specific Al modeis before quantization.
12. The system of claim 10, wherein the memory contains further instructions that, when executed by the processing circuitry, further configure the system to: employ specific algorithms within the at least one blending engine for multiplication and addition operations that are optimized for data types typically processed in Al applications.
13. The system of claim 10, wherein applying activation functions includes using a range of functions selected based on the type of neural network layer being processed and dynamically adjusting the quantization scale to optimize computational efficiency.
14. The system of claim 10, wherein executing optimized Al kernels includes using tensor cores specifically designed for accelerating deep learning tasks, with configurations being adaptable based on the complexity of the neural network model being executed.
15. The system of claim 10 wherein executing optimized Al kernels further comprises leveraging hardware optimizations that include adjusting control bits and utilizing multiplexers and shift registers within the at least one blending engine for efficient kernel operations.
16. The system of claim 10, wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:manage efficient system communications by coordinating command and data flow across GPU components, ensuring synchronized operations and optimal data transfer speeds between internal and external system interfaces, and using high-bandwidth interconnects like PCIe for robust inter-system communication.
17. The system of claim 16, wherein managing efficient system communications includes using a designated control unit to prioritize data packets based on real-time processing needs and adjust communication protocols dynamically to prevent bottlenecks.