Step-wise quantization for iterative neural networks

By generating scale factors based on feature map statistics for each layer and time step, the proposed step-wise quantization technique addresses the inefficiencies of conventional methods for iterative neural networks, enhancing processing efficiency and maintaining model precision.

WO2025129568A1PCT designated stage expired Publication Date: 2025-06-26INTEL CORP +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/140630
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Conventional model quantization techniques are ineffective for iterative neural networks, such as stable diffusion models, due to difficulties in selecting appropriate precision levels across time steps, leading to performance degradation and increased memory requirements.

Method used

The proposed solution involves generating feature map statistics at each layer and time step during a calibration phase, using these statistics to create scale factors for tensors, and applying these scale factors for quantization at each layer and time step during inference.

Benefits of technology

This step-wise quantization approach improves processing efficiency and maintains model precision by adapting quantization parameters per layer and time step, overcoming the limitations of prior methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140630_26062025_PF_FP_ABST
    Figure CN2023140630_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Technology disclosed herein provides for generating, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, where the trained neural network has a plurality of layers and is operated iteratively over a predetermined number of time steps for each input from the calibration data set, generating, on the per layer, per time step basis, a set of scale factors based on the feature map statistics, where the set of scale factors is determined relative to a target feature map quantization, and where elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and storing the set of scale factors as metadata associated with the trained neural network. The feature map statistics can include tensor values for the respective layers and time steps.
Need to check novelty before this filing date? Find Prior Art

Description

STEP-WISE QUANTIZATION FOR ITERATIVE NEURAL NETWORKSBACKGROUND

[0001] Artificial intelligence (AI) systems use trained models to operate on input data to produce an output. Latency and throughput concerns have led to attempts to improve inference performance, including through model quantization efforts. Conventional model quantization techniques involve quantizing model weights and / or tensors from floating point representation (e.g., 32-bit floating point / FP32) to a lower bit precision such as 8-bit integer (INT8) , 4-bit floating point (FP4) etc., which are supported by AI hardware devices to provide compute acceleration. Such conventional quantization techniques are directed to single-shot inference, where the model is run only once on an input data set to generate the output. These techniques do not work well, however, with iterative models such as, e.g., stable diffusion models, where the model is run iteratively to generate output from an input data set.

[0002] Prior attempts to quantize iteratively-run models have been based on selecting between a first version of the model at FP32 precision and a second version of the model at lower precision (e.g., INT8) at different time steps during execution, providing mixed-precision inference across time steps. These prior efforts, however, present several practical disadvantages, including, e.g., difficulty in determining an appropriate criterion to select between the first and second model versions, the inherent / negative impact on performance when executing a FP32 model, and the increase in memory allocation required to hold both the first and second model versions for execution.BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The various advantages of the embodiments will become apparent to one skilled in the art by reading the following specification and appended claims, and by referencing the following drawings, in which:

[0004] FIGs. 1A-1C provide diagrams illustrating an example of an image generation system according to one or more embodiments;

[0005] FIG. 2 provides a diagram illustrating an example of a U-Net network architecture for use in an image generation system according to one or more embodiments;

[0006] FIGs. 3A-3B provide diagrams illustrating an example of model calibration for an iterative neural network model according to one or more embodiments;

[0007] FIGs. 4A-4B provide diagrams illustrating an example of operating an iterative neural network model during inference according to one or more embodiments;

[0008] FIG. 5 provides a flow diagram illustrating an example method of calibrating an iterative neural network model for quantization according to one or more embodiments;

[0009] FIG. 6 provides a flow diagram illustrating an example method of operating an iterative neural network model according to one or more embodiments;

[0010] FIGs. 7A-7B illustrate examples of comparative results of operating an image generation system according to one or more embodiments;

[0011] FIG. 8 provides a block diagram illustrating an example performance-enhanced computing system according to one or more embodiments;

[0012] FIG. 9 provides a block diagram illustrating an example semiconductor apparatus for accelerating compute tasks according to one or more embodiments;

[0013] FIG. 10 is a block diagram illustrating an example processor core according to one or more embodiments; and

[0014] FIG. 11 is a block diagram illustrating an example of a multi-processor based computing system according to one or more embodiments.DESCRIPTION OF EMBODIMENTS

[0015] Embodiments relate generally to artificial intelligence (AI) systems and deep learning (DL) technology. More particularly, embodiments relate to quantization of input, output and feature maps for each layer and time step for iterative neural networks. An iterative neural network is a trained neural network that is operated iteratively based on a single input. The technology provided herein recognizes that feature map statistics from tensors within a given iterative model typically vary over each time step. As described herein, an improved model quantization technique provides for generating feature map statistics at each layer and time step during a calibration phase, using a calibration data set, and generating, based on the statistics, scale factors for the tensors at each layer and time step. At inference, the scale factors are applied to quantize the feature maps for tensors in the iterative model at each layer and time step. Thus, the quantization technique described herein provides for per layer, per time step quantization (i.e., step-wise quantization) for neural networks that are operated iteratively over time steps. By using scale factors that vary per layer and per time step for quantization, the technology as described herein overcomes disadvantages of prior solutions and provides for improved performance both in terms of processing requirements / time and model precision / accuracy.

[0016] FIG. 1A provides a block diagram illustrating an example of an image generation system 100 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. As shown in FIG. 1A, the image generation system 100 includes an image generator network 120 that receives an input text prompt 110 and generates an output generated image 130. The image generator network 120 includes an iterative neural network that is trained to generate images (such as, e.g., the generated image 130) from a text prompt (such as, e.g., the text prompt 110) . In embodiments the image generator network 120 is a stable diffusion network. Further details regarding the image generator network 120 are provided herein with reference to FIG. 1B.

[0017] Some or all components and / or features in the system 100 can be implemented using one or more of a central processing unit (CPU) , a graphics processing unit (GPU) , an artificial intelligence (AI) accelerator, a field programmable gate array (FPGA) accelerator, an application specific integrated circuit (ASIC) , and / or via a processor with software, or in a combination of a processor with software and an FPGA or ASIC, and / or in an AI  / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow, etc. or any similar framework. More particularly, components of the system 100 can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as random access memory (RAM) , read only memory (ROM) , programmable ROM (PROM) , firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured programmable logic arrays (PLAs) , FPGAs, complex programmable logic devices (CPLDs) , and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with complementary metal oxide semiconductor (CMOS) logic circuits, transistor-transistor logic (TTL) logic circuits, or other circuits.

[0018] For example, computer program code to carry out operations by the system 100 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine  instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .

[0019] FIG. 1B provides a block diagram illustrating an example of an image generator network 120 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The example image generator network 120 shown in FIG. 1B is a stable diffusion network. Stable diffusion networks can be used in a variety of image generation applications such as, e.g., a text to image generative model that can generate images according to input text prompts (text-conditioned) , a sketch-conditioned generative model to generate the image based on a sketch input, or a low resolution image conditioned model to generate a high resolution image based on a low resolution input. As shown in FIG. 1B, the image generator network 120 includes a frozen contrastive language-image pre-training (CLIP) text encoder 121 that operates to produce text embeddings 122 based on the input text prompt 110. The CLIP text encoder 121 is a neural network trained (or pre-trained) on image-text pairs, and once trained is not further modified in calibration or inference operation of the image generator network 120 (hence, the CLIP text encoder 121 is “frozen” ) . In some embodiments, the text embeddings 122 have a size 77 x 768 elements.

[0020] The image generator network 120 further includes a latent seed generator 123 that generates a latent array 124. In embodiments the latent array 124 has a Gaussian distribution normalized in the range (0, 1) . The latent array 124 represents image information (e.g., latent space, or compressed image information) for a completely noisy image. As an example, in some embodiments, the latent array 124 has a size of 64 x 64 elements.

[0021] In operation, the text embeddings 122 and the noise latent array 124 are fed into a text-conditioned latent U-Net 125 for a first pass through the text-conditioned latent U-Net 125, which operates on the inputs and produces an output. The text-conditioned latent U-Net 125 is a convolutional neural network based on the U-Net architecture, and is trained to produce information representing less noisy images based on noisy image information input and text input. Further details regarding an example U-Net architecture for use in the text-conditioned latent U-Net 125 are provided herein with reference to FIG. 2.

[0022] The output of the text-conditioned latent U-Net 125 is conditioned latent array 127, representing image information (e.g., latent space, or compressed image information) for a less- noisy image than the prior input image information. The conditioned latent array 127 is fed back to the input of the text-conditioned latent U-Net 125 (e.g., via a scheduler) . The text-conditioned latent U-Net 125 operates on the conditioned latent array 127 and the text embeddings 122 to produce a new output as an updated conditioned latent array 127 (again, representing image information for a less-noisy image than the prior input image information) . The updated conditioned latent array 127 is fed back to the input of the text-conditioned latent U-Net 125 (e.g., via a scheduler) , which operates on the updated conditioned latent array 127 and the text embeddings 122 to produce a new output as a further updated conditioned latent array 127. Notably, the initial latent array 124 is only fed into the text-conditioned latent U-Net 125 for the first pass, and for subsequent passes the conditioned latent array 127, as updated for each pass, becomes the input latent array to the text-conditioned latent U-Net 125.

[0023] This process is repeated iteratively N times (e.g., at N time steps) , after which the final output conditioned latent array 127 is fed to a variational autoencoder decoder 128. The variational autoencoder decoder 128 is a network that upconverts (or decompresses) image information to produce the generated image 130 based on the final output conditioned latent array 127. The number of iterations (N) is a predetermined number. In some embodiments, the number of iterations (N) is set to a value such as, e.g., 15, 25, 50, or 100, or 200, etc.

[0024] There are two phases involved in operating the image generator network 120 –calibration and inference. During calibration, feature map statistics for each feature map (e.g., tensor) layer of the text-conditioned latent U-Net 125 at each time step are generated based on running a calibration data set through the image generator network 120 (each calibration data input involving N passes through the text-conditioned latent U-Net 125) . The statistics are used to generate a set of scale factors 126, representing scale factors for quantizing feature maps (e.g., tensors) of the text-conditioned latent U-Net 125 on a per layer per time step basis.

[0025] Once the scale factors 126 are generated, they are stored as metadata (e.g., in a configuration file such as a quantization configuration file) accompanying or associated with the image generator network 120. The scale factors 126 are then used at inference to quantize each of the feature maps (e.g., tensors) of the text-conditioned latent U-Net 125 on a per layer, per time step basis. Notably, the weights of the text-conditioned latent U-Net 125 are quantized (if at all) on a fixed one-time basis (not per time step) . Further details regarding calibration are provided herein with reference to FIGs. 3A-3B, and further details regarding inference are provided herein with reference to FIGs. 4A-4B.

[0026] Some or all components and / or features in the image generator network 120 can be implemented using one or more of a CPU, a GPU, an AI accelerator, an FPGA accelerator, an ASIC, and / or via a processor with software, or in a combination of a processor with software and an FPGA or ASIC, and / or in an AI  / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow, etc. or any similar framework. More particularly, components and / or features of the image generator network 120 can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.

[0027] For example, computer program code to carry out operations by the image generator network 120 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .

[0028] FIG. 1C provides a diagrams illustrating an example of an image generation system 100 operating on a particular input text according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. In the example of FIG. 1C, the sample phrase “An astronaut riding on a horse” is used as a particular input text prompt 110a to the image generator network 120, which has been calibrated as described herein. The image generator network 120 is operated iteratively for N time steps (e.g., as described herein with reference to FIG. 1B) with quantized feature maps as described herein to produce the sample output image 130a.

[0029] FIG. 2 provides a diagram illustrating an example of a U-Net network architecture 200 for use in an text-conditioned latent U-Net (such as, e.g., the text-conditioned latent U-Net 125 in FIG. 1B, already discussed) according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The U-Net network architecture 200 is a neural network architecture that includes a number of layers arranged in a conventional manner for U-Net architectures as illustrated in FIG. 2. For example, in the top row on the left side of the diagram, the layer 0 represents the input image (or image information, such as a compressed image) , the layer 1 represents a feature map produced via a 3x3 convolution applied to layer 0, and the layer 2 represents a feature map produced via a 3x3 convolution applied to layer 1. In the next row of the diagram (left side) , the layer 3 represents a feature map produced via a 2x2 max pool applied to layer 2, the layer 4 represents a feature map produced via a 3x3 convolution applied to layer 3, and the layer 5 represents a feature map produced via a 3x3 convolution applied to layer 4.

[0030] The U-Net architecture continues in similar fashion down the left side and then over to the right side, where in successive rows above the bottom row the left layer (e.g., layer M) is produced by concatenating a copy of a layer from the left side (e.g., layer 5) and a feature map produced via a 2x2 upconvert applied to a layer (e.g., layer K) in the row below, the middle layer (e.g., layer N) represents a feature map produced via a 3x3 convolution applied to the left layer (e.g., layer M) , and the right layer (e.g., layer O) represents a feature map produced via a 3x3 convolution applied to the middle layer (e.g., layer N) . The final output layer (layer R) is produced via a 1x1 convolution applied to the preceding layer R-1.

[0031] In the example U-Net architecture of FIG. 2, the various layers are, for purposes of illustration, shown as rectangles. For each layer / rectangle, the height is roughly indicative of the relative dimensionality (e.g., x by x or x by y) of the respective feature map (e.g., tensor) , and the width is roughly indicative of the relative number of channels of the respective feature map (e.g., tensor) .

[0032] For calibration, statistics are obtained by running each calibration input of the calibration data set through the U-Net network architecture 200 for N time steps. In some embodiments, the calibration data set can have a range of 10-100 calibration inputs Statistics are generated for each time step and each layer of the U-Net network architecture 200, that is, for the input layer (layer 0) and each of the feature maps produced in the network (e.g., layer 1, layer 2, layer 3, layer 4, layer 5, …, layer K, layer M, layer N, layer O, …layer R-1, and layer R) . These statistics are combined across all calibration inputs and used to generate per time  step, per layer scale factors. The scale factors are applied, during inference, to quantize the input layer 0 and the each of the feature maps produced in the network (e.g., layer 1, layer 2, layer 3, layer 4, layer 5, …, layer K, layer M, layer N, layer O, …layer R-1, and layer R) at each time step.

[0033] FIGs. 3A-3B provide diagrams illustrating an example of a model calibration process 300 (including process components 300A and 300B) for an iterative neural network model (such as, e.g., the text-conditioned latent U-Net 125 in FIG. 1B, already discussed) according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The process 300 (or aspects thereof) can be implemented using one or more of a CPU, a GPU, an AI accelerator, an FPGA accelerator, an ASIC, and / or via a processor with software, or in a combination of a processor with software and an FPGA or ASIC, and / or in an AI  / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow, etc. or any similar framework. More particularly, the process 300 can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.

[0034] For example, computer program code to carry out operations for the process 300 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .

[0035] Turning to FIG. 3A, the process component 300A of the model calibration process 300 is illustrated, showing the generation of feature map statistics. The statistics are generated from operating the iterative neural network model over a series of time steps covering a range of T intervals –e.g., intervals including t=T-1, t=n, and t=0. For purposes of illustration, only three individual time steps are shown in FIG. 3A, but it will be understood that the generation of statistics is similarly conducted for all T time steps. Of note, the labeling of time steps in “reverse” numerical order (e.g., intervals including T-1, …, n, …0) for the example stable diffusion network relates to the underlying diffusion algorithm, which models adding noise to a clean image (e.g., the clean image is the starting point at time interval t=0) .

[0036] For each time step t, the statistics are generated based on the feature map (e.g., tensor) values at each layer in the iterative neural network model for that time interval. For example, as shown in FIG. 3A, for time step t = T-1, a set of statistics 310 are generated for the iterative neural network model, where the statistics are generated for each layer individually --e.g., layer 0, layer 1, layer 2, …layer R (layer R being the last layer of the iterative neural network model) . Similarly, for time step t = n, a set of statistics 311 are generated for the iterative neural network model, where the statistics are generated for each layer individually --e.g., layer 0, layer 1, layer 2, …layer R. Likewise, for time step t = 0, a set of statistics 312 are generated for the iterative neural network model, where the statistics are generated for each layer individually --e.g., layer 0, layer 1, layer 2, …layer R. Collectively, for all T time steps the statistics are generated on a per time step, per layer basis; the statistics are typically tensors, but in some embodiments the statistics are scalar values. For each respective time step t and respective layer   the statistics are combined across results for all inputs from the calibration data set. For example, for a respective time step t and respective layer   the statistics are generated by summing the tensor values across all inputs from the calibration set on a per layer basis for that time step.

[0037] Turning now to FIG. 3B, , the process component 300B of the model calibration process 300 is illustrated, showing the generation of scale factors from operating the iterative neural network model over a series of time steps covering a range of T intervals. For purposes of illustration, only one individual time step (for t = n) is shown in FIG. 3B, but it will be understood that the generation of scale factors is similarly conducted for all T time steps. For the illustrated time step t = n, the respective statistics 311 (generated individually for each layer, e.g., layer 0, layer 1, layer 2, …layer R) are processed, layer by layer, by a calibrator 320, to generate a series of scale factors 321, denoted S [t: n] [l: 0-R] , on a per layer basis for the particular  time step n. That is, processing each individual statistics for a respective layer via the calibrator 320 produces, on a one for one basis, an individual respective scale factor for each respective layer (e.g., layer 0, layer 1, layer 2, …layer R) .

[0038] The calibrator 320 generates a scale factor for each per time step per layer statistics. For example, in some embodiments the statistics for a particular layer and time step includes the set of all tensor values at that time step and layer for all of the calibration inputs. The calibrator 320 then generates a scale factor by applying a function to the statistics. As one example, the function can include determining the maximum absolute value of the statistics for that time step and layer. The scale factor is adjusted to normalize the values such that the resulting quantized feature map values generated using the scale factor will fit within the desired quantization (e.g., within a range of 0-255 for INT8 quantization) . In embodiments, the scale factors are tensors having, e.g., the same number of channels as the feature map for the respective layer. In some embodiments, the scale factors are scalar values determined, e.g., by applying a function, such as average value, to the tensor values for a given layer and time step.

[0039] Once the scale factors for the T time steps (denoted S [t: 0-T-1] [l: 0-R] ) are generated, they are stored (e.g., in a configuration file such as a quantization configure file) accompanying or associated with the iterative neural network model. The scale factors are then used at inference to quantize each of the feature maps (e.g., tensors) of the iteratively-operated neural network model (e.g., the text-conditioned latent U-Net 125 in FIG. 1B, already discussed) on a per layer, per time step basis.

[0040] FIGs. 4A-4B provide diagrams illustrating an example of a process 400 (including process components 400A and 400B) for operating an iterative neural network model (such as, e.g., the text-conditioned latent U-Net 125 in FIG. 1B, already discussed) during inference according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The process 400 (or aspects thereof) can be implemented using one or more of a CPU, a GPU, an AI accelerator, an FPGA accelerator, an ASIC, and / or via a processor with software, or in a combination of a processor with software and an FPGA or ASIC, and / or in an AI  / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow, etc. or any similar framework. More particularly, the process 400 can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic,  or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.

[0041] For example, computer program code to carry out operations for the process 400 can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .

[0042] Turning to FIG. 4A, the process component 400A of the process 400 is illustrated, showing the application of scale factors to feature maps of the iterative neural network model during inference, as operated iteratively over a series of time steps covering a range of T intervals –e.g., intervals including t=T-1, t=n, and t=0. For purposes of illustration, only three individual time steps are shown in FIG. 4A, but it will be understood that the application of scale factors to quantize the feature maps is similarly conducted for all T time steps. The scale factors are generated on a per time step per layer basis as described herein with reference to FIGs. 3A-3B.

[0043] For example, as shown in FIG. 4A, for time step t = T-1, a set of scale factors 410 are applied to the iterative neural network model (which has been loaded into memory for execution) . The scale factors are applied to quantize the feature maps (e.g., tensors) for each layer individually --e.g., layer 0, layer 1, layer 2, …layer R (layer R being the last layer of the iterative neural network model) . Similarly, for time step t = n, a set of scale factors 411 are applied to the iterative neural network model, where the scale factors are applied to quantize the feature maps (e.g., tensors) for each layer individually --e.g., layer 0, layer 1, layer 2, …layer R (the scale factors 411, for t = n, correspond to the scale factors 321 in FIG. 3B, already discussed) . Likewise, for time step t = 0, a set of scale factors 412 are applied to the iterative  neural network model, where the scale factors are applied to quantize the feature maps (e.g., tensors) for each layer individually --e.g., layer 0, layer 1, layer 2, …layer R.

[0044] Turning to FIG. 4B, the process component 400B of the process 400 is illustrated, showing an example of how the scale factors are applied as a quantized neural network model 420 is operated iteratively over T time steps. The weights of the quantized neural network model 420 itself are quantized, where the quantized weights are frozen for all time steps. In the example shown in FIG. 4B, the quantized neural network model 420 is a stable diffusion text-to-image model (e.g., as described herein with reference to FIG. 1B) , and is operated in inference with a text prompt input relating to an astronaut on a horse. Xt refers to all tensors collectively (e.g., all feature maps in all R+1 layers) in the quantized neural network model 420 for the particular time step t. For the first pass (time step T-1) , a noisy latent array 430 (e.g., having a Gaussian noise distribution) is input to the quantized neural network model 420. The tensors XT-1 for the time step T-1 are quantized, on a per layer basis, for the R+1 layers using the scale factors S [T-1] [l: 0-R] . The resulting output of the neural network model 420 is a conditioned latent array 441.

[0045] The operation of the quantized neural network model 420 is repeated iteratively (e.g., as described herein with reference to FIG. 1B) over T time steps, where the resulting conditioned latent array is fed back into the neural network model 420 for the subsequent pass. As illustrated in FIG. 4B, at time step t, the scale factors S [t] [l: 0-R] are applied to quantize the tensors Xt in the R+1 layers of the quantized neural network model 420; the resulting output for the pass is a conditioned latent array 442. Similarly, at time step t-1, the scale factors S [t-1] [l: 0-R] are applied to quantize the tensors Xt-1 in the R+1 layers of the quantized neural network model 420; the resulting output for the pass is a conditioned latent array 443. Likewise, at time step 0, the scale factors S [t] [l: 0-R] are applied to quantize the tensors X0 in the R+1 layers of the quantized neural network model 420; the resulting output for the pass is a conditioned latent array 444. In the example shown, the time step 0 is the final pass, such that the conditioned latent array 444 can be fed to a variational autoencoder decoder (such as the variational autoencoder decoder 128 in FIG. 1B, already discussed) (not shown in FIG. 4B) to produce a final generated image.

[0046] FIG. 5 provides a flow diagram illustrating an example method 500 of calibrating an iterative neural network model for quantization according to one or more embodiments, with  reference to components and features described herein including but not limited to the figures and associated description. The method 500 can generally be implemented in the system 100 (FIG. 1A, already discussed) , via components of the image generator network 120 (Fig. 1B, already discussed) , and / or in an AI  / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow, etc. or any similar framework. More particularly, the method 500 can be implemented as one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.

[0047] For example, computer program code to carry out operations shown in the method 500 and / or functions associated therewith can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .

[0048] Illustrated processing block 510a provides for generating, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, where at block 510b the trained neural network has a plurality of layers, and where at block 510c the trained neural network is operated iteratively over a predetermined number of time steps for each input from the calibration data set. Illustrated processing block 520a provides for generating, on a per layer, per time step basis, a set of scale factors based on the feature map statistics, where at block 520b the set of scale factors is determined relative to a target feature map quantization, and where at block 520c elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps. Illustrated  processing block 530 provides for storing the set of scale factors as metadata associated with the trained neural network.

[0049] In some embodiments, the feature map statistics include tensor values for the respective layers and time steps, and elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps. In some embodiments, the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step. In some embodiments, the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0050] In some embodiments, the trained neural network is to be operated during inference, where the set of scale factors are applied, on a per layer, per time step basis, to quantize feature maps generated via operating the trained neural network, and where the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input. In some embodiments, the trained neural network is trained to generate an image based on a text prompt. In some embodiments, the trained neural network is a stable diffusion network. In some embodiments, the calibration data set comprises a series of input text prompts.

[0051] FIG. 6 provides a flow diagram illustrating an example method 600 of operating an iterative neural network model during inference according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The method 600 can generally be implemented in the system 100 (FIG. 1A, already discussed) , via components of the image generator network 120 (Fig. 1B, already discussed) , and / or in an AI  / deep learning framework, such as, e.g., OpenVINO, PyTorch, Tensorflow, etc. or any similar framework. More particularly, the method 600 can be implemented as one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.

[0052] For example, computer program code to carry out operations shown in the method 600 can be written in any combination of one or more programming languages, including an object- oriented programming language such as Java, JavaScript, Python, C#, C++, Perl, Smalltalk, or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. Additionally, program or logic instructions might include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc. ) .

[0053] Illustrated processing block 610a provides for loading a trained neural network along with a set of scale factors associated with the trained neural network, where at block 610b the trained neural network has a plurality of layers. Illustrated processing block 620a provides for applying, on a per layer, per time step basis, the set of scale factors to quantize feature maps (e.g., tensors) generated via operating the trained neural network, where at block 620b the trained neural network is operated iteratively over a predetermined number of time steps for a single input, and where at block 620c elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps. Illustrated processing block 630 provides for storing an output, where the output is generated based on operating the trained neural network using the single input.

[0054] In some embodiments, the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step. In some embodiments, the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0055] In some embodiments, the trained neural network is a stable diffusion network. In some embodiments, the trained neural network is trained to generate images or image information based on text prompts, wherein the input comprises a text prompt, and wherein the output comprises an image or image information.

[0056] FIGs. 7A-7B illustrate examples of comparative results of operating an image generation system according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The image generation system used in each example is a trained stable diffusion network using the same seed values and U-Net network configurations (except for the feature map quantization) . In each example, a text prompt was used as input to the stable diffusion network to generate an output image. In each example, three comparative image results are shown: (a)  the feature maps of the stable diffusion network are unquantized (FP32 data) ; (b) the feature maps of the stable diffusion network are quantized conventionally with INT8 quantization (i.e., feature maps at all time steps are quantized with the same quantization scale factor) ; and (c) the feature maps of the stable diffusion network are quantized with step-wise INT8 quantization using scale factors on a per layer, per time step basis (e.g., as described herein with reference to FIGs. 1A-1C, 2, 3A-3B, 4A-4B, and 5A-5B) , where the stable diffusion network used corresponds to the image generator network 120 (FIG. 1B, already discussed) . In case (b) and case (c) , the weights of the U-Net were quantized with fixed INT8 quantization.

[0057] Turning to FIG. 7A, the input text prompt “a man is coding” was used as input. The first image 702 was generated using unquantized FP32 feature maps. The second image 704 was generated using conventional INT quantization of the feature maps. Substantial artifacts appear in the image 704, such as an additional man’s arm. The third image 706 was generated using step-wise INT8 quantization using scale factors on a per layer, per time step basis (e.g., as described herein) . The image 706 bears significantly closer in appearance to the image 702 (unquantized FP32 feature maps) , with only relatively minor differences.

[0058] Turning now to FIG. 7B, the input text prompt “an apple in the dish” was used as input. The first image 712 was generated using unquantized FP32 feature maps. The second image 714 was generated using conventional INT quantization of the feature maps. Substantial artifacts appear in the image 714, such as an additional apple, where the apples also appear artificial. The third image 766 was generated using step-wise INT8 quantization using scale factors on a per layer, per time step basis (e.g., as described herein) . The image 716 bears includes only a single apple which has a more realistic appearance than the apples in the image 714.

[0059] As demonstrated by each set of examples, the neural network quantization technology described herein provides for improved performance in the appearance of generated images over conventional quantization techniques. While the per layer, per time step quantization technique (i.e., step-wise quantization) for iterative neural networks has been described herein with reference to stable diffusion networks as an example, it will be understood that that the quantization technology as described herein is also applicable to other types of iterative neural networks (including, for example, other types of diffusion networks) .

[0060] FIG. 8 shows a block diagram illustrating an example performance-enhanced computing system 10 for step-wise quantization for iterative neural networks according to one or more embodiments, with reference to components and features described herein including but not  limited to the figures and associated description. The system 10 can generally be part of an electronic device / platform having computing and / or communications functionality (e.g., a server, cloud infrastructure controller, database controller, notebook computer, desktop computer, personal digital assistant / PDA, tablet computer, convertible tablet, smart phone, etc. ) , imaging functionality (e.g., camera, camcorder) , media playing functionality (e.g., smart television / TV) , wearable functionality (e.g., watch, eyewear, headwear, footwear, jewelry, or other wearable devices) , vehicular functionality (e.g., car, truck, motorcycle) , robotic functionality (e.g., robot or autonomous robot) , Internet of Things (IoT) functionality, etc., or any combination thereof. In the illustrated example, the system 10 can include a host processor 12 (e.g., central processing unit / CPU) having an integrated memory controller (IMC) 14 that can be coupled to system memory 20. The host processor 12 can include any type of processing device, such as, e.g., microcontroller, microprocessor, RISC processor, ASIC, etc., along with associated processing modules or circuitry. The system memory 20 can include any non-transitory machine-or computer-readable storage medium such as RAM, ROM, PROM, EEPROM, firmware, flash memory, etc., configurable logic such as, for example, PLAs, FPGAs, CPLDs, fixed-functionality hardware logic using circuit technology such as, for example, ASIC, CMOS or TTL technology, or any combination thereof suitable for storing instructions 28.

[0061] The system 10 can also include an input / output (I / O) module 16. The I / O module 16 can communicate with for example, one or more input / output (I / O) devices 17, a network controller 24 (e.g., wired and / or wireless NIC) , and storage 22. The storage 22 can be comprised of any appropriate non-transitory machine-or computer-readable memory type (e.g., flash memory, DRAM, SRAM (static random access memory) , solid state drive (SSD) , hard disk drive (HDD) , optical disk, etc. ) . The storage 22 can include mass storage. In some embodiments, the host processor 12 and / or the I / O module 16 can communicate with the storage 22 (all or portions thereof) via a network controller 24. In some embodiments, the system 10 can also include a graphics processor 26 (e.g., a graphics processing unit / GPU) and / or an AI accelerator 27. In an embodiment, the system 10 can also include a vision processing unit (VPU) , not shown.

[0062] The host processor 12 and the I / O module 16 can be implemented together on a semiconductor die as a system on chip (SoC) 11, shown encased in a solid line. The SoC 11 can therefore operate as a computing apparatus for step-wise quantization for iterative neural networks. In some embodiments, the SoC 11 can also include one or more of the system  memory 20, the network controller 24, and / or the graphics processor 26 (shown encased in dotted lines) . In some embodiments, the SoC 11 can also include other components of the system 10.

[0063] The host processor 12 and / or the I / O module 16 can execute program instructions 28 retrieved from the system memory 20 and / or the storage 22 to perform one or more aspects of process the process 300, the process 400, the method 500, and / or the method 600 as described herein with reference to FIGs. 3A-3B, 4A-4B, 5, and / or 6. The system 10 can implement one or more aspects of the image generation system 100, the image generator network 120, text-conditioned latent U-Net 125, and / or the U-Net network architecture 200 as described herein with reference to FIGs. 1A-1B, 2, 3A-3B, 4A-4B, 5, and / or 6. The system 10 is therefore considered to be performance-enhanced at least to the extent that the technology provides an improved quantization technique for iterative neural networks.

[0064] Computer program code to carry out the processes described above can be written in any combination of one or more programming languages, including an object-oriented programming language such as JAVA, JAVASCRIPT, PYTHON, SMALLTALK, C++ or the like and / or conventional procedural programming languages, such as the “C” programming language or similar programming languages, and implemented as program instructions 28. Additionally, program instructions 28 can include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, state-setting data, configuration data for integrated circuitry, state information that personalizes electronic circuitry and / or other structural components that are native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, microprocessor, etc. ) .

[0065] I / O devices 17 can include one or more of input devices, such as a touchscreen, keyboard, mouse, cursor-control device, microphone, digital camera, video recorder, camcorder, biometric scanners and / or sensors; input devices can be used to enter information and interact with system 10 and / or with other devices. The I / O devices 17 can also include one or more of output devices, such as a display (e.g., touchscreen, liquid crystal display / LCD, light emitting diode / LED display, plasma panels, etc. ) , speakers and / or other visual or audio output devices. The input and / or output devices can be used, e.g., to provide a user interface.

[0066] FIG. 9 shows a block diagram illustrating an example semiconductor apparatus 30 for step-wise quantization for iterative neural networks according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The semiconductor apparatus 30 can be implemented, e.g.,  as a chip, die, or other semiconductor package. The semiconductor apparatus 30 can include one or more substrates 32 comprised of, e.g., silicon, sapphire, gallium arsenide, etc. The semiconductor apparatus 30 can also include logic 34 comprised of, e.g., transistor array (s) and other integrated circuit (IC) components) coupled to the substrate (s) 32. The logic 34 can be implemented at least partly in configurable logic or fixed-functionality logic hardware. The logic 34 can implement the system on chip (SoC) 11 described above with reference to FIG. 8. The logic 34 can implement one or more aspects of the processes described above, including , the process 300, the process 400, the method 500, and / or the method 600. The logic 34 can implement one or more aspects of the image generation system 100, the image generator network 120, text-conditioned latent U-Net 125, and / or the U-Net network architecture 200 as described herein with reference to FIGs. 1A-1B, 2, 3A-3B, 4A-4B, 5, and / or 6. The apparatus 30 is therefore considered to be performance-enhanced at least to the extent that the technology provides an improved quantization technique for iterative neural networks.

[0067] The semiconductor apparatus 30 can be constructed using any appropriate semiconductor manufacturing processes or techniques. For example, the logic 34 can include transistor channel regions that are positioned (e.g., embedded) within the substrate (s) 32. Thus, the interface between the logic 34 and the substrate (s) 32 may not be an abrupt junction. The logic 34 can also be considered to include an epitaxial layer that is grown on an initial wafer of the substrate (s) 32.

[0068] FIG. 10 is a block diagram illustrating an example processor core 40 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The processor core 40 can be the core for any type of processor, such as a micro-processor, an embedded processor, a digital signal processor (DSP) , a network processor, a graphics processing unit (GPU) , or other device to execute code. Although only one processor core 40 is illustrated in FIG. 10, a processing element can alternatively include more than one of the processor core 40 illustrated in FIG. 10. The processor core 40 can be a single-threaded core or, for at least one embodiment, the processor core 40 can be multithreaded in that it can include more than one hardware thread context (or “logical processor” ) per core.

[0069] FIG. 10 also illustrates a memory 41 coupled to the processor core 40. The memory 41 can be any of a wide variety of memories (including various layers of memory hierarchy) as are known or otherwise available to those of skill in the art. The memory 41 can include one or more code 42 instruction (s) to be executed by the processor core 40. The code 42 can  implement one or more aspects of the , the process 300, the process 400, the method 500, and / or the method 600 described above. The processor core 40 can implement one or more aspects of the image generation system 100, the image generator network 120, text-conditioned latent U-Net 125, and / or the U-Net network architecture 200 as described herein with reference to FIGs. 1A-1B, 2, 3A-3B, 4A-4B, 5, and / or 6. The processor core 40 can follow a program sequence of instructions indicated by the code 42. Each instruction can enter a front end portion 43 and be processed by one or more decoders 44. The decoder 44 can generate as its output a micro operation such as a fixed width micro operation in a predefined format, or can generate other instructions, microinstructions, or control signals which reflect the original code instruction. The illustrated front end portion 43 also includes register renaming logic 46 and scheduling logic 48, which generally allocate resources and queue the operation corresponding to the convert instruction for execution.

[0070] The processor core 40 is shown including execution logic 50 having a set of execution units 55-1 through 55-N. Some embodiments can include a number of execution units dedicated to specific functions or sets of functions. Other embodiments can include only one execution unit or one execution unit that can perform a particular function. The illustrated execution logic 50 performs the operations specified by code instructions.

[0071] After completion of execution of the operations specified by the code instructions, back end logic 58 retires the instructions of code 42. In one embodiment, the processor core 40 allows out of order execution but requires in order retirement of instructions. Retirement logic 59 can take a variety of forms as known to those of skill in the art (e.g., re-order buffers or the like) . In this manner, the processor core 40 is transformed during execution of the code 42, at least in terms of the output generated by the decoder, the hardware registers and tables utilized by the register renaming logic 46, and any registers (not shown) modified by the execution logic 50.

[0072] Although not illustrated in FIG. 10, a processing element can include other elements on chip with the processor core 40. For example, a processing element can include memory control logic along with the processor core 40. The processing element can include I / O control logic and / or can include I / O control logic integrated with memory control logic. The processing element can also include one or more caches.

[0073] FIG. 11 is a block diagram illustrating an example of a multi-processor based computing system 60 according to one or more embodiments, with reference to components and features described herein including but not limited to the figures and associated description. The  multiprocessor system 60 includes a first processing element 70 and a second processing element 80. While two processing elements 70 and 80 are shown, it is to be understood that an embodiment of the system 60 can also include only one such processing element.

[0074] The system 60 is illustrated as a point-to-point interconnect system, wherein the first processing element 70 and the second processing element 80 are coupled via a point-to-point interconnect 71. It should be understood that any or all of the interconnects illustrated in FIG. 11 can be implemented as a multi-drop bus rather than point-to-point interconnect.

[0075] As shown in FIG. 11, each of the processing elements 70 and 80 can be multicore processors, including first and second processor cores (i.e., processor cores 74a and 74b and processor cores 84a and 84b) . Such cores 74a, 74b, 84a, 84b can be configured to execute instruction code in a manner similar to that discussed above in connection with FIG. 10.

[0076] Each processing element 70, 80 can include at least one shared cache 99a, 99b. The shared cache 99a, 99b can store data (e.g., instructions) that are utilized by one or more components of the processor, such as the cores 74a, 74b and 84a, 84b, respectively. For example, the shared cache 99a, 99b can locally cache data stored in a memory 62, 63 for faster access by components of the processor. In one or more embodiments, the shared cache 99a, 99b can include one or more mid-level caches, such as level 2 (L2) , level 3 (L3) , level 4 (L4) , or other levels of cache, a last level cache (LLC) , and / or combinations thereof.

[0077] While shown with only two processing elements 70, 80, it is to be understood that the scope of the embodiments is not so limited. In other embodiments, one or more additional processing elements can be present in a given processor. Alternatively, one or more of the processing elements 70, 80 can be an element other than a processor, such as an accelerator or a field programmable gate array. For example, additional processing element (s) can include additional processors (s) that are the same as a first processor 70, additional processor (s) that are heterogeneous or asymmetric to processor a first processor 70, accelerators (such as, e.g., graphics accelerators or digital signal processing (DSP) units) , field programmable gate arrays, or any other processing element. There can be a variety of differences between the processing elements 70, 80 in terms of a spectrum of metrics of merit including architectural, micro architectural, thermal, power consumption characteristics, and the like. These differences can effectively manifest themselves as asymmetry and heterogeneity amongst the processing elements 70, 80. For at least one embodiment, the various processing elements 70, 80 can reside in the same die package.

[0078] The first processing element 70 can further include memory controller logic (MC) 72  and point-to-point (P-P) interfaces 76 and 78. Similarly, the second processing element 80 can include a MC 82 and P-P interfaces 86 and 88. As shown in FIG. 11, MC’s 72 and 82 couple the processors to respective memories, namely a memory 62 and a memory 63, which can be portions of main memory locally attached to the respective processors. While the MC 72 and 82 is illustrated as integrated into the processing elements 70, 80, for alternative embodiments the MC logic can be discrete logic outside the processing elements 70, 80 rather than integrated therein.

[0079] The first processing element 70 and the second processing element 80 can be coupled to an I / O subsystem 90 via P-P interconnects 76 and 86, respectively. As shown in FIG. 11, the I / O subsystem 90 includes P-P interfaces 94 and 98. Furthermore, the I / O subsystem 90 includes an interface 92 to couple I / O subsystem 90 with a high performance graphics engine 64. In one embodiment, a bus 73 can be used to couple the graphics engine 64 to the I / O subsystem 90. Alternately, a point-to-point interconnect can couple these components.

[0080] In turn, the I / O subsystem 90 can be coupled to a first bus 65 via an interface 96. In one embodiment, the first bus 65 can be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the embodiments are not so limited.

[0081] As shown in FIG. 11, various I / O devices 65a (e.g., biometric scanners, speakers, cameras, and / or sensors) can be coupled to the first bus 65, along with a bus bridge 66 which can couple the first bus 65 to a second bus 67. In one embodiment, the second bus 67 can be a low pin count (LPC) bus. Various devices can be coupled to the second bus 67 including, for example, a keyboard / mouse 67a, communication device (s) 67b, and a data storage unit 68 such as a disk drive or other mass storage device which can include code 69, in one embodiment. The illustrated code 69 can implement one or more aspects of the processes described above, including process , the process 300, the process 400, the method 500, and / or the method 600. The illustrated code 69 can be similar to the code 42 (FIG. 10) , already discussed. Further, an audio I / O 67c can be coupled to second bus 67 and a battery 61 can supply power to the computing system 60. The system 60 can implement one or more aspects of the image generation system 100, the image generator network 120, text-conditioned latent U-Net 125, and / or the U-Net network architecture 200 as described herein with reference to FIGs. 1A-1B, 2, 3A-3B, 4A-4B, 5, and / or 6.

[0082] Note that other embodiments are contemplated. For example, instead of the point-to-point architecture of FIG. 11, a system can implement a multi-drop bus or another such  communication topology. Also, the elements of FIG. 11 can alternatively be partitioned using more or fewer integrated chips than shown in FIG. 11.

[0083] Embodiments of each of the above systems, devices, components and / or methods, including the image generation system 100, the image generator network 120, text-conditioned latent U-Net 125, the U-Net network architecture 200, the process 300, the process 400, the method 500, and / or the method 600, and / or any other system components, can be implemented in hardware, software, or any suitable combination thereof. For example, hardware implementations can include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic include suitably configured PLAs, FPGAs, CPLDs, and general purpose microprocessors. Examples of fixed-functionality logic include suitably configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-functionality logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits. For example, embodiments of each of the above systems, devices, components and / or methods can be implemented via the system 10 (FIG. 8, already discussed) , the semiconductor apparatus 30 (FIG. 9, already discussed) , the processor 40 (FIG. 10, already discussed) , and / or the computing system 60 (FIG. 11, already discussed) .

[0084] Alternatively, or additionally, all or portions of the foregoing systems and / or devices and / or components and / or methods can be implemented in one or more modules as a set of program or logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., to be executed by a processor or computing device. For example, computer program code to carry out the operations of the components can be written in any combination of one or more operating system (OS) applicable / appropriate programming languages, including an object-oriented programming language such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages.

[0085] Additional Notes and Examples:

[0086] Example CA1 includes at least one computer readable storage medium comprising a set of instructions which, when executed by a computing device, cause the computing device to generate, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for each input from the calibration data set, generate, on a per layer, per  time step basis, a set of scale factors based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and store the set of scale factors as metadata associated with the trained neural network.

[0087] Example CA2 includes the at least one computer readable storage medium of Example CA1, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.

[0088] Example CA3 includes the at least one computer readable storage medium of Example CA1 or CA2, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0089] Example CA4 includes the at least one computer readable storage medium of any of Examples CA1-CA3, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0090] Example CA5 includes the at least one computer readable storage medium of any of Examples CA1-CA4, wherein the trained neural network is to be operated during inference, wherein the instructions, when executed, cause the computing device to apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.

[0091] Example CA6 includes the at least one computer readable storage medium of any of Examples CA1-CA5, wherein the trained neural network is trained to generate an image based on a text prompt.

[0092] Example CA7 includes the at least one computer readable storage medium of any of Examples CA1-CA6, wherein the calibration data set comprises a series of text prompts.

[0093] Example SA1 includes a performance-enhanced computing system comprising a processor, and memory coupled to the processor, the memory to store instructions which, when executed by the processor, cause the computing system to generate, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for each input from the calibration data set, generate, on a per layer, per time step basis, a set of scale factors  based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and store the set of scale factors as metadata associated with the trained neural network.

[0094] Example SA2 includes the system of Example SA1, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.

[0095] Example SA3 includes the system of Example SA1 or SA2, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0096] Example SA4 includes the system of any of Examples SA1-SA3, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0097] Example SA5 includes the system of any of Examples SA1-SA4, wherein the trained neural network is to be operated during inference, wherein the instructions, when executed, cause the computing system to apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.

[0098] Example SA6 includes the system of any of Examples SA1-SA5, wherein the trained neural network is trained to generate an image based on a text prompt.

[0099] Example SA7 includes the system of any of Examples SA1-SA6, wherein the calibration data set comprises a series of text prompts.

[0100] Example MA1 includes a method comprising generating, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is operated iteratively over a predetermined number of time steps for each input from the calibration data set, generating, on a per layer, per time step basis, a set of scale factors based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and storing the set of scale factors as metadata associated with the trained neural network.

[0101] Example MA2 includes the method of Example MA1, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.

[0102] Example MA3 includes the method of Example MA1 or MA2, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0103] Example MA4 includes the method of any of Examples MA1-MA3, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0104] Example MA5 includes the method of any of Examples MA1-MA4, wherein the trained neural network is operated during inference, wherein the method further comprises applying, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.

[0105] Example MA6 includes the method of any of Examples MA1-MA5, wherein the trained neural network is trained to generate an image based on a text prompt.

[0106] Example MA7 includes the method of any of Examples MA1-MA6, wherein the calibration data set comprises a series of text prompts.

[0107] Example AA1 includes a semiconductor apparatus comprising one or more substrates, and logic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable logic or fixed-functionality hardware logic, the logic to generate, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for each input from the calibration data set, generate, on a per layer, per time step basis, a set of scale factors based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and store the set of scale factors as metadata associated with the trained neural network.

[0108] Example AA2 includes the apparatus of Example AA1, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements  of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.

[0109] Example AA3 includes the apparatus of Example AA1 or AA2, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0110] Example AA4 includes the apparatus of any of Examples AA1-AA3, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0111] Example AA5 includes the apparatus of any of Examples AA1-AA4, wherein the trained neural network is to be operated during inference, wherein the logic is to apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.

[0112] Example AA6 includes the apparatus of any of Examples AA1-AA5, wherein the trained neural network is trained to generate an image based on a text prompt.

[0113] Example AA7 includes the apparatus of any of Examples AA1-AA6, wherein the calibration data set comprises a series of text prompts.

[0114] Example CB1 includes at least one computer readable storage medium comprising a set of instructions which, when executed by a computing device, cause the computing device to load a trained neural network along with a set of scale factors associated with the trained neural network, wherein the trained neural network has a plurality of layers, apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and store an output, wherein the output is generated based on operating the trained neural network using the single input.

[0115] Example CB2 includes the at least one computer readable storage medium of Example CB1, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0116] Example CB3 includes the at least one computer readable storage medium of Example CB1 or CB2, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0117] Example CB4 includes the at least one computer readable storage medium of any of Examples CB1-CB3, wherein the trained neural network is a stable diffusion network.

[0118] Example CB5 includes the at least one computer readable storage medium of any of Examples CB1-CB4, wherein the trained neural network is trained to generate images or image information based on text prompts, wherein the input comprises a text prompt, and wherein the output comprises an image or image information.

[0119] Example SB1 includes a performance-enhanced computing system comprising a processor, and memory coupled to the processor, the memory to store instructions which, when executed by the processor, cause the computing system to load a trained neural network along with a set of scale factors associated with the trained neural network, wherein the trained neural network has a plurality of layers, apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and store an output, wherein the output is generated based on operating the trained neural network using the single input.

[0120] Example SB2 includes the system of Example SB1, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0121] Example SB3 includes the system of Example SB1 or SB2, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0122] Example SB4 includes the system of any of Examples SB1-SB3, wherein the trained neural network is a stable diffusion network.

[0123] Example SB5 includes the system of any of Examples SB1-SB4, wherein the trained neural network is trained to generate images or image information based on text prompts, wherein the input comprises a text prompt, and wherein the output comprises an image or image information.

[0124] Example MB1 includes a method comprising loading a trained neural network along with a set of scale factors associated with the trained neural network, wherein the trained neural network has a plurality of layers, applying, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, wherein the trained neural network is operated iteratively over a predetermined number of time steps for  a single input, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and storing an output, wherein the output is generated based on operating the trained neural network using the single input.

[0125] Example MB2 includes the method of Example MB1, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0126] Example MB3 includes the method of Example MB1 or MB2, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0127] Example MB4 includes the method of any of Examples MB1-MB3, wherein the trained neural network is a stable diffusion network.

[0128] Example MB5 includes the method of any of Examples MB1-MB4, wherein the trained neural network is trained to generate images or image information based on text prompts, wherein the input comprises a text prompt, and wherein the output comprises an image or image information.

[0129] Example AB1 includes a semiconductor apparatus comprising one or more substrates, and logic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable logic or fixed-functionality hardware logic, the logic to load a trained neural network along with a set of scale factors associated with the trained neural network, wherein the trained neural network has a plurality of layers, apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps, and store an output, wherein the output is generated based on operating the trained neural network using the single input.

[0130] Example AB2 includes the apparatus of Example AB1, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.

[0131] Example AB3 includes the apparatus of Example AB1 or AB2, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.

[0132] Example AB4 includes the apparatus of any of Examples AB1-AB3, wherein the trained neural network is a stable diffusion network.

[0133] Example AB5 includes the apparatus of any of Examples AB1-AB4, wherein the trained neural network is trained to generate images or image information based on text prompts, wherein the input comprises a text prompt, and wherein the output comprises an image or image information.

[0134] Example AM1 includes an apparatus comprising means for performing the method of any of Examples MA1 to MA7.

[0135] Embodiments are applicable for use with all types of semiconductor integrated circuit ( “IC” ) chips. Examples of these IC chips include but are not limited to processors, controllers, chipset components, programmable logic arrays (PLAs) , memory chips, network chips, systems on chip (SoCs) , solid state drive (SSD)  / NAND drive controller ASICs, and the like. In addition, in some of the drawings, signal conductor lines are represented with lines. Some may be different, to indicate more constituent signal paths, have a number label, to indicate a number of constituent signal paths, and / or have arrows at one or more ends, to indicate primary information flow direction. This, however, should not be construed in a limiting manner. Rather, such added detail may be used in connection with one or more exemplary embodiments to facilitate easier understanding of a circuit. Any represented signal lines, whether or not having additional information, may actually comprise one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, optical fiber lines, and / or single-ended lines.

[0136] Example sizes / models / values / ranges may have been given, although embodiments are not limited to the same. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices of smaller size could be manufactured. In addition, well known power / ground connections to IC chips and other components may or may not be shown within the figures, for simplicity of illustration and discussion, and so as not to obscure certain aspects of the embodiments. Further, arrangements may be shown in block diagram form in order to avoid obscuring embodiments, and also in view of the fact that specifics with respect to implementation of such block diagram arrangements are highly dependent upon the platform within which the embodiment is to be implemented, i.e., such specifics should be well within purview of one skilled in the art. Where specific details (e.g., circuits) are set forth in order to describe example embodiments, it should be apparent to one skilled in the art that embodiments can be practiced without, or with variation of, these specific details. The description is thus to be regarded as illustrative instead of limiting.

[0137] The term “coupled” may be used herein to refer to any type of relationship, direct or indirect, between the components in question, and may apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections, including logical connections via intermediate components (e.g., device A may be coupled to device C via device B) . In addition, the terms “first” , “second” , etc. may be used herein only to facilitate discussion, and carry no particular temporal or chronological significance unless otherwise indicated.

[0138] As used in this application and in the claims, a list of items joined by the term “one or more of” may mean any combination of the listed terms. For example, the phrases “one or more of A, B or C” may mean A, B, C; A and B; A and C; B and C; or A, B and C.

[0139] Those skilled in the art will appreciate from the foregoing description that the broad techniques of the embodiments can be implemented in a variety of forms. Therefore, while the embodiments have been described in connection with particular examples thereof, the true scope of the embodiments should not be so limited since other modifications will become apparent to the skilled practitioner upon a study of the drawings, specification, and following claims.

Claims

1.At least one computer readable storage medium comprising a set of instructions which, when executed by a computing device, cause the computing device to:generate, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for each input from the calibration data set;generate, on the per layer, per time step basis, a set of scale factors based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps; andstore the set of scale factors as metadata associated with the trained neural network.2.The at least one computer readable storage medium of claim 1, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.3.The at least one computer readable storage medium of claim 1, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.4.The at least one computer readable storage medium of claim 1, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.5.The at least one computer readable storage medium of claim 1, wherein the trained neural network is to be operated during inference, wherein the instructions, when executed, cause the computing device to apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.6.The at least one computer readable storage medium of claim 1, wherein the trained neural network is trained to generate an image based on a text prompt.7.The at least one computer readable storage medium of any one of claims 1-6, wherein the calibration data set comprises a series of text prompts.8.A computing system comprising:a processor; andmemory coupled to the processor, the memory to store instructions which, when executed by the processor, cause the computing system to:generate, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for each input from the calibration data set;generate, on the per layer, per time step basis, a set of scale factors based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps; andstore the set of scale factors as metadata associated with the trained neural network.9.The system of claim 8, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.10.The system of claim 8, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.11.The system of claim 8, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.12.The system of claim 8, wherein the trained neural network is to be operated during inference, wherein the instructions, when executed, cause the computing system to apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.13.The system of claim 8, wherein the trained neural network is trained to generate an image based on a text prompt.14.The system of any one of claims 8-13, wherein the calibration data set comprises a series of text prompts.15.A method comprising:generating, on a per layer, per time step basis, feature map statistics from processing a calibration data set via a trained neural network, wherein the trained neural network has a plurality of layers, and wherein the trained neural network is operated iteratively over a predetermined number of time steps for each input from the calibration data set;generating, on the per layer, per time step basis, a set of scale factors based on the feature map statistics, wherein the set of scale factors is determined relative to a target feature map quantization, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps; andstoring the set of scale factors as metadata associated with the trained neural network.16.The method of claim 15, wherein the feature map statistics include tensor values for the respective layers and time steps, and wherein elements of the set of scale factors are determined based on a maximum absolute value of the tensor values for the respective layers and time steps.17.The method of claim 15, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.18.The method of claim 15, wherein the set of scale factors includes a series of scalar values, each scalar value corresponding to a specific layer and time step.19.The method of claim 15, wherein the trained neural network is operated during inference, wherein the method further comprises applying, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, and wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input.20.The method of claim 15, wherein the trained neural network is trained to generate an image based on a text prompt.21.The method of any one of claims 15-20, wherein the calibration data set comprises a series of text prompts.22.A semiconductor apparatus comprising:one or more substrates; andlogic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable logic or fixed-functionality hardware logic, the logic to:load a trained neural network along with a set of scale factors associated with the trained neural network, wherein the trained neural network has a plurality of layers;apply, on a per layer, per time step basis, the set of scale factors to quantize feature maps generated via operating the trained neural network, wherein the trained neural network is to be operated iteratively over a predetermined number of time steps for a single input, and wherein elements of the set of scale factors are associated with respective layers of the trained neural network and respective time steps; andstore an output, wherein the output is generated based on operating the trained neural network using the single input.23.The apparatus of claim 22, wherein the set of scale factors includes a series of tensors, each tensor corresponding to a specific layer and time step.24.The apparatus of claim 22 or 23, wherein the trained neural network is a stable diffusion network, and wherein the trained neural network is trained to generate images or image information based on text prompts, wherein the input comprises a text prompt, and wherein the output comprises an image or image information.25.An apparatus comprising means for performing the method of any of claims 15 to 20.

Citation Information

Patent Citations

  • Quantization range estimation for quantized training

    CN117136364A

  • Method and system for smooth training of a quantized neural network

    US20230306255A1