Method and apparatus for approximate numerical calculation in deep neural networks

By employing a method based on customizable bit-width lookup tables and wide vector softmax in deep neural networks, the inefficiency problem of computationally intensive and non-computationally constrained operators is addressed, improving the computational efficiency and accuracy of activation functions and achieving more efficient model performance.

CN120937019APending Publication Date: 2025-11-11INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380092494.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2023-05-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

The computationally intensive and non-computationally constrained operators of deep neural networks are inefficient to deploy in hardware-constrained scenarios, especially the element-wise operators, which have insufficient computation cycle and accuracy efficiency, leading to model performance bottlenecks.

Method used

We adopt an implementation method based on customizable bit-width lookup tables, combined with wide vector softmax operations, to reduce exponential calls and improve the computational efficiency of activation functions. Furthermore, we integrate lookup tables and standard arithmetic calculations with low-precision tensor inputs, providing higher accuracy and computational speed.

Benefits of technology

It significantly reduces computation time, improves the efficiency of activation function implementation in deep neural networks, and enhances the computational performance and accuracy of the model, especially in low-precision inference scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937019A_ABST
    Figure CN120937019A_ABST
Patent Text Reader

Abstract

Methods, apparatus, and systems for numerical calculation in approximate deep neural networks are disclosed. An example apparatus includes at least one memory, machine-readable instructions, and programmable circuitry to instantiate or execute at least one of the machine-readable instructions to generate a lookup table based on an input element, index the lookup table using a tensor value, the tensor value being associated with an output index, and output the output index. And outputting a vector in the target digital format based on the lookup table, the vector including the output value of the floating point representation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This patent claims the benefit of PCT application No. PCT / CN2023 / 085054, filed on March 30, 2023, entitled "Methods and Apparatus for Approximating Numerical Computations in Deep Neural Networks". The entire disclosure of PCT application No. PCT / CN2023 / 085054 is hereby incorporated herein by reference.

[0003] Cross-references to related applications

[0004] This disclosure relates generally to software processing, and more specifically to methods, systems, and apparatus for approximate numerical computation in deep neural networks. Background Technology

[0005] Deep neural networks (DNNs), such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), can be used to provide accurate solutions to problems associated with a wide range of fields, including image classification, speech recognition, medical diagnosis, and / or autonomous driving. The increasing size of datasets and the corresponding increase in the complexity of DNNs lead to increased computational intensity and memory requirements for deep learning-based tasks. Attached Figure Description

[0006] Figure 1 This is an example environment for performing numerical computations using example numerical computation approximator circuits based on the teachings disclosed in this article.

[0007] Figure 2 It means that it is possible Figure 1 A block diagram of a numerical computation approximator circuit implemented in the example environment.

[0008] Figure 3 This is a flowchart illustrating machine-readable instructions that can be executed to achieve the teachings disclosed herein. Figure 1 The example numerical approximator circuit performs numerical approximation.

[0009] Figure 4 This is a flowchart representing machine-readable instructions that can be executed to achieve... Figure 1 An example numerical computation approximator circuit is provided, which generates an array of elements (e.g., a lookup table) that tabulates the function output.

[0010] Figure 5 This is a flowchart representing machine-readable instructions that can be executed to achieve... Figure 1 The example numerical computation approximator circuit performs softmax-based lookup table (LUT) generation.

[0011] Figure 6 This is a flowchart representing machine-readable instructions that can be executed to achieve... Figure 1 An example numerical computation approximator circuit is provided, which performs softmax-based lookup table generation for wide vectors.

[0012] Figure 7 Is using Figure 1 The numerical calculation approximator circuit is based on Figure 3 and / or Figure 4 An example of an array of elements (lookup table) generated from machine-readable instructions.

[0013] Figure 8 Is using Figure 1 The numerical calculation approximator circuit is based on Figure 3 and / or Figure 5 An example of machine-readable instructions generated using a softmax-based lookup table.

[0014] Figures 9A-9B It shows the use of Figure 1 The numerical calculation approximator circuit is based on Figure 3 and / or Figure 6 The machine-readable instructions are an example of softmax-based lookup table generation for wide vectors.

[0015] Figure 10 Example results for two matrix sizes are shown, illustrating the performance degradation of the system when using non-linear function calls.

[0016] Figure 11 The illustrations show example tabulation and calculation identifiers, as well as example polynomial independent variable creation and average relative error based on quantized input bits, for implementations associated with the shown tabulation and calculation identifiers.

[0017] Figure 12 The illustration shows example results related to piecewise polynomials and range reduction based on d-degree approximation of quantized input data, in accordance with the teachings disclosed herein.

[0018] Figures 13A-13F The diagram illustrates the relationship between... Figure 11-12 A graphical representation of the relative error associated with the tabulated results.

[0019] Figures 14A-14BThis includes example performance data associated with the pure-eltwiseop operator exp, the eltwiseop-intensive non-computation-bound op softmax, and the GELU activation function.

[0020] Figure 15 This is a block diagram of an example processing platform based on the teachings disclosed herein, which is constructed to perform... Figure 3-6 Instructions to achieve Figure 1 The example numerical approximator circuit performs numerical approximation.

[0021] Figure 16 yes Figure 15 A block diagram of an example implementation of the processor circuit.

[0022] Figure 17 yes Figure 15 A block diagram of another example implementation of the processor circuit.

[0023] Figure 18 This is a block diagram of an example software distribution platform (e.g., one or more servers) used to distribute software (e.g., with...) Figure 3 The software corresponding to the example machine-readable instructions is distributed to client devices associated with end users and / or consumers (e.g., for licensing, selling and / or using), retailers (e.g., for selling, reselling, licensing and / or sublicensing), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products to be distributed to, for example, retailers and / or other end users such as direct purchase customers).

[0024] Generally, the same reference numerals will be used throughout the accompanying drawings and written description to refer to the same or similar parts. The drawings are not drawn to scale. Unless otherwise specified, descriptors such as “first,” “second,” “third,” etc., are used herein not to indicate priority, physical order, arrangement in a list, and / or any sorting, but solely as labels and / or arbitrary names to distinguish elements for ease of understanding of the disclosed examples. In some examples, the descriptor “first” may be used to refer to an element in the detailed description, while the same element may be referred to in the claims by different descriptors such as “second” or “third.” In such cases, it should be understood that such descriptors are only used to clearly identify those elements that may, for example, otherwise share the same name.

[0025] As used herein, the phrase “communication” (including its variations) covers direct communication and / or indirect communication via one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals, predetermined intervals, non-periodic intervals and / or one-off events.

[0026] As used herein, “processor circuitry” is defined to include: (i) one or more dedicated circuits configured to perform one or more specific operations and comprising one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors), and / or (ii) one or more general-purpose semiconductor-based circuits programmable by instructions to perform specific operations and comprising one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuitry include programmable microprocessors, field-programmable gate arrays (FPGAs) that can instantiate instructions, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), XPUs, or microcontrollers and integrated circuits such as application-specific integrated circuits (ASICs). For example, an XPU may be implemented by a heterogeneous computing system comprising multiple types of processor circuitry (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc., and / or combinations thereof) and one or more application programming interfaces (APIs) that can assign one or more computational tasks to the processor circuitry of the multiple types most suitable for performing the computational tasks. Detailed Implementation

[0027] Deep neural networks (DNNs) have enabled state-of-the-art accuracy for a wide range of tasks. However, the adoption of DNNs in some domains is hampered by lengthy and complex training processes and the high computational cost of the inference phase associated with deep learning. For example, while the training phase represents the process by which a machine learns and optimizes a model from data, the inference phase represents the use of the trained model to predict and / or estimate results from new observations in an efficient deployment. The inference phase of deep learning allows machines to bridge the gap between the data used for training and ambiguous information in the real world (e.g., extrapolation of features identified by DNNs). Most DNN models include two types of operators: (1) computationally intensive operators involving matrix-matrix multiplication (e.g., GEMM) and convolution (CONV), and (2) non-computationally constrained operators including element-wise operators (e.g., GELU, ReLU, etc.), element-wise reduction operators, etc. Although these non-computationally constrained operators have a much lower cost compared to CONV / GEMM kernels, their efficiency in terms of computational cycles and accuracy is crucial for the efficient deployment of DNN models in hardware-constrained scenarios.

[0028] In some examples, using lookup tables to approximate numerical values ​​is particularly noteworthy for scenarios where a small number of input bits can significantly impact the results of complex computations. In such cases, the entire computational data path can be replaced by a single lookup table driven by a small number of input bits, which may consist of several high-cost operators (e.g., basic or transcendental functions). For example, when low-precision tensor values ​​(e.g., int8 / bfloat8, 4-bit representation data types, etc.) are input to certain nonlinear computations (e.g., activation functions, softmax kernels, etc.), a single lookup table can replace the computational data path. In some examples, non-computational operators may present a critical bottleneck to model performance due to the difficulty in fusing these operators. While some architectures can maintain support for certain element-wise computational capabilities (e.g., those required to implement element-wise nonlinear functions), supporting higher-density, low-precision (e.g., int8, bf16, fp16, bf8) tensor / vector product computation capabilities is desirable when significant computational acceleration is expected for the computationally intensive parts of DNNs (e.g., including CONV and GEMM). Existing matrix acceleration techniques (e.g., AMX or "Tensor Cores") typically introduce registers different from those used by vector units (e.g., AVX512, CUDA Cores, etc.). Incorporating non-computationally constrained operators into a previous CONV / GEMM layer can benefit from transferring data between matrix and vector / core register stacks and / or adding additional input / output pressure.

[0029] The methods and apparatus disclosed herein facilitate the use of customizable bit-width lookup-based implementations in element-wise operator-intensive operators (e.g., eltwseop). Compared to existing techniques, the proposed techniques increase the range of the lookup table, resulting in higher accuracy and providing more robust representational capabilities. In the examples disclosed herein, implementations based on customizable bit-width lookup tables can be used to output values ​​directly in floating-point representations (such as FP16 (e.g., half-precision) or bfloat16). In the examples disclosed herein, several lookup-based implementations can be mixed with standard arithmetic computations used in DNN models for nonlinear operator implementations. Furthermore, in the examples disclosed herein, a novel softmax implementation specifically for wide vectors is introduced, which significantly reduces the number of exponentiation calls, thereby significantly reducing the total computation time. Additionally, in the examples disclosed herein, a novel polynomial-based implementation (i) takes tensor-formatted data as input and outputs floating-point data, and (ii) provides a trade-off between lookup cost (e.g., bit width), computation (e.g., using multiply-accumulate calls), and accuracy. In some examples, lookup reductions of 93% and 75% were achieved at the cost of a maximum of three multiply-accumulate calls, compared to lookup-only solutions. In the examples disclosed herein, an exponential function-specific implementation is introduced that (i) takes tensor-formatted data as input, (ii) proposes dequantization range reduction capabilities, and (iii) provides a trade-off between the cost of creating a lookup table (LUT), computational cost, and parallelization.

[0030] Figure 1 This is an example environment 100 that performs container verification using the example container verification circuit 110 according to the teachings disclosed herein. In some examples, the lookup-based implementation uses input tensor values ​​as table indices. For example, if the input type is unsigned int8, the array is constructed of 256 elements, indexed by values ​​in the range {0, ..., 255} (e.g., where the input tensor value i corresponds to the dequantized value j). The array entry for index i will correspond to the function output evaluated at point j. Figure 1 This section provides a detailed example of calling the same lookup table-based implementation for all elements of a C×H×W structure. Figure 1In the example, input 105 with a C×H structure is provided to numerical computation approximator circuit 110, which generates a lookup table (LUT) 115, thereby providing output 125 with a C×H×W structure. When using some known systems (e.g., AMX), tensors can be stored in a cache from a first type of register (e.g., a TMM register) and then reloaded into a second type of register (e.g., a ZMM register), and cache bottlenecks can significantly reduce the fusion benefits. The same occurs in the case of GPUs; for example, tensor cores in GPUs use different registers than those in Computational Unified Device Architecture (CUDA) cores, thus the fusion benefits are greatly affected by shared graphics memory bandwidth. Figure 10 As shown in the example, Table 1000 presents the performance measured (e.g., reported in gigafloat-per-second (GFLOPS)) when using different nonlinear functions for each output element on an SPR kernel (e.g., performing a rank-1 update of a symmetric packing matrix). Table 1000 reports numbers for two matrix sizes (N, P, Q) with A(N, P) and B(P, Q), where Size[0] = (256, 768, 768) and Size[1] = (256, 2048, 1024). The results demonstrate just how much these nonlinear function calls degrade system performance. Therefore, it is necessary to find efficient ways to implement these activation functions. The methods and apparatuses disclosed herein improve the efficiency of activation function implementations for numerical approximations in deep neural networks.

[0031] In some known systems, DNN libraries (e.g., oneDNN) can directly accept low-precision (e.g., int8) tensors as input for element-wise op-intensive operations, but internally convert them to a wider bit-width representation (e.g., softmax) or insert dequantized nodes before the op-wise op-intensive operations. Similarly, lookup-based methods have been extensively explored but rarely deployed. TensorFlow optimization tools introduced an 8-bit LUT focused on pure-eltwise op-chain operations that take int8 as input and output, but it has stalled as an experimental API. For example, oneDNN typically converts low-precision inputs to higher-precision data types, which prevents the gains from low precision and introduces additional costs for data type conversion. Using this approach prevents the full utilization of the hardware's low-precision arithmetic capabilities. Since future hardware will bring greater computational power to low precision, acceleration can be improved. For example, traditional LUTs focus on full int8 inference / training and pure eltwise operation, but cannot cover bf16 / bf8 with even lower precision floating-point representations, which may be required by (one or more) subsequent operators.

[0032] In the examples disclosed in this paper, a LUT based on the softmax function is implemented. For example, during the inference phase of deep learning, the final network layer may contain a call to the softmax function. The softmax function is used to calculate the probability (e.g., a value between 0 and 1) that a network input is classified into a bin associated with a particular network output. The mathematical definition of the softmax function is shown in Equation 1 below:

[0033]

[0034] In the example of Equation 1, the output vector Each element depends on the input vector All elements of the input vector. Equation 1 shows how to scale each output vector value using a common denominator (e.g., the sum of the exponents of the input elements of the input vector Z). Each element of the output vector is computed as element z. i The ratio between the exponential value and the common exponent and scaling factor. Therefore, the computation of the softmax function can become a bottleneck due to all exponential function calls. During the training phase of deep learning, some activation functions and their derivatives are also used (e.g., for stochastic gradient descent training methods). The GELU activation function is an example of such an example, as shown in Equation 2:

[0035]

[0036] In some examples, non-computationally restricted operators are pure element-wise operators applied to all elements of a tensor (e.g., eltwiseop), or eltwiseop-dense suboperators included in the inference and training phases. Accelerating the computation of these operators in the context of low-precision inference is crucial for model acceleration and compression. The methods and apparatuses disclosed in this paper accelerate element-wise operator-dense non-computationally restricted DNN operators while providing sufficient accuracy for these models. With increased computational power set to reduce time spent in the computationally restricted parts of the code, reducing the time required for odd-numbered element-wise arithmetic operators will significantly improve numerical computation methods for deep neural networks.

[0037] Figure 2 yes Figure 2 Block diagram 200 of an example implementation of the container verification circuit 110. Figure 2 The container verification circuit 110 can be instantiated by a processor circuit, such as a central processing unit that executes instructions (e.g., creating an instance of it, generating it over any time period, materializing it, implementing it, etc.). Additionally or alternatively, Figure 2 The container verification circuit 110 can be instantiated by an ASIC or FPGA configured to perform operations corresponding to instructions (e.g., creating an instance of it, generating it over any time period, materializing it, implementing it, etc.). It should be understood that... Figure 2 Some or all of the circuitry can therefore be instantiated at the same or different times. Some or all of the circuitry can be instantiated, for example, in one or more threads that execute concurrently on the hardware and / or serially on the hardware. Furthermore, in some examples, Figure 2 Some or all of the circuitry can be implemented by microprocessor circuitry that executes instructions to implement one or more virtual machines and / or containers.

[0038] exist Figure 2 In the example, the numerical computation approximator circuit 110 includes an example input recognizer circuit 202, an example lookup table (LUT) generator circuit 204, an example index generator circuit 206, an example output recognizer circuit 208, an example softmax LUT determiner circuit 210, an example wide vector softmax LUT determiner circuit 212, and an example data storage device 214. Figure 2 In the example, the input recognizer circuit 202, the lookup table (LUT) generator circuit 204, the index generator circuit 206, the output recognizer circuit 208, the softmax LUT determiner circuit 210, the wide vector softmax LUT determiner circuit 212, and / or the data storage device 214 communicate via the example bus 220.

[0039] Input recognizer circuit 202 determines the inputs, such as the composite function set (S) used for generation, input type (IT), quantization scheme parameters (Q), and / or output data type (OT). In some examples, input recognizer circuit 202 recognizes a lookup table (LUT) as the input and / or source vector (src-vec). For example, as combined with... Figure 4 , Figure 5 and / or Figure 6 In more detail, the input recognizer circuit 202 retrieves one or more input parameters for generating an array of elements (e.g., a LUT) that tabulates the function output, for performing softmax-based LUT generation, and / or for performing softmax-based LUT generation for wide vectors.

[0040] Lookup table (LUT) generator circuit 204 generates a lookup table. For example, LUT generator circuit 204 generates a lookup table based on the composite function set (S), input type (IT), quantization scheme parameters (Q), and / or output data type (OT) used for generation. For example, LUT generator circuit 204 performs variable setting and identifies LUT indices. In some examples, LUT generator circuit 204 outputs a lookup table for the composite function set (S) and optional offset arguments, such as combining... Figure 4 and / or Figure 7 More detailed description.

[0041] Index generator circuit 206 performs indexing of the lookup table generated by LUT generator circuit 204. In some examples, index generator circuit 206 indexes the lookup table(s) using tensor values, such as combining... Figure 4 and / or Figure 7 A more detailed description is provided. For example, tensor values ​​in the source vector can be associated with a specific output index. In some examples, the index generator circuit 206 determines sub-range indices from the quantized input by analyzing a given number of bits, such as in combination with... Figure 10 As described by A.

[0042] The output recognizer circuit 208 determines the output to be generated based on the provided input(s)(s). For example, the output recognizer circuit 208 generates a lookup table(T) as the output of a composite function set(S). In some examples, such as combining... Figure 4 and / or Figure 7 In more detail, the output recognizer circuit 208 generates an optional offset argument (O). In some examples, the output recognizer circuit 208 generates the target vector based on the exponential value, such as in combination with... Figure 12 More detailed description.

[0043] The softmax LUT determinant circuit 210 generates a hybrid LUT algorithm, such as combining... Figure 4 As described. For example, the softmax LUT determiner circuit 210 divides the operator into two parts, allowing certain operations (e.g., "exp", "log" calculations) to be performed based on the LUT for all inputs, and other "multiplication" and "division" operations to be performed in numerical methods. In some examples, the softmax LUT determiner circuit 210 determines the integer equivalent values ​​of the input type format boundaries, identifies one or more dequantized true values, and iterates over a set of composite functions using one or more dequantized values ​​to generate a lookup table, such as in combination. Figure 4 More detailed description.

[0044] The wide-vector softmax LUT determiner circuit 212 performs a softmax-based operation to generate a lookup table for the wide vector. For example, the wide-vector softmax LUT determiner circuit 212 can be used to directly output individual values ​​in the target numeric format, such as when a downstream operator is decreasing, as in combination with... Figure 3 , Figure 6 , Figure 9A and Figure 9B As described. In some examples, the wide vector softmax LUT determiner circuit 212 determines whether the size of the input vector is greater than the lookup address size. In some examples, the wide vector softmax LUT determiner circuit 212 generates an output consisting of a vector in the target numeric format. For example, the wide vector softmax LUT determiner circuit 212 can determine the exponent value of the target vector format based on the identified softmax denominator(s) and / or the reciprocal of the denominator(s), such as in combination with... Figure 6 As described.

[0045] Data storage device 214 can be used to store any information associated with input identifier circuit 202, lookup table (LIT) generator circuit 204, index generator circuit 206, output identifier circuit 208, softmax LUT determiner circuit 210 and / or wide vector softmax LUT determiner circuit 212. Figure 2 The example data storage device 214 shown can be implemented by any memory, storage device, and / or storage disk (such as flash memory, magnetic media, optical media, etc.) used for storing data. Furthermore, the data stored in the example data storage device 214 can be in any data format, such as binary data, comma-separated data, tab-separated data, Structured Query Language (SQL) structures, image data, etc.

[0046] In some examples, the device includes components for recognizing input. For example, the components for recognizing input may be implemented by an input recognizer circuit 202. In some examples, the input recognizer circuit 202 may be composed of, for example, Figure 15 The example processor circuit 1512 is instantiated. For example, the input recognizer circuit 202 can be instantiated by executing functions such as those described by... Figure 3 At least those machine-executable instructions implemented in box 305. Figure 16 The example microprocessor 1600 is used for instantiation. In some examples, the input recognizer circuit 202 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 17 The input recognizer circuit 202 may be implemented by an ASIC, XPU, or FPGA circuit 1700. Additionally or alternatively, the input recognizer circuit 202 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the input recognizer circuit 202 may be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, XPU, comparator, operational amplifier, logic circuitry, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0047] In some examples, the apparatus includes components for generating a lookup table (LUT). For example, the components for generating the lookup table may be implemented by a lookup table (LUT) generator circuit 204. In some examples, the lookup table (LUT) generator circuit 204 may be implemented by a processor circuit (such as...) Figure 15 Example processor circuit 1512) instantiates. For example, lookup table (LUT) generator circuit 204 can be instantiated via... Figure 16 The example microprocessor 1600 executes machine-executable instructions (such as those from at least...). Figure 3 The lookup table (LUT) generator circuit 204 can be instantiated by those implemented in box 310. In some examples, the lookup table (LUT) generator circuit 204 can be instantiated by hardware logic circuitry that is configured to perform operations corresponding to machine-readable instructions. Figure 17The ASIC, XPU, or FPGA circuitry 1700 is used for implementation. Additionally or alternatively, the lookup table (LUT) generator circuitry 204 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the lookup table (LUT) generator circuitry 204 can be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, XPU, comparator, operational amplifier, logic circuitry, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware; however, other configurations are equally suitable.

[0048] In some instances, the apparatus includes components for generating indexes. For example, the components for generating indexes may be implemented by index generator circuitry 206. In some examples, index generator circuitry 206 may be implemented by processor circuitry (such as...) Figure 15 Example processor circuit 1512) instantiated. For example, index generator circuit 206 can be instantiated via... Figure 16 Example microprocessor 1600 performs tasks such as those performed by Figure 3 At least those machine-executable instructions implemented in box 315 are instantiated. In some examples, the index generator circuit 206 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 17 The index generator circuit 206 can be implemented using an ASIC, XPU, or FPGA circuit 1700. Additionally or alternatively, the index generator circuit 206 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the index generator circuit 206 can be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, XPU, comparators, operational amplifiers, logic circuitry, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0049] In some instances, the device includes components for identifying the output. For example, the components for identifying the output may be implemented by an output identifier circuit 208. In some examples, the output identifier circuit 208 may be implemented by, for example,... Figure 15 The example processor circuit 1512 is instantiated. For example, the output recognizer circuit 208 can be instantiated via... Figure 16 Example microprocessor 1600 performs tasks such as those performed by Figure 3At least those machine-executable instructions implemented in block 320 are instantiated. In some examples, the output recognizer circuit 208 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 17 The output recognizer circuit 208 may be implemented by an ASIC, XPU, or FPGA circuit 1700. Additionally or alternatively, the output recognizer circuit 208 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the output recognizer circuit 208 may be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, XPU, comparator, operational amplifier, logic circuitry, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0050] In some examples, the apparatus includes components for determining the softmax (LUT). For example, the components for determining the softmax LUT may be implemented by a softmax LUT determiner circuit 210. In some examples, the softmax LUT determiner circuit 210 may be implemented by, for example... Figure 15 The example processor circuit 1512 is instantiated. For example, the softmax LUT determiner circuit 210 can be instantiated by... Figure 16 The example microprocessor 1600 executes machine-executable instructions (such as those from at least...). Figure 3 The softmax LUT determiner circuit 210 can be instantiated using those implemented in box 335. In some examples, the softmax LUT determiner circuit 210 can be instantiated by hardware logic circuitry that is configured to perform operations corresponding to machine-readable instructions. Figure 17 The softmax LUT determiner circuit 210 may be implemented by an ASIC, XPU, or FPGA circuit 1700. Additionally or alternatively, the softmax LUT determiner circuit 210 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the softmax LUT determiner circuit 210 may be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, XPU, comparator, operational amplifier, logic circuitry, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0051] In some examples, the apparatus includes components for determining a wide vector softmax LUT. For example, the components for determining the wide vector softmax LUT can be implemented by a wide vector softmax LUT determiner circuit 212. In some examples, the wide vector softmax LUT determiner circuit 212 can be implemented by, for example... Figure 15 The example processor circuit 1512 is instantiated. For example, the wide vector softmax LUT determiner circuit 212 can be instantiated. Figure 16 The example microprocessor 1600 executes machine-executable instructions (such as those from at least...). Figure 3 The wide vector softmax LUT determiner circuit 212 can be instantiated by those implemented in box 340. In some examples, the wide vector softmax LUT determiner circuit 212 can be instantiated by hardware logic circuitry that is configured to perform operations corresponding to machine-readable instructions. Figure 17 The ASIC, XPU, or FPGA circuit 1700 is implemented. Additionally or alternatively, the wide-vector softmax LUT determiner circuit 212 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the wide-vector softmax LUT determiner circuit 212 can be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, XPU, comparators, operational amplifiers, logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.

[0052] Although Figure 2 The diagram illustrates the implementation. Figure 1 The numerical approximator circuit 110 is an example of such a circuit, but... Figure 2 One or more of the elements, processes, and / or devices shown can be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. Furthermore, the example input recognizer circuit 202, the example lookup table (LIT) generator circuit 204, the example index generator circuit 206, the example output recognizer circuit 208, the example softmax LUT determiner circuit 210, the example wide vector softmax LUT determiner circuit 212, and / or more generally, Figure 1The example numerical computation approximator circuit 110 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, there are example input recognizer circuits 202, example lookup table (LIT) generator circuits 204, example index generator circuits 206, example output recognizer circuits 208, example softmax LUT determiner circuits 210, example wide vector softmax LUT determiner circuits 212, and / or more generally... Figure 1 Any of the example numerical computation approximator circuits 110 can be implemented by processor circuitry, analog circuitry, digital circuitry, logic circuitry, programmable processors, programmable microcontrollers, graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field-programmable logic devices (FPLDs) (such as field-programmable gate arrays (FPGAs)). Furthermore, additional to or replacing... Figure 2 Those shown, Figure 1 The example numerical computation approximator circuit 110 may include one or more elements, processes and / or devices, and / or may include more than one of any or all of the elements, processes and devices shown.

[0053] Figure 3-6 The diagram shows a flowchart illustrating example machine-readable instructions that can be executed to configure processor circuitry to implement... Figure 1 The numerical computation approximator circuit 110. Machine-readable instructions may be one or more executable programs or portions thereof for execution by processor circuitry, such as those described below. Figure 15 The processor circuitry 1512 shown in the example processor platform 1500 discussed below and / or in conjunction with the processor circuitry 1512 shown below. Figure 16 and / or Figure 17The discussion focuses on example processor circuitry. Programs may be embodied in software stored on one or more non-transitory computer-readable storage media associated with processor circuitry located in one or more hardware devices, such as optical discs (CDs), floppy disks, hard disk drives (HDDs), solid-state drives (SSDs), digital versatile discs (DVDs), Blu-ray discs, volatile memory (e.g., any type of random access memory (RAM), etc.) or non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), FLASH memory, HDDs, SSDs, etc.). However, the entire program and / or portions thereof may alternatively be executed by one or more hardware devices other than the processor circuitry and / or embodied in firmware or dedicated hardware. Machine-readable instructions may be distributed across multiple hardware devices and / or executed by two or more hardware devices (e.g., server and client hardware devices). For example, client hardware devices may be implemented by endpoint client hardware devices (e.g., hardware devices associated with a user) or intermediate client hardware devices (e.g., radio access network (RAN)) gateways that facilitate communication between the server and endpoint client hardware devices. Similarly, non-transitory computer-readable storage media may include one or more media located in one or more hardware devices. Furthermore, although references... Figure 3-6 The flowchart shown illustrates the example program, but an implementation can be used instead. Figure 1 Many other methods exist for the example numerical computation approximator circuit 110. For example, the execution order of the blocks can be changed, and / or some of the blocks described can be altered, eliminated, or combined. Additionally or alternatively, any or all blocks can be implemented by one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, ASIC, comparator, operational amplifier, logic circuitry, etc.) configured to perform the corresponding operation without executing software or firmware. The processor circuitry can be distributed across different network locations and / or local to one or more hardware devices (e.g., a single-core processor (e.g., a single-core central processing unit (CPU), a multi-core processor in a single machine (e.g., a multi-core CPU, XPU, etc.), multiple processors distributed across multiple servers in a server rack, multiple processors distributed across one or more server racks, a CPU and / or FPGA located in the same package (e.g., the same integrated circuit (IC) package or two or more separate housings, etc.)).

[0054] The machine-readable instructions described herein can be stored in one or more of the following formats: compressed format, encrypted format, segmented format, compiled format, executable format, and packaged format. Machine-readable instructions as described herein can be stored as data or data structures (e.g., as part of instructions, code, code representations, etc.) that can be used to create, manufacture, and / or produce machine-executable instructions. For example, machine-readable instructions can be segmented and stored on one or more storage devices and / or computing devices (e.g., servers) located in the same or different locations (e.g., in the cloud, on edge devices, etc.) within a network or network set. Machine-readable instructions may require one or more of the following to be installed, modified, adapted, updated, combined, supplemented, configured, decrypted, decompressed, unpacked, distributed, redistributed, compiled, etc., so that they can be directly read, interpreted, and / or executed by computing devices and / or other machines. For example, machine-readable instructions can be stored in multiple parts that are individually compressed, encrypted, and / or stored on separate computing devices, wherein these parts, when decrypted, decompressed, and / or combined, form a set of machine-executable instructions that implement one or more operations, which together can form a program such as that described herein.

[0055] In another example, machine-readable instructions may be stored in a state where they can be read by processor circuitry, but libraries (e.g., dynamic link libraries (DLLs)), software development kits (SDKs), application programming interfaces (APIs), etc., need to be added to enable the machine-readable instructions to be executed on a specific computing device or other device. In yet another example, the machine-readable instructions (e.g., storage settings, data input, recorded network addresses, etc.) may need to be configured before they can be fully or partially executed. Therefore, as used herein, machine-readable media may include machine-readable instructions and / or / one or more programs, regardless of the specific format or state in which they are stored or otherwise stored or transmitted.

[0056] The machine-readable instructions described in this article can be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, machine-readable instructions can be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0057] As mentioned above, Figure 3-6Example operations can be implemented using executable instructions (e.g., computer and / or machine-readable instructions) stored on one or more non-transitory computer and / or machine-readable media, such as optical storage devices, magnetic storage devices, HDDs, flash memory, read-only memory (ROM), CDs, DVDs, caches, any type of RAM, registers, and / or any other storage device or disk in which information is stored for any duration (e.g., extended time period, permanent, transient, temporary buffered, and / or cached information). As used herein, the terms non-transitory computer-readable medium, non-transitory computer-readable storage medium, non-transitory machine-readable medium, and non-transitory machine-readable storage medium are explicitly defined to include any type of computer-readable storage device and / or disk, excluding propagating signals and transmission media. As used herein, the terms "computer-readable storage device" and "machine-readable storage device" are defined to include any physical (mechanical and / or electrical) structure for storing information, but excluding propagating signals and transmission media. Examples of computer-readable and machine-readable storage devices include any type of random access memory, any type of read-only memory, solid-state memory, flash memory, optical disk, hard disk, disk drive, and / or redundant array of independent disks (RAID) system. As used herein, the term "device" refers to a physical structure, such as mechanical and / or electrical equipment, hardware and / or circuitry, which may or may not be configured to, and / or manufactured to, execute computer-readable instructions, machine-readable instructions, etc.

[0058] "Comprising" and "including" (and all their forms and tenses) are used herein as open-ended terms. Therefore, whenever a claim uses any form of "comprising" or "including" (e.g., including, comprising, having, etc.) as a preamble or in any kind of claim statement, it should be understood that additional elements, terms, etc., may be present without exceeding the scope of the corresponding claim or statement. As used herein, when the phrase "at least" is used as a transitional term, for example, in the preamble of a claim, it is open-ended in the same way that the terms "comprising" and "including" are open-ended. When used, for example, in the form of A, B, and / or C, the term "and / or" refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, or (7) A and B and C. As used herein in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A and B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A or B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein in the context of describing the execution or operation of processes, instructions, actions, activities, and / or steps, the phrase “at least one of A and B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing the execution or operation of a process, instruction, action, activity and / or step, the phrase “at least one of A or B” is intended to refer to the implementation of any of the following: (1) at least one A, (2) at least one B or (3) at least one A and at least one B.

[0059] As used herein, singular references (e.g., "a," "an," "first," "second," etc.) do not exclude plurals. As used herein, the term "a" or "an" refers to one or more of that object. The terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein. Furthermore, although listed separately, multiple components, elements, or methodological actions can be implemented by, for example, the same entity or object. Additionally, although individual features may be included in different examples or claims, these features may be combined, and inclusion in different examples or claims does not imply that the combination of features is impractical and / or disadvantageous.

[0060] Figure 3This indicates that it can be executed and / or instantiated by processor circuitry to achieve [the desired result]. Figure 1 The flowchart shows an example of machine-readable instructions and / or operations 300 for an example numerical computation approximator circuit 110. Figure 3 The machine-readable instructions and / or operations 300 begin at block 305, where the input recognizer circuit 202 recognizes the input element 305. (Example) Figure 3 In the middle, the lookup table (LUT) generator circuit 204 generates an element array (lookup table) tabulation function output (box 310), such as combined with Figure 4 As described. In some examples, the index generator circuit module 206 indexes the lookup table using tensor values ​​(box 315). In some examples, the output recognizer circuit 208 determines whether the output value should be output directly in the target numeric format (box 320), for example, when a downstream operator is performing a reduction operation. If the output recognizer circuit 208 determines that the output value does not need to be in the target numeric format, the lookup table (LUT) generator circuit 204 outputs a lookup table for the composite function set (S) (box 325). In some examples, the input recognizer circuit 202 determines that the size of the input vector is greater than the lookup address size (box 330). If the size of the input vector is not greater than the lookup address size, the softmax LUT determiner circuit 210 performs softmax-based LUT generation (box 335), as in combination with Figure 5 As stated above. If the size of the input vector is greater than the lookup address size, the wide vector softmax LUT determiner circuit 212 performs softmax-based LUT generation for the wide vector (box 340), as combined with... Figure 6 The wide vector softmax LUT determiner circuit 212 outputs the vector in the target digital format (box 345).

[0061] Figure 4 This indicates that it can be executed and / or instantiated by processor circuitry to achieve [the desired result]. Figure 2 The flowchart of example machine-readable instructions and / or operations 310 for example lookup table generator circuit 204. Figure 4 Machine-readable instructions and / or operations 310 begin at box 405, where the lookup table generator circuit 204 retrieves one or more input parameters (e.g., a set of composite functions (S), an input type (IT), etc.) used to generate the lookup table. The lookup table generator circuit 204 determines one or more integer equivalents of the input type (IT) format boundaries based on the input parameters (one or more) (box 410). In some examples, such as in combination... Figure 7As described, function calls such as `get_precise_{min|max}_bin` are used to determine one or more equivalent values ​​of the integer. In some examples, the lookup table generator circuit 204 uses type information (e.g., input type) (e.g., using the function `get_dequantized_real(i,IT,Q)`) to determine one or more dequantized real values ​​to convert the binary representation of `i` (box 415). For example, quantization scheme information associated with the variable `Q` can be used to determine one or more dequantized real values. The lookup table generator circuit 204 uses the determined dequantized real values ​​to iterate over the set of composite functions (S) to compute the output of each function (box 420). While not all input points have been processed (box 425), the lookup table generator circuit 204 continues to iterate over the set of composite functions. Once all input points have been processed, the lookup table generator circuit 204 outputs a lookup table with offset input values ​​(box 430), as combined with... Figure 7 More detailed description.

[0062] Figure 5 This indicates that it can be executed and / or instantiated by processor circuitry to achieve [the desired result]. Figure 2 The flowchart of example machine-readable instructions and / or operations 335 executed by the example softmax LUT determiner circuit 210. Figure 5 Machine-readable instructions and / or operations 335 begin at box 505, where the softmax LUT determiner circuit 210 retrieves one or more relevant inputs (e.g., the generated lookup table). In some examples, the softmax LUT determiner circuit 210 performs a maximum reduction on one or more input tensors (box 510). For example, a maximum reduction can be performed on an 8-bit tensor, which combines... Figure 8 As shown. In some examples, the softmax LUT determiner circuit 210 performs a lookup-based exponential implementation to avoid the complex data paths associated with the numerical computation of the softmax function (box 515). The softmax LUT determiner circuit 210 further performs summation reduction (e.g., producing a 16-bit width for the reduced elements using a single instruction) (box 520). Once all input points have been processed (box 525), the softmax LUT determiner circuit 210 outputs a softmax-based lookup table (box 530), combined with... Figure 8 As shown.

[0063] Figure 6 This indicates that it can be executed and / or instantiated by processor circuitry to achieve [the desired result]. Figure 2 The flowchart of example machine-readable instructions and / or operations 340 for example wide vector softmaxLUT determiner circuit 212. Figure 6Machine-readable instructions and / or operations 340 begin at box 605, where the wide vector softmax LUT determiner circuit 212 identifies the source vector input. In some examples, the wide vector softmax LUT determiner circuit 212 identifies compile-time constants (box 610). For example, the length of the input vector may be identified as a constant and included in the list of compile-time constraints. Subsequently, the wide vector softmax LUT determiner circuit 212 determines the frequency of each input vector element within the input vector (box 615), as in combination with... Figures 9A-9B As shown and described. In some examples, the wide vector softmax LUT determiner circuit 212 locates the maximum element of the input vector (box 620). For example, iterating backward over a previously filled array of counting vectors and stopping at the first non-zero element can be used to identify the maximum value of the input array. The wide vector softmax LUT determiner circuit 212 uses coarse-scale normalization to perform scaling to check that no exponent value is too large to cause overflow, where a table lookup is used to obtain one or more scaling values ​​(box 625). In some examples, the wide vector softmax LUT determiner circuit 212 uses subsequent multiplication to determine subsequent exponent values ​​(box 630), such as in combination with Figures 9A-9B In more detail, the wide vector softmax LUT determiner circuit 212 determines the softmax denominator and the reciprocal of the denominator output (box 635) (e.g., based on a dot product call between a new scaled exponent vector and a corresponding count vector). Based on the softmax denominator and the reciprocal of the denominator output, the wide vector softmax LUT determiner circuit 212 also determines the exponent value in the target vector (box 640), which is used to output the vector in the target numeric format.

[0064] Figure 7 It is based on Figure 3 and Figure 4 Machine-readable instructions 310 are used Figure 2 An example of an array of elements (e.g., a lookup table) generated by the lookup table generator circuit 204. Figure 7 In the example, lookup table generator circuit 204 receives input at block 705 and identifies the output at block 710. Lookup table generator circuit 204 performs variable settings based on the input, then determines the dequantized true value at block 720 and iterates over the set of composite functions. Lookup table generator circuit 204 in... Figure 7 The lookup table generation process is completed at box 725 by assigning the calculated values ​​to the corresponding table entries, as described in more detail below.

[0065] For example, a lookup table-based technique can be used to map the input and output elements of a computation. With 8-bit input elements (e.g., corresponding to a typical tensor bit width), a 2^8 = 256-element array can be constructed to tabulate the function's output. This table is then indexed using tensor values ​​(e.g., 8-bit integers). Figure 7 An example algorithm for a lookup table generation process, executed by lookup table generator circuit 204, is presented. For example, the algorithm input S is the set of functions to be composed to compute a given expression. For example, given the function g(x) = log( 3 / 4+exp(1-x)), the set of functions stored in S for function g is S={1-x,exp(x), 3 / 4+x,log(x)}, where S[0]=1-x, S[1]=exp(x), etc. The algorithm also inputs IT, which represents the quantization input type, representing anything from INT4 to UINT8 and up to FP16. Next to IT, the algorithm inputs Q, which represents the quantization scheme information. In some examples, for affine transformations, this information can consist of scaling parameters and zero offset parameters. Finally, the algorithm inputs the output type OT. This information is used to encode the function output value in the desired format.

[0066] exist Figure 7In the example, lookup table generator circuit 204 starts two variables, m and M (lines 02-03). The function call get_precise_{min|max}_bin is used to identify the integer equivalent values ​​of the IT format boundaries. For INT8, although the supported range is [-128, 127], the values ​​of the variables can be m = 0 and M = 255, while for bfloat16, the values ​​can be m = 0 and M = (2^16) - 1. In practice, additional knowledge that allows tightening these boundaries can be obtained. Thus, for INT8, iterations can be performed from m = 47 to M = 212. This can happen when a symmetric quantization scheme is used, where a portion of the output mapping is not mapped to valid (e.g., reachable) inputs. For example, the range [-2, 4] can be mapped to [-128, 128] using a symmetric (e.g., zero-reserved) mapping. However, due to the input range limitation, the value [-128, -64) that would map to [-4, -2) will not be used. Subsequently, at lines 08-20, the lookup table generator circuit 204 performs iterations on the valid entries of the numerical values ​​corresponding to the quantization format. At line 09, the dequantization of the value is initiated. The lookup table generator circuit 204 uses a function (e.g., get_dequantized_real(i,IT,Q)) to convert the binary representation of i using the type information IT. Then, using the quantization scheme information provided in Q, the lookup table generator circuit 204 determines the dequantized value. Furthermore, the return value x is dequantized to the “real” format, which corresponds to a highly accurate format. Using the dequantized real value x, at lines 11-14, the lookup table generator circuit 204 performs iterations on the functions stored in S to compute the output of each function one by one. For the first function corresponding to S[0], the input is the dequantized value x. The output of this function is then written back to x, becoming the input of the function S[1], and so on. The function g (e.g., g(x) = log( 3 Example of evaluating / 4+exp(1-x)):

[0067] y0 = S[0](z); / / y0 = 1 - z

[0068] y1=S[1](y0); / / y1=exp(y0)=exp(1-z)

[0069] y2=S[2](y1); / / y2= 3 / 4+y1= 3 / 4+exp(1-z)

[0070] y3=S[3](y2); / / y3=log(y2)=log(3 / 4+exp(1-z));

[0071] exist Figure 7In the example, the input z corresponds to the variable x, and the outputs y0 through y3 are also written back to the same "real" variable x. Finally, in line 16 of the algorithm, the lookup table generator circuit 204 converts the computed function value into the closest value that can be represented on the output type. The result is then written to table T at index idx (e.g., these can be offset relative to boundaries m and M). Once all input points have been processed, the lookup table generator circuit 204 returns the lookup table T along with an offset O set to m, thus allowing sufficient information to potentially offset the input values ​​used when addressing the table.

[0072] In most non-computational operators, there are still many expensive computations. In some examples, exponentiation computation takes the most time in the softmax operation. As described in conjunction with Equation 2, GELU is a non-computationally restricted operator found in natural language processing (NLP) models. Although GELU can use significantly fewer multiply-accumulate (MAC) operations compared to the computationally intensive Conv / GEMM kernel, it is still a significant part of the model inference time. For example, although the special function call "erf" is quite expensive, the methods and apparatuses disclosed in this paper allow a portion of the erf computation to be replaced with a lookup table (e.g., or a combination of lookup tables and arithmetic operators, as described in more detail below). In some examples, the quantized input x is used to address the elements of the table. The table stores the floating-point value corresponding to the evaluation erf(dequantize(i) / sqrt(2)) at index i, thus absorbing the multiplication with the reciprocal of sqrt(2). A new formula for Equation 2 is shown in conjunction with Equation 3:

[0073] GELU(x)=x / 2(1+LUT(x)) Equation 3

[0074] Since erf(x) => -1 when x => -inf, for small input values ​​such as -2, the number of correct bits obtained in FP16 is only approximately 5 (out of 11). Alternatively, the same technique can be used to perform tabulation on (1 + erf(x / sqrt(2))) / 2. In this case, the accuracy is significantly improved because a single rounding error is performed when storing the tabulated value. An increasingly better approach for this function is to extend the tabulation to the entire body of the function, as shown in Equation 4:

[0075] Equation 4: GELU(x) = LUT_{GELU,HP}(x)

[0076] Tabulation absorbs all computations, improving accuracy and reducing operation counts. As shown in more detail below, methods using fewer tables and performing more arithmetic operations are also possible. This can be of increasing importance when tabulation cost becomes a bottleneck. Furthermore, operators with more than one input, in addition to operators like GELU (e.g., which has only one input), can also benefit from the methods and apparatus disclosed herein. For example, focus loss identifies and focuses on incorrect model predictions, rather than examples that the model confidently predicts, thus ensuring that predictions for inaccurate examples improve over time. Sigmoid focus loss can be considered an example as follows:

[0077]

[0078] For example, an operator can be divided into two parts: "exp" and "log" computations performed on both inputs based on a LUT, followed by "multiplication" and "division" in a numerical method. Compared to a pure LUT, this hybrid approach of LUT and numerical computation based on a Single Instruction Multiple Data (SIMD) architecture is expected to result in a smaller decrease in accuracy because the LUT only simulates the simpler process rather than the entire mapping.

[0079] Figure 8 It is based on Figure 3 and Figure 5 Machine-readable instructions 335 used Figure 1 An example of softmax-based lookup table generation for the softmax LUT determiner circuit 210. Figure 8 In the example, the softmax LUT determiner circuit 210 receives the input and determines the output format at block 805. The softmax LUT determiner circuit 210 then performs maximum reduction on the input tensor(s) and lookup-based exponent implementation at blocks 810, 815, and then performs summation reduction at block 820 to generate a softmax-based LUT, as described in more detail below.

[0080] As mentioned earlier, most lookup-based methods are typically limited to one or more 8-bit inputs and outputs. Since hardware can support some low-precision floating-point arithmetic (fp16 / bf16), wider lookup tables can be used to directly output the target numeric format and / or any type of format expected downstream in tabulation operations. This is particularly noteworthy when downstream operators are performing reductions (such as max / sum / mean), as these operators can be accelerated via high-level 16-bit floating-point SIMD instructions. For example, in a softmax inference process, a hybrid LUT algorithm (e.g., based on the Intel AVX512-bf16 ISA) can be executed as follows: Figure 8In the first while loop, the softmax LUT determiner circuit 210 performs maximum reduction on the 8-bit input tensor (e.g., faster than traditional 32-bit). In some examples, the softmax LUT determiner circuit 210 uses a lookup-based exponential implementation to avoid the complex data paths(s) required for the fp16 / bfp16 numerical computation of softmax. Figure 8 In the second while loop, when the softmax LUT determiner circuit 210 performs summation reduction, the bit width of the reduced element is 16 bits. Furthermore, compared to oneDNN, the dot product and summation operations can be performed with only one instruction (e.g., vdpbf16ps) at a smaller loop cost.

[0081] Figures 9A-9B It shows that according to Figure 3 and Figure 6 Machine-readable instructions 340 used Figure 2 The example of the wide vector softmax LUT determiner circuit 212 for generating a softmax-based lookup table for wide vectors. Figures 9A-9B In the example, the wide vector softmax LUT determiner circuit 212 identifies the source vector input and compiles the time constant at block 905, creates a zero-count vector at block 910, and determines the frequency of each input vector element within the input vector at block 915. Additionally, the wide vector softmax LUT determiner circuit 212 locates the maximum element of the input vector at block 920, performs scaling using coarse-scale normalization at block 925, determines the subsequent exponent value using subsequent multiplication at block 930, and outputs the exponent value in the target vector at block 935, as described in more detail below.

[0082] In some examples, the softmax implementation can be sped up if the size of the input vector is larger than the lookup address size (e.g., if the vector size is 512 elements and quantized INT8 data is being used). The softmax function takes an N-element vector as input and returns an N-element vector. Each element j of the output array is calculated as the ratio between the exponent of the corresponding element j in the input array and the sum of the exponents of all elements in the input array. Figures 9A-9B This paper presents a method for calculating the softmax function when the number of elements in the input vector is at least twice the dynamic range corresponding to the width of the input type. For example, if the input type is INT8 (which corresponds to 2^8 = 256 elements), then for the case N>=512, the wide vector softmax LUT determiner circuit 212 uses... Figures 9A-9BThe algorithm is shown in the figure. This algorithm has some unique features, including that instead of calculating the exponent for each element of the input array, the algorithm calculates each exponent the number of times it would need to be computed. Furthermore, the exponent is calculated for all possible tensor values. The process is incremental, and given the exponent (exp(dequantize(i))) of tensor value i, we can obtain the exponent of i+1 = exp(dequantize(i+1)) = exp(dequantize(i) + dequantize(1)) =

[0083] exp(dequantize(i))*exp(dequantize(1)). For example, multiplying by the constant exp(dequantize(1)) can be used to obtain subsequent exponential values.

[0084] exist Figure 9A In the example, the wide vector softmax LUT determiner circuit 212 sets N to the length of the input vector (line 00). If this value is known to be constant beforehand, it can be part of the compile-time constant. The softmax LUT determiner circuit 210 then processes the counting of the number of times each input vector element is found within the input vector (lines 01 to 10). Knowing that the input element type (e.g., tensor type) will be a relatively short type (e.g., such as UINT8), this implementation is simplified. A vector (e.g., count_vec) with 2^w elements (e.g., where w is the input type width, and for UINT8, w = 8 => 256 elements) is created and initialized with 0. Next, the softmax LUT determiner circuit 210 performs a loop on the elements of the input vector. For element j = src_vec[i], the counting vector is implemented as count_vec[j] = count_vec[j] + 1. Subsequently, the softmax LUT determiner circuit 210 locates the maximum element of the input vector (lines 12 to 19). This is achieved by iterating backwards over the previously filled count_vec array and stopping at the first non-zero element. Therefore, the element at index j to which count_vec[j]! = 0 corresponds is the maximum value of the input array.

[0085] The softmax LUT determiner circuit 210 continues to populate the exp_vec array with #seed exponent values ​​(e.g., evenly spaced) by identifying exponent values ​​in a pre-computed table (lines 21-28). The distance between two seed values ​​is identified as seed_stride. The seed values ​​are scaled down, which represents integer powers of 2. Scaling completes coarse-scale normalization so that no exponent value is too large to cause overflow. In some examples, scaling is implemented via a call to ldexp. For example, the softmax LUT determiner circuit 210 obtains the scaling value via a table lookup. Each #seed_stride has a single scale stored for the entire range. Starting with the seed value and the constant exp_stride (e.g., pre-computed and static for a given quantization scheme), the softmax LUT determiner circuit 210 computes subsequent seed_stride-1 exponent values ​​through subsequent multiplications. In some examples, the outer loop (line 31) has no data dependencies between iterations and can therefore be executed in parallel. In some examples, the softmax denominator (sum of exponents) is computed via a dot product call between a new scaled exponent vector (exp_vec) and the corresponding count (count_vec), ensuring that the dot product does not overflow the range of the floating-point format. The softmax LUT determiner circuit 210 determines the reciprocal of the denominator (line 39), and the reciprocal of the denominator is multiplied by all elements of the exp_vec array (line 40). Figure 9B In the example, the final loop (lines 41-43) produces the final exponent value in the target vector by copying values ​​from exp_vec_norm. In some examples, the size of the target vector is N (e.g., matching the size of the input vector), which typically does not match the size of the exp_vec array, which is assumed to be shorter. Therefore, for the output index i, the softmax LUT determiner circuit 210 obtains the tensor value from the source vector, and this value is used to extract the output value from the exp_vec_norm array.

[0086] In some examples, function implementations based on lookup-based techniques can be performed. As previously mentioned, a single lookup table can be used to tabulate all possible function outputs for a given input type. The peculiarity of this approach lies in the fact that the table input is a quantized (e.g., tensor) value with a low bit width (e.g., 8 bits for INT8), while the table output can be any value from quantized INT8 to FP16, bfloat16, or even FP32. In the examples disclosed herein, the reduction in overall tab size can be achieved at the cost of some simple arithmetic operations (e.g., typically available in silicon in modern processors). This has the potential to alleviate the potential cache pressure caused by large array sizes that need to be retrieved from memory. This technique is introduced in scenarios where one input to the implementation is in a quantized format and the output is in a floating-point format (e.g., FP16, bfloat16, FP32, etc.). In particular, the methods and apparatus disclosed herein involve replacing a single table with several tables, each storing coefficient data. Despite using more tables, the number of entries in each table is significantly reduced, which ultimately reduces the number of tab bits in the implementation. The main difference from techniques based on “typical” polynomial approximations lies in the fact that (1) with polynomial approximations, the input is not dequantized, and (2) the polynomial approximation is tuned so that creating a floating-point input from an INT8 / UINT8 quantized input is trivial and requires no special conversion hardware. In the example disclosed herein, an exponential function is used, where the input before quantization is [-4, 4], and the quantization uses an 8-bit format. This approach is general and applicable to other functions (e.g., implementations such as erf / erfc).

[0087] Figure 10 Example results for two matrix sizes are shown, illustrating the performance degradation of the system when using nonlinear function calls, such as when combined with... Figure 1 More detailed descriptions are provided. For example, Table 1000 presents the performance measured on the SPR kernel (e.g., performing a rank-1 update of a symmetric packing matrix) for each output element using different nonlinear functions (reported in GFLOPS). Table 1000 reports the numbers for the two matrices 1005 (e.g., matrix sizes (N, P, Q) with A(N, P) and B(P, Q), where Size[0] = (256, 768, 768) and Size[1] = (256, 2048, 1024)). Figure 10 The functions used in the example shown include Pure IP 101, GELU 1015, Hardswish 1020, Logsimoid 1025, and SQRT 1030.

[0088] Figure 11The illustration shows example tabulation and computation identifiers, along with the average relative error for the implementation associated with the shown tabulation and computation identifiers, and the creation of example polynomial arguments based on quantized input bits. Considering range subdivision, the initial input range can be divided into 32 uniform sub-intervals. The size of the interval, represented by the stride, is defined as stride = 8 / 32 = 1 / 4. For sub-interval index i, the resulting input range can be defined as [i*stride, (i+1)*stride]. When approximating f(x) over interval index i via polynomial P, instead of addressing the polynomial using the absolute value x, an input relative to the interval [0, stride] can be used. Therefore, when approximating f(x) over [i*stride, (i+1)*stride], an approximation of f(y+i*stride) is performed for y in [0, stride]. Furthermore, converting INT8 to the corresponding value in the interval [0, stride] while working with quantized inputs can be expensive. Therefore, instead of approximating f(y+i*step) for y in [0, step], we perform an approximation of f(z-1+i*step) for z in [1, 1+step]. This transformation allows z to be identified from the quantized format using only basic operations, as previously shown. Figure 11 Table 1100 summarizes the tabulation and calculation identifiers, as well as the average relative error for several instances of the method for the input interval [-4, 4]. Figure 11 In the example, Table 1100 includes a degree list 1105, an interval list 1110, a calculation step size 1115, tab positions 1120, multiply-accumulate calls 1125, and an average relative error 1130. Table 1100 illustrates several possible trade-offs between calculations (e.g., the number of multiply-accumulate operations) and tab size (e.g., the number of tab positions), and each of these implementations exposes a different relative error boundary (e.g., the lower the value, the more accurate the implementation). Table 1150 of Figure 1100 includes a degree list 1105, an interval list 1100, and a z-value list 1155. Table 1150 further illustrates the creation of polynomial arguments based on quantized input bits for the above implementations (e.g., where different numbers of subintervals are used).

[0089] Figure 12 The illustration shows example results of a d-th order approximation related to piecewise polynomials and range reduction, based on quantized input data, according to the teachings disclosed herein. Figure 12In the example, Table 1200 includes (one or more) interval index values ​​1202, (one or more) intervals 1205, (one or more) interval / log(2) calculations 1210, (one or more) E values ​​1215, (one or more) Y = interval -log(2)E calculations 1220, and (one or more) offset values ​​1224, as described in more detail below. In particular, piecewise polynomials and range reduction for d-th order approximations are performed based on quantized input data. For example, range reduction based on quantized data can be integrated into the method disclosed herein. This allows for an increase in the range of supported data and provides more accurate implementation results (e.g., specific to exponential functions). The proposed range reduction is based on the following: taking x as input to an exponential function and representing x as x = E * log2 + y, where E is the integer closest to x / log(2), and y will be some value in the interval [-log 2 / 2, log 2 / 2]. Calculating the exponent in floating-point results in exp(x) = exp(E*log2 + y) = exp(E*log2)*exp(y) = 2^E*exp(y). Since y will be relatively small (e.g., approximately [-0.34, 0.34]), exp(y) will also have a narrow range (e.g., approximately [0.7, 1.41]), which may require an exponent update (e.g., decrementing) if exp(y) is less than 1. In the example shown below, the same input interval [-4, 4] is used, divided into 16 subintervals, each with a size of 0.5 (e.g., calculated as the range (4 - (-4)) / 16).

[0090] exist Figure 12In the example, the first column of Table 1200 shows the interval index 1202 (e.g., the range from 0 to 15). The second column shows the input range corresponding to the interval index 1202 on the left (e.g., interval 1205). The third column of Table 1200 shows the quotient produced by dividing the interval 1205 in the second column by the constant log(2) (e.g., (one or more) intervals / log(2) calculates 1210), while the fourth column shows (one or more) custom E values ​​1215. For example, for a negative interval, E can be calculated using ceil(sup(I / log(2)) and for a positive interval, E can be calculated using floor(inf(I / log(2)). The fifth column of Table 1200 shows the range 1220 of I-log(2)E (e.g., for a range reduction input for the exponent). Finally, the sixth column shows the offset 1225, which is the left boundary of the interval calculated in the previous column. For each interval, a predetermined value of E is identified and tabulated. Since the range of E is [-5, 5], it has 4 integer bits. The format is sufficient. A wider supported input range can produce larger E values, so the range can be increased. Regarding the calculation of exp(y), where y is shown in the fifth column, the following identity can be used: e^(a+b)=e^(a)*e^(b). In the calculation of exp(y), the above independent variable can be expressed as e^(offset+[0,0.5])=e^(offset)*e^([0,0.5]). The values ​​of e^(offset) can be tabulated for all 16 input intervals, where the exponent is calculated within the interval [0,0.5].

[0091] Therefore, the proposed method may require 4 bits * 16 = 64 bits for the tabulated E value, 16 bits * 16 = 256 bits for the tabulated exp(offset) value, and 3 (coefficients) * 16 = 48 bits for the quadratic polynomial coefficients, which are used to approximate exp(z) for z in [0, 0.5]. In total, the architecture uses (4 + 16) * 16 + 3 * 16 = 320 + 48 = 368 bits for tabulation. Additionally, the polynomial evaluation uses two multiply-add operations (e.g., for evaluating the quadratic polynomial using Horner's scheme), one LDEXP operation (e.g., for adding E to the tabulated offset exponent), and one multiplication to assemble the final result. The high 4 bits are used for address decoding: address = 0xF & (input >> 4). The function exp(x) is approximated for x in [0, 0.5]. In the example disclosed herein, the function exp(z-1) is approximated for z in [1, 1.5]. This transformation allows z to be created by padding the lower 4 bits of the input into a floating-point value, as shown below:

[0092] S E E E E E F F F F F F F F F F 0 0 1 1 1 1 0 i3 i2 i1 i0 0 0 0 0 0

[0093] In some examples, the exponent value is obtained from the lookup, where E = e_table[address] is used to identify integer values ​​(e.g., 4 digits). The offset_exponential value is calculated as the exponent in the rightmost column of the range reduction table 1200 above, as follows: offset_exp = offset_table[address] represents a floating-point (e.g., half-precision) value. Finally, a single quadratic polynomial is obtained: p = (a2*z + a1)*z + a0. The polynomial p is 0.6435546875*x^2 + -0.3154296875*x + 0.6728515625, and the final value is calculated as out = LDEXP(offset_exp, E)*p.

[0094] Figures 13A-13F The diagram illustrates the relationship between... Figure 11-12 Example graphical representations of the relative errors associated with the tabulated results are shown in tables 1300, 1310, 1320, 1330, 1340, and 1350. The accompanying figures illustrate the relative errors for the 2nd and 3rd degree implementations in the tables above, for the corresponding number of subintervals given in the tables (e.g., based on ULP ranging from -8 last digits (ULP) to 1 ULP, as shown in tables 1305, 1315, 1325, 1335, 1345, and 1355). In some examples, piecewise polynomial approximations of degree d can be performed based on quantized input data. For example, for a first-degree approximation, two coefficients are used, and the polynomial has the form P(x) = a0 + a1*x. For an input interval subdivided into 32 subintervals, Figure 13A The relative error of this implementation is shown, highlighting that the relative error is typically defined by 4 ULPs, and the average relative error is 2.52 ULPs. The sub-interval index is simply obtained from the quantized input by analyzing the high 5 bits using address = 0x1F&(input >> 3). The coefficient tables table_a0[] and table_a1[] output data in floating-point format (e.g., half-precision). The corresponding coefficients are obtained by indexing these tables using address variables.

[0095] a0 = table_a0[address]; a1 = table_a1[address]. Next, the floating-point polynomial input can be obtained by manipulating the input data as follows: z = 0x3C00 | ((input & 0x7) << 5). This expression first creates a floating-point fraction from the lower bits of the input by concatenating five zeros on the right and two zeros on the left. Next, the exponent corresponding to the power of 2^0 (e.g., binary “01111”) is concatenated to the left of the fraction. The constant (e.g., in hexadecimal representation) 0x3C00 can be used to concatenate the exponent of “01111” and the sign “0” to the newly created fraction using a bitwise OR operation. The resulting input (e.g., half-precision) is shown in the example below:

[0096] S E E E E E F F F F F F F F F F 0 0 1 1 1 1 0 0 i2 i1 i0 0 0 0 0 0

[0097] Therefore, a simple multiply-accumulate operation with half precision (e.g., p = a1 * z + a0) is used to perform polynomial evaluation.

[0098] Figures 14A-14B This includes example computation graph 1400 used in low-precision training, and example performance data graphs 1450, 1460, and 1470 associated with the pure element-wise operator exp, the element-wise operator-intensive non-computationally restricted operator softmax, and the GELU activation function. In the examples disclosed herein, to verify performance and accuracy, the pure element-wise operator exp is compared with the element-wise operator-intensive non-computationally restricted operator softmax from oneDNN. Both implementations can be benchmarked (e.g., using Intel Xeon Sapphire Rapids), where the configuration is a 1024x1024 input shape, single-core, single-threaded. The associated computation graph 1400 is a widely used low-precision training / inference scenario, which takes int8 as example input 1405, applies example dequantization boxes 1410 (e.g., converting INT8 to floating-point format), and applies example non-computationally restricted operator boxes 1415 (e.g., exp, softmax, etc.) to produce example output 1420. In some examples, both the dequantized box 1410 and the exp box 1415 are purely element-wise operations. Therefore, operator fusion optimization is applied in oneDNN to benchmark the exp operator, as shown in the computational graph 1400. The benchmark results shown below demonstrate that the LUT-based implementation of the exponent disclosed in this paper can improve performance without sacrificing accuracy. Regarding accuracy, mean squared error (MSE) is used to model the loss function, and the results shown below indicate that the method disclosed in this paper achieves the same accuracy as oneDNN:

[0099] Operator MSE (oneDNN vs. the method disclosed in this paper) GELU 0 softmax 0

[0100] Figure 1450 shows the performance data for the pure element-wise operator exp, Figure 1460 shows the performance data for the element-wise operator-intensive, non-computationally constrained operator softmax, and Figure 1470 shows the performance data for the GELU activation function under different implementations (e.g., oneDNN and the methods disclosed herein). For example, Figure 1450 shows the time-varying performance data 1456 of dequantization using exp(fusion) 1452 compared to LUT_exp 1454; Figure 1460 shows the time-varying performance data 1456 of oneDNN softmax 1462 compared to LUT hybrid softmax 1464; and Figure 1470 shows the time-varying performance data 1456 of oneDNN(dequant+GELU+quant) 1472 compared to LUT(int8-LUT-GELU) 1474. For the exponential implementation, when comparing the proposed lookup-based solution (e.g., INT8 input, BFP16 output) with a oneDNN dequantization call (e.g., INT8->gt BFP16) followed by a BFP16 exponential call, the implementation disclosed herein indicates a 1.3x speedup (e.g., 1.3x less computation time), as shown in conjunction with Figure 1450. On the softmax benchmark, the performance gain widens for the lookup-based implementation (e.g., 1.56x), as shown in conjunction with Figure 1460. This performance improvement can be attributed to the efficient use of the low-precision correlation ISA in conjunction with the lookup-based exponential implementation disclosed herein. The lookup-based approach is also benchmarked against the GELU activation function, a key operator in BERT models. Consider the following operator chain: dequantization, followed by a GELU call, followed by (re)quantization. The input data type is INT8, dequantization converts this data to FP32 floating-point format for GELU computation, and finally, the FP32 output of the GELU call is converted back (e.g., quantized to) INT8 data type. In this process, the lookup-based method shows a significant improvement, reducing the total computation time by approximately 8 times, as illustrated in Figure 1470.

[0101] Figure 15 It is constructed to execute and / or instantiate Figure 3-6This is a block diagram of an example processor platform 1500 that provides machine-readable instructions and / or operations to implement the example numerical computation approximator circuit 110. The processor platform 1500 may be, for example, a server, personal computer, workstation, self-learning machine (e.g., neural network), mobile device (e.g., cellular phone, smartphone, tablet computer such as iPad™), personal digital assistant (PDA), internet device, DVD player, CD player, digital video recorder, Blu-ray player, game console, personal video recorder, set-top box, head-mounted device (e.g., augmented reality (AR) head-mounted device, virtual reality (VR) head-mounted device, etc.) or other wearable device, or any other type of computing device.

[0102] The processor platform 1500 shown in the example includes processor circuitry 1512. Processor circuitry 1512 in the example shown is hardware. For example, processor circuitry 1512 can be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. Processor circuitry 1512 can be implemented by one or more semiconductor-based (e.g., silicon-based) devices. In this example, processor circuitry 1512 implements input recognizer circuitry 202, lookup table (LUT) generator circuitry 204, index generator circuitry 206, output recognizer circuitry 208, softmax LUT determiner circuitry 210, and wide vector softmax LUT determiner circuitry 212.

[0103] The processor circuitry 1512 shown in the example includes local memory 1513 (e.g., cache, registers, etc.). Figure 15 In this example, the local memory 1513 is implemented Figure 2 Example data storage device 214. However, any of the example memories 1514 and 1516 can be implemented Figure 2 The example data storage device 214 may be all or part of it. The processor circuitry 1512 of the example shown communicates via bus 1518 with main memory, which includes volatile memory 1514 and non-volatile memory 1516. Volatile memory 1514 may be synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), etc.

[0104] Dynamic Random Access Memory And / or any other type of RAM device. The non-volatile memory 1516 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1514, 1516 of the illustrated example is controlled by the memory controller 1517.

[0105] The processor platform 1500 shown in the example also includes interface circuitry 1520. Interface circuitry 1520 can be implemented in hardware according to any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB) interface, etc. Interfaces include Near Field Communication (NFC) interfaces, Peripheral Component Interconnect (PCI) interfaces, and / or Fast Peripheral Component Interconnect (PCIe) interfaces.

[0106] In the example shown, one or more input devices 1522 are connected to interface circuitry 1520. The input devices 1522 allow a user to input data and / or commands into processor circuitry 1512. The input devices 1522 may be implemented as, for example, audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, touchpads, trackballs, isotope devices, and / or voice recognition systems.

[0107] One or more output devices 1524 are also connected to the interface circuitry 1520 of the illustrated example. The output devices 1524 may be implemented, for example, by display devices (e.g., light-emitting diode (LED), organic light-emitting diode (OLED), liquid crystal display (LCD), cathode ray tube (CRT) display, in-place switching (IPS) display, touchscreen, etc.), haptic output devices, printers, and / or speakers. Therefore, the interface circuitry 1520 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or graphics processor circuitry such as a GPU.

[0108] The interface circuit 1520 of the example shown also includes communication devices, such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces, to facilitate the exchange of data with external machines (e.g., any kind of computing device) via network 1526. Communication can be achieved through, for example, Ethernet connections, digital subscriber line (DSL) connections, telephone line connections, coaxial cable systems, satellite systems, line-of-sight wireless systems, cellular telephone systems, optical connections, etc.

[0109] The processor platform 1500 shown in the example also includes one or more mass storage devices 1528 for storing software and / or data. Examples of such mass storage devices 1528 include magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disc drives, redundant array of independent disks (RAID) systems, solid-state storage devices (such as flash memory devices), and DVD drives.

[0110] can be Figure 3-6The machine-executable instructions 1532 implemented by the machine-readable instructions can be stored in a mass storage device 1528, a volatile memory 1514, a non-volatile memory 1516, and / or a removable non-transient computer-readable storage medium such as a CD or DVD.

[0111] Figure 16 yes Figure 15 A block diagram of an example implementation of the processor circuit 1512. In this example, Figure 15 The processor circuit 1512 is implemented by the microprocessor 1600. For example, the microprocessor 1600 may be a general-purpose microprocessor (e.g., a general-purpose microprocessor circuit). The microprocessor 1300 executes... Figure 3-6 The flowchart contains some or all of the machine-readable instructions to effectively instantiate... Figure 2 The logic circuitry of the circuit, thereby executing operations corresponding to those machine-readable instructions. In some such examples, Figure 2 The circuitry is instantiated by the hardware circuitry of the microprocessor 1600 in conjunction with instructions. For example, the microprocessor 1600 can implement multi-core hardware circuitry, such as a CPU, DSP, GPU, XPU, etc. Although it can include any number of example cores 1602 (e.g., one core), the example microprocessor 1600 is a multi-core semiconductor device including N cores. The cores 1602 of the microprocessor 1600 can operate independently or can cooperate to execute machine-readable instructions. For example, machine code corresponding to firmware, embedded software programs, or software programs can be executed by one of the cores 1602, or can be executed by multiple cores 1602 at the same or different times. In some examples, the machine code corresponding to firmware, embedded software programs, or software programs is divided into threads and executed in parallel by two or more cores 1602. Software programs can correspond to... Figure 3-6 The flowchart represents part or all of the machine-readable instructions and / or operations.

[0112] Core 1602 can communicate via example bus 1604. In some examples, bus 1604 can implement a communication bus to enable communication with one or more cores 1602. For example, bus 1604 can implement at least one of an Inter-Integrated Circuit (I2C) bus, a Serial Peripheral Interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, bus 1604 can implement any other type of computing or electrical bus. Core 1602 can obtain data, instructions, and / or signals from one or more external devices via example interface circuitry 1606. Core 1602 can output data, instructions, and / or signals to one or more external devices via interface circuitry 1606. Although core 1602 in this example includes example local memory 1620 (e.g., an L1 cache that can be split into a Level 1 (L1) data cache and an L1 instruction cache), microprocessor 1600 also includes example shared memory 1610 (e.g., a Level 2 (L2) cache) that can be shared by the various cores for high-speed access to data and / or instructions. Data and / or instructions can be transferred (e.g., shared) by writing to and / or reading from shared memory 1610. The local memory 1620 and shared memory 1610 of each core 1602 can be a combination of multi-level cache memory and main memory (e.g., ...). Figure 15 The cache is part of a hierarchy of storage devices (main memory 1514, 1516). Typically, higher-level memories in the hierarchy exhibit lower access times and have smaller storage capacities compared to lower-level memories. Variations at each level of the cache hierarchy are managed by cache coherence policies (e.g., reconciliation).

[0113] Each core 1602 can be referred to as a CPU, DSP, GPU, or any other type of hardware circuitry. Each core 1602 includes a control unit circuitry 1614, an arithmetic and logic (AL) circuitry (sometimes called an ALU) 1616, multiple registers 1618, an L1 cache 1620, and an example bus 1622. Other structures may exist. For example, each core 1602 may include vector unit circuitry, single instruction multiple data (SIMD) unit circuitry, load / store unit (LSU) circuitry, branch / jump unit circuitry, floating-point unit (FPU) circuitry, etc. The control unit circuitry 1614 includes semiconductor-based circuitry configured to control (e.g., coordinate) the movement of data within the corresponding core 1602. The AL circuitry 1616 includes semiconductor-based circuitry configured to perform one or more mathematical and / or logical operations on the data within the corresponding core 1602. Some examples of the AL circuitry 1616 perform integer-based operations. In other examples, the AL circuitry 1616 also performs floating-point operations. In other examples, AL circuit 1616 may include a first AL circuit performing integer-based operations and a second AL circuit performing floating-point operations. In some examples, AL circuit 1616 may be referred to as an Arithmetic Logic Unit (ALU). Register 1618 is a semiconductor-based structure used to store data and / or instructions, such as the results of one or more operations performed by AL circuit 1616 of the corresponding core 1602. For example, register 1618 may include one or more vector registers, one or more SIMD registers, one or more general-purpose registers, one or more flag registers, one or more segment registers, one or more machine-specific registers, one or more instruction pointer registers, one or more control registers, one or more debug registers, one or more memory management registers, one or more machine check registers, etc. Register 1618 may be arranged in a manner such as Figure 16 The memory bank shown. Alternatively, register 1618 can be organized in any other arrangement, format, or structure, including being distributed throughout core 1602 to reduce access time. The second bus 1622 can be implemented by at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.

[0114] Each core 1602 and / or more generally, the microprocessor 1600 may include additional and / or alternative structures to the structures shown and described above. For example, one or more clock circuits, one or more power supplies, one or more power gates, one or more cache home agents (CHAs), one or more aggregation / common grid sites (CMS), one or more shifters (e.g., one or more barrel shifters), and / or other circuitry may be present. The microprocessor 1600 is a semiconductor device fabricated to include a plurality of interconnected transistors to implement the above-described structures in one or more integrated circuits (ICs) contained in one or more packages. The processor circuitry may include one or more accelerators and / or cooperate with one or more accelerators. In some examples, accelerators are implemented by logic circuitry to perform certain tasks faster and / or more efficiently than a general-purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. GPUs or other programmable devices may also be accelerators. Accelerators may be on the processor circuitry, in the same chip package as the processor circuitry, and / or in one or more packages separate from the processor circuitry.

[0115] Figure 17 yes Figure 15 A block diagram illustrating another example implementation of the processor circuitry. In this example, processor circuitry 1512 is implemented using FPGA circuitry 1700. For example, FPGA circuitry 1700 can be implemented using an FPGA. FPGA circuitry 1700 can be used, for example, to execute instructions that would otherwise be executed by a processor that executes corresponding machine-readable instructions. Figure 16 The example microprocessor 1600 performs the operations shown. However, once configured, the FPGA circuitry 1700 instantiates machine-readable instructions in hardware and can therefore typically perform operations faster than a general-purpose microprocessor executing the corresponding software.

[0116] More specifically, with the above Figure 16 The microprocessor 1600 (which is a general-purpose device that can be programmed to execute commands) is a microprocessor that... Figure 3-6 The flowchart represents some or all of the machine-readable instructions, but its interconnections and logic circuitry are fixed once manufactured. Figure 17 The example FPGA circuit 1700 includes interconnects and logic circuitry, which can be configured and / or interconnected in different ways after manufacturing to instantiate, for example, by... Figure 3-6The flowchart represents some or all of the machine-readable instructions. Specifically, the FPGA 1700 can be considered an array of logic gates, interconnects, and switches. Switches can be programmed to change how logic gates are interconnected, effectively forming one or more dedicated logic circuits (unless and until the FPGA circuit 1700 is reprogrammed). The configured logic circuits enable logic gates to cooperate in different ways to perform different operations on data received from the input circuitry. These operations can correspond to... Figure 3-6 The flowchart represents some or all of the software. Therefore, the FPGA circuit 1700 can be constructed to effectively utilize... Figure 3-6 The flowchart instantiates some or all of the machine-readable instructions into special-purpose logic circuits to perform the operations corresponding to those software instructions in a specialized manner similar to that of an ASIC. Therefore, the FPGA circuit 1700 can perform operations that a general-purpose microprocessor can perform. Figure 3-6 The operations corresponding to some or all of the machine-readable instructions can be executed faster.

[0117] exist Figure 17 In the example, the FPGA circuit 1700 is configured to be programmed (and / or reprogrammed once or multiple times) by an end user using a hardware description language (HDL) such as Verilog. Figure 17 The FPGA circuit 1700 includes example input / output (I / O) circuitry 1702 for obtaining and / or outputting data to / from example configuration circuitry 1704 and / or external hardware 1706. For example, configuration circuitry 1704 may implement interface circuitry that can obtain machine-readable instructions to configure FPGA circuitry 1700 or portions thereof. In some such examples, configuration circuitry 1704 may obtain machine-readable instructions from a user, a machine (e.g., hardware circuitry (e.g., programming or dedicated circuitry) that can implement an artificial intelligence / machine learning (AI / ML) model to generate instructions), etc. In some examples, external hardware 1706 may be implemented by external hardware circuitry. For example, external hardware 1706 may be implemented by... Figure 17 The microprocessor 1700 is implemented. The FPGA circuit 1700 also includes an array of example logic gates 1708, multiple example configurable interconnects 1710, and example memory circuits 1712. The logic gates 1708 and configurable interconnects 1710 are configurable to instantiate corresponding to... Figure 3-6 One or more operations and / or other desired operations in at least some of the machine-readable instructions. Figure 17The logic gate circuit 1708 shown is fabricated in groups or blocks. Each block includes a semiconductor-based electrical structure that can be configured into a logic circuit. In some examples, the electrical structure includes logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that provide basic building blocks for the logic circuit. Electrically controllable switches (e.g., transistors) are present within each logic gate circuit 1708 to enable the configuration of the electrical structure and / or logic gates to form a circuit that performs the desired operation. The logic gate circuit 1708 may include other electrical structures such as lookup tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.

[0118] The configurable interconnect 1710 in the example shown is a conductive path, trace, via, etc., which may include electrically controllable switches (e.g., transistors). The state of the electrically controllable switches can be changed by programming (e.g., using an HDL instruction language) to activate or deactivate one or more connections between one or more logic gates 1708 to program the desired logic circuit.

[0119] The storage circuit 1712 in the example shown is configured to store the results of one or more operations performed by the corresponding logic gates. The storage circuit 1712 can be implemented using registers, etc. In the example shown, the storage circuit 1712 is distributed among the logic gates 1708 to facilitate access and improve execution speed.

[0120] Figure 17 The example FPGA circuit 1700 also includes example dedicated operation circuitry 1714. In this example, dedicated operation circuitry 1714 includes dedicated circuitry 1716, which can be invoked to implement common functions, thus avoiding the need for field programming of those functions. Examples of such dedicated circuitry 1716 include memory (e.g., DRAM) controller circuitry, PCIe controller circuitry, clock circuitry, transceiver circuitry, memory, and multiplier-accumulator circuitry. Other types of dedicated circuitry may be present. In some examples, FPGA circuitry 1700 may also include example general-purpose programmable circuitry 1718, such as example CPU 1720 and / or example DSP 1722. Other general-purpose programmable circuitry 1718, such as GPUs, XPUs, etc., may additionally or alternatively exist and can be programmed to perform other operations.

[0121] although Figure 16 and Figure 17 The diagram shows... Figure 15 Two example implementations of the processor circuit 1512 are provided, but many other approaches are conceivable. For example, as mentioned above, modern FPGA circuits may include an onboard CPU, such as... Figure 17 One or more of the example CPUs 1720. Therefore, Figure 15The processor circuit 1512 can be additionally combined Figure 16 Example microprocessor 1600 and Figure 17 The example FPGA circuit 1700 is used for implementation. In some such hybrid examples, it is implemented by... Figure 3-6 The first part of the machine-readable instruction represented by the flowchart can be derived from... Figure 16 Executed by one or more 1602 cores, by Figure 3-6 The second part of the machine-readable instruction represented by the flowchart can be derived from... Figure 17 The FPGA circuit 1700 executes, and / or is powered by Figure 3-6 The third part of the machine-readable instructions shown in the flowchart can be executed by the ASIC. It should be understood that... Figure 15 Some or all of the circuits can therefore be instantiated at the same or different times. Some or all of the circuits can be instantiated, for example, in one or more threads that execute simultaneously and / or serially. Furthermore, in some examples, Figure 15 Some or all of the circuitry can be implemented within one or more virtual machines and / or containers that execute on the microprocessor.

[0122] In some examples, Figure 15 The processor circuitry 1512 can be housed in one or more packages. For example, Figure 16 The processor circuitry 1600 and / or Figure 17 The FPGA circuitry 1700 can be housed in one or more packages. In some examples, the XPU can be comprised of components that can be housed in one or more packages. Figure 15 The processor circuitry 1512 is implemented. For example, the XPU may include a CPU in one package, a DSP in another package, a GPU in yet another package, and an FPGA in yet another package.

[0123] Figure 18 The diagram shows a block diagram illustrating the use of, for example, Figure 15 Example software distribution platform 1505 distributes software such as example machine-readable instructions 1532 to hardware devices owned and / or operated by a third party. Example software distribution platform 1805 can be implemented by any computer server, data facility, cloud service, etc., capable of storing software and transferring it to other computing devices. The third party can be a customer of the entity that owns and / or operates software distribution platform 1805. For example, the entity owning and / or operating software distribution platform 1805 can be the software (such as...) Figure 15The example machine-readable instructions 1532 refer to the developer, seller, and / or licensor. Third parties can be consumers, users, retailers, OEMs, etc., who purchase and / or license the software for use and / or resell and / or sublicense. In the example shown, the software distribution platform 1805 includes one or more servers and one or more storage devices. The storage devices store... Figure 15 The machine-readable instruction 1532, which can correspond to the above-described instructions. Figure 3-6 Example machine-readable instructions 300, 310, 335, and 340. One or more servers of example software distribution platform 1805 communicate with network 1810, which may correspond to the Internet and / or any one or more of the example networks described above. In some examples, one or more servers respond to a request to transfer software as part of a commercial transaction to a requesting party. Payment for the delivery, sale, and / or licensing of the software may be processed by one or more servers of the software distribution platform and / or by a third-party payment entity. The servers enable purchasers and / or licensors to download machine-readable instructions 1532 from software distribution platform 1805. For example, this may correspond to... Figure 3-6 Software containing example machine-readable instructions 300, 310, 335, and 340 can be downloaded to example processor platform 1500, which executes machine-readable instruction 1532 to implement numerical computation approximator circuit 110. In some examples, one or more servers of software distribution platform 1805 periodically provide, transmit, and / or force software updates (e.g., Figure 15 Example machine-readable instructions (1532) are used to ensure that improvements, patches, updates, etc., are distributed and applied to the software at the end-user device.

[0124] Based on the foregoing, it should be understood that example systems, methods, apparatuses, and articles of art have been disclosed that allow the use of customizable bit-width lookup-based implementations in element-wise operator-intensive operators. In the examples disclosed herein, implementations based on customizable bit-width lookup tables can be used to output values ​​directly in floating-point representations (such as FP16 (e.g., half-precision) or bfloat16). In the examples disclosed herein, a novel softmax implementation specifically designed for wide vectors is introduced, which significantly reduces the number of exponentiation calls, thereby significantly reducing the total computation time. In the examples disclosed herein, an exponentiation function-specific implementation is introduced that (i) takes data in tensor format, (ii) proposes dequantization range reduction capabilities, and (iii) provides a trade-off between tabulated LUT costs, computational costs, and parallelization. Therefore, the disclosed systems, methods, apparatuses, and articles of art relate to one or more improvements in the operation of machines such as computers or other electronic and / or mechanical devices.

[0125] This document discloses example methods, apparatuses, systems, and artifacts for container authentication in client-based workloads. Further examples and combinations thereof include the following:

[0126] Example 1 includes an apparatus comprising at least one memory, machine-readable instructions, and programmable circuitry for instantiating or executing at least one of the machine-readable instructions to generate a lookup table based on input elements associated with a training phase of a deep neural network, indexing the lookup table using tensor values ​​associated with output indices, and outputting a vector in a target numeric format based on the lookup table, the vector including output values ​​in floating-point representation.

[0127] Example 2 includes the apparatus of Example 1, wherein the programmable circuitry is used to generate the lookup table based on expected input elements.

[0128] Example 3 includes the apparatus of Example 1, wherein the programmable circuitry is used to regenerate the lookup table based on the current input element.

[0129] Example 4 includes the apparatus of Example 1, wherein the programmable circuitry is used to index a second lookup table based on a first output of a first lookup table.

[0130] Example 5 includes the apparatus of Example 4, wherein the programmable circuitry is used to combine a first output of a first lookup table and a second output of a second lookup table into an index of a third lookup table, the first and second outputs being combined based on a numerical function.

[0131] Example 6 includes the apparatus of Example 1, wherein the programmable circuitry is used to generate a histogram of the input elements.

[0132] Example 7 includes the apparatus of Example 6, wherein the programmable circuitry is used to identify the index of the maximum value in the histogram and to scale the lookup table based on the maximum value.

[0133] Example 8 includes the apparatus of Example 7, wherein the programmable circuitry is used to scale the lookup table to prevent overflow.

[0134] Example 9 includes the apparatus of Example 1, wherein the programmable circuitry is configured to: identify a unique input value, the total number of unique input values ​​being less than the total number of input values; and calculate a result output associated with the unique input value, the result output being associated with the input value.

[0135] Example 10 includes the apparatus of Example 1, wherein the input element includes at least one of the following: a set of composite functions, an input type, a quantization scheme parameter, or an output data type.

[0136] Example 11 includes the apparatus of Example 10, wherein the programmable circuitry is configured to generate the lookup table for the composite function set when the output data type is a target number format.

[0137] Example 12 includes the apparatus of Example 1, wherein the programmable circuitry is used to identify when the size of the input vector is greater than the lookup address size.

[0138] Example 13 includes the apparatus of Example 1, wherein the programmable circuitry is used to initiate a piecewise polynomial based on quantized input data.

[0139] Example 14 includes the apparatus of Example 1, wherein the programmable circuitry is used to perform exponential-specific range reduction based on input tensor values.

[0140] Example 15 includes the apparatus of Example 14, wherein the programmable circuitry is used to output an exponent-specific range reduction result using floating-point representation.

[0141] Example 16 includes a method comprising: generating a lookup table based on input elements associated with a training phase of a deep neural network by executing instructions using at least one processor; indexing the lookup table with tensor values ​​associated with output indices by executing instructions using at least one processor; and outputting a vector in a target numeric format based on the lookup table by executing instructions using at least one processor, the vector including output values ​​in floating-point representation.

[0142] Example 17 includes the method of Example 16, and further includes generating the lookup table based on the expected input elements.

[0143] Example 18 includes the method of Example 16, and further includes generating the lookup table based on the current input element.

[0144] Example 19 includes the method of Example 16, and also includes indexing a second lookup table based on a first output of a first lookup table.

[0145] Example 20 includes the method of Example 19, and further includes combining the first output of the first lookup table and the second output of the second lookup table into an index of the third lookup table, based on a numerical function to combine the first output and the second output.

[0146] Example 21 includes the method of Example 16, and also includes generating a histogram of the input elements.

[0147] Example 22 includes the method of Example 21, and further includes: identifying an index of the maximum value in the histogram; and scaling the lookup table based on the maximum value.

[0148] Example 23 includes the method of Example 22, and further includes scaling the lookup table to prevent overflow.

[0149] Example 24 includes the method of Example 16, wherein the programmable circuitry is configured to: identify a unique input value, the total number of unique input values ​​being less than the total number of input values; and calculate a result output associated with the unique input value, the result output being associated with the input value.

[0150] Example 25 includes the method of Example 16, wherein the input elements include a set of composite functions, an input type, quantization scheme parameters, or an output data type.

[0151] Example 26 includes the method of Example 25, and further includes: generating the lookup table for the composite function set when the output data type is a target number format.

[0152] Example 27 includes the method of Example 16, and also includes identifying when the size of the input vector is greater than the lookup address size.

[0153] Example 28 includes the method of Example 16, and also includes starting a piecewise polynomial based on quantized input data.

[0154] Example 29 includes the method of Example 16, and also includes performing exponent-specific range reduction based on the input tensor values.

[0155] Example 30 includes the method of Example 29, and also includes the result of exponent-specific range reduction output using floating-point representation.

[0156] Example 31 includes a non-transitory machine-readable storage medium including instructions that, when executed, cause processor circuitry to perform at least the following operations: generate a lookup table based on input elements associated with a training phase of a deep neural network; index the lookup table using tensor values ​​associated with output indices; and output a vector in a target numeric format based on the lookup table, the vector including output values ​​in floating-point representation.

[0157] Example 32 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause a processor to generate the lookup table based on expected input elements.

[0158] Example 33 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause the processor to regenerate the lookup table based on the current input elements.

[0159] Example 34 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause the processor to index a second lookup table based on a first output of a first lookup table.

[0160] Example 35 includes the non-transitory machine-readable storage medium defined in Example 34, wherein the instructions, when executed, cause the processor to combine a first output of a first lookup table and a second output of a second lookup table into an index of a third lookup table, the first and second outputs being combined based on a numerical function.

[0161] Example 36 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause a processor to generate a histogram of the input elements.

[0162] Example 37 includes a non-transitory machine-readable storage medium as defined in Example 36, wherein the instructions, when executed, cause a processor to perform the following operations: identify an index of the maximum value in the histogram; and scale the lookup table based on the maximum value.

[0163] Example 38 includes the non-transitory machine-readable storage medium defined in Example 37, wherein the instructions, when executed, cause the processor to scale the lookup table to prevent overflow.

[0164] Example 39 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause a processor to perform the following operations: identify a unique input value, the total number of unique input values ​​being less than the total number of input values; and compute a result output associated with the unique input value, the result output being associated with the input value.

[0165] Example 40 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the input elements include a set of composite functions, an input type, quantization scheme parameters, or an output data type.

[0166] Example 41 includes the non-transitory machine-readable storage medium defined in Example 40, wherein the instructions, when executed, cause the processor to perform the following operation: generate the lookup table for the set of composite functions when the output data type is a target number format.

[0167] Example 42 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause the processor to recognize when the size of the input vector is greater than the lookup address size.

[0168] Example 43 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause the processor to initiate a piecewise polynomial based on quantized input data.

[0169] Example 44 includes the non-transitory machine-readable storage medium defined in Example 31, wherein the instructions, when executed, cause the processor to perform an exponential-specific range reduction based on the input tensor values.

[0170] Example 45 includes the non-transitory machine-readable storage medium defined in Example 44, wherein the instructions, when executed, cause the processor to output an exponent-specific range-reduced result using a floating-point representation.

[0171] Example 46 includes an apparatus comprising: means for generating a lookup table based on input elements associated with a training phase of a deep neural network; means for indexing the lookup table using tensor values; and means for outputting a vector in a target numeric format based on the lookup table, the vector including output values ​​in floating-point representation.

[0172] Example 47 includes the apparatus of Example 46, wherein the components for generating the lookup table are used to generate the lookup table based on the expected input elements.

[0173] Example 48 includes the apparatus of Example 46, wherein the components for generating the lookup table are used to regenerate the lookup table based on the current input element.

[0174] Example 49 includes the apparatus of Example 46, wherein the component for indexing a lookup table is used to index a second lookup table based on a first output of a first lookup table.

[0175] Example 50 includes the apparatus of Example 49, wherein the component for indexing a lookup table is used to combine a first output of a first lookup table and a second output of a second lookup table into an index of a third lookup table, the first and second outputs being combined based on a numerical function.

[0176] Example 51 includes the apparatus of Example 46, and also includes a component for generating a histogram of the input elements.

[0177] Example 52 includes the apparatus of Example 51, wherein the components for indexing the lookup table are used to identify the index of the maximum value in the histogram and to scale the lookup table based on the maximum value.

[0178] Example 53 includes the apparatus of Example 51, wherein the component for indexing the lookup table is used to scale the lookup table to prevent overflow.

[0179] Example 54 includes the apparatus of Example 46, wherein the components for generating a lookup table are configured to: identify a unique input value, the total number of unique input values ​​being less than the total number of input values; and calculate a result output associated with the unique input value, the result output being associated with the input value.

[0180] Example 55 includes the apparatus of Example 46, wherein the input elements include a set of composite functions, an input type, a quantization scheme parameter, or an output data type.

[0181] Example 56 includes the apparatus of Example 55, wherein the components for generating the lookup table include: generating the lookup table for the set of composite functions when the output data type is a target number format.

[0182] Example 57 includes the apparatus of Example 46, wherein the components for generating the lookup table include identifying when the size of the input vector is greater than the lookup address size.

[0183] Example 58 includes the apparatus of Example 46, wherein the component for generating the lookup table is used to initiate a piecewise polynomial based on quantized input data.

[0184] Example 59 includes the apparatus of Example 46, wherein the components for generating the lookup table are used to perform exponent-specific range reduction based on the input tensor values.

[0185] Example 60 includes the apparatus of Example 59, wherein the component for outputting the vector is used to output the result of exponent-specific range reduction using floating-point representation.

[0186] The following claims are incorporated herein by reference. Although certain example systems, methods, apparatuses, and articles of manufacture have been disclosed herein, the scope of this patent is not limited thereto. Rather, this patent covers all systems, methods, apparatuses, and articles of manufacture that fall fully within the scope of the claims of this patent.

Claims

1. An apparatus comprising: At least one memory; Machine-readable instructions; as well as Programmable circuitry for instantiating or executing at least one of the machine-readable instructions to: A lookup table is generated based on the input elements associated with the training phase of the deep neural network; The lookup table is indexed using tensor values, which are associated with the output index; as well as Based on the lookup table, a vector is output in a target numeric format, the vector including output values ​​in floating-point representation.

2. The apparatus according to claim 1, wherein, The programmable circuitry is used to generate the lookup table based on the expected input elements.

3. The apparatus according to claim 1 or 2, wherein, The programmable circuit is used to regenerate the lookup table based on the current input element.

4. The apparatus according to claim 1, 2 or 3, wherein, The programmable circuit is used to index a second lookup table based on a first output of a first lookup table.

5. The apparatus according to claim 4, wherein, The programmable circuit is used to combine the first output of the first lookup table and the second output of the second lookup table into an index of the third lookup table, wherein the first output and the second output are combined based on a numerical function.

6. The apparatus according to claim 1, 2 or 3, wherein, The programmable circuit is used to generate a histogram of the input elements.

7. The apparatus according to claim 6, wherein, The programmable circuitry is used to identify the index of the maximum value in the histogram and to scale the lookup table based on the maximum value.

8. The apparatus according to claim 7, wherein, The programmable circuitry is used to scale the lookup table to prevent overflow.

9. A method comprising: A lookup table is generated based on input elements associated with the training phase of a deep neural network by utilizing instructions executed by at least one processor. The lookup table is indexed using tensor values ​​associated with output indexes by executing instructions using at least one processor. as well as By executing instructions using at least one processor, a vector in target numeric format is output based on the lookup table, the vector including output values ​​in floating-point representation.

10. The method of claim 9, further comprising generating the lookup table based on the expected input elements.

11. The method of claim 9 or 10, further comprising generating the lookup table based on the current input element.

12. The method of claim 9, 10 or 11, further comprising indexing a second lookup table based on a first output of the first lookup table.

13. The method of claim 12, further comprising: The first output of the first lookup table and the second output of the second lookup table are combined to form the index of the third lookup table, wherein the first output and the second output are combined based on a numerical function.

14. The method of claim 9, 10 or 11, further comprising generating a histogram of the input elements.

15. The method of claim 14, further comprising identifying an index of the maximum value in the histogram and scaling the lookup table based on the maximum value.

16. The method of claim 15, further comprising scaling the lookup table to prevent overflow.

17. A non-transitory machine-readable storage medium comprising instructions that, when executed, cause processor circuitry to at least: A lookup table is generated based on the input elements associated with the training phase of the deep neural network; The lookup table is indexed using tensor values, which are associated with the output index; as well as Based on the lookup table, a vector is output in a target numeric format, the vector including output values ​​in floating-point representation.

18. The non-transitory machine-readable storage medium of claim 17, wherein the instructions, when executed, cause the processor to generate the lookup table based on the expected input elements.

19. The non-transitory machine-readable storage medium according to claim 17 or 18, wherein, When the instruction is executed, it causes the processor to regenerate the lookup table based on the current input elements.

20. The non-transitory machine-readable storage medium of claim 17, 18 or 19, wherein the instructions, when executed, cause the processor to base a second lookup table on a first output index of a first lookup table.

21. The non-transitory machine-readable storage medium of claim 20, wherein the instructions, when executed, cause the processor to combine the first output of the first lookup table and the second output of the second lookup table into an index of the third lookup table, the first output and the second output being combined based on a numerical function.

22. The non-transitory machine-readable storage medium according to claim 17, 18 or 19, wherein, When the instruction is executed, it causes the processor to generate a histogram of the input elements.

23. The non-transitory machine-readable storage medium of claim 22, wherein the instructions, when executed, cause the processor to identify the index of the maximum value in the histogram and scale the lookup table based on the maximum value.

24. The non-transitory machine-readable storage medium of claim 23, wherein the instructions, when executed, cause the processor to scale the lookup table to prevent overflow.

25. The non-transitory machine-readable storage medium according to claim 17, 18 or 19, wherein, When the instruction is executed, it causes the processor to identify a unique input value, the total number of unique input values ​​being less than the total number of input values, and to calculate a result output associated with the unique input value, the result output being associated with the input value.