Data processing method and data processing apparatus
By using secondary indexing to determine the segmented range of data bits in neural network systems, the problem of processor performance bottleneck in the prior art is solved, and more efficient processor performance and lower design costs are achieved.
Patent Information
- Application Number
- PCT/CN2024/099928
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-06-18
- Publication Date
- 2025-05-30
AI Technical Summary
Due to the nonlinear nature of the activation function in existing neural network systems, there are performance bottlenecks in the performance of the processor. How to improve the performance of the processor has become an urgent problem.
Using the secondary indexing method, the processor core first determines the first segment range where the data bit is located based on the first index, and then determines the second segment range where the data bit is located based on the second index. The bit width of the second index is related to the slope of the objective function in the first segment range, reducing the segment range and storage space of the objective function.
By reducing the segmentation range and storage space of the objective function, the performance of the processor core is improved, and the design cost and circuit area are reduced.
Smart Images

Figure CN2024099928_30052025_PF_FP_ABST
Abstract
Description
Data processing method and data processing device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 20, 2023, with application number 202311554717.6 and application name “Data Processing Method and Data Processing Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of chip technology, and in particular to a data processing method and a data processing device. Background Art
[0003] Today, various implementations of artificial intelligence (AI) and machine learning are driving innovation in many technological fields. Among them, mainstream neural networks such as convolutional neural networks and recurrent neural networks have achieved remarkable results in areas such as image classification and image processing.
[0004] In current neural network systems, the results of each layer must be processed through an activation function to increase the nonlinearity of the neural network model. The continuous development of activation functions is a key component in the continuous improvement of neural network systems. However, neural network systems typically require a large number of parameters and computational complexity, and the nonlinear nature of activation functions creates performance bottlenecks in processors. Therefore, improving processor performance has become a pressing issue.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a data processing method and a data processing device, which improve the problem of low processor performance.
[0007] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions.
[0008] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to a data processing device, wherein the data processing device includes a processor core, an input interface, and an output interface. The method includes: the processor core obtains input data of the target function through the input interface, the target function includes multiple first segmentation ranges, each first segmentation range includes multiple second segmentation ranges, the multiple second segmentation ranges correspond to multiple polynomial coefficients one-to-one, the input data includes a first index, a second index, and a data bit, the first index corresponds to the first segmentation range where the data bit is located, the second index corresponds to the second segmentation range where the data bit is located, and the bit width of the second index is related to the slope of the target function in the first segmentation range. The processor core determines the polynomial coefficient corresponding to the data bit based on the first index and the second index. The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain the output data of the target function, and outputs the output data of the target function through the output interface.
[0009] In the data processing method provided by the embodiment of the present application, a secondary indexing method is adopted, and the processor core first determines the first segmentation range where the data bit is located based on the first index, wherein the first segmentation range can be understood as a coarse-grained segment. The processor core then determines the second segmentation range where the data bit is located based on the second index, wherein the second segmentation range can be understood as a fine-grained segment. In addition, the bit width of the second index is related to the slope of the target function in the first segmentation range, wherein if the slope of the target function in the first segmentation range is large, a second index with a larger bit width can be used, and if the slope of the target function in the first segmentation range is small, a second index with a smaller bit width can be used. Thus, the number of segmentation ranges of the target function can be reduced, the storage space for storing polynomial coefficients can be reduced, and the performance of the processor core can be improved. Moreover, the method provided by the embodiment of the present application is uniformly segmented, and there is no need to use a comparator to obtain the segmentation range, which further reduces the design cost and circuit area.
[0010] In one possible design, the processor core determines the polynomial coefficient corresponding to the data bit based on a first index and a second index of input data of the objective function, including: the processor core determines, based on the first index, a lookup table corresponding to the first index, the lookup table storing a correspondence between the second segment range and the polynomial coefficient; the processor core determines, based on the second index, the polynomial coefficient corresponding to the second segment range indicated by the second index in the lookup table, and determines the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
[0011] This design uses a secondary indexing approach. First, the lookup table is determined based on the first index. The lookup table identifier corresponds one-to-one with the first index, and therefore one-to-one with the first segment range. The polynomial coefficients stored in the lookup table are then determined based on the second index. Because the bit width of the second index is related to the slope of the objective function within the first segment range, this reduces the number of segments of the objective function, reduces the storage space occupied by the lookup table, and improves processor core performance.
[0012] In one possible design, the polynomial coefficient includes a high-order field and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, when the lookup table stores the first polynomial coefficient among the multiple polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the low-order field of the polynomial coefficient is stored.
[0013] In this design, a partial storage method of omitting the same high-order bits in the same lookup table and only storing the low-order bits is adopted, which can further reduce the storage space occupied by the lookup table and improve the performance of the processor core.
[0014] In one possible design, the polynomial coefficient includes a high-order field and a low-order field. If the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the multiple polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the high-order field of the polynomial coefficient is stored.
[0015] In this design, a partial storage method is adopted in which the low-order bits that meet the threshold in the same lookup table are omitted and only the high-order bits are stored. This can further reduce the storage space occupied by the lookup table and improve the performance of the processor core.
[0016] In one possible design, the polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of multiple polynomial coefficients are the same and the difference between the low-order fields of the multiple polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the multiple polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
[0017] In this design, a partial storage method is adopted in which the same high-order bits and low-order bits that meet the threshold in the same lookup table are omitted and only the middle bits are stored. This can further reduce the storage space occupied by the lookup table and improve the performance of the processor core.
[0018] In one possible design, the first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width; the second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width.
[0019] In this design, only data bits in the input data participate in the polynomial calculation, and the bit width of the polynomial calculation is small, which can reduce hardware overhead.
[0020] In one possible design, if the input variable of the target function is floating-point data, the processor core obtains the input data of the target function through the input interface, including: the processor core obtains the floating-point data through the input interface. The processor core separates the floating-point data according to characteristics of the target function to obtain a sign bit, an exponent value, and the input data of the target function.
[0021] In one possible design, the processor core performs polynomial calculations based on data bits and polynomial coefficients to obtain output data of the target function, including: the processor core performs polynomial calculations based on data bits and polynomial coefficients to obtain a calculation result of the target function; the processor core processes the calculation result, sign bit, and exponent value according to the characteristics of the target function to obtain the output data of the target function.
[0022] In this design, the data processing method provided in the embodiments of the present application also supports the calculation process of the objective function of floating-point data. Before determining the polynomial coefficients, floating-point preprocessing is performed, that is, separating the floating-point number to obtain the sign bit, exponent value, and input data. The calculation result of the objective function is then subjected to floating-point postprocessing, that is, combining the calculation result, sign bit, and exponent value to obtain the output data of the objective function.
[0023] In a second aspect, an embodiment of the present application provides a data processing device, which includes: a processor core, an input interface, and an output interface. The input interface is used to obtain input data of the objective function, the objective function includes multiple first segmentation ranges, each first segmentation range includes multiple second segmentation ranges, the multiple second segmentation ranges correspond one-to-one to multiple polynomial coefficients, the input data includes a first index, a second index, and a data bit, the first index corresponds to the first segmentation range where the data bit is located, the second index corresponds to the second segmentation range where the data bit is located, and the bit width of the second index is related to the slope of the objective function in the first segmentation range. The processor core is used to determine the polynomial coefficient corresponding to the input data based on the first index and the second index. The processor core is also used to perform polynomial calculations based on the data bits and the polynomial coefficients to obtain the output data of the objective function. The output interface is used to output the output data of the objective function.
[0024] In one possible design, the processor core is specifically configured to, based on the first index, determine a lookup table whose lookup table identifier corresponds to the first index, where the lookup table stores a correspondence between the second segment range and the polynomial coefficient. Based on the second index, determine, in the lookup table, the polynomial coefficient corresponding to the second segment range indicated by the second index, and determine the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
[0025] In one possible design, the polynomial coefficient includes a high-order field and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, when the lookup table stores the first polynomial coefficient among the multiple polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the low-order field of the polynomial coefficient is stored.
[0026] In one possible design, the polynomial coefficient includes a high-order field and a low-order field. If the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the multiple polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the high-order field of the polynomial coefficient is stored.
[0027] In one possible design, the polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of multiple polynomial coefficients are the same and the difference between the low-order fields of the multiple polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the multiple polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
[0028] In one possible design, the first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width; the second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width.
[0029] In one possible design, if the input variable of the objective function is floating-point data, the processor core is specifically configured to obtain the floating-point data through an input interface, separate the floating-point data according to the characteristics of the objective function, and obtain the sign bit, exponent value, and input data of the objective function.
[0030] In one possible design, the processor core is specifically used to perform polynomial calculations based on data bits and polynomial coefficients to obtain the calculation results of the target function; the calculation results, sign bits and exponent values are processed according to the characteristics of the target function to obtain the output data of the target function.
[0031] The beneficial effects of the second aspect can be found in the description of the first aspect.
[0032] In a third aspect, an embodiment of the present application provides a chip, comprising a data processing device and a memory, wherein the data processing device is configured to read and execute program instructions stored in the memory to implement the method of the first aspect.
[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the data processing method in any possible implementation of the first aspect above.
[0034] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer or processor, enables the computer or processor to execute the data processing method in the first aspect and any possible implementation thereof.
[0035] It can be understood that any of the data processing devices, chips, computer-readable storage media or computer program products provided above can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods and will not be repeated here.
[0036] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] FIG1 is a schematic diagram of a segmented error-equalizing polynomial approximation method and a uniform piecewise polynomial approximation method provided by an embodiment of the present application;
[0038] FIG2 is a hardware architecture diagram of an error equalization polynomial approximation method according to an embodiment of the present application;
[0039] FIG3 is a flow chart of a method for applying uniform piecewise polynomial approximation provided in an embodiment of the present application;
[0040] FIG4 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0041] FIG5 is a flow chart of a data processing method provided in an embodiment of the present application;
[0042] FIG6 is a schematic diagram of the structure of a lookup table provided in an embodiment of the present application;
[0043] FIG7 is a flowchart of a processing of an objective function provided in an embodiment of the present application;
[0044] FIG8 is a processing flow chart of another objective function provided in an embodiment of the present application;
[0045] FIG9 is a processing flow chart of another objective function provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] For ease of understanding, some examples of concepts related to the embodiments of this application are provided for reference as follows:
[0047] Floating point data (FP), the Institute of Electrical and Electronics Engineers (IEEE) has developed IEEE 754 as the binary floating-point arithmetic standard, which defines floating-point data representation methods such as double-precision FP64, single-precision FP32, and half-precision FP16. Among them, for FP64 data, the sign field is 1 bit, the exponent field is 11 bits, and the mantissa field is 52 bits. For FP32 data, the sign field is 1 bit, the exponent field is 8 bits, and the mantissa field is 23 bits. For FP16 data, the sign field is 1 bit, the exponent field is 5 bits, and the mantissa field is 10 bits.
[0048] A scalar computation unit (CMU) is a circuit designed for scalar calculations. A scalar, also known as a pure quantity, has only magnitude and no direction. Scalar calculations are often used for general-purpose computing. The execution units (EXUs) of the multi-stage pipelines of central processing units (CPUs) and other similar processors can incorporate floating-point arithmetic logic units (ALUs).
[0049] A vector computing unit (VCU) is a computing unit specially designed for vector computing with a certain degree of parallelism, such as a single instruction multiple data (SIMD) processor. A vector, also known as a vector, typically refers to a one-dimensional array with a length greater than 1. Vector computing units are commonly used in fields such as high performance computing (HPC) and AI machine learning, including solving mathematical problems such as linear programming, Fourier transforms, filtering calculations, and linear algebra, partial differential equations, and integration. In a vector computing accelerator unit or vector processor, an arithmetic execution unit (vector unit) based on floating-point data format can be embedded.
[0050] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0051] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0052] At present, the nonlinear characteristics of the activation function in the neural network affect the performance of the processor, so many studies are devoted to the efficient approximation of nonlinear functions. Nonlinear activation functions may include sigmoid function, tanh function, softmax function, ReLU function, ELU function, and PReLU function. Among them, nonlinear functions can be implemented by iterative methods or piecewise polynomial approximations. For example, iterative methods may include Newton iteration method and coordinate rotation digital computer (CORDIC) method, etc. However, iterative methods require a longer delay to achieve the target accuracy of nonlinear functions. The piecewise polynomial approximation has advantages in terms of delay and is more suitable for current AI processors that have high requirements for real-time processing performance.
[0053] The piecewise polynomial approximation method divides the objective function into several segments, each segment corresponds to a polynomial, and the polynomial can be of any order, such as the first-order polynomial y=a*x+b and the second-order polynomial y=a*x 2 +b*x+c, where x is the input variable, y is the output variable, and a, b, and c are the polynomial coefficients. In this method, after the software determines the segment range and polynomial coefficients, the hardware uses a lookup table (LUT) to store the polynomial coefficients for each segment. The hardware then determines the polynomial coefficients for the segment corresponding to the input variable based on the value of the input variable. The multiplication and addition unit then completes the polynomial calculation to obtain the final approximate result.
[0054] Specifically, piecewise polynomial approximation can be divided into error-equalizing polynomial approximation and uniform piecewise polynomial approximation, as shown in Figure 1. Figure 1 (a) shows a piecewise schematic diagram of the error-equalizing polynomial approximation method, and Figure 1 (b) shows a piecewise schematic diagram of the uniform piecewise polynomial approximation method. Among them, the error-equalizing polynomial approximation uses segments of unequal length to fit the objective function, making the errors of each segment the same. Its advantage is that it can minimize the number of segments through a piecewise algorithm, thereby reducing the entries in the lookup table storing the polynomial coefficients. However, the error-equalizing polynomial approximation requires an additional comparator to obtain the segment index of the input variable. The uniform piecewise polynomial approximation divides the objective function into segments of equal length and directly uses the high bits of the input variable as an index to look up the polynomial coefficients stored in the lookup table. However, the number of segments of the uniform piecewise polynomial approximation increases with the increase of the target accuracy.
[0055] The error equalization polynomial approximation and uniform piecewise polynomial approximation are further introduced below.
[0056] As shown in Figure 2, Figure 2 is a hardware architecture diagram of an error-balanced polynomial approximation method provided by an embodiment of the present application. The error-balanced polynomial approximation method divides the target function into several segments through a segmentation algorithm, ensures that the maximum absolute value error of the original function value and the polynomial approximation value of each segment is the same, and obtains the range and polynomial coefficient of each segment, thereby achieving the minimum number of segments under the target accuracy. In the hardware implementation, a lookup table is used to store the polynomial coefficients of each segment, and a group of parallel comparators are used to obtain the segment index where the input variable is located. Among them, the starting points of each segment are x2, x3, x4, ..., x n , the starting point of a segment is also the end point of the previous segment, that is, x3 is the end point of the first segment and the starting point of the second segment. The end point of the last segment is not shown in Figure 2. Among them, the input variable x needs to be compared with the starting point to obtain the index of the segment it is in. The index can be expressed as {S1, S2, ..., S n-1}, S1 is the comparison result of x and x2, S2 is the comparison result of x and x3, S n-1 For x and x n The multiplexer determines the corresponding polynomial coefficient [k i , b i ], wherein the lookup table may store n polynomial coefficients, such as [k1, b1], [k2, b2], [k3, b3], ..., [k n , b n Then input variable x and polynomial coefficient [k i , b i ] Perform multiplication and addition operations to obtain the value of y.
[0057] However, since the segments are non-uniform, a set of comparators is required to obtain the segment index, which incurs significant hardware overhead. Furthermore, the input variable x uses its full bit width for polynomial calculations, and the multiplication and addition operations have a large bit width, which also incurs significant hardware overhead.
[0058] As shown in Figure 3, Figure 3 is a flowchart of an application of a uniform piecewise polynomial approximation method provided in an embodiment of the present application. The uniform piecewise polynomial approximation method divides the target function into several equal-length segments, applies polynomial approximation to each segment, and obtains a polynomial system through a software algorithm. In hardware implementation, the polynomial coefficients are stored in a lookup table, and the high bit of the input variable x is used as an index to obtain the polynomial coefficients of the corresponding segment. Finally, the polynomial operation is performed through the multiplication and addition unit to obtain the final approximation result y.
[0059] The process includes: inputting the input variable x into the separator, obtaining the high-order index part (x_u) and the low-order data part (x_l), and using the high-order index part to look up the polynomial coefficients stored in the table, such as a, b, and c, where a is x 2 The coefficient of x, b is the coefficient of x, and c is a constant. In addition, the low-order data part is squared to obtain x_l 2 Therefore, the low-order data part (x_1), the square of the low-order data part (x_1 2 ) and the polynomial coefficients (a, b and c) into the second-order polynomial (y = a*x_l 2 +b*x_l+c) to perform polynomial operations and obtain the approximate result y corresponding to the input variable x.
[0060] However, because the objective function varies across different input ranges, more segments are required to meet the target accuracy in steeply varying regions, while fewer segments are sufficient in slowly varying regions. The uniform piecewise polynomial approximation uses the high-order bits of the input variables directly as segment indices, which results in an unnecessary number of segments in the slowly varying regions of the objective function, increasing the lookup table overhead and resulting in a large area cost.
[0061] Thus, an embodiment of the present application provides a data processing method, in which a secondary indexing method is adopted, wherein the processor core first determines the first segmentation range where the data bit is located based on the first index, wherein the first segmentation range can be understood as a coarse-grained segment. The processor core then determines the second segmentation range where the data bit is located based on the second index, wherein the second segmentation range can be understood as a fine-grained segment. In addition, the bit width of the second index is related to the slope of the target function in the first segmentation range, wherein if the slope of the target function in the first segmentation range is large, a second index with a larger bit width can be used, and if the slope of the target function in the first segmentation range is small, a second index with a smaller bit width can be used. Thus, the number of segmentation ranges of the target function can be reduced, the storage space for storing polynomial coefficients can be reduced, and the performance of the processor core can be improved. Moreover, the methods provided in the embodiments of the present application are all uniformly segmented, and there is no need to use a comparator to obtain the segmentation range, further reducing design cost and circuit area.
[0062] For ease of understanding, the data processing device applied to the data processing method provided in the embodiment of the present application is first introduced below. In the above scenario, the data processing method and device of the present application can be applied to different systems or devices, such as applied to an execution device. The execution device can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, a vehicle terminal, etc., and can also be a server cluster, etc. The data processing method provided in the present application can be applied to scenarios related to function calculations such as a central processing unit, a graphics processing unit (GPU), a neural processor unit (NPU), high-performance computing and AI in the execution device, such as a scalar computing unit and a vector computing unit.
[0063] Specifically, the data processing method can be applied to division and square root operations on fixed-point and floating-point data by a central processing unit. It can also be applied to logarithmic operations, exponential operations, square root operations, trigonometric functions, reciprocal operations, and square root reciprocal operations on floating-point data by a graphics processing unit, and to logarithmic operations, exponential operations, square root operations, division operations, and reciprocal operations on floating-point data by a neural processing unit.
[0064] In some embodiments, the data processing device proposed in the present application may be a chip, for example, the chip is a system-on-a-chip (SoC). As shown in Figure 4, Figure 4 is a structural diagram of a data processing device provided in an embodiment of the present application. The data processing device may include a processor core, an input interface (input) and an output interface (output). Among them, the data processing device may include one processor core, and the data processing device may also include multiple processor cores. In addition, the data processing device may also include a memory, and instructions or data may be stored in the memory. After the processor core loads the data and application in the memory, it processes the data, such as performing the polynomial calculation processing of the data bits and polynomial coefficients in the embodiment of the present application. In addition, in an embodiment of the present application, the input interface can be used to obtain the input variables of the objective function, and the output interface can be used to output the output data of the objective function. In one example, in a neural network system, the output data of the objective function can be used as the input data of the next convolution layer.
[0065] In addition to SoC, the chip may also be a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0066] Applied to the above-mentioned data processing device, the data processing method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0067] The present application provides a data processing method, as shown in FIG5 , which is a flow chart of a data processing method provided by the present application. The method includes the following steps:
[0068] S501. The processor core obtains input data of an objective function through an input interface. The objective function includes multiple first segment ranges, each of which includes multiple second segment ranges. The multiple second segment ranges correspond one-to-one to multiple polynomial coefficients. The input data includes a first index, a second index, and a data bit. The first index corresponds to the first segment range in which the data bit is located, the second index corresponds to the second segment range in which the data bit is located, and the bit width of the second index is related to the slope of the objective function in the first segment range.
[0069] For example, the input data of the objective function can be a scalar or a vector. If the input data of the objective function is a scalar, a polynomial operation can be performed on the input data of the objective function by a scalar calculation unit in the processor core. If the input data of the objective function is a vector, a polynomial operation can be performed on the input data of the objective function by a vector calculation unit in the processor core.
[0070] Exemplarily, the first segmentation range can be understood as a coarse-grained segment, and the second segmentation range can be understood as a fine-grained segment. Both the first segmentation range and the second segmentation range are uniform segmentations. In a specific example, assuming that the objective function is a tanh function, due to the symmetry of the tanh function, only the positive half axis of the tanh function on the x-axis can be segmented. Specifically, the first segmentation range can be [0,4), [4,8), [8,12), ..., [n,n+4). In the first segmentation range [0,4), if the slope of the objective function is large, that is, the objective function is steeper, the second segmentation range can be [0,1), [1,2), [2,3) and [3,4). If the slope of the objective function is small, that is, the objective function is relatively flat, the second segmentation range can be [0,2) and [2,4).
[0071] The input data may include a first index (represented by x_u in the embodiment of the present application), a second index (represented by x_m in the embodiment of the present application), and a data bit (represented by x_l in the embodiment of the present application). The first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width. The second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width. If the input variable is floating-point data, the input data may be (1. mantissa field), and the fixed bit is 1.
[0072] Exemplarily, the second bit width is related to the slope of the objective function in the first segmentation range. The larger the second bit width, the greater the number of second segmentation ranges in the first segmentation range, and the smaller the second bit width, the smaller the number of second segmentation ranges in the first segmentation range. Thus, the number of segmentation ranges of the objective function can be reduced. In addition, the second bit width is also related to the target accuracy of the objective function. The greater the second bit width, the higher the target accuracy of the objective function, and the smaller the second bit width, the lower the target accuracy of the objective function.
[0073] Specifically, taking a 23-bit binary stream as an example, assuming the input data is 23'b01011010010010000000010, the uniform piecewise polynomial approximation method uses a single indexing scheme. The high-order 8 bits can be used as the index, and the remaining 15 bits as the data bits for the polynomial operation. Therefore, the range of this method's segmentation can be 2^8 = 256.
[0074] In an embodiment of the present application, the first bit width may be 3 bits, and the second bit width may be 3 bits, 4 bits, or 5 bits. The first index may be a bit value starting from the highest bit and having a bit width of 3 bits, and the first index may be 3'b010. If the second bit width is 3 bits, the second index may be 3'b110, and the data bits are 17'b10010010000000010. If the second bit width is 4 bits, the second index may be 4'b1101, and the data bits are 16'b0010010000000010. If the second bit width is 5 bits, the second index may be 5'b11010, and the data bits are 15'b010010000000010.
[0075] Therefore, the embodiment of the present application includes 2^3=8 first segment ranges. Assuming that the second bit widths corresponding to the 8 first segment ranges are 5, 4, 3, 3, 3, 3, 3 and 3 respectively, the first first segment range is divided into 2^5=32 second segment ranges, the second first segment range is divided into 2^4=16 second segment ranges, and the third first segment range to the eighth first segment range are divided into 2^3=8 second segment ranges respectively. Thus, the 8 first segment ranges are divided into 32+16+8+8+8+8+8+8=96 second segment ranges in total. Compared with the 256 segment ranges used in the uniform piecewise polynomial approximation method, the data processing method used in the embodiment of the present application greatly reduces the number of segments.
[0076] S502: The processor core determines polynomial coefficients corresponding to the data bits based on the first index and the second index.
[0077] Exemplarily, the processor core first determines the first segment range of the polynomial coefficient in the lookup table based on the first index, and then determines the second segment range through the second index to determine the polynomial coefficient corresponding to the second segment range. If the target function is a first-order polynomial y=a*x+b, the polynomial coefficients are a and b. If the target function is a second-order polynomial y=a*x 2 +b*x+c, the polynomial coefficients are a, b and c. It is understandable that if the target accuracy of the objective function is higher, the objective function can also be a third-order polynomial or above.
[0078] Optionally, S502 may include: the processor core determining, based on the first index, a lookup table corresponding to the first index and a lookup table identifier, the lookup table storing a correspondence between the second segment range and the polynomial coefficient. The processor core determining, based on the first index, a polynomial coefficient in the lookup table corresponding to the second segment range indicated by the second index, and determining the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
[0079] For example, as shown in Figure 6, Figure 6 is a schematic diagram of the structure of a lookup table provided in an embodiment of the present application. Figure 6 shows 2^(x_u) lookup tables. Assuming that the first index is 3 bits, it can include 2^3=8 lookup tables, and the lookup table identifiers can be 3'b000, 3'b001, 3'b010, 3'b011, 3'b100, 3'b101, 3'b110 and 3'b111. Figure 6 also specifically shows three lookup tables, such as LUT-A, LUT-B and LUT-C. Each LUT stores 2^(x_m) entries, and each entry corresponds to a polynomial coefficient. The processor core first determines the identifier of the lookup table based on the first index, and further determines the entry in the identifier of the lookup table based on the second index.
[0080] S503: The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the target function, and outputs the output data of the target function through the output interface.
[0081] For example, since the second bit width of the second index in the input data varies, the processor core should align the data bits before performing the polynomial calculation based on the data bits and the polynomial coefficients to ensure that the data bits have the same bit width. In a specific example, if the data bits are 15'b010010000000010, 2 bits can be added to the low order bits of the data bits to obtain 17'b01001000000001000. If the data bits are 16'b0010010000000010, 1 bit can be added to the low order bits of the data bits to obtain 17'b00100100000000100.
[0082] For example, if the objective function is a first-order polynomial y=a*x+b, as shown in FIG7 , FIG7 is a processing flow chart of an objective function provided by an embodiment of the present application. The input data x of the objective function is separated by a separator to obtain a first index (x_u), a second index (x_m), and a data bit (x_l). The first index and the second index are used to search the lookup table to obtain the polynomial coefficients a and b. The data bits and the polynomial coefficients are substituted into the first-order polynomial (y=a*x_l+b) to perform a polynomial operation to obtain the output data y of the objective function.
[0083] For example, if the objective function is a second-order polynomial y=a*x 2+b*x+c, as shown in Figure 8, which is a processing flow chart of another objective function provided by an embodiment of the present application. In which, the input data x of the objective function is separated by a separator to obtain a first index (x_u), a second index (x_m) and a data bit (x_l). The first index and the second index are used to search the lookup table to obtain the polynomial coefficients a, b and c. The data bits are squared to obtain x_l 2 , substitute the data bits and polynomial coefficients into the second-order polynomial (y = a*x_l 2 +b*x_l+c) to perform polynomial operations and obtain the output result y of the objective function.
[0084] Optionally, if the input variable of the objective function is floating-point data, S501 may include: the processor core obtains the floating-point data through the input interface, and the processor core separates the floating-point data according to the characteristics of the objective function to obtain the sign bit (sign), exponent value (exponent) and input data of the objective function.
[0085] In addition, S503 may include: the processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain a calculation result of the target function. The processor core also processes the calculation result, the sign bit, and the exponent value according to the characteristics of the target function to obtain output data of the target function.
[0086] For example, if the input data is floating point data, taking floating point data FP32 as an example, the input data may be 32'b00111110001011010010000000010, then the sign bit of the floating point data is 0, the exponent value is 8'b01111100, and the input data is 23'b01011010010010000000010.
[0087] For example, as shown in FIG9 , FIG9 is a flowchart of another objective function processing provided by an embodiment of the present application. Assume that the objective function is a second-order polynomial y=a*x 2 +b*x+c, where floating-point preprocessing is performed on the input variable to obtain the sign bit, exponent value, and input data x. The input data x is separated by a separator to obtain the first index (x_u), the second index (x_m), and the data bit (x_l). The first index and the second index are used to search the lookup table to obtain the polynomial coefficients a, b, and c. The data bit is squared to obtain x_l 2 , polynomial operations are performed on the data bits and polynomial coefficients to obtain the calculation result of the objective function. The calculation result, sign bit, and exponent value are then subjected to floating-point post-processing to obtain the output data y of the objective function.
[0088] In one possible example, in order to achieve the target accuracy, for computing units with half-precision requirements, the second segmented range can use a first-order polynomial to fit the objective function; for computing units with single-precision requirements and double-precision requirements, the second segmented range can use a second-order polynomial to fit the objective function.
[0089] Optionally, the polynomial coefficient may include a high-order field and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, when the lookup table stores the first polynomial coefficient among the multiple polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the low-order field of the polynomial coefficient is stored.
[0090] In a specific example, taking a binary stream with 8-bit polynomial coefficients as an example, assume that the first polynomial coefficient is 8'b10110011, the second polynomial coefficient is 8'b10011110, ..., and the nth polynomial coefficient is 8'b10100100. The high-order 2 bits of the polynomial coefficients are the same, both 2'b10. Therefore, when storing n polynomial coefficients, the lookup table can omit the high-order 2 bits. Specifically, when storing n polynomial coefficients, the lookup table can store the first polynomial coefficient as 8'b10110011, the second polynomial coefficient as 6'b011110, ..., and the nth polynomial coefficient as 6'b100100. This reduces the storage space occupied by the lookup table.
[0091] Optionally, the polynomial coefficient may include a high-order field and a low-order field. If the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, when the lookup table stores the first polynomial coefficient among the multiple polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the high-order field of the polynomial coefficient is stored.
[0092] Continuing with the above example, the difference between the low-order 2 bits of the first and second polynomial coefficients is 2'b01, the difference between the low-order 2 bits of the first and n-th polynomial coefficients is 2'b11, and the difference between the low-order 2 bits of the second and n-th polynomial coefficients is 2'b10. It can be considered that the difference between the low-order 2 bits of the n polynomial coefficients meets the preset threshold. Therefore, when storing n polynomial coefficients in the lookup table, the storage of the low-order 2 bits can be omitted. Specifically, when storing n polynomial coefficients, the lookup table can store the first polynomial coefficient as 8'b10110011, the second polynomial coefficient as 6'b100111, ..., and the n-th polynomial coefficient as 6'b101001. This reduces the storage space occupied by the lookup table.
[0093] Optionally, the polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of multiple polynomial coefficients are the same and the difference between the low-order fields of the multiple polynomial coefficients is less than a threshold, then when the lookup table stores the first polynomial coefficient among the multiple polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
[0094] Continuing with the above example, where the upper 2 bits of multiple polynomial coefficients are identical, and the difference between the lower 2 bits of the first and second polynomial coefficients is 2'b01, the difference between the lower 2 bits of the first and nth polynomial coefficients is 2'b11, and the difference between the lower 2 bits of the second and nth polynomial coefficients is 2'b10, it can be considered that the difference between the lower 2 bits of the n polynomial coefficients meets the preset threshold. In this case, when storing the n polynomial coefficients in the lookup table, both the upper 2 bits and the lower 2 bits can be omitted. Specifically, when storing n polynomial coefficients, the lookup table can store the first polynomial coefficient as 8'b10110011, the second polynomial coefficient as 4'b1001, ..., and the nth polynomial coefficient as 4'b1010. Thus, the storage space occupied by the lookup table can be reduced.
[0095] Thus, the data processing method provided in the embodiment of the present application can reduce the segmentation range of the target function through the method of secondary indexing, reduce the storage space used to store polynomial coefficients, and improve the performance of the processor core. In addition, in the embodiment of the present application, no comparator is required to obtain the segmentation range, further reducing design costs and circuit area. In addition, when storing polynomial coefficients in the lookup table, a partial storage method is adopted, which reduces the storage space occupied by the lookup table.
[0096] An embodiment of the present application further provides an electronic device comprising one or more processors and one or more memories. The one or more memories are coupled to the one or more processors and are configured to store computer program code, the computer program code comprising computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the above-mentioned related method steps to implement the data processing method in the above-mentioned embodiment.
[0097] It is understandable that in order to implement the above functions, the electronic device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of this application.
[0098] In this embodiment, the electronic device can be divided into functional modules according to the above-mentioned method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into a single processing module. The above-mentioned integrated modules can be implemented in the form of hardware. It should be noted that the module division in this embodiment is illustrative and is only a logical functional division. In actual implementation, other division methods may be used.
[0099] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the data processing method in the above-mentioned embodiment.
[0100] An embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the data processing method executed by the electronic device in the above-mentioned embodiment.
[0101] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the data processing method performed by the electronic device in the above-mentioned method embodiments.
[0102] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.
[0103] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0105] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0107] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0108] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: The data processing method is applied to a data processing device, the data processing device includes a processor core, an input interface and an output interface, and the method includes: The processor core obtains input data of an objective function through the input interface, wherein the objective function includes a plurality of first segment ranges, each of the first segment ranges includes a plurality of second segment ranges, the plurality of second segment ranges correspond to a plurality of polynomial coefficients one-to-one, the input data includes a first index, a second index, and a data bit, the first index corresponds to the first segment range where the data bit is located, the second index corresponds to the second segment range where the data bit is located, and the bit width of the second index is related to the slope of the objective function in the first segment range; The processor core determines the polynomial coefficient corresponding to the data bit based on the first index and the second index; The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the objective function, and outputs the output data of the objective function through the output interface.
2. The method according to claim 1, characterized in that The processor core determines the polynomial coefficients corresponding to the data bits based on the first index and the second index of the input data of the target function, including: The processor core determines, according to the first index, a lookup table whose lookup table identifier corresponds to the first index, wherein the lookup table stores a correspondence between the second segment range and the polynomial coefficients; The processor core determines, according to the second index, a polynomial coefficient in the lookup table corresponding to the second segment range indicated by the second index, and determines the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
3. The method according to claim 2, characterized in that The polynomial coefficients include a high-order field and a low-order field. If the high-order fields of the plurality of polynomial coefficients are the same, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the low-order field of the polynomial coefficient is stored.
4. The method according to claim 2, characterized in that: The polynomial coefficient includes a high-order field and a low-order field. If a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold value, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the high-order field of the polynomial coefficient is stored.
5. The method according to claim 2, characterized in that: The polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of a plurality of polynomial coefficients are the same and a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
6. The method according to any one of claims 1 to 5, characterized in that The first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width. The second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width.
7. The method according to claim 1, characterized in that If the input variable of the objective function is floating point data, the processor core obtains the input data of the objective function through the input interface, including: The processor core obtains the floating-point data through the input interface; The processor core separates the floating-point data according to the characteristics of the target function to obtain the sign bit, exponent value and input data of the target function.
8. The method according to claim 7, characterized in that The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the objective function, including: The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain a calculation result of the objective function; The processor core processes the calculation result, the sign bit, and the exponent value according to the objective function characteristics to obtain the output data of the objective function.
9. A data processing device, characterized in that: The data processing device comprises: a processor core, an input interface and an output interface; The input interface is used to obtain input data of the target function, the target function includes multiple first segment ranges, each of the first segment ranges includes multiple second segment ranges, the multiple second segment ranges correspond to multiple polynomial coefficients one by one, the input data includes a first index, a second index and a data bit, the first index corresponds to the first segment range where the data bit is located, The second index corresponds to the second segment range where the data bit is located, and the bit width of the second index is related to the slope of the target function in the first segment range; The processor core is configured to determine a polynomial coefficient corresponding to the data bit based on the first index and the second index; The processor core is further configured to perform polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the objective function; The output interface is used to output the output data of the objective function.
10. The device according to claim 9, characterized in that The processor core is specifically configured to determine, according to the first index, a lookup table whose lookup table identifier corresponds to the first index, wherein the lookup table stores a correspondence between the second segment range and polynomial coefficients; According to the second index, a polynomial coefficient in the lookup table corresponding to the second segment range indicated by the second index is determined, and the polynomial coefficient is determined as the polynomial coefficient corresponding to the data bit.
11. The device according to claim 10, characterized in that The polynomial coefficients include a high-order field and a low-order field. If the high-order fields of the plurality of polynomial coefficients are the same, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the low-order field of the polynomial coefficient is stored.
12. The device according to claim 10, characterized in that The polynomial coefficient includes a high-order field and a low-order field. If a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold value, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the high-order field of the polynomial coefficient is stored.
13. The device according to claim 10, characterized in that The polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of a plurality of polynomial coefficients are the same and a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
14. The device according to any one of claims 9 to 13, characterized in that The first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width. The second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width.
15. The device according to claim 9, characterized in that If the input variable of the objective function is floating-point data, the processor core is specifically used to obtain the floating-point data through the input interface; separate the floating-point data according to the characteristics of the objective function to obtain the sign bit, exponent value and input data of the objective function.
16. The device according to claim 15, characterized in that The processor core is specifically used to perform polynomial calculation based on the data bits and the polynomial coefficients to obtain the calculation result of the objective function; and process the calculation result, the sign bit and the exponent value according to the characteristics of the objective function to obtain the output data of the objective function.
17. A chip, characterized in that: The chip includes: a data processing device and a memory; wherein the data processing device is used to read and execute program instructions stored in the memory to implement the method described in any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data processing method and data processing device
CN120020674A
Digital pre-distortion and post-distortion based on segmentwise piecewise polynomial approximation
CN105320492A
Acceleration unit for a deep learning engine
CN110197111A
Data processing method and device, processor and data searching method and device
CN113296732A
Neural network activation method and device, NPU, equipment and storage medium
CN115759217A
Cited By
Disturbance detection method, device and equipment based on trend extraction and control system
CN121637131A