Data processing method and data processing device
By using secondary indexing to determine the segmented range of data bits in neural network systems, the problem of processor performance bottleneck in the prior art is solved, and more efficient processor performance and lower design costs are achieved.
Patent Information
- Application Number
- CN202311554717.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
Due to the nonlinear nature of the activation function in existing neural network systems, there are performance bottlenecks in the performance of the processor. How to improve the performance of the processor has become an urgent problem.
Using the secondary indexing method, the processor core first determines the first segment range where the data bit is located based on the first index, and then determines the second segment range where the data bit is located based on the second index. The bit width of the second index is related to the slope of the objective function in the first segment range, reducing the segment range and storage space of the objective function.
By reducing the segmentation range and storage space of the objective function, the performance of the processor core is improved, and no need to use a comparator to obtain the segmentation range, further reducing design costs and circuit area.
Smart Images

Figure CN120020674A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of chip technology, and in particular, to a data processing method and a data processing device. Background Art
[0002] Nowadays, various implementations of artificial intelligence (AI) and machine learning are driving innovation in many technical fields. Among them, mainstream neural networks such as convolutional neural networks and recurrent neural networks have achieved remarkable results in fields such as image classification and image processing.
[0003] In the current neural network system, the results of each layer of the network need to be processed by an activation function to increase the nonlinearity of the neural network model. The continuous development of the activation function is an important link in the continuous progress and improvement of the neural network system. However, the neural network system usually requires a large number of parameters and a large amount of computation. The non-linear characteristics of the activation function cause a performance bottleneck in the performance of the processor. Therefore, how to improve the performance of the processor has become an urgent problem to be solved. Summary of the Invention
[0004] The embodiments of the present application provide a data processing method and a data processing device, which improve the problem of low processor performance.
[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions.
[0006] In a first aspect, the embodiments of the present application provide a data processing method. The data processing method is applied to a data processing device, and the data processing device includes a processor core, an input interface, and an output interface. The method includes: The processor core obtains input data of a target function through the input interface. The target function includes a plurality of first segmentation ranges, each first segmentation range includes a plurality of second segmentation ranges, the plurality of second segmentation ranges correspond to a plurality of polynomial coefficients one by one, the input data includes a first index, a second index, and data bits, the first index corresponds to the first segmentation range where the data bits are located, the second index corresponds to the second segmentation range where the data bits are located, and the bit width of the second index is related to the slope of the target function in the first segmentation range. The processor core determines the polynomial coefficient corresponding to the data bit based on the first index and the second index. The processor core performs polynomial calculation based on the data bit and the polynomial coefficient to obtain output data of the target function, and outputs the output data of the target function through the output interface.
[0007] In the data processing method provided by the embodiments of the present application, a secondary indexing method is adopted. First, the processor core determines the first segmentation range where the data bit is located based on the first index. Here, the first segmentation range can be understood as a coarse-grained segment. Then, the processor core determines the second segmentation range where the data bit is located based on the second index. Here, the second segmentation range can be understood as a fine-grained segment. In addition, the bit width of the second index is related to the slope of the objective function in the first segmentation range. Specifically, if the slope of the objective function in the first segmentation range is large, a second index with a larger bit width can be used; if the slope of the objective function in the first segmentation range is small, a second index with a smaller bit width can be used. Thus, the number of segmentation ranges of the objective function can be reduced, the storage space for storing polynomial coefficients can be decreased, and the performance of the processor core can be improved. Moreover, in the method provided by the embodiments of the present application, all segmentations are uniform, and a comparator is not required to obtain the segmentation range, further reducing the design cost and circuit area.
[0008] In a possible design, the processor core determines the polynomial coefficient corresponding to the data bit based on the first index and the second index of the input data of the objective function, including: The processor core determines the lookup table whose identifier corresponds to the first index according to the first index. The correspondence between the second segmentation range and the polynomial coefficient is stored in the lookup table. The processor core determines the polynomial coefficient corresponding to the second segmentation range indicated by the second index in the lookup table according to the second index, and determines the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
[0009] In this design, a secondary indexing method is adopted. First, the lookup table is determined according to the first index. Here, the lookup table identifier corresponds one-to-one with the first index, that is, it corresponds one-to-one with the first segmentation range. Then, the polynomial coefficient stored in the lookup table is determined according to the second index. Since the bit width of the second index is related to the slope of the objective function in the first segmentation range, the number of segments of the objective function can be reduced, the storage space occupied by the lookup table can be decreased, and the performance of the processor core can be improved.
[0010] In a possible design, the polynomial coefficient includes a high-order field and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, the high-order field and the low-order field of the first polynomial coefficient are stored. When storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, only the low-order field of the polynomial coefficient is stored.
[0011] In this design, a partial storage method of omitting the same high-order bits in the same lookup table and only storing the low-order bits is adopted, which can further reduce the storage space occupied by the lookup table and improve the performance of the processor core.
[0012] In a possible design, the polynomial coefficient includes a high-order field and a low-order field. If the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field of the polynomial coefficient.
[0013] In this design, by adopting a partial storage method of omitting the low-order bits that meet the threshold in the same lookup table and only storing the high-order bits, the storage space occupied by the lookup table can be further reduced, and the performance of the processor core can be improved.
[0014] In a possible design, the polynomial coefficient includes a high-order field, a middle-order field, and a low-order field. If the high-order fields of multiple polynomial coefficients are the same and the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the middle-order field of the polynomial coefficient.
[0015] In this design, by adopting a partial storage method of omitting the same high-order bits and the low-order bits that meet the threshold in the same lookup table and only storing the middle-order bits, the storage space occupied by the lookup table can be further reduced, and the performance of the processor core can be improved.
[0016] In a possible design, the first index starts from the highest bit except the fixed bits in the input data and has a bit width of the first bit width, and the second index starts from the highest bit except the first index in the input data and has a bit width of the second bit width.
[0017] In this design, only the data bits in the input data participate in the polynomial calculation, and the bit width of the polynomial calculation is small, which can reduce the hardware overhead.
[0018] In a possible design, if the input variable of the objective function is floating-point data, the processor core obtains the input data of the objective function through the input interface, including: the processor core obtains the floating-point data through the input interface. The processor core separates the floating-point data according to the characteristics of the objective function to obtain the sign bit, exponent value, and input data of the objective function.
[0019] In a possible design, the processor core performs polynomial calculation based on the data bits and polynomial coefficients to obtain the output data of the objective function, including: the processor core performs polynomial calculation based on the data bits and polynomial coefficients to obtain the calculation result of the objective function; the processor core processes the calculation result, sign bit, and exponent value according to the characteristics of the objective function to obtain the output data of the objective function.
[0020] In this design, the data processing method provided by the embodiments of the present application also supports the calculation process of the objective function of floating-point data. Before determining the polynomial coefficients, floating-point preprocessing is performed, that is, the floating-point number is separated to obtain the sign bit, exponent value, and input data. Then, floating-point postprocessing is performed on the calculation result of the objective function, that is, the calculation result, sign bit, and exponent value are combined to obtain the output data of the objective function.
[0021] In a second aspect, the embodiments of the present application provide a data processing device, which includes: a processor core, an input interface, and an output interface. The input interface is used to obtain the input data of the objective function. The objective function includes multiple first segmentation ranges, each first segmentation range includes multiple second segmentation ranges, and the multiple second segmentation ranges correspond to multiple polynomial coefficients one by one. The input data includes a first index, a second index, and data bits. The first index corresponds to the first segmentation range where the data bits are located, the second index corresponds to the second segmentation range where the data bits are located, and the bit width of the second index is related to the slope of the objective function in the first segmentation range. The processor core is used to determine the polynomial coefficient corresponding to the input data based on the first index and the second index. The processor core is further used to perform polynomial calculation based on the data bits and the polynomial coefficient to obtain the output data of the objective function. The output interface is used to output the output data of the objective function.
[0022] In a possible design, the processor core is specifically configured to determine, according to the first index, a lookup table whose identifier corresponds to the first index. The lookup table stores the correspondence between the second segmentation range and the polynomial coefficient. According to the second index, determine the polynomial coefficient corresponding to the second segmentation range indicated by the second index in the lookup table, and determine the polynomial coefficient as the polynomial coefficient corresponding to the data bits.
[0023] In a possible design, the polynomial coefficient includes a high-order field and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the low-order field of the polynomial coefficient.
[0024] In a possible design, the polynomial coefficient includes a high-order field and a low-order field. If the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field of the polynomial coefficient.
[0025] In a possible design, the polynomial coefficient includes a high-order field, a middle-order field, and a low-order field. If the high-order fields of multiple polynomial coefficients are the same and the difference between the low-order fields of the multiple polynomial coefficients is less than a threshold, when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the middle-order field of the polynomial coefficient.
[0026] In a possible design, the first index starts from the highest bit except the fixed bits in the input data and has a bit width of the first bit width, and the second index starts from the highest bit except the first index in the input data and has a bit width of the second bit width.
[0027] In a possible design, if the input variable of the objective function is floating-point data, the processor core is specifically configured to obtain the floating-point data through the input interface. Separate the floating-point data according to the characteristics of the objective function to obtain the sign bit, the exponent value, and the input data of the objective function.
[0028] In a possible design, the processor core is specifically configured to perform polynomial calculations based on the data bits and the polynomial coefficients to obtain the calculation result of the objective function; process the calculation result, the sign bit, and the exponent value according to the characteristics of the objective function to obtain the output data of the objective function.
[0029] For the beneficial effects of the second aspect, reference can be made to the description of the first aspect.
[0030] In a third aspect, an embodiment of the present application provides a chip, which includes a data processing device and a memory. The data processing device is configured to read and execute program instructions stored in the memory to implement the method of the first aspect.
[0031] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, which, when running on an electronic device, cause the electronic device to execute the data processing method in any possible implementation manner of the first aspect above.
[0032] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a computer or a processor, causes the computer or the processor to execute the data processing method in the first aspect and any possible implementation manner thereof.
[0033] It can be understood that any of the above-provided data processing devices, chips, computer-readable storage media, or computer program products can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, which will not be elaborated here.
[0034] These aspects or other aspects of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A segmentation schematic diagram of an error equalization polynomial approximation method and a uniform segmented polynomial approximation method provided for an embodiment of the present application;
[0036] Figure 2 A hardware architecture diagram of an application of an error equalization polynomial approximation method provided for an embodiment of the present application;
[0037] Figure 3 A flowchart of an application of a uniform segmented polynomial approximation method provided for an embodiment of the present application;
[0038] Figure 4 A structural schematic diagram of a data processing device provided for an embodiment of the present application;
[0039] Figure 5 A flowchart of a data processing method provided for an embodiment of the present application;
[0040] Figure 6 A structural schematic diagram of a look-up table provided for an embodiment of the present application;
[0041] Figure 7 A processing flowchart of an objective function provided for an embodiment of the present application;
[0042] Figure 8 Another processing flowchart of an objective function provided for an embodiment of the present application;
[0043] Figure 9 Yet another processing flowchart of an objective function provided for an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] For ease of understanding, some descriptions of concepts related to the embodiments of the present application are given as examples for reference. As follows:
[0045] For floating point (FP) data, the Institute of Electrical and Electronics Engineers (IEEE) has developed IEEE 754 as the binary floating-point arithmetic standard, which defines floating-point data representation methods such as double-precision FP64, single-precision FP32, and half-precision FP16. Among them, for FP64 data, the sign field is 1 bit, the exponent field is 11 bits, and the mantissa field is 52 bits. For FP32 data, the sign field is 1 bit, the exponent field is 8 bits, and the mantissa field is 23 bits. For FP16 data, the sign field is 1 bit, the exponent field is 5 bits, and the mantissa field is 10 bits.
[0046] Scalar calculation unit: The circuit for scalar calculation is called a scalar calculation unit. Here, a scalar, also known as a pure quantity, has only magnitude and no direction. Scalar calculation is mostly used in general computing. In the execution unit (EXU) part of the multi-stage pipeline of a central processing unit (CPU) and the scalar calculation part of other processors with similar functions, an arithmetic logic unit (ALU) based on the floating-point data format can be embedded.
[0047] Vector calculation unit: A calculation unit with a certain degree of parallelism specially designed for vector calculation, such as a single instruction multiple data (SIMD) processor. Here, a vector, also known as a vector quantity, usually refers to a one-dimensional array with a length greater than 1. Vector calculation units are mostly used in high-performance computing (HPC) and AI machine learning and other fields, including the solution of mathematical problems such as linear programming, Fourier transform, filtering calculation, and linear algebra, partial differential equations, and integrals. In a vector calculation acceleration unit or a vector processor, an arithmetic execution unit (vector unit) based on the floating-point data format can be embedded.
[0048] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.
[0049] Hereinafter, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.
[0050] Currently, the non-linear characteristics of the activation function in the neural network affect the performance of the processor. Therefore, many studies are dedicated to the efficient approximation of non-linear functions. Non-linear activation functions may include sigmoid function, tanh function, softmax function, ReLU function, ELU function, and PReLU function. Among them, non-linear functions can be implemented by iterative methods or piecewise polynomial approximation. For example, iterative methods may include Newton iterative method and coordinate rotation digital computer (CORDIC), etc. However, iterative methods require a long delay to achieve the target accuracy of non-linear functions. Piecewise polynomial approximation has an advantage in terms of delay and is more suitable for current AI processors with high requirements for real-time processing performance.
[0051] The piecewise polynomial approximation method divides the target function into several segments, and each segment corresponds to a polynomial. The polynomial can be of any order, such as the first-order polynomial y = a * x + b and the second-order polynomial y = a * x 2 + b * x + c, where x is the input variable, y is the output variable, and a, b, and c are polynomial coefficients. In this method, after the software determines the segmentation range and polynomial coefficients, the polynomial coefficients of each segment are stored in a look-up table (LUT) in the hardware, and the polynomial coefficients of the segment corresponding to the input variable are determined according to the value of the input variable. Then, the polynomial calculation is completed using a multiply-accumulate unit to obtain the final approximation result.
[0052] Specifically, piecewise polynomial approximation can be divided into error-equalized polynomial approximation and uniform piecewise polynomial approximation, as Figure 1 shown, Figure 1 in (a) shows the segmentation schematic diagram of the error-equalized polynomial approximation method, Figure 1Figure (b) in [reference] shows a sectional schematic diagram of the uniform segmented polynomial approximation method. Among them, the error-equalized polynomial approximation uses segments of unequal lengths to fit the target function, making the errors of each segment the same. Its advantage is that the minimum number of segments can be achieved through the segmentation algorithm, thereby reducing the entries in the lookup table storing the polynomial coefficients. However, the error-equalized polynomial approximation requires an additional comparator to obtain the segment index of the input variable. The uniform segmented polynomial approximation divides the target function into segments of equal length and directly uses the high-order bits of the input variable as the index to look up the polynomial coefficients stored in the lookup table. However, the number of segments of the uniform segmented polynomial approximation increases with the increase of the target accuracy.
[0053] The error-equalized polynomial approximation and the uniform segmented polynomial approximation are further introduced below.
[0054] As Figure 2 shown, Figure 2 This is a hardware architecture diagram provided by an embodiment of the present application using the error-equalized polynomial approximation method. The error-equalized polynomial approximation method divides the target function into several segments through a segmentation algorithm, ensuring that the maximum absolute value error between the original function value and the polynomial approximation value of each segment is the same, and obtaining the range and polynomial coefficients of each segment. Thus, the minimum number of segments can be obtained under the target accuracy. In hardware implementation, a lookup table is used to store the polynomial coefficients of each segment, and a set of parallel comparators is used to obtain the segment index where the input variable is located. Among them, the starting points of each segment are x 2 , x 3 , x 4 , ……, x n . The starting point of a certain segment is also the end point of the previous segment, that is, x 3 is the end point of the first segment and also the starting point of the second segment. Figure 2 The end point of the last segment is not shown in [reference]. Among them, the input variable x needs to be compared with the starting point to obtain the segment index where it is located, and this index can be expressed as {S 1 , S 2 , ……, S n-1}, where S 1 is the comparison result of x and x 2 , S 2 is the comparison result of x and x 3 , and S n-1 is the comparison result of x and x n . The multiplexer determines the corresponding polynomial coefficients [k i , b i from multiple polynomial coefficients in the lookup table according to the segment index. Among them, n polynomial coefficients can be stored in the lookup table. For example, [k 1 , b 1 , [k2 , b 2 , [k 3 , b 3 , ……, [k n , b n . Then, perform multiplication and addition operations on the input variable x and the polynomial coefficients [k i , b i to obtain the value of y.
[0055] However, since the segmentation is non-uniform, a set of comparators is required to obtain the segmentation index, which will incur a large hardware overhead. In addition, the input variable x participates in the polynomial calculation using the full bit width, and the bit width of the multiplication and addition operations is large, which will also bring a large hardware overhead.
[0056] As Figure 3 shown, Figure 3 is a flowchart of an application of the uniform segmentation polynomial approximation method provided by an embodiment of the present application. The uniform segmentation polynomial approximation method divides the objective function into several segments of equal length, approximates each segment using a polynomial, and obtains a polynomial system through a software algorithm. When implemented in hardware, the polynomial coefficients are stored in a lookup table, the high bits of the input variable x are used as an index to obtain the polynomial coefficients of the corresponding segment, and finally, a multiplication and addition unit is used to perform polynomial operations to obtain the final approximate result y.
[0057] Among them, the process includes: inputting the input variable x into a separator to obtain a high-bit index part (x_u) and a low-bit data part (x_l), obtaining the polynomial coefficients stored in the lookup table through the high-bit index part, such as a, b, and c, where a is the coefficient of x 2 , b is the coefficient of x, and c is a constant. In addition, the low-bit data part is squared (square) to obtain x_l 2 . Thus, substituting the low-bit data part (x_l), the square of the low-bit data part (x_l 2 ) and the polynomial coefficients (a, b, and c) into a second-order polynomial (y = a * x_l 2 + b * x_l + c) for polynomial operations to obtain an approximate result y corresponding to the input variable x.
[0058] However, since the trend of change of the objective function is different in different input ranges, more segments are required in the steep part of the change to meet the target accuracy, while only a small number of segments are needed in the slow part of the change to meet the target accuracy. The characteristic of the uniform segmentation polynomial approximation is to directly use the high bits of the input variable as the segmentation index, which will result in using too many unnecessary segments in the slow part of the objective function, increasing the overhead of the lookup table, and thus bringing a large area overhead.
[0059] Accordingly, an embodiment of the present application provides a data processing method. In this method, a secondary indexing method is adopted. The processor core first determines the first segmentation range where the data bit is located based on the first index. Here, the first segmentation range can be understood as a coarse-grained segment. Then, the processor core determines the second segmentation range where the data bit is located based on the second index. Here, the second segmentation range can be understood as a fine-grained segment. Additionally, the bit width of the second index is related to the slope of the objective function in the first segmentation range. Specifically, if the slope of the objective function in the first segmentation range is large, a second index with a larger bit width can be used. If the slope of the objective function in the first segmentation range is small, a second index with a smaller bit width can be used. Thereby, the number of segmentation ranges of the objective function can be reduced, the storage space for storing polynomial coefficients can be decreased, and the performance of the processor core can be improved. Moreover, in the method provided by the embodiment of the present application, all segmentations are uniform, and a comparator does not need to be used to obtain the segmentation range, further reducing the design cost and circuit area.
[0060] For ease of understanding, the data processing device to which the data processing method provided by the embodiment of the present application is applied will be introduced first. In the above scenario, the data processing method and device of the present application can be applied to different systems or devices, such as an execution device. The execution device can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, a vehicle-mounted terminal, etc., or it can also be a server cluster, etc. The data processing method provided by the present application can be applied to scenarios related to function calculations in an execution device, such as a central processing unit, a graphics processing unit (GPU), a neural processor unit (NPU), high-performance computing, and AI, for example, a scalar calculation unit and a vector calculation unit.
[0061] Specifically, this data processing method can be applied to division operations and square root operations on fixed-point data and floating-point data by a central processing unit. It can also be applied to logarithmic operations, exponential operations, square root operations, trigonometric function operations, reciprocal operations, and square root reciprocal operations on floating-point data by a graphics processing unit. It can also be applied to logarithmic operations, exponential operations, square root operations, division operations, and reciprocal operations on floating-point data by a neural processor unit.
[0062] In some embodiments, the data processing device proposed by the present application can be a chip. For example, this chip is a system-on-a-chip (SoC). As Figure 4 shown, Figure 4A structural schematic diagram of a data processing device provided by an embodiment of the present application. The data processing device may include a processor core, an input interface (input), and an output interface (output). Among them, the data processing device may include one processor core, and the data processing device may also include multiple processor cores. In addition, the data processing device may also include a memory, and instructions or data may be stored in the memory. After loading the data and application programs in the memory, the processor core may process the data, for example, perform polynomial calculation processing on data bits and polynomial coefficients in the embodiment of the present application. In addition, in the embodiment of the present application, the input interface may be used to obtain the input variables of the objective function, and the output interface may be used to output the output data of the objective function. In one example, in a neural network system, the output data of the objective function may be used as the input data of the next convolutional layer.
[0063] In addition to the SoC, the chip may also be a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), etc.
[0064] Applied to the above data processing device, the data processing method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0065] The embodiment of the present application provides a data processing method, as Figure 5 shown, Figure 5 is a flowchart of a data processing method provided by an embodiment of the present application. The method includes the following processes:
[0066] S501. The processor core obtains the input data of the objective function through the input interface. The objective function includes multiple first segmentation ranges, each first segmentation range includes multiple second segmentation ranges, and the multiple second segmentation ranges correspond to multiple polynomial coefficients one by one. The input data includes a first index, a second index, and data bits. The first index corresponds to the first segmentation range where the data bits are located, the second index corresponds to the second segmentation range where the data bits are located, and the bit width of the second index is related to the slope of the objective function in the first segmentation range.
[0067] Exemplarily, the input data of the objective function may be a scalar or a vector. If the input data of the objective function is a scalar, the input data of the objective function may be subjected to polynomial operations through the scalar calculation unit in the processor core. If the input data of the objective function is a vector, the input data of the objective function may be subjected to polynomial operations through the vector calculation unit in the processor core.
[0068] Exemplarily, the first segmented range can be understood as a coarse-grained segment, and the second segmented range can be understood as a fine-grained segment. Both the first segmented range and the second segmented range are uniform segments. In a specific example, assuming the objective function is the tanh function, due to the symmetry of the tanh function, only the positive half-axis of the x-axis of the tanh function needs to be segmented. Specifically, the first segmented range can be [0, 4), [4, 8), [8, 12), ……, [n, n + 4). In the first segmented range [0, 4), if the slope of the objective function is large, that is, the objective function is relatively steep, the second segmented range can be [0, 1), [1, 2), [2, 3), and [3, 4). If the slope of the objective function is small, that is, the objective function is relatively gentle, the second segmented range can be [0, 2) and [2, 4).
[0069] Among them, the input data may include a first index (denoted as x_u in the embodiments of the present application), a second index (denoted as x_m in the embodiments of the present application), and data bits (denoted as x_l in the embodiments of the present application). Among them, the first index starts from the highest bit except for the fixed bits in the input data, and the bit width is the bit value of the first bit width. The second index starts from the highest bit except for the first index in the input data, and the bit width is the bit value of the second bit width. Among them, if the input variable is floating-point data, then the input data can be (1. mantissa field), and the fixed bit is 1.
[0070] Exemplarily, the second bit width is related to the slope of the objective function in the first segmented range. Among them, the larger the second bit width, the more the number of the second segmented ranges in the first segmented range; the smaller the second bit width, the fewer the number of the second segmented ranges in the first segmented range. Thus, the number of segmented ranges of the objective function can be reduced. In addition, the second bit width is also related to the target accuracy of the objective function. Among them, the more the second bit width, the higher the target accuracy of the objective function; the fewer the second bit width, the lower the target accuracy of the objective function.
[0071] Specifically, taking the input data as a 23-bit binary stream as an example, assuming the input data is 23’b01011010010010000000010. In the uniform segmented polynomial approximation method, which adopts a one-time indexing method, the high 8 bits can be used as the index, and the remaining 15 bits are used as data bits for polynomial operations. Thus, the segmented range of this method can be 2^8 = 256.
[0072] In the embodiments of the present application, the first bit width may be 3 bits, and the second bit width may be 3 bits, 4 bits or 5 bits. Among them, the first index may be a bit value starting from the highest bit with a bit width of 3 bits, so the first index may be 3’b010. If the second bit width is 3 bits, the second index may be 3’b110, and at this time the data bit is 17’b10010010000000010. If the second bit width is 4 bits, the second index may be 4’b1101, and at this time the data bit is 16’b0010010000000010. If the second bit width is 5 bits, the second index may be 5’b11010, and at this time the data bit is 15’b010010000000010.
[0073] Therefore, there are 2^3 = 8 first segmentation ranges in the embodiments of the present application. Assuming that the second bit widths corresponding to the 8 first segmentation ranges are 5, 4, 3, 3, 3, 3, 3, and 3 respectively, the first first segmentation range is divided into 2^5 = 32 second segmentation ranges, the second first segmentation range is divided into 2^4 = 16 second segmentation ranges, and the third first segmentation range to the eighth first segmentation range are each divided into 2^3 = 8 second segmentation ranges. Therefore, the 8 first segmentation ranges are divided into a total of 32 + 16 + 8 + 8 + 8 + 8 + 8 + 8 = 96 second segmentation ranges. Compared with the 256 segmentation ranges used in the uniform segmentation polynomial approximation method, the data processing method used in the embodiments of the present application greatly reduces the number of segments.
[0074] S502. The processor core determines the polynomial coefficient corresponding to the data bit based on the first index and the second index.
[0075] Exemplarily, the processor core first determines the first segmentation range where the polynomial coefficient is located in the lookup table based on the first index, and then determines the second segmentation range through the second index to determine the polynomial coefficient corresponding to the second segmentation range. If the objective function is a first-order polynomial y = a*x + b, the polynomial coefficients are a and b. If the objective function is a second-order polynomial y = a*x 2 + b*x + c, the polynomial coefficients are a, b, and c. It can be understood that if the target accuracy of the objective function is relatively high, the objective function can also be a third-order polynomial or higher.
[0076] Optionally, S502 may include: The processor core determines the lookup table corresponding to the first index according to the first index, and the correspondence between the second segmentation range and the polynomial coefficient is stored in the lookup table. The processor core determines the polynomial coefficient corresponding to the second segmentation range indicated by the second index in the lookup table according to the first index, and determines the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
[0077] Exemplarily, as Figure 6 shownFigure 6 A structural schematic diagram of a lookup table provided by an embodiment of the present application. Among them, Figure 6 2^(x_u) lookup tables are shown. Assuming that the first index is 3 bits, 2^3 = 8 lookup tables can be included, and the lookup table identifiers can be 3’b000, 3’b001, 3’b010, 3’b011, 3’b100, 3’b101, 3’b110, and 3’b111. Figure 6 3 lookup tables are also specifically shown, such as LUT-A, LUT-B, and LUT-C. Among them, each LUT stores 2^(x_m) entries, and each entry corresponds to a polynomial coefficient. Among them, the processor core first determines the identifier of the lookup table based on the first index, and further determines the entry in the identifier of the lookup table based on the second index.
[0078] S503. The processor core performs polynomial calculation based on the data bit and the polynomial coefficient to obtain the output data of the objective function, and outputs the output data of the objective function through the output interface.
[0079] Exemplarily, since the second bit width of the second index in the input data is variable, the processor core should perform an alignment operation on the data bit before performing polynomial calculation based on the data bit and the polynomial coefficient to ensure that the bit widths of the data bits are the same. In a specific example, if the data bit is 15’b010010000000010, 2 bits can be added to the low bit of the data bit to obtain 17’b01001000000001000. If the data bit is 16’b0010010000000010, 1 bit can be added to the low bit of the data bit to obtain 17’b00100100000000100.
[0080] Exemplarily, if the objective function is a first-order polynomial y = a*x + b, as Figure 7 shown, Figure 7 A processing flow chart of an objective function provided by an embodiment of the present application. Among them, the input data x of the objective function is separated by a separator to obtain a first index (x_u), a second index (x_m), and a data bit (x_l), the lookup table is searched through the first index and the second index to obtain the polynomial coefficients a and b, and the data bit and the polynomial coefficients are substituted into the first-order polynomial (y = a*x_l + b) for polynomial operation to obtain the output data y of the objective function.
[0081] Exemplarily, if the objective function is a second-order polynomial y = a*x 2 + b*x + c, as Figure 8 shown, Figure 8It is a processing flow chart of another objective function provided by an embodiment of the present application. Among them, the input data x of the objective function is separated by a separator to obtain a first index (x_u), a second index (x_m), and data bits (x_l). The polynomial coefficients a, b, and c are obtained by looking up a lookup table through the first index and the second index. The data bits are squared to obtain x_l 2 , and the data bits and polynomial coefficients are substituted into a second-order polynomial (y = a * x_l 2 + b * x_l + c) for polynomial operation to obtain the output result y of the objective function.
[0082] Optionally, if the input variable of the objective function is floating-point data, S501 may include: The processor core obtains the floating-point data through the input interface. The processor core separates the floating-point data according to the characteristics of the objective function to obtain the sign bit (sign), exponent value (exponent), and input data of the objective function.
[0083] In addition, S503 may include: The processor core performs polynomial calculation based on the data bits and polynomial coefficients to obtain the calculation result of the objective function. The processor core also processes the calculation result, sign bit, and exponent value according to the characteristics of the objective function to obtain the output data of the objective function.
[0084] Exemplarily, if the input data is floating-point data, taking the floating-point data FP32 as an example, the input data may be 32’b00111110001011010010010000000010, then the sign bit of the floating-point data is 0, the exponent value is 8’b01111100, and the input data is 23’b01011010010010000000010.
[0085] Exemplarily, as Figure 9 shown, Figure 9 It is a processing flow chart of yet another objective function provided by an embodiment of the present application. Assume that the objective function is a second-order polynomial y = a * x 2 + b * x + c. Among them, floating-point preprocessing is performed on the input variable to obtain the sign bit, exponent value, and input data x. The input data x is separated by a separator to obtain a first index (x_u), a second index (x_m), and data bits (x_l). The polynomial coefficients a, b, and c are obtained by looking up a lookup table through the first index and the second index. The data bits are squared to obtain x_l 2 , and polynomial operation is performed on the data bits and polynomial coefficients to obtain the calculation result of the objective function. Then, floating-point post-processing is performed on the calculation result, sign bit, and exponent value to obtain the output data y of the objective function.
[0086] In a possible example, to achieve the target accuracy, for a computing unit with a half-precision requirement, the second segmented range can use a first-order polynomial to fit the objective function, and for a computing unit with a single-precision requirement and a double-precision requirement, the second segmented range can use a second-order polynomial to fit the objective function.
[0087] Optionally, the polynomial coefficient can include a high-order field and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, when storing the first polynomial coefficient among multiple polynomial coefficients in the lookup table, store the high-order field and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among multiple polynomial coefficients in the lookup table, store the low-order field of the polynomial coefficient.
[0088] In a specific example, taking the polynomial coefficient as an 8-bit binary stream as an example, assume the first polynomial coefficient is 8’b10110011, the second polynomial coefficient is 8’b10011110,..., and the nth polynomial coefficient is 8’b10100100. Among them, the high-order 2 bits of the polynomial coefficients are the same, both 2’b10. Then the lookup table can omit storing the high-order 2 bits when storing n polynomial coefficients. Specifically, when storing n polynomial coefficients in the lookup table, it can store the first polynomial coefficient as 8’b10110011, the second polynomial coefficient as 6’b011110,..., and the nth polynomial coefficient as 6’b100100. Thus, the storage space occupied by the lookup table can be reduced.
[0089] Optionally, the polynomial coefficient can include a high-order field and a low-order field. If the difference between the low-order fields of multiple polynomial coefficients is less than a threshold, when storing the first polynomial coefficient among multiple polynomial coefficients in the lookup table, store the high-order field and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among multiple polynomial coefficients in the lookup table, store the high-order field of the polynomial coefficient.
[0090] Continuing to refer to the above example, among them, the difference between the low-order 2 bits of the first polynomial coefficient and the second polynomial coefficient is 2’b01, the difference between the low-order 2 bits of the first polynomial coefficient and the nth polynomial coefficient is 2’b11, and the difference between the low-order 2 bits of the second polynomial coefficient and the nth polynomial coefficient is 2’b10. It can be considered that the difference between the low-order 2 bits of the n polynomial coefficients meets the preset threshold. Then the lookup table can omit storing the low-order 2 bits when storing n polynomial coefficients. Specifically, when storing n polynomial coefficients in the lookup table, it can store the first polynomial coefficient as 8’b10110011, the second polynomial coefficient as 6’b100111,..., and the nth polynomial coefficient as 6’b101001. Thus, the storage space occupied by the lookup table can be reduced.
[0091] Optionally, the polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of multiple polynomial coefficients are the same, and the difference between the low-order fields of the multiple polynomial coefficients is less than a threshold, then when storing the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient, and when storing the polynomial coefficients other than the first polynomial coefficient among the multiple polynomial coefficients in the lookup table, store the middle-order field of the polynomial coefficient.
[0092] Continuing to refer to the above example, where the high 2 bits of multiple polynomial coefficients are the same, and the difference between the low 2 bits of the first polynomial coefficient and the second polynomial coefficient is 2'b01, the difference between the low 2 bits of the first polynomial coefficient and the nth polynomial coefficient is 2'b11, and the difference between the low 2 bits of the second polynomial coefficient and the nth polynomial coefficient is 2'b10. It can be considered that the difference between the low 2 bits of the n polynomial coefficients satisfies the preset threshold. Then, when storing the n polynomial coefficients in the lookup table, the high 2 bits and the low 2 bits can be omitted simultaneously. Specifically, when the lookup table stores the n polynomial coefficients, it can store the first polynomial coefficient as 8'b10110011, the second polynomial coefficient as 4'b1001,..., and the nth polynomial coefficient as 4'b1010. Thereby, the storage space occupied by the lookup table can be reduced.
[0093] Therefore, the data processing method provided by the embodiments of the present application can reduce the segmented range of the objective function, reduce the storage space for storing polynomial coefficients, and improve the performance of the processor core through the secondary indexing method. And in the embodiments of the present application, a comparator is not required to obtain the segmented range, further reducing the design cost and circuit area. In addition, when storing polynomial coefficients in the lookup table, a partial storage method is adopted, reducing the storage space occupied by the lookup table.
[0094] The embodiments of the present application further provide an electronic device, including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code. The computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device is caused to execute the above-related method steps to implement the data processing method in the above embodiments.
[0095] It can be understood that, in order to implement the above functions, the electronic device includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.
[0096] In this embodiment, the electronic device can be divided into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0097] An embodiment of the present application further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on the electronic device, the electronic device is enabled to execute the above-related method steps to implement the data processing method in the above embodiment.
[0098] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-related steps to implement the data processing method executed by the electronic device in the above embodiment.
[0099] In addition, an embodiment of the present application further provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected thereto; wherein, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory to enable the chip to execute the data processing method executed by the electronic device in each of the above method embodiments.
[0100] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0101] From the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0102] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0103] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0105] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs and other various media that can store program codes.
[0106] The above content is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A data processing method, characterized in that: The data processing method is applied to a data processing device, the data processing device includes a processor core, an input interface and an output interface, and the method includes: The processor core obtains input data of an objective function through the input interface, wherein the objective function includes a plurality of first segment ranges, each of the first segment ranges includes a plurality of second segment ranges, the plurality of second segment ranges correspond to a plurality of polynomial coefficients one-to-one, the input data includes a first index, a second index, and a data bit, the first index corresponds to the first segment range where the data bit is located, the second index corresponds to the second segment range where the data bit is located, and the bit width of the second index is related to the slope of the objective function in the first segment range; The processor core determines the polynomial coefficient corresponding to the data bit based on the first index and the second index; The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the objective function, and outputs the output data of the objective function through the output interface.
2. The method according to claim 1, characterized in that The processor core determines the polynomial coefficients corresponding to the data bits based on the first index and the second index of the input data of the target function, including: The processor core determines, according to the first index, a lookup table whose lookup table identifier corresponds to the first index, wherein the lookup table stores a correspondence between the second segment range and the polynomial coefficients; The processor core determines, according to the second index, a polynomial coefficient in the lookup table corresponding to the second segment range indicated by the second index, and determines the polynomial coefficient as the polynomial coefficient corresponding to the data bit.
3. The method according to claim 2, characterized in that The polynomial coefficients include a high-order field and a low-order field. If the high-order fields of the plurality of polynomial coefficients are the same, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the low-order field of the polynomial coefficient is stored.
4. The method according to claim 2, characterized in that: The polynomial coefficient includes a high-order field and a low-order field. If a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold value, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the high-order field of the polynomial coefficient is stored.
5. The method according to claim 2, characterized in that: The polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of a plurality of polynomial coefficients are the same and a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
6. The method according to any one of claims 1 to 5, characterized in that The first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width. The second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width.
7. The method according to claim 1, characterized in that If the input variable of the objective function is floating point data, the processor core obtains the input data of the objective function through the input interface, including: The processor core obtains the floating-point data through the input interface; The processor core separates the floating-point data according to the characteristics of the target function to obtain the sign bit, exponent value and input data of the target function.
8. The method according to claim 7, characterized in that The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the objective function, including: The processor core performs polynomial calculation based on the data bits and the polynomial coefficients to obtain a calculation result of the objective function; The processor core processes the calculation result, the sign bit, and the exponent value according to the objective function characteristics to obtain the output data of the objective function.
9. A data processing device, characterized in that: The data processing device comprises: a processor core, an input interface and an output interface; The input interface is used to obtain input data of an objective function, wherein the objective function includes a plurality of first segment ranges, each of the first segment ranges includes a plurality of second segment ranges, the plurality of second segment ranges correspond to a plurality of polynomial coefficients one-to-one, the input data includes a first index, a second index, and a data bit, the first index corresponds to the first segment range where the data bit is located, the second index corresponds to the second segment range where the data bit is located, and the bit width of the second index is related to the slope of the objective function in the first segment range; The processor core is configured to determine a polynomial coefficient corresponding to the data bit based on the first index and the second index; The processor core is further configured to perform polynomial calculation based on the data bits and the polynomial coefficients to obtain output data of the objective function; The output interface is used to output the output data of the objective function.
10. The device according to claim 9, characterized in that The processor core is specifically configured to determine, according to the first index, a lookup table whose lookup table identifier corresponds to the first index, wherein the lookup table stores a correspondence between the second segment range and polynomial coefficients; According to the second index, a polynomial coefficient in the lookup table corresponding to the second segment range indicated by the second index is determined, and the polynomial coefficient is determined as the polynomial coefficient corresponding to the data bit.
11. The device according to claim 10, characterized in that The polynomial coefficients include a high-order field and a low-order field. If the high-order fields of the plurality of polynomial coefficients are the same, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the low-order field of the polynomial coefficient is stored.
12. The device according to claim 10, characterized in that The polynomial coefficient includes a high-order field and a low-order field. If a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold value, when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the high-order field of the polynomial coefficient is stored.
13. The device according to claim 10, characterized in that The polynomial coefficients include a high-order field, a middle-order field, and a low-order field. If the high-order fields of a plurality of polynomial coefficients are the same and a difference value of the low-order fields of the plurality of polynomial coefficients is less than a threshold, then when the lookup table stores a first polynomial coefficient among the plurality of polynomial coefficients, the high-order field, the middle-order field, and the low-order field of the first polynomial coefficient are stored, and when the lookup table stores polynomial coefficients other than the first polynomial coefficient among the plurality of polynomial coefficients, the middle-order field of the polynomial coefficient is stored.
14. The device according to any one of claims 9 to 13, characterized in that The first index starts from the highest bit in the input data except the fixed bit, and the bit width is the bit value of the first bit width. The second index starts from the highest bit in the input data except the first index, and the bit width is the bit value of the second bit width.
15. The device according to claim 9, characterized in that If the input variable of the objective function is floating-point data, the processor core is specifically used to obtain the floating-point data through the input interface; separate the floating-point data according to the characteristics of the objective function to obtain the sign bit, exponent value and input data of the objective function.
16. The device according to claim 15, characterized in that The processor core is specifically used to perform polynomial calculation based on the data bits and the polynomial coefficients to obtain the calculation result of the objective function; and process the calculation result, the sign bit and the exponent value according to the characteristics of the objective function to obtain the output data of the objective function.
17. A chip, characterized in that: The chip includes: a data processing device and a memory; wherein the data processing device is used to read and execute program instructions stored in the memory to implement the method described in any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Data processing method and data processing apparatus
EP4797037A1
Data processing method and data processing apparatus
WO2025107602A1