Data conversion method, data conversion device, data conversion circuit and storage and calculation integrated equipment

By using mantissa cutoff processing technology during data conversion, the problems of large calculation amount, high hardware resource utilization and high energy consumption when converting floating-point data into integer data in the prior art are solved, and more efficient data conversion and lower resource consumption are achieved.

CN119995610APending Publication Date: 2025-05-13INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411310096.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When converting floating-point data into integer data, the prior art has a large amount of calculation, a large amount of hardware resources and a high energy consumption.

Method used

By performing data format conversion calculation based on the target quantization coefficient after mantissa cutoff processing and floating-point data to be converted, the calculation amount of data conversion is reduced, efficiency is improved, and hardware resource occupation and energy consumption are reduced.

Benefits of technology

It realizes the reduction of data conversion calculation, improves data conversion efficiency, and reduces hardware resource usage and energy consumption overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995610A_ABST
    Figure CN119995610A_ABST
Patent Text Reader

Abstract

The invention discloses a data conversion method, a data conversion device, a data conversion circuit and storage and calculation integrated equipment. The method comprises the following steps: acquiring floating point type data to be converted and a target quantization coefficient; the mantissa of the to-be-converted floating point type data and the mantissa of the target quantization coefficient are subjected to bit cutting processing based on the data bits of the target integer data, so that the effective mantissa of the to-be-converted floating point type data and the effective mantissa of the target quantization coefficient are obtained, and the bits of the effective mantissa are one bit less than the data bits of the target integer data; and performing format conversion on the to-be-converted floating point type data after the bit cutting processing based on the target quantization coefficient after the bit cutting processing so as to obtain target integer data corresponding to the to-be-converted floating point type data. According to the method, data format conversion calculation is carried out based on the target quantization coefficient after mantissa truncation processing and the to-be-converted floating point type data, the calculation amount is greatly reduced, and the data conversion efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data conversion technology, and in particular to a data conversion method, a computer-readable storage medium, a data conversion device, a data conversion circuit and a storage-computing integrated device. Background Art

[0002] In the data conversion technology solution for converting floating-point data into integer data, a FP16, FP32 or FP64 multiplier is usually used to multiply the FP data with the corresponding FP quantization coefficient, and then round it to an integer, thereby converting the floating-point data into integer data. However, this technical solution has a large amount of calculation, occupies a large amount of hardware resources, and has a large energy consumption overhead. Summary of the invention

[0003] The present invention aims to solve one of the technical problems in the related art at least to a certain extent. To this end, the first purpose of the present invention is to propose a data conversion method, which performs data format conversion calculation based on the target quantization coefficient after mantissa truncation and the floating-point data to be converted, greatly reducing the amount of data conversion calculation, improving the data conversion efficiency, and reducing the occupation of hardware resources and energy consumption.

[0004] A second object of the present invention is to provide a computer-readable storage medium.

[0005] The third objective of the present invention is to provide a data conversion device.

[0006] A fourth objective of the present invention is to provide a data conversion circuit.

[0007] The fifth objective of the present invention is to provide a storage and computing integrated device.

[0008] To achieve the above-mentioned purpose, a first aspect of the present invention proposes a data conversion method, which includes: obtaining floating-point data to be converted and a target quantization coefficient; truncating the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the number of data bits of the target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data; format conversion of the truncated floating-point data to be converted based on the truncated target quantization coefficient, so as to obtain the target integer data corresponding to the floating-point data to be converted.

[0009] According to the data conversion method of the embodiment of the present invention, first, the floating-point data to be converted and the target quantization coefficient are obtained, and the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient are truncated based on the data bit number of the target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit number of the target integer data, and then the format of the truncated floating-point data to be converted is converted based on the target quantization coefficient after the truncated processing, so as to obtain the target integer data corresponding to the floating-point data to be converted. Thus, the method performs data format conversion calculation based on the target quantization coefficient after the mantissa truncated processing and the floating-point data to be converted after the truncated processing, which greatly reduces the amount of data conversion calculation, improves the data conversion efficiency, and reduces the occupation of hardware resources and energy consumption.

[0010] In addition, the data conversion method according to the above embodiment of the present invention may also have the following additional technical features:

[0011] According to one embodiment of the present invention, the mantissa of the floating-point data to be converted is truncated based on the number of data bits of the target integer data to obtain the effective mantissa of the floating-point data to be converted, including: determining the high-order retained mantissa of the floating-point data to be converted based on the number of data bits of the target integer data, and rounding the remaining mantissa of the floating-point data to be converted to obtain the effective mantissa of the floating-point data to be converted.

[0012] According to one embodiment of the present invention, the mantissa of the target quantization coefficient is truncated based on the number of data bits of the target integer data to obtain the effective mantissa of the target quantization coefficient, including: determining the high-order retained mantissa of the target quantization coefficient based on the number of data bits of the target integer data, and rounding the remaining mantissa of the target quantization coefficient to obtain the effective mantissa of the target quantization coefficient.

[0013] According to one embodiment of the present invention, the format of the floating-point data to be converted after truncation is converted based on the target quantization coefficient after truncation to obtain the target integer data corresponding to the floating-point data to be converted, including: determining the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted; determining the data of the target integer data corresponding to the floating-point data to be converted based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted; and outputting the target integer data according to the sign of the target integer data and the data of the target integer data.

[0014] According to one embodiment of the present invention, the sign of the target integer data corresponding to the floating-point data to be converted is determined based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, including: performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted; or, the sign of the target integer data corresponding to the floating-point data to be converted is determined based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, including: when the sign of the target quantization coefficient is negative, performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted, and when the sign of the target quantization coefficient is positive, using the sign of the floating-point data to be converted as the sign of the target integer data corresponding to the floating-point data to be converted.

[0015] According to one embodiment of the present invention, the data of the target integer data corresponding to the floating-point data to be converted is determined based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted, including: respectively filling the effective mantissa of the target quantization coefficient and the effective mantissa of the floating-point data to be converted with one at the high bit to obtain the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted; obtaining the product of the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted to obtain the first data; obtaining the sum of the exponent of the target quantization coefficient and the exponent of the floating-point data to be converted to obtain the decimal point shift amount, and shifting the decimal point of the first data based on the decimal point shift amount; rounding the data after the decimal point of the shifted first data to obtain the second data; truncating the second data based on the data range of the target integer data to obtain the data of the target integer data corresponding to the floating-point data to be converted.

[0016] To achieve the above-mentioned purpose, a second aspect of the present invention provides a computer storage medium on which a data conversion program is stored. When the data conversion program is executed by a processor, the above-mentioned data conversion method is implemented.

[0017] According to the computer storage medium of an embodiment of the present invention, the above-mentioned data conversion method is implemented when the data conversion program is executed by the processor. Based on the above-mentioned data conversion method, the amount of data conversion calculation is greatly reduced, the data conversion efficiency is improved, and the hardware resource occupancy and energy consumption overhead are reduced.

[0018] To achieve the above-mentioned purpose, a third aspect of the present invention proposes a data conversion device, which includes: a data input module for receiving floating-point data to be converted and a target quantization coefficient; a preprocessing module for truncating the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the number of data bits of the target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data; a conversion module for performing format conversion on the truncated floating-point data to be converted based on the truncated target quantization coefficient, so as to obtain the target integer data corresponding to the floating-point data to be converted.

[0019] According to the data conversion device of the embodiment of the present invention, the floating-point data to be converted and the target quantization coefficient are received through the data input module, and the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient are truncated respectively based on the data bit number of the target integer data through the preprocessing module to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit number of the target integer data, and the conversion module performs format conversion on the truncated floating-point data to be converted based on the target quantization coefficient after the truncated processing to obtain the target integer data corresponding to the floating-point data to be converted. Thus, the device performs data format conversion calculation based on the target quantization coefficient after the mantissa truncated processing and the floating-point data to be converted after the mantissa truncated processing, which greatly reduces the amount of data conversion calculation, improves the data conversion efficiency, and reduces the occupation of hardware resources and energy consumption.

[0020] To achieve the above-mentioned purpose, a fourth aspect of the present invention proposes a data conversion circuit, which includes: a preprocessing unit, which is used to truncate the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the number of data bits of the received target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data; a data conversion unit, which is used to perform format conversion on the floating-point data to be converted after truncation based on the target quantization coefficient after truncation, so as to obtain the target integer data corresponding to the floating-point data to be converted.

[0021] According to the data conversion circuit of the embodiment of the present invention, the preprocessing unit performs truncation processing on the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the data bit number of the received target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit number of the target integer data, and the data conversion unit performs format conversion on the floating-point data to be converted after the truncation processing based on the target quantization coefficient after the truncation processing, so as to obtain the target integer data corresponding to the floating-point data to be converted. Thus, the circuit performs data format conversion calculation based on the target quantization coefficient after the mantissa truncation processing and the floating-point data to be converted after the mantissa truncation processing, which greatly reduces the amount of data conversion calculation, improves the data conversion efficiency, and reduces the occupation of hardware resources and energy consumption.

[0022] In addition, the data conversion circuit according to the above embodiment of the present invention may also have the following additional technical features:

[0023] According to one embodiment of the present invention, the preprocessing unit includes: a first rounding processing module, which is used to determine the high-order retained mantissa of the floating-point data to be converted based on the number of data bits of the target integer data, and round the remaining mantissa of the floating-point data to be converted to obtain the effective mantissa of the floating-point data to be converted; and determine the high-order retained mantissa of the target quantization coefficient based on the number of data bits of the target integer data, and round the remaining mantissa of the target quantization coefficient to obtain the effective mantissa of the target quantization coefficient.

[0024] According to one embodiment of the present invention, the data conversion unit includes: an XOR module, which is used to perform an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, so as to obtain the sign of the target integer data corresponding to the floating-point data to be converted; a one-complement processing module, which is used to perform one-complement on the effective mantissa of the target quantization coefficient and the effective mantissa of the floating-point data to be converted, so as to obtain the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted; a multiplier, which is used to obtain the product of the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted, so as to obtain the first data; an adder, which is used to obtain The exponent of the target quantization coefficient and the exponent of the floating-point data to be converted are summed to obtain the decimal point shift quantity; a shifter is used to shift the decimal point of the first data based on the decimal point shift quantity; a second rounding processing module is used to round the data after the decimal point of the shifted first data to obtain second data; a truncation processing module is used to truncate the second data based on the data range of the target integer data to obtain the data of the target integer data corresponding to the floating-point data to be converted; wherein the target integer data corresponding to the floating-point data to be converted is composed of the sign of the target integer data and the data of the target integer data.

[0025] In order to achieve the above-mentioned purpose, the fifth aspect of the present invention proposes a storage and computing integrated device, including the above-mentioned data conversion circuit.

[0026] According to the storage and computing integrated device of the embodiment of the present invention, based on the above-mentioned data conversion circuit, the amount of calculation is greatly reduced, the data conversion efficiency is improved, and the hardware resource occupancy and energy consumption are reduced.

[0027] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flow chart of a data conversion method according to an embodiment of the present invention;

[0029] Figure 2 is a flow chart of a data conversion method according to a specific embodiment of the present invention;

[0030] Figure 3 A connection diagram of a data conversion device according to an embodiment of the present invention;

[0031] Figure 4 is a connection diagram of a data conversion circuit according to an embodiment of the present invention;

[0032] Figure 5 is a connection diagram of a data conversion circuit according to a specific embodiment of the present invention;

[0033] Figure 6 A block diagram of a storage-computing integrated device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0035] The following describes the data conversion method, computer-readable storage medium, data conversion device, data conversion circuit and storage-computing integrated device proposed in the embodiments of the present invention with reference to the accompanying drawings.

[0036] Figure 1 4 is a flow chart of a data conversion method according to an embodiment of the present invention.

[0037] like Figure 1 As shown, the data conversion method of the embodiment of the present invention may include:

[0038] S1, obtaining floating-point data to be converted and a target quantization coefficient, wherein the target quantization coefficient is determined based on a data format of the floating-point data to be converted;

[0039] S2, truncating the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient respectively based on the number of data bits of the target integer data to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data;

[0040] S3, performing format conversion on the floating-point data to be converted after the truncation process based on the target quantization coefficient after the truncation process, so as to obtain target integer data corresponding to the floating-point data to be converted.

[0041] Specifically, the following takes the floating point data to be converted as FP32, the target integer data as INT8, and the target quantization coefficient as FP32 quantization coefficient as an example to explain the data conversion method in detail. Among them, FP32 and FP32 quantization coefficients include 1-bit sign + 8-bit exponent + 23-bit mantissa, and INT8 includes 1-bit sign + 7-bit data. Figure 5 shown.

[0042] The 23-bit mantissa of the FP32 to be converted is truncated according to the 7-bit data of INT8 to obtain the 6-bit effective mantissa of the FP32 to be converted, and the 23-bit mantissa of the FP32 quantization coefficient is truncated according to the 7-bit data of INT8 to obtain the 6-bit effective mantissa of the FP32 quantization coefficient, thereby realizing the truncation and shortening of the 23-bit mantissa of the FP32 to be converted and the FP32 quantization coefficient.

[0043] The FP32 to be converted after truncation includes 1-bit sign + 8-bit exponent + 6-bit valid mantissa, and the FP32 quantization coefficient after truncation includes 1-bit sign + 8-bit exponent + 6-bit valid mantissa. Data conversion calculation is performed based on the FP32 quantization coefficient after truncation and the FP32 to be converted after truncation to convert the digital format of the FP32 to be converted into INT8 output.

[0044] This embodiment performs data format conversion calculation based on the target quantization coefficient after mantissa truncation and the floating-point data to be converted after truncation, which greatly reduces the amount of data conversion calculation, improves data conversion efficiency, and reduces hardware resource occupancy and energy consumption.

[0045] In one embodiment of the present invention, the mantissa of the floating-point data to be converted is truncated based on the number of data bits of the target integer data to obtain the effective mantissa of the floating-point data to be converted, including: determining the high-order retained mantissa of the floating-point data to be converted based on the number of data bits of the target integer data, and rounding the remaining mantissa of the floating-point data to be converted to obtain the effective mantissa of the floating-point data to be converted.

[0046] Specifically, according to the 7-bit data of INT8, the high 6-bit mantissa of the FP32 to be converted is determined as the high-order retained mantissa, and the low 17-bit mantissa of the FP32 to be converted is rounded, and the first 6 bits after rounding are taken as the valid bits, thereby obtaining the 6-bit valid mantissa of the FP32 to be converted.

[0047] In one embodiment of the present invention, the mantissa of the target quantization coefficient is truncated based on the number of data bits of the target integer data to obtain the effective mantissa of the target quantization coefficient, including: determining the high-order retained mantissa of the target quantization coefficient based on the number of data bits of the target integer data, and rounding the remaining mantissa of the target quantization coefficient to obtain the effective mantissa of the target quantization coefficient.

[0048] Specifically, according to the 7-bit data of INT8, the high 6-bit mantissa of the FP32 quantization coefficient is determined as the high-order reserved mantissa, and the low 17-bit mantissa of the FP32 quantization coefficient is rounded, and the first 6 bits after rounding are taken as the valid bits, thereby obtaining the 6-bit valid mantissa of the FP32 quantization coefficient.

[0049] In one embodiment of the present invention, the format of the floating-point data to be converted after truncation is converted based on the target quantization coefficient after truncation to obtain the target integer data corresponding to the floating-point data to be converted, including: determining the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted; determining the data of the target integer data corresponding to the floating-point data to be converted based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted; and outputting the target integer data according to the sign of the target integer data and the data of the target integer data.

[0050] Specifically, the 1-bit sign of the converted INT8 is determined according to the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted, and the 7-bit data of the converted INT8 is determined based on the 8-bit exponent and 6-bit significant mantissa of the FP32 quantization coefficient and the 8-bit exponent and 6-bit significant mantissa of the FP32 to be converted. According to the 1-bit sign of INT8 and the 7-bit data of INT8, the INT8 corresponding to the FP32 to be converted is obtained, that is, INT8 is 1-bit sign + 8-bit data.

[0051] In addition, in the process of determining the 1-bit sign of the converted INT8 according to the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted, the 1-bit sign of the INT8 can be calculated by the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted. It is also possible to directly use the sign of the FP32 to be converted as the 1-bit sign of the INT8 when the FP32 quantization coefficient is positive to reduce the calculation process, and when the FP32 quantization coefficient is negative, the 1-bit sign of the INT8 is calculated by the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted.

[0052] In one embodiment of the present invention, the sign of the target integer data corresponding to the floating-point data to be converted is determined based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, including: performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted; or, the sign of the target integer data corresponding to the floating-point data to be converted is determined based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, including: when the sign of the target quantization coefficient is negative, performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted, and when the sign of the target quantization coefficient is positive, using the sign of the floating-point data to be converted as the sign of the target integer data corresponding to the floating-point data to be converted.

[0053] That is, when one of the signs of the FP32 quantization coefficient and the sign of the FP32 to be converted is 0 and the other is 1, the sign of the converted INT8 is 1; when the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted are both 0 or both 1, the sign of the converted INT8 is 0. Among them, 0 represents a positive number and 1 represents a negative number.

[0054] In another possible implementation of the present application, the sign of the FP32 quantization coefficient may be judged first. For example, when the sign of the FP32 quantization coefficient is positive, that is, 0, the sign of the FP32 to be converted is directly used as the sign of INT8 without the need for logical operations; when the sign of the FP32 quantization coefficient is negative, that is, 1, an XOR logical operation is performed on the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted to obtain the sign of INT8.

[0055] In one embodiment of the present invention, the data of the target integer data corresponding to the floating-point data to be converted is determined based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted, including: respectively filling the effective mantissa of the target quantization coefficient and the effective mantissa of the floating-point data to be converted with one at the high bit to obtain the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted; obtaining the product of the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted to obtain the first data; obtaining the sum of the exponent of the target quantization coefficient and the exponent of the floating-point data to be converted to obtain the decimal point shift amount, and shifting the decimal point of the first data based on the decimal point shift amount; rounding the data after the decimal point of the shifted first data to obtain the second data; truncating the second data based on the data range of the target integer data to obtain the data of the target integer data corresponding to the floating-point data to be converted.

[0056] Specifically, the first bit of the 6-bit effective mantissa of the FP32 quantization coefficient is padded with 1 to obtain the 7-bit target effective mantissa of the FP32 quantization coefficient. At the same time, the first bit of the 6-bit effective mantissa of the FP32 to be converted is padded with 1 to obtain the 7-bit target effective mantissa of the FP32 to be converted. The 7-bit target effective mantissa of the FP32 quantization coefficient is multiplied by the 7-bit target effective mantissa of the FP32 to be converted using a 7-bit multiplier to obtain the first data. The 8-bit exponent of the FP32 quantization coefficient is added to the 8-bit exponent of the FP32 to be converted by an 8-bit adder, and the decimal point of the first data is shifted based on the output result of the 8-bit adder to ensure that the decimal point position is correct. Then the decimal point data of the shifted first data is rounded to form integer data to obtain the second data. The second data is truncated according to the data range of INT8, where the data range of INT8 is [-128, 127]. Assuming that the second data is 11001000 (corresponding to 200 in decimal), the 7-bit data of INT8 obtained by truncating is 1111111; assuming that the second data is 1101001, the 7-bit data of INT8 obtained by truncating is 1101001.

[0057] As a specific embodiment of the present invention, taking the floating point data to be converted as FP32, the target integer data as INT8, and the target quantization coefficient as FP32 quantization coefficient as an example, Figure 2 As shown, the data conversion method may include the following steps:

[0058] S101, obtaining FP32 to be converted and FP32 quantization coefficients. Execute steps S102, S104 and S106 respectively.

[0059] S102, truncate the 23-bit mantissa of the FP32 to be converted based on the 7-bit data of INT8 to obtain a 6-bit effective mantissa of the FP32 to be converted.

[0060] S103, the high-order bits of the 6-bit effective mantissa of the FP32 to be converted are padded with one to obtain the 7-bit target effective mantissa of the FP32 to be converted. Execute step S107.

[0061] S104, truncate the 23-bit mantissa of the FP32 quantization coefficient based on the 7-bit data of INT8 to obtain a 6-bit effective mantissa of the FP32 quantization coefficient.

[0062] S105, padding the high-order bits of the 6-bit effective mantissa of the FP32 quantized coefficient with one to obtain the 7-bit target effective mantissa of the FP32 quantized coefficient. Execute step S107.

[0063] S106, perform an XOR logic operation on the sign of the FP32 quantization coefficient and the sign of the FP32 to be converted to obtain a 1-bit sign of INT8. Execute step S111.

[0064] S107, obtain the product of the 7-bit target significant mantissa of the FP32 quantization coefficient and the 7-bit target significant mantissa of the FP32 to be converted to obtain first data.

[0065] S108, obtaining the sum of the 8-bit exponent of the FP32 quantization coefficient and the 8-bit exponent of the FP32 to be converted, obtaining a decimal point shift quantity, and shifting the decimal point of the first data based on the decimal point shift quantity.

[0066] S109, rounding off the data after the decimal point of the shifted first data to obtain second data.

[0067] S110, truncate the second data based on the data range of INT8 to obtain 8-bit data of INT8.

[0068] S111, output INT8 according to the 1-bit symbol of INT8 and the 8-bit data of INT8.

[0069] Compared with the FP32 multiplier in the related art, this embodiment can convert FP32 to INT8 data format through an 8-bit adder, a 7-bit multiplier and a shifter, which can greatly reduce the energy consumption, area and delay overhead of the conversion circuit. At the same time, since the 7-bit target effective mantissa has the same resolution as the 7-bit data in INT8, it can effectively ensure that the calculation accuracy is lossless and consistent with the accuracy of INT8 itself.

[0070] In summary, according to the data conversion method of the embodiment of the present invention, first, the floating-point data to be converted and the target quantization coefficient are obtained, and the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient are truncated based on the data bit number of the target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit number of the target integer data, and then the floating-point data to be converted after the truncated processing is format-converted based on the target quantization coefficient after the truncated processing, so as to obtain the target integer data corresponding to the floating-point data to be converted. Thus, the method performs data format conversion based on the target quantization coefficient after the mantissa truncated processing and the floating-point data to be converted after the truncated processing, reduces the amount of data conversion calculation, improves the data conversion efficiency, and reduces the occupation of hardware resources and energy consumption.

[0071] Corresponding to the above embodiment, the present invention also proposes a computer storage medium.

[0072] The computer storage medium of the embodiment of the present invention stores a data conversion program, and the data conversion program implements the above-mentioned data conversion method when executed by a processor.

[0073] According to the computer storage medium of an embodiment of the present invention, the above-mentioned data conversion method is implemented when the data conversion program is executed by the processor. Based on the above-mentioned data conversion method, the amount of data conversion calculation is reduced, the data conversion efficiency is improved, and the hardware resource occupancy and energy consumption overhead are reduced.

[0074] Corresponding to the above embodiment, the present invention also proposes a data conversion device.

[0075] like Figure 3 As shown, the data conversion device of the embodiment of the present invention may include: a data input module 10, a preprocessing module 20 and a conversion module 30.

[0076] The data input module 10 is used to receive the floating-point data to be converted and the target quantization coefficient. The preprocessing module 20 is used to perform truncation processing on the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the number of data bits of the target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data. The conversion module 30 is used to perform format conversion on the floating-point data to be converted after the truncation processing based on the target quantization coefficient after the truncation processing, so as to obtain the target integer data corresponding to the floating-point data to be converted.

[0077] According to one embodiment of the present invention, the preprocessing module 20 truncates the mantissa of the floating-point data to be converted based on the number of data bits of the target integer data to obtain the effective mantissa of the floating-point data to be converted, and is specifically used to: determine the high-order retained mantissa of the floating-point data to be converted based on the number of data bits of the target integer data, and round off the remaining mantissa of the floating-point data to be converted to obtain the effective mantissa of the floating-point data to be converted.

[0078] According to one embodiment of the present invention, the preprocessing module 20 truncates the mantissa of the target quantization coefficient based on the number of data bits of the target integer data to obtain the effective mantissa of the target quantization coefficient, and is specifically used to: determine the high-order retained mantissa of the target quantization coefficient based on the number of data bits of the target integer data, and round the remaining mantissa of the target quantization coefficient to obtain the effective mantissa of the target quantization coefficient.

[0079] According to one embodiment of the present invention, the conversion module 30 performs format conversion on the floating-point data to be converted after truncation based on the target quantization coefficient after truncation to obtain target integer data corresponding to the floating-point data to be converted, and is specifically used for: determining the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted; determining the data of the target integer data corresponding to the floating-point data to be converted based on the exponent and the effective mantissa of the target quantization coefficient, and the exponent and the effective mantissa of the floating-point data to be converted; and outputting the target integer data according to the sign of the target integer data and the data of the target integer data.

[0080] According to one embodiment of the present invention, the conversion module 30 determines the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, and is specifically used for: performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted; or, the conversion module 30 determines the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, and is specifically used for: when the sign of the target quantization coefficient is negative, performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted, and when the sign of the target quantization coefficient is positive, using the sign of the floating-point data to be converted as the sign of the target integer data corresponding to the floating-point data to be converted.

[0081] According to one embodiment of the present invention, the conversion module 30 determines the target integer data corresponding to the floating-point data to be converted based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted, and is specifically used for: respectively filling the effective mantissa of the target quantization coefficient and the effective mantissa of the floating-point data to be converted with one at the high bit to obtain the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted; obtaining the product of the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted to obtain first data; obtaining the sum of the exponent of the target quantization coefficient and the exponent of the floating-point data to be converted to obtain the decimal point shift amount, and shifting the decimal point of the first data based on the decimal point shift amount; rounding the data after the decimal point of the shifted first data to obtain second data; truncating the second data based on the data range of the target integer data to obtain the target integer data corresponding to the floating-point data to be converted.

[0082] It should be noted that for details not disclosed in the data conversion device of the embodiment of the present invention, please refer to the details disclosed in the data conversion method of the above embodiment of the present invention, and the details will not be repeated here.

[0083] According to the data conversion device of the embodiment of the present invention, the floating-point data to be converted and the target quantization coefficient are received through the data input module, and the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient are truncated respectively based on the data bit number of the target integer data through the preprocessing module to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit number of the target integer data, and the conversion module performs format conversion on the truncated floating-point data to be converted based on the target quantization coefficient after the truncated processing to obtain the target integer data corresponding to the floating-point data to be converted. Thus, the device performs data format conversion based on the target quantization coefficient after the mantissa truncated processing and the floating-point data to be converted after the truncated processing, which reduces the amount of data conversion calculation, improves the data conversion efficiency, and reduces the occupation of hardware resources and energy consumption.

[0084] Corresponding to the above embodiment, the present invention further proposes a data conversion circuit.

[0085] like Figure 4 As shown, the data conversion circuit 100 according to the embodiment of the present invention includes: a pre-processing unit 110 and a data conversion unit 120 .

[0086] The preprocessing unit 110 is used to perform truncation processing on the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the data bit number of the received target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit of the target integer data. The data conversion unit 120 is used to perform format conversion on the floating-point data to be converted after the truncation processing based on the target quantization coefficient after the truncation processing, so as to obtain the target integer data corresponding to the floating-point data to be converted.

[0087] Combination Figure 5 As shown, according to one embodiment of the present invention, the preprocessing unit 110 includes: a first rounding processing module 111 (rounding I), which is used to determine the high-order retained mantissa of the floating-point data to be converted based on the number of data bits of the target integer data, and round off the remaining mantissa of the floating-point data to be converted to obtain the effective mantissa of the floating-point data to be converted; and determine the high-order retained mantissa of the target quantization coefficient based on the number of data bits of the target integer data, and round off the remaining mantissa of the target quantization coefficient to obtain the effective mantissa of the target quantization coefficient.

[0088] According to one embodiment of the present invention, the data conversion unit 120 includes: an XOR module 121, which is used to perform an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, so as to obtain the sign of the target integer data corresponding to the floating-point data to be converted; a one-complement processing module 122, which is used to perform one-complement on the effective mantissa of the target quantization coefficient and the effective mantissa of the floating-point data to be converted, so as to obtain the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted; a multiplication module 123, which is used to obtain the product of the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted, so as to obtain the first data; an adder 124, which is used to obtain The sum of the exponent of the target quantization coefficient and the exponent of the floating-point data to be converted obtains the decimal point shift quantity; the shifter 125 is used to shift the decimal point of the first data based on the decimal point shift quantity; the second rounding processing module 126 (rounding II) is used to round the data after the decimal point of the shifted first data to obtain the second data; the truncation processing module 127 is used to truncate the second data based on the data range of the target integer data to obtain the target integer data corresponding to the floating-point data to be converted; wherein the target integer data corresponding to the floating-point data to be converted is composed of the sign of the target integer data and the data of the target integer data.

[0089] It should be noted that Figure 5The circuit diagram corresponding to the FP32 conversion to the INT8 data format is shown in FIG. 1. The data conversion circuit 100 mainly includes an 8-bit adder and a 7-bit multiplier. Therefore, compared with the FP32 multiplier used in the related art, the circuit area, energy consumption, and delay overhead can be greatly reduced. At the same time, since the FP32 data mantissa is intercepted by 6 bits and supplemented by 1, the 7-bit input data, i.e., the target effective mantissa, is obtained. Therefore, the resolution of the input data is consistent with that of the INT8 output data, which greatly reduces the circuit overhead without introducing additional precision loss. In addition, the floating-point multiplier built into the CPU or other digital circuits can also be used for calculations.

[0090] According to the data conversion circuit of the embodiment of the present invention, the preprocessing unit performs truncation processing on the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient based on the data bit number of the received target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of the effective mantissa is one less than the number of the data bit number of the target integer data, and the data conversion unit performs format conversion on the floating-point data to be converted after the truncation processing based on the target quantization coefficient after the truncation processing, so as to obtain the target integer data corresponding to the floating-point data to be converted. Thus, the circuit performs data format conversion based on the target quantization coefficient after the mantissa truncation processing and the floating-point data to be converted after the truncation processing, which reduces the amount of data conversion calculation, improves the data conversion efficiency, and reduces the occupation of hardware resources and energy consumption.

[0091] Corresponding to the above-mentioned embodiment, the present invention also proposes a storage and computing integrated device.

[0092] The Compute-In-Memory (CIM) crossbar array (abbreviated as CIM) is an efficient analog computing device that can be used to implement large-scale matrix-vector multiplication and addition operations, and is often widely used in artificial intelligence and neural network calculations. The basic computing unit of its circuit is usually a conductance or charge modulated circuit device, such as a memristor, resistive random access memory, phase change memory, magnetic memory, floating gate transistor, dynamic random access memory, static random access memory, etc. By gating the row and column switches, the current is accumulated and sampled in the analog domain to achieve efficient matrix multiplication operations.

[0093] Since CIM devices use DAC (Digital to Analog Converter) and ADC (Analog to Digital Converter) for data acquisition and quantization, they can only represent calculation data linearly, so they are not compatible with floating point format data (such as FP32). Instead, they usually use integer data for calculation (such as INT8), which will cause data scaling. Therefore, every time FP (Floating Point) data is transferred from the digital computing core to the CIM computing array, the FP data needs to be quantized and converted to INT (Integer) type. This will generate additional computing delay and energy overhead.

[0094] In related technologies, FP data (such as FP32, FP64) is quantized into INT8 data directly on the digital computing core such as the CPU (Central Processing Unit), and then sent to the CIM computing array. This process usually uses FP16, FP32 or FP64 multipliers to multiply two FP data and then round them to integer bits. However, this process will occupy a lot of hardware resources and energy consumption, and the delay is high.

[0095] To solve the above technical problems, Figure 6 As shown, the storage and computing integrated device 1000 of an embodiment of the present invention includes the above-mentioned data conversion circuit 100.

[0096] For example, the data conversion circuit 100 combines the INT8 data type commonly used in CIM computing arrays with the FP32 data type commonly used in digital circuits. Compared with ordinary FP32 multipliers, the data conversion circuit 100 only includes a 7-bit multiplier instead of a 24-bit multiplier, which can greatly reduce the energy consumption, area, and delay overhead of the circuit. At the same time, since the 7-bit target effective mantissa has the same resolution as the 7-bit data in INT8, the calculation accuracy can be effectively guaranteed to be lossless and consistent with the accuracy of INT8 itself.

[0097] According to the storage and computing integrated device of an embodiment of the present invention, based on the above-mentioned data conversion circuit, the amount of data conversion calculation is reduced, the data conversion efficiency is improved, and at the same time the hardware resource occupancy and energy consumption overhead are reduced.

[0098] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0099] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0100] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0101] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0102] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0103] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A data conversion method, characterized in that: The method comprises: Get the floating point data to be converted and the target quantization coefficient; Based on the number of data bits of the target integer data, the mantissa of the to-be-converted floating-point data and the mantissa of the target quantization coefficient are truncated to obtain the effective mantissa of the to-be-converted floating-point data and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data; The format of the floating-point data to be converted after the truncation process is converted based on the target quantization coefficient after the truncation process, so as to obtain the target integer data corresponding to the floating-point data to be converted.

2. The data conversion method according to claim 1, characterized in that: The truncation of the mantissa of the to-be-converted floating-point data based on the number of data bits of the target integer data to obtain the effective mantissa of the to-be-converted floating-point data includes: The high-order reserved mantissa of the floating-point data to be converted is determined based on the number of data bits of the target integer data, and the remaining mantissa of the floating-point data to be converted is rounded to obtain the effective mantissa of the floating-point data to be converted.

3. The data conversion method according to claim 1, characterized in that: The truncation of the mantissa of the target quantization coefficient based on the number of data bits of the target integer data to obtain the effective mantissa of the target quantization coefficient includes: The high-order retained mantissa of the target quantization coefficient is determined based on the number of data bits of the target integer data, and the remaining mantissa of the target quantization coefficient is rounded to obtain the effective mantissa of the target quantization coefficient.

4. The data conversion method according to claim 1, characterized in that: The method of performing format conversion on the floating-point data to be converted after the truncation process based on the target quantization coefficient after the truncation process to obtain target integer data corresponding to the floating-point data to be converted includes: Determine the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted; Determine the target integer data corresponding to the floating-point data to be converted based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted; The target integer data is output according to the sign of the target integer data and the data of the target integer data.

5. The data conversion method according to claim 4, characterized in that: The step of determining the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted comprises: Performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted; or The step of determining the sign of the target integer data corresponding to the floating-point data to be converted based on the sign of the target quantization coefficient and the sign of the floating-point data to be converted comprises: When the sign of the target quantization coefficient is negative, the sign of the target quantization coefficient is subjected to an XOR logic operation with the sign of the floating-point data to be converted to obtain the sign of the target integer data corresponding to the floating-point data to be converted; and when the sign of the target quantization coefficient is positive, the sign of the floating-point data to be converted is used as the sign of the target integer data corresponding to the floating-point data to be converted.

6. The data conversion method according to claim 4, characterized in that: The method of determining the target integer data corresponding to the floating-point data to be converted based on the exponent and the effective mantissa of the target quantization coefficient and the exponent and the effective mantissa of the floating-point data to be converted comprises: Respectively padding the significant mantissa of the target quantization coefficient and the significant mantissa of the floating-point data to be converted with one at the high bit to obtain the target significant mantissa of the target quantization coefficient and the target significant mantissa of the floating-point data to be converted; Obtaining the product of a target effective mantissa of the target quantization coefficient and a target effective mantissa of the floating-point data to be converted, to obtain first data; Obtaining a sum of an exponent of the target quantization coefficient and an exponent of the floating-point data to be converted, obtaining a decimal point shift amount, and shifting a decimal point of the first data based on the decimal point shift amount; Rounding off the data after the decimal point of the shifted first data to obtain second data; The second data is truncated based on the data range of the target integer data to obtain the target integer data corresponding to the floating-point data to be converted.

7. A computer storage medium, characterized in that: A data conversion program is stored thereon, and when the data conversion program is executed by a processor, the data conversion method according to any one of claims 1-6 is implemented.

8. A data conversion device, characterized in that: The device comprises: A data input module, used for receiving floating point data to be converted and target quantization coefficients; A preprocessing module, configured to perform truncation processing on the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient respectively based on the number of data bits of the target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data; The conversion module is used to perform format conversion on the floating-point data to be converted after truncation based on the target quantization coefficient after truncation, so as to obtain the target integer data corresponding to the floating-point data to be converted.

9. A data conversion circuit, characterized in that: The circuit comprises: A preprocessing unit, configured to perform truncation processing on the mantissa of the floating-point data to be converted and the mantissa of the target quantization coefficient respectively based on the number of data bits of the received target integer data, so as to obtain the effective mantissa of the floating-point data to be converted and the effective mantissa of the target quantization coefficient, wherein the number of bits of the effective mantissa is one less than the number of bits of the target integer data; The data conversion unit is used to perform format conversion on the floating-point data to be converted after truncation based on the target quantization coefficient after truncation, so as to obtain the target integer data corresponding to the floating-point data to be converted.

10. The data conversion circuit according to claim 9, characterized in that: The pre-processing unit comprises: A first rounding processing module is used to determine the high-order reserved mantissa of the floating-point data to be converted based on the number of data bits of the target integer data, and round off the remaining mantissa of the floating-point data to be converted to obtain the effective mantissa of the floating-point data to be converted; and The high-order retained mantissa of the target quantization coefficient is determined based on the number of data bits of the target integer data, and the remaining mantissa of the target quantization coefficient is rounded to obtain the effective mantissa of the target quantization coefficient.

11. The data conversion circuit according to claim 9, characterized in that: The data conversion unit comprises: An XOR module, used for performing an XOR logic operation on the sign of the target quantization coefficient and the sign of the floating-point data to be converted, so as to obtain the sign of the target integer data corresponding to the floating-point data to be converted; A one-fill processing module, used to fill the high-order bits of the effective mantissa of the target quantization coefficient and the effective mantissa of the floating-point data to be converted by one, respectively, to obtain the target effective mantissa of the target quantization coefficient and the target effective mantissa of the floating-point data to be converted; a multiplier, configured to obtain a product of a target effective mantissa of the target quantization coefficient and a target effective mantissa of the floating-point data to be converted, to obtain first data; An adder, used for obtaining the sum of the exponent of the target quantization coefficient and the exponent of the floating-point data to be converted to obtain the decimal point shift amount; a shifter, configured to shift a decimal point of the first data based on the decimal point shift amount; A second rounding processing module is used to round off the data after the decimal point of the shifted first data to obtain second data; A truncation processing module, used for performing truncation processing on the second data based on the data range of the target integer data to obtain data of the target integer data corresponding to the floating-point data to be converted; The target integer data corresponding to the floating-point data to be converted is composed of the sign of the target integer data and the data of the target integer data.

12. A storage and computing integrated device, characterized in that: Comprising the data conversion circuit according to any one of claims 9-11.

Citation Information

Patent Citations

  • Data format conversion device, processor, electronic equipment and model operation method

    CN111796870A

  • Floating point to fixed point device and method, electronic equipment and storage medium

    CN113625990A

  • Data processing method and device, chip and electronic equipment

    CN117707472A

  • Operational circuit and related method

    CN117742655A

  • Farrow filter based on logic circuit and implementation method thereof

    WO2013060138A1