Converter, chip, electronic device and method for converting data types
By adopting a converter including the first conversion stage and the second conversion stage in the AI chip, the problems of low data type conversion efficiency and poor scalability are solved, and more efficient data processing and lower power consumption are achieved.
Patent Information
- Application Number
- CN201911025769.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-11-07
AI Technical Summary
In the prior art, data type conversion efficiency is low and scalable, resulting in performance bottlenecks and low production efficiency problems in AI chips.
A converter is adopted, including a first conversion stage and a second conversion stage, which receives data and description information, converts it into an intermediate result, and the second conversion stage converts the intermediate result into a target data type.
It improves the efficiency of data type conversion in AI chips, reduces the computing burden and circuit area, and enhances the scalability of the system.
Smart Images

Figure CN112711441B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and more particularly, to the conversion of data types. Background Art
[0002] For traditional arithmetic units, when implementing instructions (arithmetic units), there is generally only the conversion between fixed-precision floating-point and integer numbers, and the functions are single. In an artificial intelligence (AI) chip, the number of data type conversion instructions executed is much larger than that of traditional processing units, and the demand of programmers for the conversion function has increased significantly. Therefore, a larger number of software calculation behaviors make the weaknesses of low operation efficiency, large memory access overhead, and high calculation power consumption in implementing data type conversion through software more prominent, and its operation speed will become a performance bottleneck of the entire processor core.
[0003] At the same time, traditional arithmetic units implemented through instructions are all single-function implementations. If a processor core needs to implement a new data type conversion function, it is necessary to increase the logical expression according to the new function according to the multiplication principle, and its scalability is poor. Once a new function requirement appears, it will cause a multiple increase in the area of the arithmetic unit in the chip, there is a large amount of repetitive calculation logic, which affects the overall performance of the processor.
[0004] For example, when there are M input data types and N output data types, usually M*N data conversion paths are required. Therefore, the corresponding circuit design will be relatively complex, with high power consumption. Moreover, whenever a new data type appears, it is necessary to redesign the conversion device, which increases the workload and reduces the production efficiency.
[0005] Therefore, the traditional method for data type conversion has a poor application effect in AI chips, and we cannot refer to the traditional implementation method to implement the arithmetic unit in an AI chip. Summary of the Invention
[0006] An object of the present disclosure is to overcome the defects of low data conversion efficiency and poor scalability in the prior art.
[0007] According to a first aspect of the present disclosure, there is provided a converter for converting data types, including: a first conversion stage configured to receive first-type data and description information about the first-type data and second-type data, and convert the first-type data into an intermediate result according to the description information; and a second conversion stage configured to convert the intermediate result into second-type data.
[0008] According to a second aspect of the present disclosure, there is provided a chip including the above-mentioned converter.
[0009] According to a third aspect of the present disclosure, there is provided an electronic device including the above-mentioned chip.
[0010] According to a fourth aspect of the present disclosure, there is provided a method for converting data types, including: receiving first-type data and description information about the first-type data and second-type data, and converting the first-type data into an intermediate result according to the description information; and converting the intermediate result into second-type data.
[0011] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: one or more processors; and a memory storing computer-executable instructions, which, when run by the one or more processors, cause the electronic device to execute the method as described above.
[0012] According to a sixth aspect of the present disclosure, there is provided a computer-readable storage medium including computer-executable instructions, which, when run by one or more processors, execute the method as described above.
[0013] At least one beneficial effect of the technical solution provided by the present disclosure is that it can improve the efficiency of data type conversion in an AI chip, reduce the computing burden, and reduce the required circuit area. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] By referring to the drawings and reading the detailed description below, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0015] Figure 1 A converter for converting data types according to a first aspect of the present disclosure is shown.
[0016] Figure 2 A flowchart of a method for converting data types according to another aspect of the present disclosure is shown.
[0017] Figure 3 A schematic block diagram of a first converter L1 according to an embodiment of the present disclosure is shown.
[0018] Figure 4a The specific structure of a first computing unit C1 and the data structure of an intermediate result according to an embodiment of the present disclosure are shown.
[0019] Figure 4b The specific structure of a first computing unit C1 and the data structure of an intermediate result according to another embodiment of the present disclosure are shown.
[0020] Figure 5aShows a schematic block diagram of an absolute value calculation circuit C11 according to an embodiment of the present disclosure.
[0021] Figure 5b Shows a schematic block diagram of an absolute value calculation circuit C11 according to an embodiment of the present disclosure.
[0022] Figure 6 Shows a schematic block diagram of a second conversion stage L2 according to an embodiment of the present disclosure.
[0023] Figure 7a Shows a schematic block diagram of a pre-output calculation unit P2 according to an embodiment of the present disclosure.
[0024] Figure 7b Shows a schematic block diagram of a pre-output calculation unit P2 according to another embodiment of the present disclosure.
[0025] Figure 8 Shows a structural schematic diagram of a data recovery unit R2 according to an embodiment of the present disclosure.
[0026] Figure 9a Shows a schematic block diagram of a pre-output processing circuit R21 according to an embodiment of the present disclosure.
[0027] Figure 9b Shows a schematic block diagram of a pre-output processing circuit R21 according to another embodiment of the present disclosure.
[0028] Figure 10 Is a structural diagram showing a combined processing device according to an embodiment of the present disclosure.
[0029] Figure 11 Is a structural schematic diagram showing a board card according to an embodiment of the present disclosure. Specific embodiments
[0030] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.
[0031] It should be understood that in the claims, the description, and the drawings of the present disclosure, terms such as "first", "second", "third", and "fourth" are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the description and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0032] It should also be understood that the terms used in the description of the present disclosure herein are merely for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in the description and claims of the present disclosure, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the description and claims of the present disclosure refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0033] As used in the description and claims of this specification, the term "if" can be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0034] Figure 1 Shown is a converter for converting data types according to a first aspect of the present disclosure. Figure 2 Shown is a flowchart of a method for converting data types according to another aspect of the present disclosure.
[0035] As Figure 1 shown, the converter includes: a first conversion stage L1 configured to receive data of a first type and description information about the data of the first type and data of a second type, and convert the data of the first type into an intermediate result according to the description information; and a second conversion stage L2 configured to convert the intermediate result into data of the second type.
[0036] As Figure 2 shown, the method of the present disclosure may include: a first operation S1 of receiving data of a first type and description information about the data of the first type and data of a second type, and converting the data of the first type into an intermediate result according to the description information; and a second operation S2 of converting the intermediate result into data of the second type.
[0037] It should be understood that the above-mentioned "first type of data" can be the original first type of data, or the first type of data after transformation, splicing, or splitting. In other words, the variations of the first type of data at each stage are also included within the scope of the first type of data.
[0038] In the present disclosure, when converting data types, it can be first converted into an intermediate result, which is applicable to all data types. This intermediate result can effectively represent the data to be converted (the first type of data in the above text), and can be converted into any type of data required (the second type of data in the above text) according to this intermediate result. In other words, this intermediate result has common content and / or structure with respect to all types of data, and thus can be converted to other data types through this intermediate result.
[0039] The beneficial effects brought by converting the first type into an intermediate result and then converting the intermediate result into the second type of data include but are not limited to: In a traditional hardware structure, if there are M types of input data and N types of output data, separate circuits need to be designed for each conversion. Thus, the complexity of the circuit is approximately M * N, which will greatly increase the workload of circuit design, increase the circuit area, and further cause adverse effects such as increased power consumption and increased cost. However, in the technical solution provided by the present disclosure, in the case of the same number of data type conversions, the complexity of the circuit is only approximately M + N, which can greatly reduce the complexity of circuit design, reduce the circuit area, and thus reduce circuit power consumption and save costs.
[0040] The number of bits of the above-mentioned first type of data and second type of data can be in various situations. For example, it can be 1 bit, 2 bits, 4 bits, 8 bits, 16 bits, 32 bits, etc. In the present disclosure, the processing bit number of the converter used (such as the bit width of registers, memories, buses, etc.) may be other bit numbers, such as 32 bits. Therefore, according to an embodiment of the present disclosure, the first conversion stage L1 is further configured to determine the quantity of the received first type of data, and splice the quantity of the first type of data to form first spliced data. The first conversion stage L1 converts the first spliced data into an intermediate result according to the description information.
[0041] For example, when the input data is 8 bits, the output data is 8 bits, and the processing bit number of the converter (such as the bit width of the register) is 32 bits, then 4 input data can be received simultaneously at one time, that is, 4 input data are spliced to form 32-bit data.
[0042] When the input data is 8 bits, the output data is 16 bits, and the processing bit width of the converter is 32 bits, 2 input data can be received simultaneously at one time, that is, the 2 input data are concatenated to form 32-bit data. In this case, two 8-bit data can be extended to two 16-bit data, and then the two extended 16-bit data are concatenated.
[0043] For another example, when the input data is 16 bits, the output data is 8 bits, and the processing bit width of the converter is 32 bits, 2 input data can be received simultaneously at one time, that is, the 2 input data are concatenated to form 32-bit data. In this case, the two 16-bit output data contain the information of two 8-bit output data.
[0044] According to an embodiment of the present disclosure, the number of the received first type of data can be determined by dividing the processing bit width of the converter by the bit width of the higher one of the bit widths of the first type of data and the second type of data.
[0045] Taking the input of two 8-bit hexadecimal numbers 81 and 82 and the output of two 16-bit numbers as an example, two data can be received at one time. In this embodiment, the binary representations of the hexadecimal numbers 81 and 82 are "1000 0001" and "1000 0010" respectively, and they can be extended to two 16-bit numbers, that is, "xxxx xxxx 1000 0001" and "yyyy yyyy 1000 0010". The actual data of the 8-bit number is placed in the lower 8 bits of the 16-bit number, and the high bits of the 16-bit number are filled with zeros or other specified numbers (represented by x here). The concatenated data can be 00008182, and the binary representation is "xxxx xxxx yyyyyyyy 1000 0001 1000 0010". That is, in the 32-bit concatenated data, the first input data "81" occupies the lower 8 bits (0 to 7), and the second input data "82" occupies the middle 8 bits (8 to 15). The high bits (16 to 31) of the 32 bits are filled with x and y, where x and y are set according to the actual situation, and the two can be the same or different. Details will be explained below.
[0046] It should be understood that the above concatenation method is only an example, and those skilled in the art can set the concatenated data in the required format according to their own needs. For example, the first received data can be placed in the lower 16 bits of the 32-bit concatenated data, and the second received data can be placed in the high 16 bits of the 32-bit concatenated data. Still taking the above hexadecimal numbers 81 and 82 as an example, the form of the concatenated data can also be, for example, xxxx xxxx 1000 0001 yyyy yyyy 1000 0010, where x and y can be the same or different.
[0047] According to another embodiment of the present disclosure, splicing can be performed with a preset first fixed value. For example, the first fixed value can be 2 or other numbers.
[0048] Through the splicing operation shown in the above embodiments, the throughput of data can be increased and the processing efficiency can be improved. Of course, those skilled in the art can understand that the above data splicing is not necessary, but only a preferred method. For example, when the number of bits of at least one of the input data and the output data is the same as the number of bits processed by the converter, splicing is not required; in addition, other specified formats (such as using the method of marking valid bits, that is, specifying in advance which bits are valid bits and which bits are invalid bits) can also be used, so that even when the number of bits of at least one of the input data and the output data is different from the number of bits processed by the converter, splicing is not required. For example, in the case where the input data is 8 bits, the output data is 16 bits, and the register is 32 bits, the 8-bit input data can be directly extended to 32-bit data (for example, by adding 0s to specific bits of the original 8-bit input data), and then the 32-bit data is restored to 16-bit data when output.
[0049] The case where the number of bits of the first type of data is shorter than the number of bits of the register is described above. In another case, if the number of bits of the input data is greater than the number of bits processed by the converter, for example, when the input data is 64 bits and the number of bits processed by the converter is 32 bits, the following processing can be performed.
[0050] One processing method can be to truncate the 64-bit data, leave the required 32-bit data, discard the other 32-bit data, and process the remaining 32-bit data. This method may cause certain data loss and errors.
[0051] According to another embodiment of the present disclosure, the first conversion stage L1 is further configured to determine the number of splits of the received first type of data to be split, and split the first type of data into the number of split data, and the first conversion stage L1 converts the split data into an intermediate result according to the description information.
[0052] In this embodiment, the 64-bit data can be split into two 32-bit data, the two split 32-bit data are processed, and finally the two output data are spliced to form the required output data.
[0053] According to an embodiment of the present disclosure, the number of splits of the received first type of data to be split can be determined in the following manner: dividing the number of bits of the higher of the number of bits of the first type of data and the second type of data by the number of bits processed by the converter.
[0054] For example, when the input data is 64 bits, the output data is 64 bits, and the register is 32 bits, the input data can be split into two 32-bit data; after processing, the two 32-bit data are re-concatenated at the output end to form 64-bit output data.
[0055] For another example, when the input data is 64 bits, the output data is 16 bits, and the register is 32 bits, the input data can be split into two 32-bit data. After processing, the valid data part is intercepted from the two 32-bit data at the output end and re-concatenated into 16-bit output data.
[0056] For yet another example, when the input data is 16 bits and the output data is 64 bits, the 16-bit input data can be extended to two 32-bit data. One of the 32-bit data contains valid information, while the other 32-bit data contains invalid information (such as all 0s), and the two 32-bit data are concatenated into 64-bit output data when output.
[0057] According to another embodiment of the present disclosure, it can be split by a preset second fixed value. For example, the fixed value can be set to 2 or other numbers.
[0058] Splitting and concatenating the data is beneficial to aligning the timing in the input data and the output data, avoiding or reducing additional design in the timing control part of the circuit; in addition, this embodiment is beneficial to parallel processing of the data and improving resource utilization.
[0059] The corresponding splitting and concatenating functions can be added to the above-mentioned first conversion stage L1 and second conversion stage L2, and this function can be implemented in ways such as software and / or hardware.
[0060] It can be seen that the present disclosure does not limit the number of bits of the input, output, and converter (such as a register). Through methods such as data splitting and concatenation, the present disclosure can process data of any number of bits.
[0061] Figure 3 A schematic block diagram of a first converter L1 according to an embodiment of the present disclosure is shown.
[0062] As Figure 3 shown, the first conversion stage L1 includes a first data parsing unit P1 and a first arithmetic unit C1.
[0063] The first data parsing unit P1 is configured to generate a transition sign bit Tsign, a transition data bit Tdata, and a transition exponent bit Tshift according to the first type of data and the description information. The first arithmetic unit C1 is configured to generate an intermediate result according to the transition sign bit Tsign, the transition data bit Tdata, and the transition exponent bit Tshift.
[0064] The description information can be input manually, or can be input into the first data parsing unit P1 in the form of a file or a signal.
[0065] According to an embodiment of the present disclosure, the above description information may include: first description information for describing the data type of the first type of data and the first exponent bit of the first type of data; second description information for describing the data type of the second type of data and the second exponent bit of the second type of data.
[0066] The data types described in the above first description information and second description information can be various, including but not limited to FIX4, FIX8, FIX16, FIX32, UFIX8, UFIX16, UFIX32, FP16, FP32, BFLOAT, and any other existing or custom data types. It should be understood that only the highest 32 bits are taken as an example here for illustration. For 64 bits or higher, more data types can be included.
[0067] In addition, in this embodiment, the first exponent bit indicating the shift value of the first type of data and the second exponent bit indicating the shift value of the second type of data can also be received separately by the first data parsing unit P1, and then P1 calculates the difference between the first exponent bit and the second exponent bit.
[0068] Alternatively, according to another embodiment of the present disclosure, the description information may include the first data type of the first type of data; the second data type of the second type of data; and a differential exponent bit for indicating the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data.
[0069] Different from the differential exponent bit being calculated by the first data parsing unit P1 in the previous embodiment, in this embodiment, the differential exponent bit can be directly input into the first data parsing unit P1 without subsequent calculation.
[0070] It should be noted that the "difference" mentioned above not only indicates the magnitude of the shift but also represents the direction of the shift. The difference described in the present disclosure can be the first exponent bit minus the second exponent bit, or the second exponent bit minus the first exponent bit. This is clear to those skilled in the art, so it will not be elaborated here.
[0071] When the first data parsing unit P1 calculates or directly receives the differential exponent bit, the above transition exponent bit Tshift can be calculated according to the differential exponent bit, and the transition exponent bit Tshift is equal to the differential exponent bit.
[0072] Although the description information and data are interpreted as two different message carriers above, it should be understood that in practice, there may not be an obvious boundary between the two. For example, when both the first type of data and the second type of data are of the Fix type, the shift values of the first type of data and the second type of data can be specified in a separate description information, and the differential data bits can be calculated based on these two shift values. When the first type of data is, for example, of the Float type, the first shift value is included in the Float type data itself, and thus P1 can extract the first shift value from the first type of data. Thus, the first type of data and its first description information, as well as the second type of data and its description information, can be mixed or discrete.
[0073] It should be understood that the term "equivalent" here indicates a substantial sameness, but may be different in form. For example, for a certain 8-bit number 0000 0001, when it is transformed into 0000 0000 0000 0001, it is essentially another representation of the previous 8-bit number, but may not be exactly equal. In addition, it should be understood that in addition to the change in the number of bits, different forms of representation such as the complement code, offset code, binary, decimal, hexadecimal, etc. of a number are also within the scope of "equivalent" described in this article. In other words, as long as the valid information is not lost, any form of change can be regarded as equivalent.
[0074] For example, when the first type of data is of the Float type and the second type of data is of the Fix type, the second shift value extracted from the Float type data may be represented in the offset code format, while the shift value describing the Fix type data may be represented in the original code format. At this time, when calculating the difference between the two, it is necessary to uniformly transform them into the same code type before calculating the difference. It can be uniformly transformed into the offset code, the original code, or other types of code such as the complement code. The present invention will not describe the transformation of the code type in detail.
[0075] According to an embodiment of the present disclosure, the description information further includes a rounding type, and the rounding type includes at least one of the following: TO_ZERO, OFF_ZERO, UP, DOWN, ROUNDING_OFF_ZERO, ROUNDING_TO_EVEN, random rounding.
[0076] TO_ZERO means rounding towards zero, in other words, rounding towards the direction of smaller absolute value;
[0077] OFF_ZERO means rounding away from zero, in other words, rounding towards the direction of larger absolute value);
[0078] UP means rounding towards positive infinity;
[0079] DOWN means rounding towards negative infinity;
[0080] ROUNDING_OFF_ZERO means rounding;
[0081] ROUNDING_TO_EVEN means that, on the basis of rounding, exactly half of the values are rounded to an even number.
[0082] It should be understood that the above rounding types are only some examples, and those skilled in the art can set various desired rounding methods.
[0083] Figure 4a Shows the specific structure of the first computing unit C1 and the data structure of the intermediate result according to an embodiment of the present disclosure.
[0084] According to an embodiment of the present disclosure, the intermediate result can be divided into an intermediate data bit ABS, an intermediate sign bit Sign, and an intermediate exponent bit EXP. The following details how to obtain the above intermediate result from the transitional exponent bit Tshift, the transitional sign bit Tsign, and the transitional data bit Tdata. In other words, all input data can be converted into this intermediate data with a common structure.
[0085] As Figure 4a shown, the first arithmetic unit C1 includes: an absolute value calculation circuit C11 configured to calculate the intermediate data bit ABS according to the transitional data bit Tdata.
[0086] Figure 5a Shows a schematic block diagram of the absolute value calculation circuit C11 according to an embodiment of the present disclosure.
[0087] As Figure 5a shown, the absolute value calculation circuit C11 includes a second selector configured to determine whether the transitional data bit Tdata is less than zero; a first complement calculator configured to calculate the complement of the transitional data bit as the intermediate data bit ABS if the transitional data bit Tdata is less than zero; otherwise, use the transitional data bit Tdata as the intermediate data bit ABS. Calculating the complement is actually inverting all bits except the sign bit and adding 1. Therefore, the first complement calculator may include a first inverter and a first adder. And if the transitional data bit Tdata is greater than or equal to zero (not negative), then the intermediate data bit ABS is equal to the transitional data bit Tdata.
[0088] Figure 5b Shows a schematic block diagram of the absolute value calculation circuit C11 according to another embodiment of the present disclosure.
[0089] AsFigure 5b As shown, the absolute value calculation circuit C11 further includes a first selector and a first normalizer. The first selector receives the transition data bit Tdata and determines whether the data type of the transition data bit Tdata is the first type or the second type.
[0090] The above-mentioned first type can be, for example, the Fix type, and the second type can be, for example, the Float type. In the following description and in the accompanying drawings, Fix will be used as an example of the first type and Float will be used as an example of the second type for description. It should be understood that the first type and the second type of data can also be any other suitable number types.
[0091] If the transition data bit Tdata is of the Fix type, it enters the second selector. In the second selector, it is determined whether the transition data bit Tdata is less than zero. If the transition data bit Tdata is less than zero (negative), the complement of the Tdata is calculated in the first complement calculator and used as the intermediate data bit ABS. Calculating the complement is actually inverting all bits except the sign bit and adding 1. Therefore, the first complement calculator can include a first inverter and a first adder. If the transition data bit Tdata is greater than or equal to zero (non-negative), the intermediate data bit ABS is equal to the transition data bit Tdata.
[0092] If the transition data bit Tdata is of the Float type, it enters the first normalizer. In the first normalizer, the Tdata is normalized, and the normalized data is used as the intermediate data bit ABS.
[0093] Normalization is an operation performed on Float type numbers. Float type numbers have several types in the definition of the IEEE754 standard, including normalized numbers, denormalized numbers, zero, positive and negative infinity, and NaN. In this operation, all normalized numbers can be pre-complemented with 1, and denormalized numbers can be post-complemented with 0 to form the actual original code representation result of the number. This result has one more bit than the normalized / denormalized representation result in the Float type.
[0094] Furthermore, as Figure 4a shown, the first arithmetic unit C1 further includes an exponent bit calculation circuit C12 configured to calculate an intermediate exponent bit EXP according to the transition exponent bit Tshift. According to an embodiment of the present disclosure, the above-mentioned intermediate exponent bit (EXP) is equal to the transition exponent bit Tshift.
[0095] Furthermore, as Figure 4aAs shown, according to an embodiment of the present disclosure, the sign bit calculation circuit C13 may be a direct connection line. The first arithmetic unit C1 further includes a sign bit calculation circuit C13 configured to calculate an intermediate sign bit Sign based on the transitional sign bit Tsign. It should be understood that the sign does not change, so the intermediate sign bit Sign can be calculated based on the transitional sign bit Tsign through a direct connection line.
[0096] Further as Figure 4b shown, according to an embodiment of the present disclosure, the intermediate result may further include an intermediate rounding bit STK. To calculate this rounding bit STK, the first calculation circuit C1 may further include: a rounding bit calculation circuit C14.
[0097] According to an embodiment of the present disclosure, the rounding bit calculation circuit C14 may be configured to calculate the intermediate rounding bit based on the intermediate data bit ABS and the intermediate sign bit Sign.
[0098] According to another embodiment of the present disclosure, the rounding bit calculation circuit C14 may be configured to calculate the intermediate rounding bit based on the intermediate data bit ABS, the intermediate exponent bit EXP, and the intermediate sign bit Sign.
[0099] In the above two embodiments of calculating the intermediate rounding bit STK, the intermediate exponent bit EXP may or may not be used. For example, when the intermediate rounding bit STK is in the form of an array (e.g., all rounding contents need to be retained), the intermediate exponent bit EXP may not be used; while if a specific bit or several bits of the intermediate rounding bit need to be specified, the intermediate exponent bit EXP may be used.
[0100] According to an embodiment of the present disclosure, the rounding bit calculation circuit C14 may be implemented through AND-OR logic. For example, for rounding, STK = ABS, and for rounding towards positive infinity, STK[n] = |ABS[n:x1] && ~SIGN, etc.
[0101] As Figure 4a shown, through the above converters and methods, all types of data can be converted into intermediate results with the same content. That is, according to an embodiment of the present disclosure, the intermediate result may include an intermediate sign bit Sign, an intermediate exponent bit EXP, and an intermediate data bit ABS.
[0102] As Figure 4b shown, according to another embodiment of the present disclosure, the intermediate result may include an intermediate sign bit Sign, an intermediate exponent bit EXP, an intermediate data bit ABS, and an intermediate rounding bit STK.
[0103] Figure 4a and Figure 4bThe rounding bit calculation circuit C14 in [[ ]] can also be provided in the second conversion stage L2, that is, the second conversion stage L2 can receive an intermediate result including an intermediate sign bit Sign, an intermediate exponent bit EXP, and an intermediate data bit ABS, and calculate an intermediate rounding bit STK based on this intermediate result.
[0104] Furthermore, according to another embodiment of the present disclosure, the rounding bit calculation circuit can also be a separate module, which can exist independently of the first conversion stage L1 and the second conversion stage L2.
[0105] Although described above in connection with Figure 4a , Figure 4b , Figure 5a and Figure 5b , those skilled in the art can understand that the circuits, units and other components in these figures can exist separately, can be combined together, and can exist in combination with other conversion stages.
[0106] The intermediate result can be converted into the required data type by the second conversion stage L2.
[0107] Figure 6 FIG. [[ ]] shows a schematic block diagram of the second conversion stage L2 according to an embodiment of the present disclosure.
[0108] As Figure 6 shown, the second conversion stage L2 can include a pre-output calculation unit P2 and a data recovery unit R2. The pre-output calculation unit P2 is configured to calculate a pre-output data bit Pdata and a pre-output sign bit Psign according to the intermediate data bit ABS, the intermediate sign bit Sign, the intermediate exponent bit EXP, and the intermediate rounding bit STK. The data recovery unit R2 is configured to generate a second type of data according to the pre-output data bit Pdata and the pre-output sign bit Psign.
[0109] It should be understood that although Figure 6 does not show that the second conversion stage L2 includes a rounding bit calculation circuit C14, Figure 6 the intermediate rounding bit STK in [[ ]] can come from the first conversion stage L1 or from the rounding bit calculation circuit C14 included in L2 itself. In addition, the pre-output calculation unit P2 here receives four inputs, namely ABS, Sign, EXP, and STK. However, it should be understood that as described above, the calculation of STK can be completed in the first conversion stage L1, can be completed in the second conversion stage L2, or can also be integrated in the pre-output calculation unit P2. The four inputs shown here are only for convenience of understanding and description, and do not limit the content of the present disclosure in any way.
[0110] Figure 7aShows a schematic block diagram of a pre-output calculation unit P2 according to an embodiment of the present disclosure.
[0111] As Figure 7a shown, the pre-output calculation unit P2 includes a shift operator P21 and an adder P22, configured to generate a temporary output data bit ABS' and a pre-output sign bit Psign. The shift operator P21 is configured to shift the intermediate data bit ABS by the intermediate exponent bit EXP to obtain a shift result; the adder P22 receives the shift result of the shift operator P21 and the intermediate rounding bit STK to generate a temporary data bit ABS'; the pre-output sign bit Psign is equal to the intermediate sign bit SIGN.
[0112] First, in the pre-output calculation unit P2, the received intermediate data bit ABS is shifted, and the amount and direction of the shift are determined by the intermediate exponent bit EXP. The obtained shift result is input to the next adder.
[0113] The output of the adder is ABS' = the output result of the shift operator + STK[-EXP - 1]. And if STK is out of range, then STK takes zero. It should be explained that STK is an array, such as a 32-bit array STK[31:0]. Here, STK[0] is the lowest bit element, and STK
[31] is the highest bit element. We calculate -EXP - 1. If it is between 0 and 31, we take the corresponding value. If it is less than 0, we take 0. If it is greater than 0, we perform special processing (take 0 or 31 according to the different types of STK).
[0114] In a specific case, for example, when ABS' does not overflow, then this ABS' can be directly used as the output of the pre-output calculation unit P2.
[0115] Figure 7b Shows a schematic block diagram of a pre-output calculation unit P2 according to another embodiment of the present disclosure.
[0116] As Figure 7b shown, the pre-output calculation unit P2 further includes a selector P23. In the selector P23, it is judged whether the generated ABS' overflows. If it overflows, saturation processing is performed on ABS'. If it does not overflow, then Pdata = ABS'.
[0117] Saturation processing is a special case handling that exists in various arithmetic units. During operations including rotation, there will be a situation where the result obtained from the input data is different from the value range of the output data: if the absolute value of the result that should be obtained is larger than the absolute value upper limit of the output data representation range, an upper overflow occurs; if the absolute value of the result that should be obtained is smaller than the absolute value lower limit of the output data representation range, a lower overflow occurs; generally, there are several ways to handle overflow situations: taking the saturation value, high-bit truncation, taking infinity or special values. Any method can be adopted for saturation processing in this disclosure.
[0118] In addition, SIGN is output as Psign through a direct connection line, that is, the sign does not change.
[0119] In addition, in Figure 7a and Figure 7b the pre-output exponent bit Pshift is not shown. When all data shifting has been completed, Pshift = 0.
[0120] Figure 7a and Figure 7b For the output data in
[0121] Figure 7a and Figure 7b In certain specific situations (such as when both the input and output are of Fix type and the signs are both positive), for example, the temporary output data bit ABS’, the pre-output data bit Pdata, and the pre-output sign bit Psign can directly become the second output data without further processing. Figure 7a and Figure 7b
[0122] Figure 8
[0123] Figure 8 As
[0124] shown, the data recovery unit R2 is used to obtain the second output data according to the output pre-output data Pdata and the pre-output sign PSign. Figure 8 As shown, the data recovery unit R2 may include a pre-output processing circuit R21. Preferably, it may further include a data assembly circuit R22. Data assembly and the data splicing introduced above may be an inverse operation, restoring the spliced data into the required second type of data. Those skilled in the art can determine whether to add this assembly circuit according to the actual data type. For example, for unspliced data, the data assembly circuit R22 may not be needed. Therefore, the data assembly circuit R22 is only preferred rather than necessary.
[0125] For example, the input is a 32-bit Float type number, and the output is a 32-bit Fix type number. At this time, there is no splicing or splitting during input. Therefore, in terms of length, the data assembly circuit R22 is not needed.
[0126] As Figure 8 shown, the pre-output processing circuit R21 in the data recovery unit R2 receives Figure 7a the temporary output data bits ABS’ and the pre-shown sign bit Psign in Figure 7b or receives the pre-output data bits Pdata and the pre-output sign bit Psign in
[0127] to obtain the output data bit representation Data_out.
[0128] Considering that there are other data types such as Float in the data type, the pre-output processing circuit R21 in the present disclosure is further configured to generate the floating-point decimal point position representation SHIFT_FP.
[0129] Further as Figure 8 shown, the data assembly circuit R22 obtains the final second type of data according to the output data bit representation Data_out, the floating-point decimal point digit representation SHIFT_FP, and the pre-output sign bit Psign. It should be understood that in Figure 8 the floating-point decimal point digit representation SHIFT_FP is shown by a dashed line, indicating that the SHIFT_FP may not exist in a specific case. In this case, the data assembly circuit R22 is configured to obtain the second type of data according to the data output bit representation Data_out and the pre-output sign bit Psign.
[0130] Figure 9a Shows a schematic block diagram of the pre-output processing circuit R21 according to an embodiment of the present disclosure.
[0131] As Figure 9aAs shown, the pre-output processing circuit R21 of the present disclosure includes: a fourth selector and a two's complement calculator.
[0132] In Figure 9a it, the Pdata is received in the fourth selector, and the pre-output sign bit Psign is received. It is judged whether the PSign is a negative number or a non-negative number, that is, it is judged whether Psign is equal to 1 or equal to 0.
[0133] If Psign = 1, it enters the two's complement calculator. The two's complement calculator includes a second inverter and a second adder. The second inverter first inverts all bits except the sign bit, and then the second adder adds 1. Next, the two's complement calculator outputs the result as the output data bit representing Data_out.
[0134] If Psign = 0, the Pdata is directly output as the output data bit representing Data_out.
[0135] Considering that there are various types of data, the pre-output data bit Pdata can be judged in advance to determine how to perform further processing subsequently.
[0136] Figure 9b Fig. shows a schematic block diagram of the pre-output processing circuit R21 according to another embodiment of the present disclosure.
[0137] As Figure 9b shown, the pre-output processing circuit R21 further includes: a third selector, a second normalizer, and a floating-point decimal point determiner.
[0138] Among them, the third selector receives the pre-output data bit Pdata, judges whether the data type of the pre-output data bit Pdata is Fix or Float. If the data type of the pre-output data bit Pdata is Fix, the pre-output data bit Pdata is sent to the fourth selector. If the data type of the pre-output data bit Pdata is Float, the pre-output data bit Pdata is sent to the second normalizer.
[0139] The second normalizer can normalize the pre-output data bit Pdata and output it as the data output bit representing Data_out.
[0140] In the definition of normalized numbers, normalization and denormalization are distinguished by simple magnitude comparison. An absolute value greater than the maximum representable absolute value (positive and negative saturation values) cannot be represented, resulting in an upper overflow and saturation processing; an absolute value less than the saturation value but greater than the normalization threshold undergoes normalization; an absolute value less than the normalization threshold but greater than the minimum representable value undergoes denormalization; and an absolute value less than the minimum representable value results in a lower overflow and saturation processing (taking 0 or the minimum representable value or a special value). Normalization in the second conversion stage L2 is to remove the leading 1, and denormalization is to shift right by one bit, which is an inverse operation to the normalization operation in the previous first conversion stage L1.
[0141] The floating-point decimal point position determiner can determine the floating-point decimal point digit representation SHIFT_FP according to the output of the second normalizer.
[0142] It should be noted that the data at various stages above can maintain the same number of bits at each stage. For example, if the first type of data is concatenated (e.g., two 16-bit data are concatenated into a 32-bit data), then the transitional data bit Tdata is also two concatenated data. Similarly, the intermediate results (e.g., Sign, ABS, EXP, STK), the pre-output data (e.g., the pre-output data bit Pdata, the pre-output sign bit Psign), the output data bit representation Data_out, and the floating-point decimal point digit representation SHIFT_FP can all be two concatenated data. The concatenated form can be set according to the user's needs.
[0143] For the data assembly circuit R22, there may be multiple situations.
[0144] For example, for a 32-bit converter, if the input is a 16-bit Fix type number and the output is a 32-bit Fix type number, the input 16-bit number can be simply converted to a 32-bit number by adding zeros at the high bits, and the final output can be directly a 32-bit number without any data assembly, etc.
[0145] Another example is that for a 32-bit converter, if the input is a 32-bit Fix type number and the output is a 16-bit Fix type number, the input is normally converted in the first conversion stage, and the data obtained after conversion can be obtained by truncating the high 16 bits to get the final 16-bit Fix type number.
[0146] It can be seen that the above data assembly circuit R22 may not function in some cases, and thus is not essential for this disclosure.
[0147] In addition, since the output data bits output by the pre-output processing circuit R21, i.e., Data_out, and the floating-point decimal point number of bits, i.e., SHIFT_FP, may be multiple pieces of data spliced together, the data assembly circuit R22 can be used to transform or assemble the data into the final required data form. For example, the spliced-together data can be split, and the various parts of the data (such as the valid data part and the sign part) can be assembled.
[0148] For example, the data of Data_out may be {0000 0000 0000 0000 0101 0011 00011010}, and the sign bit of the data is {0001}. At this time, the number to be output is Fix8. Then, the data assembly circuit R22 can extract two final second-type data from the above data, which are {0101 0011} and {0001 1010} respectively, and the signs of the data are 0 and 1 respectively. Thus, the array assembly circuit can extract the final data from Data_out.
[0149] The first conversion stage L1 of the present disclosure can also receive constraint information, which can be used to indicate whether the converter supports a specific standard and / or whether it supports compilation optimization. The specific standard can be any known or unknown standard suitable for the present disclosure, such as IEEE754; compilation optimization can be, for example, the support for compiler behaviors -o0, -o1, etc.
[0150] It should be understood that only specific instances are described above, and these instances are only for convenience of illustration and do not form any limitation to the protection scope of the present disclosure. The data types of the first-type data and the second-type data of the present disclosure, as well as the content of the constraint information, can be extended in any way. Any existing or newly developed data type in the future can be implemented by the technical solution of the present disclosure.
[0151] In the above, when the intermediate data passes through the second conversion stage L2, there may be multiple states. For example Figure 7a the output of the adder, ABS', Figure 7b the output of the selector, Pdata, Figure 8 、 Figure 9a and Figure 9bThe outputs of the front-end output processing circuit such as Data_out, etc. These data (optionally, plus other auxiliary data) can all be equivalent to the second type of data. For example, ABS’ can be equivalent to the second type of data, and ABS’ + Pdata can also be equivalent to the second type of data; similarly, Pdata can be equivalent to the second type of data, and Pdata + Psign can also be equivalent to the second type of data, and the difference between the two is only in the sign bit; again, for example, Data_out can be equivalent to the second type of data, and Data_out + Shift_FP can also be equivalent to the second type of data. It should be understood that although the data at these different stages are represented by different symbols, for some data, they may be the same or equivalent data. In other words, the “second type of data” referred to in this article may be any one of the above data, but only the representation methods in each figure are different. For example, when the input number is of Fix16 type, is a positive number, and is extended to a 32-bit number, and the output is Fix32, then after Pdata passes through the fourth selector (as Figure 9a shown), it is assigned to directly output Data_out. The data of Data_out itself conforms to the format of Fix32, so it can also be directly output as the second type of data without further processing.
[0152] The various units, circuits, and components described above will be explained below with specific examples.
[0153] Example 1
[0154] Example 1 gives an example of converting Fix8 to Float16.
[0155] Let the input numbers be 81 and 82, the data type be fix8, and the output data type be Float16. Then the 16-bit number formed by concatenating the two numbers is DATA = 32’h 00008182 (0000 0000 0000 0000 1000 0001 1000 0010), the exponent bit Shift is 9 bits, for example, -1 (1 1111 1111), and the rounding method is rounding. Among them, in the above expression, 32’ represents 32 bits, and h represents hexadecimal.
[0156] As Figures 1 - 3 shown, after concatenation, a 32-bit number is formed, that is, the output after passing through the first data parsing unit P1 is:
[0157] The transitional data bit Tdata is 32’h ff81 ff82.
[0158] The shift after concatenation, that is, the transitional exponent bit Tshifit is -1 (1 1111 1111), which is the same as the original input.
[0159] The extracted Sign is 0011. Among them, only two numbers are valid (i.e., 11, which are the signs of 82 and 81 respectively), and the invalid positions are 0; the valid numbers are two negative numbers, so the value is 1. That is, the transitional sign bit Tsign is 0011.
[0160] It should be understood that the above description is based on the spliced data. If a single data (such as 81) is used as the object and described by the actual value (the data before splicing), then the transitional data bit is 81, the transitional exponent bit is -1, and the transitional sign bit is 1.
[0161] As Figure 3 shown, after calculation, that is, after passing through the first arithmetic unit C1, the following can be obtained:
[0162] ABS = 32’h 007f 007e, the input is of Fix type, and the complement code is taken through the selector.
[0163] EXP = -1 (1 1111 1111), which is the same as the transitional exponent bit.
[0164] SIGN = 0011 (directly equal).
[0165] STK = 32’h 007f 007e (when rounding, STK = ABS).
[0166] Next, the intermediate results ABS, EXP, SIGN, and STK are input to the second conversion stage L2 (as Figures 6 - 9b shown):
[0167] Through the shift arithmetic unit P21, since EXP = -1, it is shifted one bit to the right, and the shift result = 32’h 003f003f;
[0168] Through the adder P22, for example, the addition this time adds STK[-EXP - 1] = STK[0]. When there are two numbers, it corresponds to STK
[16] = 1 and STK[0] = 0: the high 16 bits [31:16] of the adder output = 16’h 003f + STK
[16] = 16’h 0040; the low 16 bits [15:0] of the adder output = 16’h 003f + STK[0] = 16’h 003f. Therefore, the adder output = 32’h0040 003f.
[0169] Through the selector P23, obviously the output of the adder P22 is smaller, there is no overflow, and no exception is reported. And Pdata = the adder output = 32’h 0040 003f = 0000 0000 0100 0000 0000 0000 0011 1111.
[0170] Next, the data enters the pre-output processing circuit R21, as Figure 8 shown.
[0171] The output type is Float16, so the normalization of Pdata is performed, DATA_out = 32’h 0000 001f
[0172] SHIFT_FP = {6-15,5-15} = {-9,-10} = {10111,10110}
[0173] Next, the data passes through the data assembly circuit R22, as Figure 8 shown.
[0174] SIGN, SHIFT_FP DATA_out are assembled into two Float 16 type data.
[0175] The second type of data = {1,10111,0000000000,1,10110,0000011111}
[0176] = 32’h dc00 d81f
[0177] Example 2
[0178] Example 2 gives an example of converting Float16 to Fix8 with SHIFT = -3.
[0179] Let the input DATA = 32’h c001 4401 (1100 0000 0000 0001 0100 0100 0000 0001),
[0180] SHIFT = -3
[0181] The rounding method is rounding towards positive infinity
[0182] As Figures 1 - 3 shown,
[0183] Tdata = 32’h 0401 0401 (0000 0100 0000 0001 0000 0100 0000 0001) (only two significant digits, each with 11 bits, and the remaining bits are sign bit extended. Since the fp is in original code representation, the sign bit is filled with 0)
[0184] Tshifit = {16,17} (10000 10001). The input type is Float, and taking several middle bits is directly equal.
[0185] Tsign = 0010 (Only two numbers are valid, with invalid positions set to 0; the valid numbers are one negative and one positive, so it is set to 10)
[0186] As Figure 3 shown, after calculation, that is, after passing through the first arithmetic unit C1, we can obtain
[0187] ABS = 32’h 0401 0401, the input is Float, and the direct original code output is ABS = Tdata.
[0188] EXP = {16 - 15 - (3), 17–15–(3)} = {-2, -1} (The input is of Float type. First, take the excess code -15, and then subtract it from the output shift) = {11110 11111}
[0189] SIGN = 0010 (Directly equal)
[0190] STK = 32’h 0000ffff. When rounding towards positive infinity, in this example, the data representation bits are ABS[31:16], ABS[15:0]), then STK[n] = |ABS[n:x1] && ~SIGN, where x2 >= n >= x1). For the high 16 bits of a 32 - bit number, x2 = 31, x1 = 16; for its low 16 bits, x2 = 15, x1 = 0.
[0191] Next, the intermediate results ABS, EXP, SIGN, and STK are input to the second conversion stage L2 (as Figures 6 - 9b shown)
[0192] Through the shift arithmetic unit P21, since EXP = {-2, -1}, shift right by 2 and 1 bits respectively, and the shift result = 32’h0008 0010
[0193] Through the adder P22, for example, the numbers added this time are STK[-EXP - 1] = STK[2], STK[1]. When there are two numbers, it corresponds to STK
[18] = 0, STK[1] = 1: The high 16 bits [31:16] of the adder output = 16’h 0008+STK
[18] = 16’h0008; the low 16 bits [15:0] of the adder output = 16’h 0010+STK[1] = 16’h 0011. Therefore, the adder output = 32’h 0008 0011.
[0194] Through the selector P23, obviously the adder output is smaller, there is no overflow, and no exception is reported. And Pdata = adder output = 32’h 0008 0011 = 0000 0000 0000 1000 0000 0000 0001 0001.
[0195] Next, the data enters the pre-output processing circuit R21, as Figure 8 shown.
[0196] The output type is Fix, so the two's complement of Pdata is taken, DATA_out = 32'h fff8 0011
[0197] Next, the data passes through the data assembly circuit R22, as Figure 8 shown.
[0198] The obtained DATA_out is converted into two Fix8-type data, placed in the lower bits and the high 16 invalid bits are set to zero.
[0199] The second type of data obtained = 32'h 0000 f811.
[0200] The present disclosure also provides a method based on the above device, as Figure 2 shown. Other operations and steps of the method in the disclosure are not shown in the drawings for the purpose of simplification. The operations of the method of the present disclosure can be based on the specific devices, units and circuits described in the present disclosure, but can also be based on other software, hardware, firmware, etc., and are not limited to the above specific structures.
[0201] According to another aspect of the present disclosure, there is also provided an electronic device, including: one or more processors; and a memory, wherein computer-executable instructions are stored in the memory, and when the computer-executable instructions are run by the one or more processors, the electronic device executes the method as described above.
[0202] According to still another aspect of the present disclosure, there is also provided a computer-readable storage medium, including computer-executable instructions, and when the computer-executable instructions are run by one or more processors, the method as described above is executed.
[0203] In traditional actual calculations, there are few conversion types and few constraints in data type conversion work. Most of them can be completed with simple software behaviors and instructions in a few clock cycles. More importantly, the frequency of data type conversion instructions is very low.
[0204] In an AI chip, due to different precision requirements, the need for data type conversion is likely to occur in each step of the calculation. Once it occurs, it is not just a small number of calculations, but very intensive large-scale calculations, and its data organization is very regular. If the traditional data type conversion method is used, large-scale intensive calculations will cause a large memory access delay. Since the frequency of data type conversion instructions is relatively high, this part of the bottleneck will affect the overall calculation performance of the processor core.
[0205] In addition, for simple stack transfer instructions, a large amount of logical redundancy will occur in the transfer module, resulting in excessive local area, overly dense wiring, and affecting the local performance of the processor. The following is an example to illustrate the problem of logical redundancy: During the data type conversion from Fix4 to fp16, the Fix4 input needs to be converted into an absolute value form, the rounding bit is calculated based on this absolute value form, and at the final stage of data conversion, the same numerical data is represented in fixed-point form and then converted into a 10-bit mantissa representation of a floating-point number in normal or denormal form, and finally, the output number is assembled by the sign bit, exponent, and mantissa. In fact, during the conversion from Fix4 to fp16, there is also exactly the same first half of the logic: converting the Fix4 input into an absolute value form and calculating the rounding bit based on this absolute value form; during the conversion from Fix8 to fp16, there is also exactly the same second half of the logic: representing the same numerical data in fixed-point form and then converting it into a 10-bit mantissa representation of a floating-point number in normal or denormal form, and finally, the output number is assembled by the sign bit, exponent, and mantissa. If the instruction set is simply expanded, a large number of hardware operations with repeated logic and repeated calculations will occur (if the compiler behavior software controls the calculation of this part of the logic, then this part of the redundant calculation does not disappear but is repeated in the software implementation), which will affect the performance of the processor.
[0206] The main purpose of the structural design of the intermediate result in this disclosure is to reduce the repeated calculation logic, reduce the memory access latency overhead compared to software implementation, and at the same time have better scalability and portability. For example, as long as an intermediate result that can represent any data type is obtained, the intermediate result can be flexibly processed, rather than necessarily adopting the specific circuits and structures described in this disclosure. The content described in this disclosure can also be easily transplanted to other processing units, such as traditional CPUs and GPUs.
[0207] In the above embodiments of this disclosure, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not conflict, it should be considered as the scope described in this specification.
[0208] The foregoing can be better understood in accordance with the following clauses:
[0209] Clause A1. A converter for converting data types, comprising:
[0210] A first conversion stage (L1) configured to receive first type data and description information regarding the first type data and second type data, and convert the first type data into an intermediate result according to the description information; and
[0211] A second conversion stage (L2) configured to convert the intermediate result into second type data.
[0212] Clause A2. The converter according to Clause A1, wherein the first conversion stage (L1) includes a first data parsing unit (P1) and a first arithmetic unit (C1),
[0213] The first data parsing unit (P1) is configured to generate a transitional sign bit (Tsign), transitional data bits (Tdata), and a transitional exponent bit (Tshift) according to the first type data and the description information;
[0214] The first arithmetic unit (C1) is configured to generate an intermediate result according to the transitional sign bit (Tsign), transitional data bits (Tdata), and transitional exponent bit (Tshift).
[0215] Clause A3. The converter according to Clause A1 or A2, wherein the intermediate result includes intermediate data bits (ABS), an intermediate sign bit (Sign), and an intermediate exponent bit (EXP), and the first arithmetic unit (C1) includes:
[0216] An absolute value calculation circuit (C11) configured to calculate the intermediate data bits (ABS) according to the transitional data bits (Tdata);
[0217] An exponent bit calculation circuit (C12) configured to calculate the intermediate exponent bit (EXP) according to the transitional exponent bit (Tshift);
[0218] A sign bit calculation circuit (C13) configured to calculate the intermediate sign bit (Sign) according to the transitional sign bit (Tsign).
[0219] Clause A4. The converter according to any one of Clauses A1 - A3, wherein the intermediate result further includes an intermediate rounding bit (STK), and the first arithmetic unit (C1) further includes: a rounding bit calculation circuit (C14) configured to calculate the intermediate rounding bit (STK) according to the intermediate data bits (ABS), intermediate exponent bit (EXP), and intermediate sign bit (Sign).
[0220] Clause A5. The converter according to claim 3, wherein the intermediate result further includes an intermediate rounding bit (STK), and the first arithmetic unit (C1) further includes:
[0221] A rounding bit calculation circuit (C14), configured to calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS) and the intermediate sign bit (Sign).
[0222] Clause A6. The converter according to any one of Clauses A1 - A5, wherein the absolute value calculation circuit (C11) includes:
[0223] A second selector, configured to determine whether the transition data bit (Tdata) is less than zero;
[0224] A first two's complement calculator, configured to calculate the two's complement of the transition data bit as the intermediate data bit (ABS) if the transition data bit (Tdata) is less than zero; otherwise
[0225] Use the transition data bit Tdata as the intermediate data bit (ABS).
[0226] Clause A7. The converter according to any one of Clauses A1 - A6, wherein the absolute value calculation circuit (C11) further includes a first selector and a first normalizer,
[0227] The first selector is configured to determine whether the data type of the transition data bit (Tdata) is the first type or the second type;
[0228] If the data type of the transition data bit (Tdata) is the first type, select the second selector for processing;
[0229] If the data type of the transition data bit (Tdata) is the second type, select the first normalizer for processing;
[0230] The first normalizer is configured to normalize the transition data bit (Tdata) as the intermediate data bit (ABS) when the data type of the transition data bit (Tdata) is the second type.
[0231] Clause A8. The converter according to any one of Clauses A1 - A7, wherein the output intermediate exponent bit (EXP) of the exponent bit calculation circuit (C12) is equal to the transition exponent bit (Tshift).
[0232] Clause A9. The converter according to any one of Clauses A1 - A8, wherein,
[0233] The sign bit calculation circuit (C13) is a direct connection line.
[0234] Clause A10. The converter according to any one of Clauses A1 - A9, wherein the first conversion stage (L1) is further configured to determine the quantity of the received first type of data, splice the quantity of the first type of data to form first spliced data, and convert the first spliced data into an intermediate result according to the description information.
[0235] Clause A11. The converter according to any one of Clauses A1 - A10, wherein the quantity of the received first type of data is determined by the following method:
[0236] With a preset first fixed value; or
[0237] By dividing the processing bit number of the converter by the bit number of the higher median of the first type of data and the second type of data.
[0238] Clause A12. The converter according to any one of Clauses A1 - A11, wherein the first conversion stage (L1) is further configured to determine the quantity of the received first type of data to be split, split the first type of data into the quantity of split data, and convert the split data into an intermediate result according to the description information.
[0239] Clause A13. The converter according to any one of Clauses A1 - A12, wherein the quantity of the received first type of data to be split is determined by the following method:
[0240] With a preset second fixed value; or
[0241] By dividing the bit number of the higher median of the first type of data and the second type of data by the processing bit number of the converter.
[0242] Clause A14. The converter according to any one of Clauses A1 - A13, wherein the description information includes:
[0243] First description information, used to describe the data type of the first type of data and the first exponent bit of the first type of data;
[0244] Second description information, used to describe the data type of the second type of data and the second exponent bit of the second type of data;
[0245] The transition exponent bit (Tshift) is equal to the difference between the first exponent bit and the second exponent bit.
[0246] Clause A15. The converter according to any one of Clauses A1 - A14, wherein the description information includes:
[0247] The first data type of the first type of data;
[0248] The second data type of the second type of data; and
[0249] A differential exponent bit, which is used to indicate the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data;
[0250] The transition exponent bit (Tshift) is equal to the differential exponent bit.
[0251] Clause A16. The converter according to any one of Clauses A1 - A15, wherein the description information further includes a rounding type, and the rounding type includes at least one of the following: TO_ZERO, OFF_ZERO, UP, DOWN, ROUNDING_OFF_ZERO, ROUNDING_TO_EVEN, random rounding.
[0252] Clause A17. The converter according to any one of Clauses A1 - A16, wherein the second conversion stage (L2) includes a rounding bit calculation circuit (C14), configured to calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS) and the intermediate sign bit (Sign).
[0253] Clause A18. The converter according to any one of Clauses A1 - A17, wherein the second conversion stage (L2) includes a rounding bit calculation circuit (C14), configured to calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS), the intermediate exponent bit (EXP) and the intermediate sign bit (Sign).
[0254] Clause A19. The converter according to any one of Clauses A1 - A18, wherein the second conversion stage (L2) is configured to generate the second type of data according to the intermediate data bit (ABS), the intermediate sign bit (Sign), the intermediate exponent bit (EXP) and the intermediate rounding bit (STK).
[0255] Clause A20. The converter according to any one of Clauses A1 - A19, wherein the rounding bit calculation circuit (C14) is implemented by AND - OR logic.
[0256] Clause A21. The converter according to any one of Clauses A1 - A20, wherein the second conversion stage (L2) includes: a pre-output calculation unit (P2) and a data recovery unit (R2), and the pre-output calculation unit (P2) is configured to calculate a pre-output data bit (Pdata) and a pre-output sign bit (Psign) based on the intermediate data bit (ABS), the intermediate sign bit (Sign), the intermediate exponent bit (EXP), and the intermediate rounding bit (STK);
[0257] The data recovery unit (R2) is configured to generate a second type of data based on the pre-output data bit (Pdata) and the pre-output sign bit (Psign).
[0258] Clause A22. The converter according to any one of Clauses A1 - A21, wherein the pre-output calculation unit (P2) includes: a shift arithmetic unit (P21) and an adder (P22), configured to generate a temporary output data bit (ABS') and a pre-output sign bit (Psign), where
[0259] The shift arithmetic unit (P21) is configured to shift the intermediate data bit (ABS) by the intermediate exponent bit (EXP) to obtain a shift result;
[0260] The adder (P22) is configured to generate a temporary data bit (ABS') based on the shift result and the intermediate rounding bit (STK);
[0261] The pre-output sign bit (Psign) is the same as the intermediate sign bit.
[0262] Clause A23. The converter according to any one of Clauses A1 - A22, the pre-output calculation unit (P2) further includes a selector (P23), and the selector (P23) is configured to detect whether the temporary data bit (ABS') is greater than a saturation value,
[0263] If it is greater, perform saturation processing on the temporary data bit (ABS') to obtain the pre-output data bit (Pdata);
[0264] If it is not greater, output the temporary data bit (ABS') as the pre-output data bit (Pdata).
[0265] Clause A24. The converter according to any one of Clauses A1 - A23, wherein the data recovery unit (R2) includes a pre-output processing circuit (R21) and a data assembly circuit (R22):
[0266] The pre-output processing circuit (R21) is configured to receive the pre-output data bit (Pdata) and the pre-output symbol bit (Psign) to generate an output data bit representation (Data_out);
[0267] The data assembly circuit (R22) is configured to generate a second type of data based on the output data bit representation (Data_out) and the pre-output symbol bit (Psign).
[0268] Clause A25. The converter according to any one of Clauses A1 - A24, wherein the pre-output processing circuit (R21) is further configured to generate a floating-point decimal point position representation (SHIFT_FP), and the data assembly circuit (R22) is configured to generate the second type of data based on the data output bit representation (Data_out), the floating-point decimal point position representation (Shift_FP), and the pre-output symbol bit (Psign).
[0269] Clause A26. The converter according to any one of Clauses A1 - A25, wherein the pre-output processing circuit (R21) includes: a fourth selector and a two's complement calculator,
[0270] The fourth selector is configured to receive the pre-output data bit (Pdata) and the pre-output symbol bit (Psign). If the pre-output symbol bit (Psign) is negative, the pre-output data bit is output to the two's complement calculator. If the pre-output symbol bit (Psign) is non-negative, the pre-output data bit is output as the data output bit representation (Data_out);
[0271] The two's complement calculator is configured to calculate the two's complement of the pre-output data bit (Pdata).
[0272] Clause A27. The converter according to any one of Clauses A1 - A25, wherein the pre-output processing circuit (R21) further includes: a third selector, a second normalizer, and a floating-point decimal point determiner, where
[0273] The third selector is configured to receive the pre-output data bit (Pdata) and determine whether the data type of the pre-output data bit (Pdata) is the first type or the second type. If the data type of the pre-output data bit (Pdata) is the first type, the pre-output data bit (Pdata) is sent to the fourth selector. If the data type of the pre-output data bit (Pdata) is the second type, the pre-output data bit (Pdata) is sent to the second normalizer;
[0274] The second normalizer is configured to normalize the pre-output data bits (Pdata) and output them as data output bit representation (Data_out);
[0275] The floating-point decimal point position determiner is configured to determine the floating-point decimal point digit representation (SHIFT_FP) according to the output of the second normalizer.
[0276] Clause A28. For the converter according to any one of Clauses A1 - A27, the first conversion stage (L1) is further configured to receive constraint information, where the constraint information is used to indicate whether a specific standard is supported and / or whether compilation optimization is supported.
[0277] Clause A29. For the converter according to any one of Clauses A1 - A28, wherein the data types of the first type of data and the second type of data are extensible.
[0278] Clause A30. A chip, including the converter according to any one of Clauses A1 - A29.
[0279] Clause A31. A computing device, including the converter according to any one of Clauses A1 - A29 or the chip according to Clause A30.
[0280] Clause A32. A method for converting data types, including:
[0281] Receiving the first type of data and description information about the first type of data and the second type of data, and converting the first type of data into an intermediate result according to the description information; and
[0282] Converting the intermediate result into the second type of data.
[0283] Clause A33. For the method according to Clause A32, wherein converting the first type of data into an intermediate result includes:
[0284] Generating a transition sign bit (Tsign), transition data bits (Tdata), and a transition exponent bit (Tshift) according to the first type of data and the description information;
[0285] Generating an intermediate result according to the transition sign bit (Tsign), transition data bits (Tdata), and the transition exponent bit (Tshift).
[0286] Clause A34. The method according to Clause A32 or A33, wherein the intermediate result includes an intermediate data bit (ABS), an intermediate sign bit (Sign), and an intermediate exponent bit (EXP), and generating the intermediate result according to the transition sign bit (Tsign), the transition data bit (Tdata), and the transition exponent bit (Tshift) includes:
[0287] Calculating the intermediate data bit (ABS) according to the transition data bit (Tdata);
[0288] Calculating the intermediate exponent bit (EXP) according to the transition exponent bit (Tshift);
[0289] Calculating the intermediate sign bit (Sign) according to the transition sign bit (Tsign).
[0290] Clause A35. The method according to any one of Clauses A32 - A34, wherein the intermediate result further includes an intermediate rounding bit (STK), and generating the intermediate result further according to the transition sign bit (Tsign) and the transition exponent bit (Tshift) includes:
[0291] Calculating the intermediate rounding bit (STK) according to the intermediate data bit (ABS), the intermediate exponent bit (EXP), and the intermediate sign bit (Sign).
[0292] Clause A36. The method according to any one of Clauses A32 - A35, wherein the intermediate result further includes an intermediate rounding bit (STK), and generating the intermediate result further according to the transition sign bit (Tsign), the transition data bit (Tdata), and the transition exponent bit (Tshift) includes:
[0293] Calculating the intermediate rounding bit (STK) according to the intermediate data bit (ABS), the intermediate exponent bit (EXP), and the intermediate sign bit (Sign).
[0294] Clause A37. The method according to any one of Clauses A32 - A36, wherein calculating the intermediate data bit (ABS) according to the transition data bit (Tdata) includes:
[0295] Determining whether the transition data bit (Tdata) is less than zero;
[0296] If the transition data bit (Tdata) is less than zero, calculating the two's complement of the transition data bit as the intermediate data bit (ABS); otherwise using the transition data bit Tdata as the intermediate data bit (ABS).
[0297] Clause A38. For the method according to any one of Clauses A32 - A37, wherein calculating the intermediate data bit (ABS) based on the transition data bit (Tdata) further includes:
[0298] Determine whether the data type of the transition data bit (Tdata) is the first type or the second type,
[0299] If the data type of the transition data bit (Tdata) is the first type, then
[0300] Determine whether the transition data bit (Tdata) is less than zero;
[0301] If the transition data bit (Tdata) is less than zero, calculate the complement code of the transition data bit as the intermediate data bit (ABS); otherwise, use the transition data bit Tdata as the intermediate data bit (ABS);
[0302] If the data type of the transition data bit (Tdata) is the second type, then
[0303] Normalize the transition data bit (Tdata) to be the intermediate data bit (ABS).
[0304] Clause A39. For the method according to any one of Clauses A32 - A38, wherein the intermediate exponent bit (EXP) is equal to the transition exponent bit (Tshift).
[0305] Clause A40. For the method according to any one of Clauses A32 - A39, wherein calculating the intermediate rounding bit (STK) is implemented through AND - OR logic.
[0306] Clause A41. For the method according to any one of Clauses A32 - A40, receiving the first - type data and the description information about the first - type data and the second - type data includes:
[0307] Determine the quantity of the received first - type data, and splice the first - type data of the quantity together to form the first spliced data, and the first spliced data is converted into an intermediate result.
[0308] Clause A42. For the method according to any one of Clauses A32 - A41, wherein the quantity of the received first - type data is determined by the following methods:
[0309] With a preset first fixed value; or
[0310] By dividing the processing bit number of the converter adopted by the method by the bit number of the higher of the bit numbers of the first - type data and the second - type data.
[0311] Clause A43. The method according to any one of Clauses A32 - A42, wherein receiving the first type of data and the description information about the first type of data and the second type of data includes:
[0312] Determining the number of splits for the received first type of data, and splitting the first type of data into that number of split data, the split data being converted into intermediate results.
[0313] Clause A44. The method according to any one of Clauses A32 - A43, wherein the number of splits for the received first type of data is determined by:
[0314] A preset second fixed value; or
[0315] The number of digits of the higher of the medians of the first type of data and the second type of data divided by the processing digits of the converter used in the method.
[0316] Clause A45. The method according to any one of Clauses A32 - A44, wherein the description information includes:
[0317] First description information for describing the data type of the first type of data and the first exponent bits of the first type of data;
[0318] Second description information for describing the data type of the second type of data and the second exponent bits of the second type of data;
[0319] The transition exponent bit (Tshift) is equal to the difference between the first exponent bit and the second exponent bit.
[0320] Clause A46. The method according to any one of Clauses A32 - A45, wherein the description information includes:
[0321] The first data type of the first type of data;
[0322] The second data type of the second type of data; and
[0323] A differential exponent bit for indicating the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data;
[0324] The transition exponent bit (Tshift) is equal to the differential exponent bit.
[0325] Clause A47. The method according to any one of Clauses A32 - A46, wherein the description information further includes a rounding type, and the rounding type includes at least one of the following: TO_ZERO, OFF_ZERO, UP, DOWN, ROUNDING_OFF_ZERO, ROUNDING_TO_EVEN, random rounding.
[0326] Clause A48. The method according to any one of Clauses A32 - A47, wherein converting the intermediate result to the second type of data includes:
[0327] Generating the second type of data according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP), and intermediate rounding bit (STK).
[0328] Clause A49. The method according to any one of Clauses A32 - A48, wherein converting the intermediate result to the second type of data includes:
[0329] Calculating a pre-output data bit (Pdata) and a pre-output sign bit (Psign) according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP), and intermediate rounding bit (STK); and
[0330] Generating the second type of data according to the pre-output data bit (Pdata) and the pre-output sign bit (Psign).
[0331] Clause A50. The method according to any one of Clauses A32 - A49, wherein calculating the pre-output data bit (Pdata) and the pre-output sign bit (Psign) according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP), and intermediate rounding bit (STK) includes:
[0332] Shifting the intermediate data bit (ABS) by the intermediate exponent bit (EXP) to obtain a shifted result;
[0333] Generating a temporary data bit (ABS’) according to the shifted result and the intermediate rounding bit (STK);
[0334] The pre-output sign bit (Psign) is the same as the intermediate sign bit.
[0335] Clause A51. The method according to any one of Clauses A21 - A50, calculating the pre-output data bit (Pdata) and the pre-output sign bit (Psign) according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP), and intermediate rounding bit (STK) further includes:
[0336] Check whether the temporary data bit (ABS’) is greater than the saturation value,
[0337] If it is greater, perform saturation processing on the temporary data bit (ABS’) to obtain the pre-output data bit (Pdata);
[0338] If it is not greater, output the temporary data bit (ABS’) as the pre-output data bit (Pdata).
[0339] Clause A52. The method according to any one of Clauses A32 - A51, wherein generating the second type of data according to the pre-output data bit (Pdata) and the pre-output sign bit (Psign) includes:
[0340] Receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign) to generate an output data bit representation (Data_out);
[0341] Obtaining the second type of data according to the data output bit representation (Data_out) and the pre-output sign bit (Psign).
[0342] Clause A53. The method according to any one of Clauses A32 - A52, wherein generating the second type of data according to the pre-output data bit (Pdata) and the pre-output sign bit (Psign) further includes: generating a floating-point decimal point bit representation (SHIFT_FP) according to the pre-output data bit (Pdata) and the pre-output sign bit (Psign);
[0343] Obtaining the second type of data according to the data output bit representation (Data_out), the floating-point decimal point bit representation (Shift_FP), and the pre-output sign bit (Psign).
[0344] Clause A54. The method according to any one of Clauses A32 - A53, wherein receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign) to generate an output data bit representation (Data_out) includes:
[0345] Receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign), if the pre-output sign bit (Psign) is negative, taking the two's complement of the pre-output data bit (Pdata);
[0346] If the pre-output sign bit (Psign) is positive, outputting the pre-output data bit as the data output bit representation (Data_out).
[0347] Clause A55. The method according to any one of Clauses A32 - A54, wherein receiving the pre - output data bits (Pdata) and the pre - output symbol bits (Psign) to generate the output data bit representation (Data_out) further includes:
[0348] Receiving the pre - output data bits (Pdata), and determining whether the data type of the pre - output data bits (Pdata) is of the first type or the second type,
[0349] If the data type of the pre - output data bits (Pdata) is of the first type, then
[0350] If the pre - output symbol bit (Psign) is negative, then take the two's complement of the pre - output data bits (Pdata);
[0351] If the pre - output symbol bit (Psign) is non - negative, then output the pre - output data bits as the data output bit representation (Data_out);
[0352] If the data type of the pre - output data bits (Pdata) is of the second type, then
[0353] Normalize the pre - output data bits (Pdata) and output them as the data output bit representation (Data_out);
[0354] The floating - point decimal point position determiner is configured to determine the floating - point decimal point digit representation (SHIFT_FP) according to the output of the second normalizer.
[0355] Clause A56. The method according to any one of Clauses A32 - A55, further includes receiving constraint information, where the constraint information is used to indicate whether a specific standard is supported, and / or whether compilation optimization is supported.
[0356] Clause A57. The method according to any one of Clauses A32 - A56, wherein the data types of the first - type data and the second - type data are extensible.
[0357] Clause A58. An electronic device, comprising: one or more processors; and a memory, where computer - executable instructions are stored in the memory, and when the computer - executable instructions are run by the one or more processors, the electronic device executes the method according to any one of Clauses A32 - A57.
[0358] Clause A59. A computer-readable storage medium includes computer-executable instructions that, when run by one or more processors, perform the method described in any one of Clauses A32 - A57.
[0359] This disclosure also discloses a combined processing device 1000, which includes the above-described computing device 1002, a general interconnect interface 1004, and other processing devices 1006. The computing device according to this disclosure interacts with other processing devices to jointly complete the operations specified by the user. Figure 10 It is a schematic diagram of the combined processing device.
[0360] Other processing devices include one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processor among general-purpose / special-purpose processors. The number of processors included in other processing devices is not limited. Other processing devices serve as an interface for machine learning computing devices to external data and control, including data transfer, and complete basic controls such as starting and stopping the machine learning computing device; other processing devices can also cooperate with the machine learning computing device to jointly complete computing tasks.
[0361] The general interconnect interface is used to transfer data and control instructions between the computing device (including, for example, a machine learning computing device) and other processing devices. The computing device obtains the required input data from other processing devices and writes it into the storage device on the chip of the computing device; it can obtain control instructions from other processing devices and write them into the control cache on the chip of the computing device; it can also read the data in the storage module of the computing device and transfer it to other processing devices.
[0362] Optionally, this structure may further include a storage device 1008, which is respectively connected to the computing device and the other processing devices. The storage device is used to store data in the computing device and the other processing devices, especially suitable for data that cannot be fully stored in the internal storage of the computing device or other processing devices for the required operations.
[0363] The combined processing device can be used as a system-on-chip (SOC) for devices such as mobile phones, robots, drones, and video surveillance devices, effectively reducing the core area of the control part, increasing the processing speed, and reducing the overall power consumption. In this case, the general interconnect interface of the combined processing device is connected to certain components of the device. Certain components such as cameras, displays, mice, keyboards, network cards, and Wi-Fi interfaces.
[0364] In some embodiments, this disclosure also discloses a chip that includes the above-described computing device or combined processing device.
[0365] In some embodiments, this disclosure also discloses a chip packaging structure that includes the above-described chip.
[0366] In some embodiments, the present disclosure also discloses a board card, which includes the above chip packaging structure. Refer to Figure 11 , which provides an exemplary board card. In addition to including the above chip 1102, the board card may further include other supporting components, and the supporting components include but are not limited to: a storage device 1104, an interface device 1106, and a control device 1108;
[0367] The storage device is connected to the chip in the chip packaging structure through a bus and is used for storing data. The storage device may include multiple groups of storage units 1110. Each group of the storage units is connected to the chip through a bus. It can be understood that each group of the storage units may be DDR SDRAM (English: Double Data Rate SDRAM, double data rate synchronous dynamic random access memory).
[0368] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read on both the rising edge and the falling edge of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device may include 4 groups of the storage units. Each group of the storage units may include multiple DDR4 chips. In one embodiment, the chip may internally include 4 72-bit DDR4 controllers. 64 bits of the 72-bit DDR4 controllers are used for data transmission, and 8 bits are used for ECC verification. In one embodiment, each group of the storage units includes multiple double data rate synchronous dynamic random access memories arranged in parallel. DDR can transmit data twice within one clock cycle. A controller for controlling DDR is provided in the chip to control the data transmission and data storage of each of the storage units.
[0369] The interface device is electrically connected to the chip in the chip packaging structure. The interface device is used to implement data transmission between the chip and an external device 1112 (such as a server or a computer). For example, in one embodiment, the interface device may be a standard PCIE interface. For example, the data to be processed is transmitted from the server to the chip through the standard PCIE interface to achieve data transfer. In another embodiment, the interface device may also be other interfaces. The present disclosure does not limit the specific form of the above other interfaces, as long as the interface unit can achieve the transfer function. In addition, the calculation result of the chip is still transmitted back to the external device (such as a server) by the interface device.
[0370] The control device is electrically connected to the chip. The control device is used to monitor the state of the chip. Specifically, the chip and the control device can be electrically connected through an SPI interface. The control device may include a microcontroller unit (MCU). For example, the chip may include multiple processing chips, multiple processing cores, or multiple processing circuits, and can drive multiple loads. Therefore, the chip can be in different operating states such as multi-load and light-load. Through the control device, the operating states of multiple processing chips, multiple processes, and / or multiple processing circuits in the chip can be regulated.
[0371] In some embodiments, the present disclosure also discloses an electronic device or apparatus, which includes the above-mentioned board.
[0372] The electronic device or apparatus includes a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, headphones, a mobile storage device, a wearable device, a vehicle, a household appliance, and / or a medical device.
[0373] The vehicle includes an airplane, a ship, and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, a rice cooker, a humidifier, a washing machine, a light, a gas stove, an oil fume extractor; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument, and / or an electrocardiogram instrument.
[0374] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.
[0375] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0376] In several embodiments provided by this disclosure, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, optical, acoustic, magnetic or other forms.
[0377] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0378] In addition, in each embodiment of this disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software program modules.
[0379] If the above integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, when the technical solution of this disclosure can be embodied in the form of a software product, the computer software product is stored in a memory, including several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this disclosure. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0380] The above has introduced the embodiments of this disclosure in detail. Specific examples are used in this article to elaborate on the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only used to help understand the method and its core idea of this disclosure; at the same time, for those of ordinary skill in the art, according to the idea of this disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this disclosure.
Claims
1. A converter for converting data types, comprising: The first conversion stage (L1) is configured to receive the first type of data and the description information about the first type of data and the second type of data, and convert the first type of data into an intermediate result according to the description information; And The second conversion stage (L2) is configured to convert the intermediate result into the second type of data; Wherein, the first conversion stage (L1) includes a first data parsing unit (P1) and a first arithmetic unit (C1), The first data parsing unit (P1) is configured to generate a transition exponent bit (Tshift) according to the first type of data and the description information; the transition exponent bit (Tshift) is equal to the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data; The first arithmetic unit (C1) is configured to generate an intermediate result according to the transition exponent bit (Tshift).
2. The converter according to claim 1, wherein, The first data parsing unit (P1) is configured to generate a transition sign bit (Tsign) and a transition data bit (Tdata) according to the first type of data and the description information; The first arithmetic unit (C1) is configured to generate an intermediate result according to the transition sign bit (Tsign), the transition data bit (Tdata), and the transition exponent bit (Tshift).
3. The converter according to claim 2, wherein, The intermediate result includes an intermediate data bit (ABS), an intermediate sign bit (Sign), and an intermediate exponent bit (EXP), and the first arithmetic unit (C1) includes: An absolute value calculation circuit (C11) configured to calculate the intermediate data bit (ABS) according to the transition data bit (Tdata); An exponent bit calculation circuit (C12) configured to calculate the intermediate exponent bit (EXP) according to the transition exponent bit (Tshift); A sign bit calculation circuit (C13) configured to calculate the intermediate sign bit (Sign) according to the transition sign bit (Tsign).
4. The converter according to claim 3, wherein, The intermediate result further includes an intermediate rounding bit (STK), and the first arithmetic unit (C1) further includes: A rounding bit calculation circuit (C14) configured to calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS) and the intermediate sign bit (Sign).
5. The converter according to claim 3, wherein, The intermediate result further includes an intermediate rounding bit (STK), and the first arithmetic unit (C1) further includes: A rounding bit calculation circuit (C14) configured to calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS), the intermediate exponent bit (EXP), and the intermediate sign bit (Sign).
6. The converter according to any one of claims 3-5, wherein, The absolute value calculation circuit (C11) includes: A second selector configured to determine whether the transition data bit (Tdata) is less than zero; A first two's complement calculator configured to calculate the two's complement of the transition data bit as the intermediate data bit (ABS) if the transition data bit (Tdata) is less than zero; otherwise Use the transition data bit (Tdata) as the intermediate data bit (ABS).
7. The converter according to claim 6, wherein, The absolute value calculation circuit (C11) further includes a first selector and a first normalizer, The first selector is configured to determine whether the data type of the transition data bit (Tdata) is the first type or the second type; If the data type of the transition data bit (Tdata) is the first type, select the second selector for processing; If the data type of the transition data bit (Tdata) is the second type, select the first normalizer for processing; The first normalizer is configured to normalize the transition data bit (Tdata) as the intermediate data bit (ABS) when the data type of the transition data bit (Tdata) is the second type.
8. The converter according to any one of claims 3-5, wherein, The output intermediate exponent bit (EXP) of the exponent bit calculation circuit (C12) is equal to the transition exponent bit (Tshift).
9. The converter according to any one of claims 3-5, wherein, The sign bit calculation circuit (C13) is a direct connection line.
10. The converter according to any one of claims 1-5, wherein, The first conversion stage (L1) is further configured to determine the number of the received first type of data, splice the number of the first type of data to form first spliced data, and convert the first spliced data into an intermediate result according to the description information.
11. The converter according to claim 10, wherein, The number of the received first type of data is determined by the following methods: Using a preset first fixed value; or Dividing the processing bit number of the converter by the bit number of the higher one of the first type of data and the second type of data.
12. The converter according to any one of claims 1-5, wherein, The first conversion stage (L1) is further configured to determine the number of splits of the received first type of data, split the first type of data into the number of split data, and convert the split data into an intermediate result according to the description information.
13. The converter according to claim 12, wherein, The number of splits of the received first type of data is determined by the following methods: Using a preset second fixed value; or Dividing the bit number of the higher one of the first type of data and the second type of data by the processing bit number of the converter.
14. The converter according to any one of claims 1-5, wherein, The description information includes: First description information, which is used to describe the data type of the first type of data and the first exponent bit of the first type of data; Second description information, which is used to describe the data type of the second type of data and the second exponent bit of the second type of data.
15. The converter according to any one of claims 1-5, wherein, The description information includes: The first data type of the first type of data; The second data type of the second type of data; and A differential exponent bit, which is used to indicate the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data; The transition exponent bit (Tshift) is equal to the differential exponent bit.
16. The converter according to claim 14, wherein, The description information further includes a rounding type, and the rounding type includes at least one of the following: TO_ZERO, OFF_ZERO, UP, DOWN, ROUNDING_OFF_ZERO, ROUNDING_TO_EVEN, random rounding.
17. The converter according to claim 4, wherein, The second conversion stage (L2) includes a rounding bit calculation circuit (C14), which is configured to calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS) and the intermediate sign bit (Sign).
18. The converter according to claim 4, wherein, The second conversion stage (L2) includes a rounding bit calculation circuit (C14) configured to calculate the intermediate rounding bit (STK) based on the intermediate data bit (ABS), the intermediate exponent bit (EXP), and the intermediate sign bit (Sign).
19. The converter according to claim 4, 5, 17 or 18, wherein, The second conversion stage (L2) is configured to generate a second type of data based on the intermediate data bit (ABS), the intermediate sign bit (Sign), the intermediate exponent bit (EXP), and the intermediate rounding bit (STK).
20. The converter according to claim 4, 5, 17 or 18, wherein, The rounding bit calculation circuit (C14) is implemented by AND-OR logic.
21. The converter according to claim 19, wherein, The second conversion stage (L2) includes: a pre-output calculation unit (P2) and a data recovery unit (R2), The pre-output calculation unit (P2) is configured to calculate a pre-output data bit (Pdata) and a pre-output sign bit (Psign) based on the intermediate data bit (ABS), the intermediate sign bit (Sign), the intermediate exponent bit (EXP), and the intermediate rounding bit (STK); The data recovery unit (R2) is configured to generate a second type of data based on the pre-output data bit (Pdata) and the pre-output sign bit (Psign).
22. The converter according to claim 21, wherein, The pre-output calculation unit (P2) includes: a shift operator (P21) and an adder (P22), configured to generate a temporary output data bit (ABS’) and a pre-output sign bit (Psign), where The shift operator (P21) is configured to shift the intermediate data bit (ABS) by the intermediate exponent bit (EXP) to obtain a shift result; The adder (P22) is configured to generate a temporary data bit (ABS’) based on the shift result and the intermediate rounding bit (STK); The pre-output sign bit (Psign) is identical to the intermediate sign bit.
23. The converter according to claim 22, wherein the pre-output calculation unit (P2) further includes a selector (P23), and the selector (P23) is configured to detect whether the temporary data bit (ABS’) is greater than the saturation value. If it is greater, perform saturation processing on the temporary data bit (ABS’) to obtain the pre-output data bit (Pdata). If it is not greater, output the temporary data bit (ABS’) as the pre-output data bit (Pdata).
24. The converter according to claim 21, wherein The data recovery unit (R2) includes a pre-output processing circuit (R21) and a data assembly circuit (R22): The pre-output processing circuit (R21) is configured to receive the pre-output data bit (Pdata) and the pre-output sign bit (Psign) to generate an output data bit representation (Data_out); The data assembly circuit (R22) is configured to generate a second type of data based on the output data bit representation (Data_out) and the pre-output sign bit (Psign).
25. The converter according to claim 24, wherein The pre-output processing circuit (R21) is further configured to generate a floating-point decimal point digit representation (Shift_FP), and the data assembly circuit (R22) is configured to generate the second type of data based on the output data bit representation (Data_out), the floating-point decimal point digit representation (Shift_FP), and the pre-output sign bit (Psign).
26. The converter according to claim 24, wherein The pre-output processing circuit (R21) includes: a fourth selector and a two's complement calculator, The fourth selector is configured to receive the pre-output data bit (Pdata) and the pre-output sign bit (Psign). If the pre-output sign bit (Psign) is negative, it outputs the pre-output data bit to the two's complement calculator. If the pre-output sign bit (Psign) is non-negative, it outputs the pre-output data bit as the output data bit representation (Data_out); The two's complement calculator is configured to complement the pre-output data bit (Pdata).
27. The converter according to claim 26, wherein The pre-output processing circuit (R21) further includes: a third selector, a second normalizer, and a floating-point decimal point position determiner, where The third selector is configured to receive the pre-output data bit (Pdata), and determine whether the data type of the pre-output data bit (Pdata) is the first type or the second type. If the data type of the pre-output data bit (Pdata) is the first type, it sends the pre-output data bit (Pdata) to the fourth selector. If the data type of the pre-output data bit (Pdata) is the second type, it sends the pre-output data bit (Pdata) to the second normalizer; The second normalizer is configured to normalize the pre-output data bit (Pdata) and output it as the output data bit representation (Data_out); The floating-point decimal point position determiner is configured to determine the floating-point decimal point digit representation (Shift_FP) according to the output of the second normalizer.
28. The converter according to any one of claims 1-5, wherein the first conversion stage (L1) is further configured to receive constraint information, and the constraint information is used to indicate whether a specific standard is supported and / or whether compilation optimization is supported.
29. The converter according to any one of claims 1-5, wherein The data types of the first type of data and the second type of data are extensible.
30. A chip, comprising the converter according to any one of claims 1-29.
31. A computing device, comprising the converter according to any one of claims 1-29 or the chip according to claim 30.
32. A method for converting data types, comprising: Receiving the first type of data and the description information about the first type of data and the second type of data, and converting the first type of data into an intermediate result according to the description information; And Converting the intermediate result into the second type of data; Among them, converting the first type of data into an intermediate result includes: Generating a transitional exponent bit (Tshift) according to the first type of data and the description information; the transitional exponent bit (Tshift) is equal to the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data; Generating an intermediate result according to the transitional exponent bit (Tshift).
33. The method according to claim 32, wherein, Converting the first type of data into an intermediate result includes: Generating a transitional sign bit (Tsign) and a transitional data bit (Tdata) according to the first type of data and the description information; Generating an intermediate result according to the transitional sign bit (Tsign) and the transitional data bit (Tdata).
34. The method according to claim 33, wherein, The intermediate result includes an intermediate data bit (ABS), an intermediate sign bit (Sign), and an intermediate exponent bit (EXP). Generating an intermediate result according to the transitional sign bit (Tsign), the transitional data bit (Tdata), and the transitional exponent bit (Tshift) includes: Calculating the intermediate data bit (ABS) according to the transitional data bit (Tdata); Calculate the intermediate exponent bit (EXP) according to the transition exponent bit (Tshift); Calculate the intermediate sign bit (Sign) according to the transition sign bit (Tsign).
35. The method according to claim 34, wherein, The intermediate result further includes an intermediate rounding bit (STK). Generating the intermediate result according to the transition sign bit (Tsign), transition data bit (Tdata), and transition exponent bit (Tshift) further includes: Calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS) and the intermediate sign bit (Sign).
36. The method according to claim 34, wherein, The intermediate result further includes an intermediate rounding bit (STK). Generating the intermediate result according to the transition sign bit (Tsign), transition data bit (Tdata), and transition exponent bit (Tshift) further includes: Calculate the intermediate rounding bit (STK) according to the intermediate data bit (ABS), intermediate exponent bit (EXP), and intermediate sign bit (Sign).
37. The method according to any one of claims 34-36, wherein, Calculating the intermediate data bit (ABS) according to the transition data bit (Tdata) includes: Determine whether the transition data bit (Tdata) is less than zero; If the transition data bit (Tdata) is less than zero, calculate the two's complement of the transition data bit as the intermediate data bit (ABS); otherwise, use the transition data bit (Tdata) as the intermediate data bit (ABS).
38. The method according to claim 37, wherein, Calculating the intermediate data bit (ABS) according to the transition data bit (Tdata) further includes: Determine whether the data type of the transition data bit (Tdata) is the first type or the second type. If the data type of the transition data bit (Tdata) is the first type, then Determine whether the transition data bit (Tdata) is less than zero; If the transition data bit (Tdata) is less than zero, calculate the two's complement of the transition data bit as the intermediate data bit (ABS); otherwise, use the transition data bit (Tdata) as the intermediate data bit (ABS); If the data type of the transition data bit (Tdata) is the second type, then Normalize the transition data bit (Tdata) to be the intermediate data bit (ABS).
39. The method according to any one of claims 34-36, wherein, The intermediate exponent bit (EXP) is equal to the transition exponent bit (Tshift).
40. The method according to claim 36, wherein, Calculating the intermediate rounding bit (STK) is implemented through AND-OR logic.
41. The method according to any one of claims 32-36, receiving the first type of data and the description information about the first type of data and the second type of data includes: Determine the number of received first-type data, and splice the first-type data of the number to form a first spliced data, which is converted into an intermediate result.
42. The method according to claim 41, wherein, Determine the number of received first-type data in the following way: With a preset first fixed value; or Divide the processing bit number of the converter used by the method by the bit number of the higher one of the first-type data and the second-type data.
43. The method according to any one of claims 32-36, wherein,Receiving the first-type data and the description information about the first-type data and the second-type data includes: Determine the number of splits of the received first-type data to be split, and split the first-type data into the split data of the number, which is converted into an intermediate result.
44. The method according to claim 43, wherein the number of splits of the received first type of data is determined by the following means: With a preset second fixed value; or By dividing the number of bits of the higher one of the median of the first type of data and the second type of data by the number of bits processed by the converter used in the method.
45. The method according to any one of claims 32-36, wherein, The description information includes: The first description information is used to describe the data type of the first type of data and the first exponent bit of the first type of data; The second description information is used to describe the data type of the second type of data and the second exponent bit of the second type of data.
46. The method according to any one of claims 32-36, wherein, The description information includes: The first data type of the first type of data; The second data type of the second type of data; and The differential exponent bit, which is used to indicate the difference between the first exponent bit of the first type of data and the second exponent bit of the second type of data; The transition exponent bit (Tshift) is equal to the differential exponent bit.
47. The method according to claim 45, wherein, The description information further includes a rounding type, and the rounding type includes at least one of the following: TO_ZERO, OFF_ZERO, UP, DOWN, ROUNDING_OFF_ZERO, ROUNDING_TO_EVEN, random rounding.
48. The method according to claim 35 or 36, wherein, Converting the intermediate result to the second type of data includes: Generating the second type of data according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP) and intermediate rounding bit (STK).
49. The method according to claim 48, wherein, Converting the intermediate result to the second type of data includes: Calculating a pre-output data bit (Pdata) and a pre-output sign bit (Psign) according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP) and intermediate rounding bit (STK); and Generating the second type of data according to the pre-output data bit (Pdata) and the pre-output sign bit (Psign).
50. The method according to claim 49, wherein, Calculating the pre-output data bit (Pdata) and the pre-output sign bit (Psign) according to the intermediate data bit (ABS), intermediate sign bit (Sign), intermediate exponent bit (EXP) and intermediate rounding bit (STK) includes: Shifting the intermediate data bit (ABS) by the intermediate exponent bit (EXP) to obtain a shifted result; Generating a temporary data bit (ABS’) according to the shifted result and the intermediate rounding bit (STK); The pre-output sign bit (Psign) is the same as the intermediate sign bit.
51. The method according to claim 50, further comprising calculating the pre-output data bit (Pdata) and the pre-output sign bit (Psign) according to the intermediate data bit (ABS), the intermediate sign bit (Sign), the intermediate exponent bit (EXP), and the intermediate rounding bit (STK): Detecting whether the temporary data bit (ABS’) is greater than the saturation value, If it is greater, performing saturation processing on the temporary data bit (ABS’) to obtain the pre-output data bit (Pdata); If it is not greater, outputting the temporary data bit (ABS’) as the pre-output data bit (Pdata).
52. The method according to claim 49, wherein, Generating the second type of data according to the pre-output data bit (Pdata) and the pre-output sign bit (Psign) includes: Receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign) to generate an output data bit representation (Data_out); Obtaining the second type of data according to the output data bit representation (Data_out) and the pre-output sign bit (Psign).
53. The method according to claim 52, wherein, Generating the second type of data based on the pre-output data bit (Pdata) and the pre-output sign bit (Psign) further includes: generating a floating-point decimal point position representation (Shift_FP) based on the pre-output data bit (Pdata) and the pre-output sign bit (Psign); Obtaining the second type of data based on the output data bit representation (Data_out), the floating-point decimal point position representation (Shift_FP), and the pre-output sign bit (Psign).
54. The method according to claim 52, wherein, Receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign) to generate the output data bit representation (Data_out) includes: Receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign), If the pre-output sign bit (Psign) is negative, taking the two's complement of the pre-output data bit (Pdata); If the pre-output sign bit (Psign) is positive, outputting the pre-output data bit as the output data bit representation (Data_out).
55. The method according to claim 54, wherein, Receiving the pre-output data bit (Pdata) and the pre-output sign bit (Psign) to generate the output data bit representation (Data_out) further includes: Receiving the pre-output data bit (Pdata) and determining whether the data type of the pre-output data bit (Pdata) is the first type or the second type, If the data type of the pre-output data bit (Pdata) is the first type, then If the pre-output sign bit (Psign) is negative, taking the two's complement of the pre-output data bit (Pdata); If the pre-output sign bit (Psign) is non-negative, outputting the pre-output data bit as the output data bit representation (Data_out); If the data type of the pre-output data bit (Pdata) is the second type, then Normalizing the pre-output data bit (Pdata) and outputting it as the output data bit representation (Data_out); The floating-point decimal point position determiner is configured to determine the floating-point decimal point position representation (Shift_FP) according to the output of the second normalizer.
56. The method according to any one of claims 32 - 36, further comprising receiving constraint information for indicating whether a specific standard is supported, and / or whether compilation optimization is supported.
57. The method according to any one of claims 32 - 36, wherein, The data types of the first type of data and the second type of data are extensible.
58. An electronic device, comprising: One or more processors; And A memory storing computer-executable instructions that, when run by the one or more processors, cause the electronic device to execute the method according to any one of claims 32-57.
59. A computer - readable storage medium, comprising computer - executable instructions which, when run by one or more processors, perform the method according to any one of claims 32 - 57.
Citation Information
Patent Citations
Register circuit realizing grouping addressing and read write control method for register files
CN101930355A
Data type conversion circuit unit and apparatus
CN108055041A