EXTENDED FORMATTING OF BINARY FLOATTING POINT NUMBERS WITH LOWER ACCURACY
Patent Information
- Application Number
- DE112019001799
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-30
- Filing Date
- 2019-05-30
- Publication Date
- 2026-10-01
- Estimated Expiration
- 2039-05-30
AI Technical Summary
Existing deep learning applications face inefficiencies and prolonged convergence times due to the use of double and single precision representations, with lower precision formats like 1/5/10 format leading to reduced accuracy and failure in training algorithms.
Implementing an extended floating-point number format with a 1/6/9 format, utilizing a 6-bit exponent and 9-bit mantissa, which avoids subnormal numbers and simplifies handling of zero and special values, allowing for faster convergence and reduced error rates in machine and deep learning tasks.
The extended floating-point format enables efficient hardware usage, reducing complexity and resource requirements while maintaining accuracy, as demonstrated by improved convergence and performance in speech, image, and text generation tasks compared to half-precision and full-precision formats.
Abstract
Description
BACKGROUND
[0001] The present disclosure relates to a formatting of floating-point numbers. SUMMARY
[0002] The following is a summary to provide a basic understanding of one or more embodiments of the disclosed subject matter. This summary is not intended to identify key or critical elements or to limit the scope of the individual embodiments or the scope of the claims. Its sole purpose is to present concepts in simplified form as an introduction to the more detailed description that will follow. In one or more embodiments described herein, systems, units, structures, computer-implemented methods, devices, and / or computer program products are provided that enable the formation of electronic units having spirally conductive structures.
[0003] According to one embodiment, a system may include a memory that stores computer-executable components; and a processor that is functionally connected to the memory and executes computer-executable components. The computer-executable components may include a computational component that enables the processor to perform operations on binary floating-point numbers and calculate them according to a defined floating-point number format in conjunction with the execution of an application, wherein the defined floating-point number format uses six bits in an exponent array.
[0004] Another embodiment relates to a computer-implemented method that can include generating corresponding numeric fields in a defined floating-point number format by a system functionally connected to a processor, wherein the corresponding numeric fields comprise a sign field, an exponent field, and a mantissa field, the defined floating-point number format using six bits in the exponent field. The method can further include calculating binary floating-point numbers according to the defined floating-point number format by the system in conjunction with the execution of an application.
[0005] Another embodiment relates to a computer program product that enables the calculation of floating-point numbers, wherein the computer program product comprises a computer-readable storage medium containing program instructions. The program instructions can be executed by a processor to cause the processor to generate corresponding fields in a defined floating-point number format, wherein the corresponding fields comprise a sign field, an exponent field, and a fraction field, the defined floating-point number format containing six bits for the exponent in the exponent field. The program instructions can further be executed by the processor to cause the processor to calculate the floating-point numbers according to the defined floating-point number format in conjunction with the use of an application.
[0006] Yet another embodiment relates to a computer-implemented method that can include generating corresponding numeric fields in a defined floating-point number format by a system functionally connected to a processor, wherein the corresponding numeric fields comprise a sign field, an exponent field, and a fraction field. The method can also include calculating binary floating-point numbers by the system according to the defined floating-point number format in conjunction with the execution of an application, wherein the defined floating-point number format uses a binary to represent zero and normal numbers, the binary corresponding to bit values of bits of the exponent field, which consists of zeros.
[0007] These and other features will become apparent from the following detailed description of illustrative embodiments, which should be read in conjunction with the accompanying drawings. List of characters Fig. Figure 1 illustrates in a block diagram an exemplary, non-restrictive system that can be used according to various aspects and embodiments of the disclosed subject matter to perform operations on, generate and / or compute floating-point numbers using an extended floating-point number format. Fig. Figure 2 shows a block diagram of an exemplary extended bit structure of the extended floating-point number format according to various aspects and embodiments of the disclosed subject matter. Fig.Figure 3 presents a diagram of an exemplary number line that can illustrate the positions of denormal numbers and normal numbers along the number line when denormal numbers are used in a floating-point number format according to different aspects and embodiments of the disclosed subject matter. Fig. Figure 4 shows an exemplary diagram of the performance results in relation to speech recognition according to different aspects and embodiments of the disclosed subject matter. Fig. Figure 5 illustrates an exemplary diagram of the performance results in relation to image recognition according to different aspects and embodiments of the disclosed subject matter. Fig. Figure 6 shows an exemplary diagram of the performance results in relation to Shakespearean text generation according to different aspects and embodiments of the disclosed subject matter. Fig.Figure 7 shows a block diagram of an exemplary, non-restrictive system which, according to various aspects and embodiments of the disclosed subject matter, can use lower-precision computational control components for a first part of calculations to generate or compute floating-point numbers with an extended floating-point number format, and can use higher-precision computational control components for a second part of calculations. Fig. Figure 8 illustrates a flowchart of an exemplary, non-restrictive procedure for performing operations, including calculations, on data with an extended floating-point number format according to various aspects and embodiments of the disclosed subject matter. Fig.Figure 9 shows a flowchart of another exemplary, non-restrictive method for performing operations, including calculations, on data with an extended floating-point number format according to various aspects and embodiments of the disclosed subject matter. Fig. Figure 10 illustrates a block diagram of an exemplary, non-restrictive operating environment in which one or more of the embodiments described herein may be enabled. DETAILED DESCRIPTION
[0008] The following detailed description serves only for illustration and is not intended to limit embodiments and / or applications or uses of embodiments. Furthermore, express or implied information contained in the preceding sections "Background" or "Summary" or in the section "Detailed Description" is not to be construed as binding.
[0009] The following describes one or more embodiments with reference to the drawings, in which the same reference numerals are used throughout to denote identical elements. Numerous specific details are provided in the following description to facilitate a better understanding of the one or more embodiments. However, it is evident in several cases that the one or more embodiments can be implemented without these specific details.
[0010] Certain types of applications, such as deep learning applications, can have resource-intensive workloads (e.g., computationally demanding workloads). For example, training deep networks can be time-consuming (e.g., several days or weeks) even on systems with multiple graphics processing units (GPUs). It can take a very long time (e.g., weeks) for training algorithms to converge for some deep learning reference values on systems with multiple GPUs.
[0011] To significantly accelerate these types of applications, specialized accelerators can be useful. Such specialized accelerators offer a relatively high throughput density for floating-point calculations, both in terms of area (e.g., throughput per millimeter). 2 (mm 2)) as well as the performance (throughput / watts) can be quite useful for future deep learning systems.
[0012] Using double-precision (e.g., 64-bit) and single-precision (e.g., 32-bit) representations for cognitive computing workloads can be unnecessary and inefficient. One way to improve metrics for both space and power consumption with respect to high data processing loads such as cognitive computing workloads is to use smaller-bit floating-point representations for most calculations. For example, a relatively small subset of calculations, such as those that may be comparatively prone to rounding errors, can be performed using the single-precision format, while most calculations (e.g., those not particularly susceptible to rounding errors) can be performed using a lower-precision format (e.g., 16-bit).Such a division of calculations between the single-precision format and the lower-precision format can make it possible to use a desirable number (e.g., many) of lower-precision control components (e.g., by using a lower-precision format) and a relatively small number of higher-precision control components (e.g., by using a higher (e.g., single) precision format).
[0013] One way to structure a lower-precision format (e.g., 16 bits) is a 1 / 5 / 10 format (e.g., IEEE 754, half-precision format (1 / 5 / 10)), which can have a single sign bit, a 5-bit exponent, and a 10-bit mantissa, and can be used as a data exchange format. However, the 1 / 5 / 10 format can be unsuitable for applications with a significant number of computations, especially for training deep neural networks, because this format can have a relatively limited dynamic range. If the more critical computations (e.g., computations prone to rounding errors) are performed using the single-precision format, and most computations are performed using the 1 / 5 / 10 format, the quality of the trained network will be significantly lower compared to the baseline performance of a number of comparable applications.For example, if a 1 / 5 / 10 format was used for most calculations during training, the accuracy in the case of Watson Natural Language training fell below a reference value of 60% in terms of suitability (this means, for example, that the accuracy was not suitable), and in the case of AlexNet Image Classification training, the calculations did not converge at all.
[0014] For improved throughput density, it may be desirable to use lower-precision floating-point arithmetic (e.g., 16-bit). However, for a lower-precision floating-point format to be used (e.g., in applications with a significant number of calculations, such as deep neural network training and other deep learning training), it may be desirable that the lower-precision floating-point format allows for sufficiently fast convergence of programs and algorithms associated with computationally intensive applications and ensures sufficiently low error rates for computationally intensive applications such as machine training and deep learning training tasks.
[0015] The various embodiments described herein relate to performing operations on, generating, and calculating floating-point numbers (e.g., binary floating-point numbers) using an extended floating-point number format (e.g., a lower-precision extended floating-point number format). The extended floating-point number format can be sufficient for a wide range of machine training and deep learning tasks and enables an area- and performance-efficient implementation of a 16-bit fused multiply-add floating-point unit (a 16-bit fused multiply-add FPU). The extended floating-point number format (e.g., the extended lower-precision floating-point number format (e.g., 16-bit)) can have a 1 / 6 / 9 format, which can have a single sign bit, a 6-bit exponent, and a 9-bit mantissa, which can be used as both an arithmetic calculation format and a data exchange format.Compared to a 1 / 5 / 10 format, the extended floating-point format can have one more exponent bit, and the fractional part can be one bit less. The six bits for the exponent, combined with the additional (e.g., 6th) exponent bit of the extended floating-point format, can provide an additional exponent range, which may be desirable for machine training and deep learning algorithms to converge sufficiently quickly. The extended floating-point format can also exhibit desirablely low error rates for computationally intensive applications such as machine training, deep learning, and related tasks.
[0016] The extended floating-point format can also use a fixed definition for the lowest binad (e.g., data points whose exponent field consists entirely of zeros). Other floating-point formats (e.g., lower-precision floating-point formats) such as the 1 / 5 / 10 format can use the lowest binad for subnormal numbers and zero. According to the fixed definition of the lowest binad in the extended floating-point format, however, the lowest binad can be used for zero and normal numbers. Because the exponent field of the extended floating-point format is one bit wider than that of the 1 / 5 / 10 format, the extended floating-point format (e.g., the 1 / 6 / 9 format) can be closer to zero in the lowest binad, even for normal numbers, than the 1 / 5 / 10 format can be for subnormal numbers. Hardware support for subnormal numbers can be relatively expensive, and the depth (e.g.,The total gate delay of a typical fused multiply-added floating-point unit (FPU) is increased by using the specified definition for the lowest binad. This avoids the use of subnormal numbers, which can make the disclosed subject more efficient in terms of hardware utilization and support, and it avoids increasing the depth (e.g., the total gate delay) of a 16-bit fused multiply-added FPU compared to other formats such as the 1 / 5 / 10 format.
[0017] The extended floating-point number format can furthermore use a fixed definition for the highest binad (e.g., data points whose exponent field consists entirely of ones). Certain other floating-point number formats use the highest binad only for special values such as not-a-number (NaN) and infinity. In contrast, in some embodiments, according to the fixed definition of the highest binad in the extended floating-point number format, a desired portion (e.g., the largest part) of the highest binad can be used for finite numbers, and a data point of the highest binad can be used for special values such as NaN and infinity (e.g., merged NaN / infinity). By using this fixed definition of the highest binad, the extended floating-point number format can accommodate a comparatively larger data range and, compared to other floating-point number formats (e.g.,Lower-precision floating-point number formats, such as the 1 / 5 / 10 format, use comparatively less complex hardware because the logic used for the extended floating-point number format does not have to distinguish between NaN and infinity values, among other things.
[0018] In some implementations, the extended floating-point number format can define the sign of zero as a "don't-care" term, which can make handling the sign of zero less complex compared to other floating-point number formats such as the 1 / 5 / 10 format. Other formats, such as the 1 / 5 / 10 format, can have comparatively complicated rules for how to obtain the sign of zero. However, for many applications, such as deep learning applications, the sign of zero is irrelevant. The extended floating-point number format can accommodate that the sign of zero is irrelevant in such applications, where the sign of zero can be a "don't-care" term. For example, if a floating-point number is zero, any value can be generated for the sign field (e.g.,from a processor component or computer component) to represent the term or symbol for the sign of a value from zero. Any values or symbols that are practical or efficient for the system (e.g., the processor component or computer component) to generate can be created and inserted into the sign field. This allows the hardware and logic used in connection with the extended floating-point number format to be comparatively less complex, and it can use less hardware and hardware resources (and fewer logic resources) than in the case of hardware used for other floating-point number formats (e.g., the 1 / 5 / 10 format).
[0019] In certain implementations, the extended floating-point number format can define the sign of the merged NaN / infinity symbol as a "don't care" term, which can make handling the sign of the merged NaN / infinity symbol comparatively less complex compared to how the sign of NaN and infinity is handled in other formats (e.g., the 1 / 5 / 10 format). Certain other formats, for example, may have comparatively complicated rules for how to obtain the sign of NaN and infinity, especially for addition / subtraction and multiplication / addition operations. However, for many applications, including deep learning applications, the sign of NaN and infinity is irrelevant. In such cases, the sign of the merged NaN / infinity symbol in the extended floating-point number format can be a "don't care" term.If a floating-point number is, for example, NaN or infinity, any value can be generated for the sign field (e.g., by a processor or computer component) to represent the term or symbol for the sign of the merged NaN / infinity symbol. Any values or symbols that are practical or efficient for the system (e.g., the processor or computer component) to generate can be created and inserted into the sign field associated with the merged NaN / infinity symbol. This allows the hardware and logic used in connection with the extended floating-point number format to be comparatively less complex, requiring less hardware and hardware resources (and fewer logic resources) than the hardware used for other formats (e.g., the 1 / 5 / 10 format).
[0020] The extended floating-point number format can also use a rounding mode that may be less complex to implement than the rounding modes used by other types of formats, such as the 1 / 5 / 10 format. For example, the extended floating-point number format can use a single rounding mode that rounds up to the nearest decimal place. Some other types of formats (e.g., the 1 / 5 / 10 format) may use more complex rounding modes, such as rounding up to the nearest decimal place, rounding down to the nearest decimal place, rounding to even, rounding to 0, rounding to +infinity, and rounding to -infinity.The use of a single and comparatively less complex rounding mode, such as the "round up to the nearest decimal place" rounding mode by the extended floating-point number format, may have no or at least virtually no impact on workload performance, including the speed of performance and the quality of training systems (e.g., machine training or training for deep learning), while at the same time significantly reducing the amount of hardware and logic used for rounding floating-point numbers.
[0021] These and other aspects and embodiments of the disclosed subject matter are now described below with reference to the drawings.
[0022] Fig. Figure 1 illustrates an exemplary, non-restrictive system in a block diagram. 100, which, according to various aspects and embodiments of the disclosed subject matter, can be used to perform operations on, generate, and / or calculate floating-point numbers using an extended floating-point number format. The extended floating-point number format of the system 100 It can be sufficient for a wide range of machine learning or deep learning training tasks and enables an area- and power-efficient implementation of a 16-bit fused multiply-add floating-point unit (a 16-bit fused multiply-add FPU). Depending on various embodiments, the system can 100 have, be part of, or belong to one or more 16-bit fused multiply-add FPUs.
[0023] The extended floating-point number format (e.g., the lower-precision extended floating-point number format (e.g., 16-bit)) can have a 1 / 6 / 9 format, which can have a single sign bit, a 6-bit exponent, and a 9-bit mantissa, and can be used as both an arithmetic calculation format and a data exchange format. By using six bits for the exponent (e.g., as opposed to five bits), the extended floating-point number format can provide an additional exponent range, which can be desirable for machine learning algorithms, deep learning training algorithms, and / or other computationally intensive algorithms to converge sufficiently quickly. The extended floating-point number format can also exhibit desirablely low error rates for computationally intensive applications (e.g., machine training or deep learning applications).Other aspects and implementations of the extended floating-point number format are described in more detail herein.
[0024] The system 100 can a processor component 102 exhibit which a data storage 104 and a computer component 106 may be part of the processor component. 102 can be used together with the other components (e.g. the data storage) 104 , the computer component 106 etc.) work so that the various functions of the system can 100 can be carried out. The processor component 102It may use one or more processors, microprocessors, or control units that can process data such as information related to the extended floating-point number format, perform operations on floating-point numbers, generate and / or calculate them, as well as applications (e.g., machine learning, deep learning, or cognitive computing applications), machine or system training, parameters related to number formatting or calculations, data traffic flows (e.g., between components or units and / or across one or more networks), algorithms (e.g., application-specific algorithms, algorithms for extended floating-point number formatting, algorithms for calculating floating-point numbers, algorithms for rounding numbers, etc.), protocols, policies, interfaces, tools, and / or other information to govern the operation of the system. 100as explained in more detail herein, to enable and facilitate the data flow between components of the system 100 , the data flow between the system 100 and to control other components or units (e.g., computers, computer network units, data sources, applications, etc.) that are part of the system 100 are related. In some embodiments, the processor component 102 have one or more FPUs, e.g. one or more 16-bit fused multiply-add FPUs.
[0025] The data storage 104It can store data structures (e.g., user data, metadata), code structure(s) (e.g., modules, objects, hash values, classes, procedures) or instructions, information related to the extended floating-point number format, performing operations on, generating and / or calculating floating-point numbers, applications (e.g., machine learning, deep learning, or cognitive computing applications), machine or system training, parameters related to number formatting or calculations, data traffic flows, algorithms (e.g., application-specific algorithms, algorithms for extended floating-point formatting, algorithms for calculating floating-point numbers, algorithms for rounding numbers, etc.), protocols, policies, interfaces, tools, and / or other information to enable operations related to the system. 100 can be controlled. In one aspect, the processor component can be affected. 102 functionally compatible with the data storage104 be connected (e.g., via a memory bus or another bus) to store and retrieve information that is desired to operate the computer component 106 and / or other components of the system 100 and / or essentially every other functional aspect of the system 100 to operate and / or to transfer at least some of its functionality.
[0026] The processor component 102 can also be a program memory component 108 (e.g., via one or more buses) be associated (e.g., for the purpose of data transmission and / or be functionally connected), whereby the program memory component 108 the computer component 106 can exhibit (e.g., store) the program memory component. 108 can store machine-executable (e.g., computer-executable) components or instructions to which the processor component can respond. 102to be executed by the processor component 102 can access the processor component 102 can, for example, refer to the computer component 106 in the program memory component 108 access and the computer component 106 in conjunction with the processor component 102 to use in order to perform operations on data (e.g., binary floating-point numbers) according to the extended floating-point number format, as described in more detail herein.
[0027] The processor component 102 and the computer component 106They can, for example, work together (e.g., in conjunction with each other) to perform operations on data, including performing operations (e.g., carrying out mathematical calculations), generating and / or calculating floating-point numbers (e.g., binary floating-point numbers) according to (e.g., using) the extended floating-point number format (e.g., the 1 / 6 / 9 floating-point number format) of the component. 110 for an extended format. In some embodiments, the processor component 102 and / or the computer component 106 with the component's extended floating-point number format 110 For an extended format, perform operations on floating-point numbers, generate and / or calculate them to enable machine training, deep learning training, cognitive computing, and / or other computationally intensive tasks, operations, or applications. The processor component 102 and / or the computer component 106For example, the extended floating-point number format can be used to perform operations on floating-point numbers, generate and / or calculate them, in order to obtain results that relate to solving a problem (e.g., a cognitive computing problem) that is to be solved in conjunction with machine training or deep learning training. A problem to be solved might relate to machine training, deep learning, cognitive computing, artificial intelligence, neural networks, and / or other computationally intensive tasks, problems, or applications. To give a few non-limiting examples, the problem might relate to image recognition to detect or identify objects (e.g., people, places, and / or things) in an image or video; or to speech recognition to detect or identify words and / or voice identities in audio content (e.g.,Speech, a broadcast with sound, a song and / or a program with sound), or text recognition to detect or identify text data (e.g. words, alphanumeric characters) in text content (e.g. book, manuscript, email or a series of documents, etc.).
[0028] The processor component 102 and the computer component 106 can belong to one or more applications (e.g., applications for machine training or deep learning), e.g., the application 112 , to perform operations on data in connection with the application(s) 112 to execute. The processor component 102 and the computer component 106 For example, data from an application can 112receive, generate or determine (e.g. calculate) numerical results (e.g. binary floating-point numbers) that are at least partially based on the data, in accordance with the extended floating-point number format, and can process the application's numerical results. 112 and / or provide to another desired destination.
[0029] With reference to briefly Fig. 2 (together with Fig. 1) shows Fig. 2 a block diagram of an exemplary, extended bit structure 200 of the extended floating-point number format according to various aspects and embodiments of the disclosed subject matter. The extended bit structure 200 can a sign field 202 exhibiting a data bit that can contain a data bit that can represent the sign of the floating-point number (e.g., the binary floating-point number).
[0030] The extended bit structure 200 Furthermore, an exponential field can 204contained, which is located next to the sign field 202 can be located. The exponential field 204 It can have six data bits that can represent the exponent of the floating-point number. The extended floating-point number format, by using the additional (e.g., 6th) exponent bit (e.g., compared to the 1 / 5 / 10 format), can provide an extended (e.g., better or additional) exponent range, which may be desirable for machine training or deep learning algorithms (e.g., single-precision (e.g., 32-bit) machine training or deep learning algorithms running in half-precision (e.g., 16-bit) mode) to converge quickly enough while achieving comparatively small error rates for computationally intensive applications, such as machine training and deep learning applications and related machine training and deep learning tasks.
[0031] The extended bit structure 200 can also be a fractional field (f) 206 contained, which is located next to the exponential field 204 can be located. The fractional field 206It can have nine data bits that can represent the fraction (e.g., the fractional value) of the floating-point number. The nine fractional bits for the extended floating-point number format are sufficient to provide the desired fractional precision for machine training algorithms, deep learning algorithms, and / or other computationally intensive algorithms (e.g., single-precision (e.g., 32-bit) machine training, deep learning, or other computationally intensive single-precision (e.g., 32-bit) machine training, deep learning, or other computationally intensive algorithms running in half-precision (e.g., 16-bit) mode), even though the nine fractional bits are one less than in the 1 / 5 / 10 format. It is further noted that if there were only eight fractional bits (e.g., eight fractional bits with seven exponents), the eight fractional bits would not provide sufficient fractional precision for use with machine training or deep learning algorithms (e.g.,Algorithms for machine training or for single-precision deep learning (running in half-precision mode) would provide without at least some modifications to such algorithms, if such modifications to the algorithms were even possible.
[0032] The value of a number x=(s,e,f) can be defined as follows W e r t ( x ) = ( − 1 ) s ∗ 2 [ e ] ∗ ( i . ) , where [e] can be the value of the biased exponent and (if) can be the binary mantissa or the signifier, where f is the fractional part and i is the implied bit. According to the extended floating-point format, the value of i can be 1 for all finite non-zero numbers (and can be 0 for a true zero). With the extended floating-point format using a biased exponent, it is possible for the exponent to be represented as an unsigned binary integer by the shift.
[0033] Table 1 illustrates various differences between the extended floating-point number format (also referred to herein as the extended 1 / 6 / 9 format) and the 1 / 5 / 10 format. TABLE 1 1 / 5 / 10 format Extended 1 / 6 / 9 format exponential field 5 bits 6 bits Exponential shift 15 31 Smallest exponent -14 -31 Largest exponent +15 +32 Smallest positive number 2 -24 2 -31 * (1 + 2 -9 ) Largest positive number 2 15 (2 - 2 -10 ) 2 32 * (2 - 2 -9 )
[0034] As can be seen from Table 1, the smallest representable positive number is 2 -31 * (1 + 2 -9 ), if the component uses the extended floating-point number format 110 is used for an extended format, which is a significantly smaller number than the smallest positive subnormal number (e.g., 2). -24 ), which can be achieved with the 1 / 5 / 10 format. Therefore, the processor component 102 and the computer component 106Even without using subnormal numbers in the extended floating-point number format, the extended floating-point number format can be used to generate or calculate numbers that are significantly closer to 0 than when using a 1 / 5 / 10 format. This can improve the convergence of machine learning training runs (e.g., deep learning).
[0035] With respect to floating-point numbers, there can be a set of binads, where a binad can be a set of binary floating-point values, each of which can have the same exponent. The extended floating-point number format can have 64 binads. The binads of the extended floating-point number format can contain a first binad and a last binad. For the first (e.g., lowest) binad, the exponent field can be... 204consist exclusively of zeros (e.g., each of the 6 bits is a 0), and for the last binary, the exponent field can be used. 204 consist exclusively of ones (e.g., each of the 6 bits is a 1). The component's extended floating-point number format. 110 For an extended format, it may use fixed definitions for the data points in the first binad and the last binad, whereby such fixed definitions for the data points in the first binad and the last binad may differ from definitions used with respect to data points for the first and last binad of other formats such as the 1 / 5 / 10 format.
[0036] The other binads, which can have non-extreme exponents (e.g., exponents that do not consist exclusively of ones or zeros), can be used to represent normal numbers. The implied bit can be one (i=1), which can lead to a mantissa of 1.f. The exponent value of the exponent array 204 can be derived as [e] = e - shift. As already disclosed, the extended floating-point number format can differ from the 1 / 5 / 10 format in part in that the exponent is in the exponent field. 204 is one bit wider than in the 1 / 5 / 10 format, the fraction in the fraction field 206 is one bit shorter than in the 1 / 5 / 10 format, and the shift for the extended floating-point number format may also differ from the 1 / 5 / 10 format.
[0037] The use of subnormal (e.g., denormal) numbers in a floating-point number format, especially a half-precision floating-point number format (e.g., 16-bit), can be inefficient and / or relatively expensive, as explained in more detail here. With reference to briefly Fig. 3 (together with the Fig. 1 and Fig. 2) represents Fig. 3. A diagram of an example number line 300 This can illustrate the positions of denormal and normal numbers along the number line when denormal numbers are used in a floating-point number format according to various aspects and embodiments of the disclosed subject matter. As in the number line 300 As represented, in a floating-point number format such as a 32-bit floating-point number format (e.g. 1 / 8 / 23), a first subset of the numbers can be displayed. 302, which may be denormal numbers (denormous) (also referred to herein as subnormal numbers), from greater than 0 to 2 -126 suffice. In the 1 / 8 / 23 format, there can be a single bit for the sign field, 8 bits for the exponent field, and 23 bits for the fraction field. The first set of numbers 302 It can be in a first binad. Just like in the number line. 300 As represented, a second set of numbers can be shown. 304 , which can be normal numbers, consist of numbers greater than 2 -126 are.
[0038] As revealed herein, the smallest representable positive number can be 2 -31 * (1 + 2 -9 ) (e.g. (2 -31 * (1 + 2 -9 )) → 0×0001) with the extended floating-point number format, which is a significantly smaller number than the smallest positive subnormal number (e.g. 2 -24), which can be achieved with the 1 / 5 / 10 format. By using 6 bits in the exponent field, the extended floating-point number format can represent a range of numerical values that can be used when running an application. 112 (e.g., in a machine learning or deep learning application) can be expected or predicted without using subnormal number modes. According to various implementations, the extended floating-point number format can therefore eliminate the need for subnormal (e.g., denormal) numbers. Compared to the 1 / 5 / 10 format, the extended floating-point number format can actually represent a wider range of numbers while exhibiting lower logical complexity.
[0039] With regard to the lowest (e.g., first) binary, the extended floating-point number format can use a fixed definition for lowest binary data points (e.g., data points where the exponent field consists entirely of zeros). According to the fixed definition of the lowest binary in the extended floating-point number format, the processor components 102 and / or the computer component 106 Use the lowest binadi for zero and normal numbers, while discarding subnormal numbers. For a zero fraction, the value can be zero. For a non-zero fraction, the value can be a normal non-zero number. If the smallest represented positive number (2 -31 * (1 + 2 -9 )) with regard to the extended floating-point number format, the processor component 102 and / or the computer component 106According to the extended floating-point number format, rounding of numbers smaller than (2 -31 *(1 + 2 -9 )) are, either on (2 -31 * (1 + 2 -9 )) or perform 0.
[0040] According to the extended floating-point number format, the sign can also be ignored, which can make hardware implementations less complex and more efficient than other systems or units that use other floating-point number formats (e.g., the 1 / 5 / 10 format). If the processor component 102 and / or the computer component 106 For example, when generating or calculating a zero result, the value of the sign bit can be undefined or implementation-specific. Zero can be, for example, 0x0000 or 0x8000, where the sign bit can be represented as a "don't care" term. If the processor component 102 and / or the computer component 106When performing operations on a non-zero fraction, generating or calculating it, the value can be a normal number, where the implied bit can be, for example, one (i=1), which can lead to a mantissa of 1.f, and where the exponent value [e] can be e - shift = 0 - shift.
[0041] As previously stated, the extended floating-point number format can define the sign of zero as a "don't care" term, which can make handling the sign of zero less complex compared to other formats such as the 1 / 5 / 10 format. For example, certain other floating-point number formats, such as the 1 / 5 / 10 format, can have comparatively complicated rules for how to obtain the sign of zero. However, for many applications, including deep learning applications, the sign of zero is irrelevant. The processor component 102 and / or the computer component 106Thus, according to the extended floating-point number format, the sign of zero can be represented as a "don't-care" term. The "don't-care" term can, for example, represent a defined value that, according to the extended floating-point number format, indicates that the sign of zero is irrelevant to the value of zero. If the floating-point number is zero, the processor components 102 and / or the computer component 106 generate any value or symbol in the sign field of the extended floating-point number format, which is used by the processor component. 102 or the computer component 106 to generate practically or efficiently (e.g., most practically or efficiently) and insert it into the sign field. The processor component 102 and / or the computer component 106For example, it can generate any value or enable the generation of any value that can utilize the smallest amount of resources (e.g., processing or data processing resources), or at least fewer resources than usual, to determine and generate a non-arbitrary value for the sign field that represents the sign of a value of zero. By using the extended floating-point number format, including the handling of the sign associated with a value of zero, the disclosed subject matter can enable the hardware used by the system to 100 When used in conjunction with the extended floating-point number format, it is comparatively less complex and requires less hardware and hardware resources (and fewer logic resources) than in the case of hardware used for other floating-point number formats (e.g., the 1 / 5 / 10 format).
[0042] Unlike the extended floating-point format, certain other floating-point formats, such as the 1 / 5 / 10 format, use the lowest binadi for subnormal (e.g., denormal) numbers and zero. Because the exponent field of the extended floating-point format, as revealed here, is one bit wider than in the 1 / 5 / 10 format, the extended floating-point format can be closer to zero in the lowest binadi, even for normal numbers, than the 1 / 5 / 10 format can be for subnormal numbers.
[0043] Hardware support for subnormal numbers can be comparatively expensive and increase the depth (e.g., the total gate delay) of a typical fused multiply-added FPU, and further support for subnormal numbers can increase the complexity of the logic to an undesirable degree. By using the specified definition for the lowest binadi, the disclosed subject can avoid the use of subnormal numbers, which can make the disclosed subject more efficient with respect to hardware utilization and support (e.g., the scope of hardware and hardware resources used can be reduced), and it can avoid increasing the depth (e.g., the total gate delay) of a 16-bit fused multiply-added FPU compared to other formats such as the 1 / 5 / 10 format.
[0044] With regard to the highest (e.g., last) binary, the extended floating-point number format can use a fixed definition for the highest binary (e.g., data points where the exponent field consists exclusively of ones). According to the fixed definition of the highest binary in the extended floating-point number format, the processor component 102 and / or the computer component 106 Use a desired part (e.g., the largest part) of the highest binary for finite numbers and use a data point of the highest binary for special values such as a non-number (NaN) and infinity (e.g., concatenated NaN / infinity).
[0045] Table 2 illustrates the use of most of the data (e.g., all but one data point) of the highest binary for finite numbers and one data point of the highest binary for the special value for combined NaN / infinity according to some embodiments of the disclosed subject matter as follows: TABLE 2 exponent mantissa 111111 000...000 to 111...110 Normal numbers with mantissa 1.f and [e] = 32 111111 111...111 Combined NaN / infinity (sign bit can be a "don't care" term)
[0046] It is understood and should be noted that, according to various other embodiments, more mantissa codes are reserved for different variants of a NaN and are accordingly assigned by the processor component. 102 and / or the computer component 106 can be used. If and to the extent desired, for example, two or more mantissa codes can be used for corresponding (e.g., different) variants of a NaN according to the extended floating-point number format (e.g., by the processor component). 102 and / or the computer component 106If and to the extent desired, as a further exemplary embodiment, the data points with the exponent 111111 and with a mantissa in the range of 000...000 to 111...101 can be normal numbers, with a mantissa of 1.f and [e] = 32; for a mantissa of 111...110, the data point can represent + or -infinity, depending on the value of the sign bit in the sign field. 202 ; and for a mantissa consisting exclusively of 1 values, the data point can represent NaN, where the sign bit is in the sign field. 202 It can be a "don't care" term.
[0047] The extended floating-point number format can further define the sign of the merged NaN / infinity symbol as a "don't care" term, which can make handling the sign of the merged NaN / infinity symbol comparatively less complex compared to handling the sign of NaN and infinity in other floating-point number formats (e.g., the 1 / 5 / 10 format). Certain other floating-point number formats, for example, may have comparatively complicated rules for how to preserve the sign of NaN and infinity, especially for addition / subtraction and multiplication / addition operations. However, for many applications, including deep learning applications, the sign of NaN and infinity is irrelevant.
[0048] The processor component 102 and / or the computer component 106Thus, according to the extended floating-point number format, the sign of the merged NaN / infinity symbol can be represented as a "don't care" term. The "don't care" term can, for example, represent a defined value that, according to the extended floating-point number format, indicates that the sign of the merged NaN / infinity symbol is irrelevant to the value of the merged NaN / infinity symbol. If infinity or NaN is the result of a floating-point number, the processor component can 102 and / or the computer component 106The processor component can generate or enable the generation of any value in the sign field associated with the value(s) for infinity and / or NaN (e.g., a merged NaN / infinity symbol, an individual NaN value, or an individual infinity value) to represent the term or symbol for the sign of the value(s) for infinity and / or NaN. According to the extended floating-point number format, the processor component can 102 and / or the computer component 106 generate or enable the generation of any value or symbol in the sign field corresponding to the value(s) for infinity and / or NaN, which is required by the processor component 102 or the computer component 106 to generate practically or efficiently (e.g., most practically or efficiently) and insert it into such a sign field. The processor component 102 and / or the computer component106 For example, it can generate or enable the generation of any value that uses the least amount of resources (e.g., processing or data processing resources), or at least fewer resources than usual, to determine and generate a non-arbitrary value for the sign field, representing the sign of the value(s) for infinity and / or NaN. By using the extended floating-point number format, including the handling of infinity, NaN, and the sign associated with infinity and NaN, the disclosed subject matter can enable the hardware used by the system to 100 When used in conjunction with the extended floating-point number format, it is comparatively less complex and requires less hardware and hardware resources (and fewer logic resources) than in the case of hardware used for other floating-point number formats (e.g., the 1 / 5 / 10 format).
[0049] Furthermore, certain other floating-point number formats, such as the 1 / 5 / 10 format, unlike the fixed definition of the highest binad in the extended floating-point number format, use the highest binad exclusively for the special values of NaN and infinity, where for a zero fraction, the value can be + / -infinity depending on the sign bit, and for a non-zero fraction, the data point can represent NaN. These special values can be useful in some ways for handling corner conditions. Infinity can be used to indicate that the calculation has exceeded the valid data range. In arithmetic operations, infinity can obey the rules of algebra, such as infinity - x = infinity for any finite number x.In certain other floating-point formats, NaN can be used to represent mathematically undefined scenarios, such as the square root of a negative number or the difference between two infinite numbers with the same sign. It should be noted that in floating-point number formats such as the 1 / 5 / 10 format, there can be multiple representations of NaN. If an arithmetic operation has at least one NaN operand, the result must be one of the NaN operands; for example, its payload must be passed on.
[0050] While using a binade exclusively for these special values of NaN and infinity may not cause any significant impairment with 8 or more bits for the exponent, such as with higher-precision floating-point number formats (e.g., 32-bit or 64-bit floating-point formats), using a binade exclusively for these special values of NaN and infinity may be unacceptably expensive with smaller bit widths, such as those associated with lower-precision floating-point number formats (e.g., 16-bit floating-point formats such as the extended 1 / 6 / 9 format and the 1 / 5 / 10 format).
[0051] The extended floating-point number format, by using such a fixed definition for the highest binary described herein, can enable a comparatively larger data range and, compared to other floating-point number formats such as the 1 / 5 / 10 format, utilize comparatively less complex hardware, since the system requires less hardware for the extended floating-point number format. 100 (e.g., through the processor component) 102 and / or computer component 106 The logic used does not need to distinguish between NaN and infinity values, among other things.
[0052] The extended floating-point number format can, in some implementations, use a rounding mode that may be less complex to implement than the rounding modes used by other types of formats. For example, the extended floating-point number format can use a rounding mode to round up to the nearest decimal place without using any other rounding modes. Some other types of formats (e.g., the 1 / 5 / 10 format) use more complex rounding modes such as rounding up to the nearest decimal place, rounding down to the nearest decimal place, and rounding to even, as well as rounding to 0, rounding to +infinity, and rounding to -infinity. Using the less complex rounding mode (e.g.,Rounding up to the nearest digit) without using other rounding modes through the extended floating-point number format may have no or at least virtually no impact on workload performance, including the speed of performance and the quality of training systems (e.g., machine training or training for deep learning), while at the same time significantly reducing the amount of hardware and logic used for rounding floating-point numbers.
[0053] By using the extended floating-point format, the disclosed subject matter can reduce the size of the hardware, hardware resources, logic resources, and other computational resources used, for example, in computationally intensive applications (e.g., machine training or deep learning applications), as described in detail herein. In some embodiments, the disclosed subject matter can also use a 16-bit FPU, which may be about 25 times smaller than certain double-precision FPUs (e.g., 64-bit). Furthermore, it is expected (e.g., estimated or predicted) that the disclosed subject matter, by using the extended floating-point format, will use a 16-bit FPU, which may be about 25% to 30% smaller than a comparable half-precision FPU (e.g., a 16-bit FPU).Half-precision FPU in 1 / 5 / 10 format).
[0054] The Fig. 4, Fig. 5 and Fig. 6. Exemplary performance results can be used to illustrate the performance of the extended floating-point number format (e.g., extended half-precision floating-point number format (e.g., 16-bit)) relative to the performance results of a half-precision floating-point number format (e.g., 16-bit) (e.g., 1 / 5 / 10 format) and the performance results of a 32-bit floating-point number format (e.g., full-precision). With reference to briefly Fig. 4 shows Fig. 4. An example diagram 400 The performance results regarding speech recognition according to various aspects and embodiments of the disclosed subject matter. The diagram 400The diagram illustrates the respective speech recognition results for the extended floating-point format, the half-precision floating-point format, and the 32-bit floating-point format for a speech-deep neural network (speech DNN) on a 50-hour message dataset. 400 represents the phoneme error rate (%) on the y-axis in relation to the training period along the x-axis.
[0055] As shown in the diagram 400 The performance results can be seen, the performance results 402 The speech recognition for the extended floating-point number format (extended format) (and the associated 16-bit hardware and software) converge quite well and provide comparatively favorable results that match the performance results. 404The speech recognition of the 32-bit floating-point number format (full-precision format) (and the associated 32-bit hardware and software) comes relatively close. As shown in the diagram. 400 Furthermore, the performance results are evident. 406 The speech recognition for the half-precision floating-point number format (half-precision format) (and the associated 16-bit hardware and software) is unable to converge at all and therefore cannot provide meaningful speech recognition results.
[0056] Fig. Figure 5 illustrates an example diagram. 500 The performance results with regard to image recognition according to various aspects and embodiments of the disclosed subject matter. The diagram 500The diagram illustrates the respective results of image recognition for the extended floating-point format, the half-precision floating-point format, and the 32-bit floating-point format with respect to AlexNet on a 2012 ImageNet dataset with a thousand output classes. 500 represents the test error (%) on the y-axis in relation to the training period along the x-axis.
[0057] As shown in the diagram 500 The performance results can be seen, the performance results 502 The image recognition for the extended floating-point number format (and the associated 16-bit hardware and software) converge quite well and provide results that compare favorably with the performance results. 504 Image recognition of the 32-bit floating-point number format (and the associated 32-bit hardware and software) can be quite advantageous (e.g., results that can come comparatively close to these). As shown in the diagram 500As can be seen, the performance results are 506 The image recognition for the half-precision floating-point number format (and the associated 16-bit hardware and software) is unable to converge at all and therefore cannot provide meaningful image recognition results.
[0058] With reference to Fig. 6 shows Fig. Figure 6 shows an example diagram. 600 the performance results relating to Shakespearean text production according to various aspects and embodiments of the disclosed subject matter. The diagram 600The diagram illustrates the respective text generation performance results for Shakespearean text using the extended floating-point format, the half-precision floating-point format, and the 32-bit floating-point format, employing a character recurrent neural network (char-RNN) with two stacked long short-term memories (LSTMs) (512 units each). 600 represents the training error on the y-axis in relation to the training epoch along the x-axis.
[0059] As shown in the diagram 600 The performance results can be seen from the performance results 602 Text generation for the extended floating-point number format (and the associated 16-bit hardware and software) converges quite well and provides results that compare favorably with the performance results. 604Text generation using the 32-bit floating-point number format (and the associated 32-bit hardware and software) can be quite advantageous (e.g., results that can come comparatively close to these). As shown in the diagram 600 Furthermore, the performance results can be seen 606 The text generation for the half-precision floating-point number format (and the associated 16-bit hardware and software) do converge, but not nearly as well as the line results. 602 text generation for the extended floating-point number format or the performance results 604 the text generation of the 32-bit floating-point number format.
[0060] With reference to Fig. 7 shows Fig. 7 a block diagram of an exemplary, non-restrictive system 700, which, according to various aspects and embodiments of the disclosed subject matter, can use lower-precision computational control components for a first part of calculations to generate or compute floating-point numbers (e.g., binary floating-point numbers) using an extended floating-point number format, and which can use higher-precision computational control components for a second part of calculations. A repeated description of similar elements used in other embodiments described herein is omitted for the sake of brevity or may be omitted for the sake of brevity.
[0061] The system 700 can a processor component 702 , a data storage device 704 and a computer component 706 exhibit the processor component 702 can the data storage 704 and a program memory component 708(e.g., via one or more buses) belong to (e.g., connected for the purpose of data transmission), whereby the program memory component 708 the computer component 706 can exhibit. The computer component 706 can a component 710 for an extended format that allows the processor component 702 and the computer component 706 As described in more detail herein, it enables operations to be performed on floating-point numbers (e.g. binary floating-point numbers), to generate and / or calculate them in the extended floating-point number format (e.g. mathematical calculations).
[0062] The processor component 702 and the computer component 706 can belong to one or more applications (e.g., applications for machine training or deep learning), e.g., the application 712, to perform operations on data (e.g., numerical data) that are available to the application(s) 712 are related to the processor component. 702 and the computer component 706 For example, data from an application can 712 receive, perform operations on numeric values or results (e.g., binary floating-point numbers), generate or determine (e.g., calculate) them that are at least partially based on the data, in accordance with the extended floating-point number format, and process the numeric values or results of the application. 712 and / or provide it to another desired destination (e.g., another application, component, or unit).
[0063] In some embodiments, the processor component 702a set of lower-precision computation control components such as the lower-precision computation control component 1 714 (LPCE1 714), the lower-precision computation control component 2 716 (LPCE2 716) up to the computation control component M 718 with lower accuracy (LPCE) M 718), where M can be virtually any desired number. The lower-precision computation control components (e.g., 714, 716, 718) of the set of lower-precision computation control components can be or include 16-bit computation control components (e.g., 16-bit FPUs) capable of performing calculations or other operations on data (e.g., numeric data) to generate floating-point numbers (e.g., binary floating-point numbers) in accordance with the extended floating-point number format (e.g., that are compliant with, consistent with, structured according to, and / or representable in the extended floating-point number format).
[0064] The processor component 702 can also include a set of higher-precision computation control components such as the higher-precision computation control component 1 720 (HPCE1 720), the higher-precision computation control component 2 722 (HPCE2 722), up to the computation control component N 724 with higher accuracy (HPCE) N724), where N can be practically any desired number. The higher-precision computation control components of the set of lower-precision computation control components (e.g., 720, 722, 724) can be or include 32-bit (and / or 64-bit) computation control components (e.g., FPUs) capable of performing calculations or other operations on data (e.g., numerical data) to generate floating-point numbers (e.g., binary floating-point numbers) according to a desired (e.g., suitable, acceptable, or optimal) higher-precision floating-point number format (e.g., that can be conformed to, match, structured according to, and / or represented in).
[0065] In connection with the execution or use of an application 712(e.g., an application related to or connected with machine training or deep learning training) there can be a variety of different types of calculations or other operations that can be performed by the processor component. 702 and / or the computer component 706 can be performed. Some types of calculations, for example, may be comparatively more prone to errors such as rounding errors than other types of calculations. In many or even most cases, most calculations required by an application can be performed. 710 These are the types of calculations that are not particularly prone to errors, and a relatively small portion of the calculations may be the types of calculations that can be comparatively more prone to errors, such as rounding errors.
[0066] To enable the efficient execution of operations, including calculations, by the processor component702 , the computer component 706 and other components of the system 700 To enable this, the system can 700 an operations management component 726 exhibit, which the processor component 702 , the data storage 704 and the program memory 708 may be associated with it (e.g., it may be connected for the purpose of data transmission). The operations management component 726 can operations of the system's components 700 control, including the processor component 702 , the data storage 704 , the computer component 706 , the program memory component 708 , the set of lower-precision computational control components (e.g. 714, 716, 718) and the set of higher-precision computational control components (e.g. 720, 722, 724).
[0067] For example, to enable the efficient execution of calculations with data (e.g., numerical data), the operations management component can be used. 726 Assign a first part of operations (e.g., arithmetic operations) and associated data to the set of lower-precision arithmetic control components (e.g., 714, 716, 718), where the set of lower-precision arithmetic control components can perform such operations on such data. The first part of operations can include performing calculations or other operations on data, where the calculations or other operations are not particularly prone to errors such as rounding errors. The operations management component 726Furthermore, a second set of operations (e.g., arithmetic operations) and associated data can be assigned to the set of higher-precision arithmetic control components (e.g., 720, 722, 724), whereby the set of higher-precision arithmetic control components can perform such operations on such data. The second set of operations may include performing calculations or other operations on data, where these calculations or other operations may be the type of calculations that are comparatively more prone to errors such as rounding errors, and therefore the use of the higher-precision arithmetic control components to perform these more prone calculations or other operations may be desirable.
[0068] Since there are usually (e.g., frequently) significantly more calculations that are not particularly error-prone (e.g., the first part of operations) compared to the number of calculations that can be error-prone (e.g., the second part of operations), in some embodiments the number (e.g., M) of computation control components in the set of lower-precision computation control components (e.g., 714, 716, 718) can be greater than the number (e.g., N) of computation control components in the set of higher-precision computation control components (e.g., 720, 722, 724). Depending on requirements (e.g., depending on the type of application or the training tasks to be performed), in other embodiments there can be an equal number of computation control components in the set of lower-precision computation control components (e.g., 714, 716, 718) and in the set of higher-precision computation control components (e.g.,720, 722, 724) or the set of higher-precision computational control components (e.g., 720, 722, 724) can have a higher number of computational control components than the set of lower-precision computational control components (e.g., 714, 716, 718).
[0069] It goes without saying and should be noted that the system 700 According to various embodiments, it may include one or more processor components, computer components, lower-precision computational control components, higher-precision computational control components, graphics processing units (GPUs), accelerators, field-programmable gate arrays (FPGAs) and / or other processing units for performing or enabling operations on data, including performing calculations on data (e.g., numerical data).
[0070] Fig. Figure 8 illustrates a flowchart of an exemplary, non-restrictive procedure.800 for performing operations, including calculations, on data with an extended floating-point number format according to various aspects and embodiments of the disclosed subject matter. The method 800 This can be performed, for example, by the processor component and / or the computer component using the component for an extended format. Repeated descriptions of similar elements used in other embodiments described herein are omitted for brevity or may be omitted for brevity.
[0071] At 802Corresponding numeric fields can be generated in a defined floating-point number format (e.g., an extended floating-point number format), where the corresponding numeric fields can have a sign field, an exponent field, and a mantissa field, and where the defined floating-point number format can use six bits in the exponent field. The processor component and / or the computer components can generate the corresponding numeric fields in the defined floating-point number format or enable their generation.
[0072] At 804Binary floating-point numbers can be calculated according to the defined floating-point number format in connection with the execution of an application. The processor component and / or the computer component can calculate the binary floating-point numbers according to the defined floating-point number format in connection with the execution of an application or otherwise perform operations on them. For example, the processor component and / or the computer component can calculate the binary floating-point numbers according to the defined floating-point number format or otherwise perform operations on them to enable the convergence of a program, an algorithm (e.g., an algorithm associated with the program), and / or data points (e.g., associated with the program or the algorithm) to obtain a result that relates to a solution to a problem to be solved in connection with the execution of an application.The problem to be solved can relate to machine training, deep learning, cognitive computing, artificial intelligence, neural networks, and / or other computationally intensive tasks, problems, or applications. The problem can involve image recognition to detect or identify objects (e.g., people, places, and / or things) in an image; speech recognition to detect or identify words and / or voice identities in audio content; or text recognition to detect or identify text data (e.g., words, alphanumeric characters) in text content (e.g., book, manuscript, email, etc.).
[0073] Fig. Figure 9 shows a flowchart of another exemplary, non-restrictive procedure. 900 for performing operations, including calculations, on data with an extended floating-point number format according to various aspects and embodiments of the disclosed subject matter. The method900 This can be performed, for example, by the processor component and / or the computer component using the component for an extended format. Repeated descriptions of similar elements used in other embodiments described herein are omitted for brevity or may be omitted for brevity.
[0074] At 902Corresponding fields can be generated according to an extended floating-point number format, where the corresponding fields can contain a sign field, an exponent field, and a mantissa field. The processor component and / or the computer component can generate the corresponding fields in the defined floating-point number format or enable their generation. The defined floating-point number format can be a 16-bit format, where the sign field can contain a single bit, the exponent field can contain six bits, and the mantissa field can contain nine bits.
[0075] At 904Operations can be performed on binary floating-point numbers according to the extended floating-point number format. The processor component and / or the compute component can perform operations on binary floating-point numbers according to the extended floating-point number format, generate them, and / or calculate them to, for example, enable a program belonging to an application to converge according to defined convergence criteria. The defined convergence criteria can relate, for example, to a time interval in which the program (e.g., the algorithm and / or data points of the program) must converge to obtain a result for a problem to be solved in connection with the execution of the application, the reachability of the program to converge in order to obtain the result, and / or an error rate associated with the result.
[0076] At 906A first binad can be used to represent zero and normal numbers according to the extended floating-point number format, where the first binad belongs to the exponential field consisting entirely of zeros. The processor component and / or the compute component can use the first binad to represent zero and normal numbers according to the extended floating-point number format. The first binad can, for example, be the lowest binad. The first binad can have a set of data points, where one data point of the set, containing a fraction of all zeros, can represent zero. The other data points of the set of data points can represent normal numbers.
[0077] Normal numbers can be finite non-zero floating-point numbers with a size greater than or equal to a minimum value that can be determined as a function of a base and a minimum exponent in conjunction with the extended floating-point number format. For example, a normal number can be a finite non-zero floating-point number with a size greater than or equal to r min(e) where r can be the base and min(e) the minimum exponent. The extended floating-point number format can be structured so that subnormal numbers are not used, not included, or removed.
[0078] At 908In accordance with the extended floating-point number format, a reduced subset of a second binary can be used to represent an infinity value and a NaN value. The processor component and / or the compute component can use this reduced subset of the second binary to represent the infinity value and the NaN value, in accordance with the extended floating-point number format. The second binary can, for example, be the highest binary, where the second binary can be assigned such that all bits of the exponent array have values of one.
[0079] The extended floating-point number format can be structured such that a reduced (e.g., smaller) subset of data points from a second binad set of data points is used to represent the infinity value and the NaN value. The reduced set of data points can have fewer data points than a set of data points belonging to a whole of the second binad. In some embodiments, the extended floating-point number format can be structured such that a single data point from the second binad set of data points is used to represent the infinity value and the NaN value as a single merged symbol. That is, the reduced subset of data points can have a single data point that can represent the single merged symbol for the infinity value and the NaN value.In other embodiments, the reduced subset of data points may include one data point that can represent the infinity value and another data point that can represent the NaN value. The other data points of the second binad's data point set that are not in the reduced set of data points may be used to represent finite numbers, which may further improve or extend the range of floating-point numbers that can be represented using the extended floating-point number format.
[0080] If a floating-point number is zero at 910, then according to the extended floating-point number format, a sign of zero can be represented as a "don't-care" term. If a floating-point number is zero, then according to the extended floating-point number format, the processor component and / or the computer component can represent the sign of zero as a "don't-care" term. The "don't-care" term can be a term or a symbol that indicates that the sign is irrelevant with respect to the zero value.
[0081] If a floating-point number is zero at 912, any value can be generated in the sign field for a value of zero to represent the term for the sign of zero. The processor component and / or the compute component can generate any value in the sign field to represent the term or symbol for the sign of the value of zero, for example, when zero is the floating-point number. The processor component and / or the compute component can generate any value or symbol in the sign field of the extended floating-point number format that is practical or efficient (e.g., most practical or efficient) for the processor component or the compute component to generate and insert into the sign field. For example, the processor component and / or the compute component can generate any value that requires the least amount of resources (e.g.,processing or data processing resources) or at least fewer resources than usual to determine and generate a non-arbitrary value for the sign field that represents the sign of the value of zero.
[0082] If a floating-point number in 914 is infinity and / or a non-number, the sign of infinity and / or NaN can be represented as a "don't-care" term according to the extended floating-point number format. The processor component and / or the computer component can represent the sign of infinity and / or NaN as a "don't-care" term according to the extended floating-point number format. The "don't-care" term can be a term or a symbol that indicates that the sign with respect to the infinity and / or NaN value is irrelevant. In some embodiments, the infinity value and the NaN value can be represented as a combined symbol. In other embodiments, the infinity value and the NaN value can be represented separately.
[0083] If, at 916, a floating-point number is infinity and / or a NaN, any value can be generated in the sign field associated with the infinity and / or NaN value(s) to represent the sign term(s) of the infinity and / or NaN value(s). The processor component and / or the compute component can generate the arbitrary value in the sign field associated with the infinity and / or NaN value(s) to represent the sign term or symbol of the infinity and / or NaN value(s), for example, if infinity or NaN is the result of the floating-point number. As described herein, in some embodiments, the infinity value and the NaN value can be represented as a single, combined symbol, and in other embodiments, the infinity value and the NaN value can be represented separately.
[0084] According to the extended floating-point number format, the processor component and / or the compute component can generate any value or symbol in the sign field that corresponds to the value(s) for infinity and / or NaN, provided that it is practical or efficient for the processor component or the compute component to generate and insert into such a sign field. For example, the processor component and / or the compute component can generate any value that uses the least amount of resources (e.g., processing or data processing resources), or at least fewer resources than usual, to determine and generate a non-arbitrary value for the sign field that represents the sign of the value(s) for infinity and / or NaN.
[0085] If a floating-point number is to be rounded at 918, a single rounding mode can be used to round binary floating-point numbers according to the extended floating-point number format. When rounding numbers is desired, the processor component and / or the compute component can perform rounding of binary intermediate floating-point numbers (e.g., preliminary calculated numerical values) using the single rounding mode to enable operations on or calculations of binary floating-point numbers according to the defined floating-point number format. In some embodiments, the single rounding mode may be a round-up mode. By using only the round-up mode to perform rounding floating-point numbers, the disclosed subject matter can improve efficiency by reducing the need for the hardware (e.g.,the scope and / or type of hardware or hardware resources) used to run the application and to perform operations on or computation of binary floating-point numbers is reduced.
[0086] For the sake of simplicity, the procedures and / or the computer-implemented procedures are presented and described as a series of steps. It is understood and should be noted that the disclosed subject matter is not limited by the presented steps and / or the sequence of steps; for example, steps may be performed in different sequences and / or simultaneously, and / or with other steps not presented and described herein. Furthermore, not all of the presented steps may be necessary to implement the computer-implemented procedures according to the disclosed subject matter. Moreover, those skilled in the art will understand and note that the computer-implemented procedures can alternatively be represented as a series of interconnected states via a state diagram or events.Furthermore, it should be noted that the computer-implemented methods disclosed below and throughout this description can be stored in a manufactured article to enable the forwarding and transfer of such computer-implemented methods to other computers. The term "manufactured article," as used herein, is intended to encompass a computer program that can be accessed by computer-readable units or storage media.
[0087] In order to create a context for the various aspects of the revealed subject matter, Fig. 10 and the following consideration provide a general description of a suitable environment in which the various aspects of the disclosed subject matter can be implemented. Fig.Section 10 illustrates a block diagram of an exemplary, non-restrictive operating environment in which one or more of the embodiments described herein may be enabled. Repeated descriptions of similar elements used in other embodiments described herein are omitted for brevity or may be omitted for brevity. With reference to Fig. 10 can provide a suitable operating environment 1000 To implement various aspects of this revelation, a computer is also needed. 1012 included. The computer 1012 can also be a processing unit 1014 , a system memory 1016 and a system bus 1018 included. The system bus 1018 connects the system components, including the system memory. 1016 with the processing unit 1014 , without being limited to that. At the processing unit 1014It can be any of the various available processors. Dual microprocessors and other multiprocessor architectures can also be used as a processing unit. 1014 to be used in the system bus 1018 It can be any of several types of bus structure(s), including a memory bus or memory control unit, a peripheral bus or external bus, and / or a local bus, utilizing a variety of available bus architectures, including Industrial Standard Architecture (ISA), Micro-Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), and FireWire (IEEE). 1394 ) and Small Computer Systems Interface (SCSI), without being limited to these. The system memory 1016 can also be a volatile memory 1020and a non-volatile memory 1022 Included is the basic input / output system (BIOS), which contains the fundamental routines for transferring information between components within the computer. 1012 such as during startup, it is stored in non-volatile memory. 1022 stored. For example, and this is not limited to, non-volatile memory. 1022 It may contain a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), flash memory, or non-volatile random-access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). The volatile memory 1020It can also include random access memory (RAM) that acts as an external cache. For example, and this is not limited to, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), synchronous SDRAM with double data rate (DDR SDRAM), extended SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct dynamic Rambus RAM (DRDRAM), and dynamic Rambus RAM.
[0088] The computer system 1012 It may also contain removable / non-removable, volatile / non-volatile computer storage media. Fig. Figure 10 illustrates, for example, a disk storage device. 1024 The disk storage 1024It can also include, but is not limited to, units such as a magnetic disk drive, a floppy disk drive, a tape drive, a Jaz drive, a Zip drive, an LS-100 drive, a flash memory card, or a memory stick. The disk storage 1024 It may also include storage media separately or in conjunction with other storage media, including, but not limited to, an optical drive such as a compact storage disk-ROM unit (CD-ROM), a recordable CD drive (CD-R drive), a rewritable CD drive (CD-RW drive), or a digital versatile ROM drive (DVD-ROM). To connect the disk storage 1024 with the system bus 1018 To enable this, a switchable or non-switchable interface is usually used, such as the interface 1026 used. Fig.Section 10 further illustrates software that acts as an interface between users and the basic computer resources available in the appropriate operating environment. 1000 are described. This software can, for example, also be an operating system. 1028 included. The operating system 1028 , which is on the disk storage 1024 It can be stored and is used to control and allocate computer resources. 1012 The system applications 1030 take advantage of the operating system's resource management 1028 using the program modules 1032 and program data 1034 , which, for example, are either in system memory 1016 or in disk storage 1024 are stored. It should be noted that this disclosure can be implemented with various operating systems or combinations of operating systems. A user inputs via the input device(s) 1036commands or information into the computer 1012 one. To the input units 1036 These include, but are not limited to, pointing devices such as a mouse, trackball, stylus, touchpad, keyboard, microphone, joystick, gamepad, satellite dish, scanner, TV card, digital camera, digital video camera, webcam, and similar devices. These and other input devices are connected via the system bus. 1018 through the interface port(s) 1038 with the processing unit 1014 connected. The interface port(s) 1038 It includes, for example, a serial port, a parallel port, a game port, and a Universal Serial Bus (USB). The output unit(s) 1040 uses some of the same connection types as the input unit(s) 1036For example, a USB port can be used to connect to the computer. 1012 to provide input and information from the computer 1012 to an output unit 1040 to output. The output adapter 1042 is provided to illustrate that it is among other output units 1040 some output units 1040 There are devices like monitors, speakers, and printers that require special adapters. For example, and this is not an exhaustive list, output adapters include... 1042 Video and sound cards that provide a method for connecting the output unit 1040 and the system bus 1018 provide. It should be noted that other units and / or systems of units provide both input and output capabilities, such as the remotely located computer(s). 1044 .
[0089] The computer 1012can be used in a networked environment with logical connections to one or more remotely located computers, such as the remotely located computer(s). 1044 be operated. Regarding the remotely located computer(s) 1044 It could be a computer, a server, a router, a network PC, a workstation, a microprocessor-controlled device, a peer unit, or another shared network node, and the like; these can usually also include many or all of the components associated with the computer. 1012 The described elements are included. For the sake of brevity, only one storage unit is used. 1046 with the remotely located computer(s) 1044 depicted. The remotely positioned computer(s). 1044 is / are logically via a network interface 1048 with the computer 1012 connected and then physically via the data transmission link1050 connected. The network interface 1048 This includes wired and / or wireless data transmission networks such as local area networks (LANs), wide area networks (WANs), mobile networks, etc. LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet, Token Ring, and similar technologies. WAN technologies include, but are not limited to, point-to-point connections, circuit-switched networks such as Integrated Services Digital Networks (ISDN) and their variants, packet-switched networks, and Digital Subscriber Lines (DSL). The data transmission connection(s) 1050 refers to the hardware / software used to establish the network interface 1048 with the system bus 1018 to connect. During the data transmission connection 1050 for illustration on the computer 1012 As depicted, it can also occur outside the computer. 1012The hardware / software for connecting to the network interface is located there. 1048 May also include, for purely illustrative purposes, internal and external technologies such as modems including standard telephone modems, cable modems and DSL modems, ISDN adapters and Ethernet cards.
[0090] One or more embodiments may be a system, a method, a device, and / or a computer program product at any possible level of technical integration. The computer program product may include a computer-readable storage medium (or media) on which computer-readable program instructions are stored to instruct a processor to execute aspects of the one or more embodiments. The computer-readable storage medium may be a physical unit capable of retaining and storing instructions for use by an execution unit.A computer-readable storage medium can be, for example, an electronic storage unit, a magnetic storage unit, an optical storage unit, an electromagnetic storage unit, a semiconductor storage unit, or any suitable combination thereof, without being limited to these. A non-exhaustive list of more specific examples of computer-readable storage media may also include the following: a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), and erasable programmable read-only memory (EPROM).Flash memory), static random-access memory (SRAM), portable compact storage disk-read-only memory (CD-ROM), DVD (digital versatile disc), USB flash drive, floppy disk, a mechanically coded unit such as punched cards or raised structures in a groove on which instructions are stored, and any suitable combination thereof. A computer-readable storage medium shall not, in its use herein, be understood as volatile signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses traveling through an optical fiber), or electrical signals transmitted by a wire.
[0091] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to individual data processing units or, via a network such as the internet, a local area network, a wide area network, and / or a wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission lines, wireless transmission, routing computers, firewalls, switching units, gateway computers, and / or edge servers. A network adapter card or network interface in each data processing unit receives computer-readable program instructions from the network and forwards them for storage on a computer-readable storage medium within the respective data processing unit.Computer-readable program instructions for performing operations of the disclosed subject matter may be assembly instructions, ISA (Instruction Set Architecture) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., as well as procedural programming languages such as the programming language "C" or similar programming languages.The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, via the internet using an internet service provider).In some embodiments, electronic circuits, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute computer-readable program instructions by using state information from the computer-readable program instructions to personalize the electronic circuits to perform aspects of the disclosed subject matter.
[0092] Aspects of the disclosed subject matter are described herein with reference to flowcharts and / or block diagrams or charts of methods, devices (systems), and computer program products according to embodiments of the disclosure. It is noted that each block of the flowcharts and / or block diagrams or charts, as well as combinations of blocks in the flowcharts and / or block diagrams or charts, can be executed by means of computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, a specialized computer, or another programmable data processing device to create a machine such that the instructions executed by the processor of the computer or other programmable data processing device constitute a method for implementing the function specified in the block or chart.The functions / steps specified in the blocks of the flowcharts and / or block diagrams or charts are generated. These computer-readable program instructions can also be stored on a computer-readable storage medium capable of controlling a computer, programmable data processing device, and / or other units to operate in a specific manner, such that the computer-readable storage medium on which instructions are stored has a manufacturing item, including instructions that implement aspects of the function / step specified in the block(s) of the flowchart and / or block diagrams or charts. The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other unit to execute a series of process steps on the computer or other unit.to cause the other programmable device or other unit to generate a process executed on a computer, such that the instructions executed on the computer, another programmable device or other unit implement the functions / steps specified in the block(s) of the flowcharts and / or block diagrams or charts.
[0093] The flowcharts and block diagrams or charts in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, processes, and computer program products according to various embodiments of the disclosed subject matter. In this context, each block in the flowcharts or block diagrams or charts can represent a module, segment, or part of instructions that includes one or more executable instructions for performing the specific logical function(s). In some alternative implementations, the functions specified in the block may occur in a different order than shown in the figures. For example, two blocks shown consecutively may in reality be executed essentially simultaneously, or the blocks may sometimes be executed in reverse order depending on the corresponding functionality.It should also be noted that each block of the block diagrams or charts and / or flowcharts, as well as combinations of blocks in the block diagrams or charts and / or flowcharts, can be implemented by special hardware-based systems that perform the specified functions or steps, or execute combinations of special hardware and computer instructions.
[0094] Although the subject matter above has been described in the general context of computer-executable instructions for a computer program product running on one or more computers, those skilled in the art will recognize that this disclosure can also be implemented in combination with other program modules. In general, program modules can include routines, programs, components, data structures, etc., that perform specific tasks and / or implement specific abstract data types. Furthermore, it is obvious to those skilled in the art that the computer-implemented methods described herein can be executed with other computer system configurations, including single- or multi-processor computer systems, mini data processing units, mainframe computers, as well as computers, portable data processing units (e.g., PDAs, telephones), microprocessor-based or programmable consumer or industrial electronics, and the like.The aspects described can also be used in distributed data processing environments, where tasks are performed by remotely located processing units connected via a data transmission network. However, some, if not all, aspects of this revelation can be implemented on standalone computers. In a distributed data processing environment, program modules can reside on both local and remote storage units.
[0095] As used in this application, the terms "component," "system," "platform," "interface," and similar terms can refer to and / or include a computer-related entity or an entity associated with a working machine, possessing one or more specific functionalities. The entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. For illustration, an application running on a server, and the server being a component, may also be an application running on a server.One or more components can reside within a single process and / or execution thread, and a component can reside on a single computer and / or be distributed across two or more computers. In another example, corresponding components can be executed from different computer-readable media containing different data structures. The components can exchange data via local and / or remote processes, for example, in accordance with a signal containing one or more data packets (e.g., data from one component interacting with another component in a local system, a distributed system, and / or with other systems via a network such as the internet).In another example, a component could be a device with specific functionality provided by mechanical parts operated by electrical or electronic circuits, which in turn were operated by a software or firmware application run by a processor. In this case, the processor could be located inside or outside the device and could execute at least part of the software or firmware application. In yet another example, a component could be a device that provides specific functionality through electronic components without any mechanical parts, where the electronic components could include a processor or other method for executing software or firmware that provides at least some of the functionality of the electronic components.With regard to one aspect, a component can emulate an electronic component via a virtual machine, e.g. in a cloud computing system.
[0096] Furthermore, the term "or" should be understood as an inclusive "or" and not an exclusive "or". That is to say, unless otherwise stated or clearly indicated from the context, "X uses A or B" should apply to all natural inclusive exchanges. That is, if X uses A; X uses B or X uses A and B, then "X uses A or B" is satisfied in each of the foregoing cases. In addition, the articles "a / an / an / a / one" and "a / an / one" as used in this description and in the accompanying drawings are generally to be understood as meaning "one or more", unless otherwise stated or these articles clearly refer to a singular form from the context. As used herein, the terms "example" and / or "exemplary" are used to serve as an example, case, or illustration.To avoid any doubt, the subject matter disclosed herein is not limited by such examples. Furthermore, any aspect or design described herein as an "example" and / or "exemplary" is not necessarily to be interpreted as preferential or advantageous over other aspects or designs, nor is it intended to exclude equivalent exemplary structures and techniques known to experts.
[0097] As used in the present description, the term "processor" can refer essentially to any data processing unit or unit that includes, but is not limited to, single-core processors; single-core processors with software multithreading capability; multi-core processors; multi-core processors with software multithreading capability; multi-core processors with hardware multithreading technology; parallel platforms; and parallel platforms with distributed shared memory.Furthermore, a processor can refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. In addition, processors can utilize nanoscale architectures such as molecular and quantum dot-based transistors, switches, and gates, without limitation, to optimize space utilization or improve the performance of user equipment. A processor can also be implemented as a combination of data processing units.In this disclosure, terms such as "store," "memory," "store data," "data storage," "database," and essentially any other information storage component essential to the operation and functionality of a component are used to refer to "memory components," entities contained within a "memory," or components that have memory. It should be noted that the memories and / or memory components described herein may be either volatile memory or non-volatile memory, or may include both. For example, and without limitation, the non-volatile memory may be read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or non-volatile random-access memory (RAM) (e.g.,The volatile memory may include ferroelectric RAM (FeRAM). The volatile memory may contain RAM that can function, for example, as an external buffer. For example, and without limitation, RAM is available in many forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), synchronous SDRAM with double data rate (DDR SDRAM), extended SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct dynamic Rambus RAM (DRDRAM), and dynamic Rambus RAM. Furthermore, the disclosed memory components of systems or computer-implemented methods contained herein are intended to include, but are not limited to, these and all other suitable types of memory.
[0098] What has been described above includes only examples of systems and computer-implemented methods. It is, of course, not possible to present every conceivable combination of components or computer-implemented methods to describe this disclosure, but a person skilled in the art can recognize that many further combinations and modifications of this disclosure are possible. Furthermore, where the terms "contain / comprise," "include," "possess," and similar terms are used in the detailed description, claims, annexes, and drawings, these terms are to be understood as inclusive in a similar way to how "include" is interpreted when used as a transitional word in a claim. The descriptions of the various embodiments have been presented for illustrative purposes but are not intended to be exhaustive or limited to the embodiments.It is obvious to those skilled in the art that many modifications and adaptations are possible without deviating from the scope and inventive concept of the described embodiments. The terminology used herein has been chosen to best explain the basic ideas of the embodiments, their practical application, or technical improvements over technologies already on the market, or to enable those skilled in the art to understand the embodiments described herein.
Claims
[1] A computer-implemented method that features: Generating corresponding numeric fields in a defined floating-point number format by a system functionally connected to a processor, wherein the corresponding numeric fields comprise a sign field, an exponent field, and a mantissa field, the defined floating-point number format using six bits in the exponent field; and Calculating binary floating-point numbers according to the defined floating-point number format by the system in conjunction with the execution of an application. [2] A computer-implemented method according to claim 1, wherein the defined floating-point number format uses a first binad to represent zero and normal numbers, the first binad belonging to the exponent field which contains only zeros, and wherein a normal number of the normal numbers is a finite non-zero floating-point number with a size greater than or equal to a minimum value determined as a function of a base and a minimum exponent in conjunction with the defined floating-point number format. [3] A computer-implemented method according to claim 2, wherein the first binary has a data point and other data points, wherein the data point has a fraction of all zeros and represents zero, and wherein the other data points represent the normal numbers. [4] A computer-implemented method according to claim 2, wherein the defined floating-point number format uses a second binad belonging to the exponent field which contains only ones, wherein a smaller set of data points in the second binad is used to represent an infinity value and a non-number value according to the defined floating-point number format, and wherein the smaller set of data points has fewer data points than the set of data points belonging to the entire second binad. [5] A computer-implemented method according to claim 4, wherein the defined floating-point number format uses only one data point of the second binad to represent both the infinity value and the non-number value. [6] A computer-implemented method according to claim 1, wherein the six bits of the exponent field have a minimum number of bits to represent a range of numerical values whose occurrence is predicted during the execution of the application without using subnormal numbers, and wherein the defined floating-point number format does not include the subnormal numbers in order to improve efficiency by reducing the hardware used to run the application and compute the binary floating-point numbers. [7] A computer-implemented method according to claim 1, further comprising: according to the defined floating-point number format Representing a sign of zero by the system as a term indicating that the sign is irrelevant with respect to zero; and Generating an arbitrary value in the sign field by the system to represent the term, using fewer resources in generating the arbitrary value than determining and generating a non-arbitrary value for the sign field. [8] A computer-implemented method according to claim 1, further comprising: in accordance with the defined floating-point number format Representing a non-number value and an infinity value together as a combined symbol by the system; Representing the sign of the merged symbol by the system as a term indicating that the sign is irrelevant with respect to the merged symbol; and Generating an arbitrary value in the sign field by the system to represent the term, using fewer resources in generating the arbitrary value than determining and generating a non-arbitrary value for the sign field. [9] A computer-implemented method according to claim 1, wherein the defined floating-point number format uses only a single rounding mode, which is rounding up to the nearest digit, and wherein the method further comprises: Rounding binary intermediate floating-point numbers by the system using the single rounding mode to perform a rounding of the binary intermediate floating-point numbers to enable the computation of the binary floating-point numbers according to the defined floating-point number format, whereby the use of the single rounding mode enables an improvement in efficiency by reducing the hardware used to run the application and compute the binary floating-point numbers. [10] System comprising a means suitable for carrying out all steps of the method according to any of the preceding method claims. [11] Computer program comprising instructions for performing all steps of the method according to any of the preceding method claims when the computer program is executed on a computer system.
Citation Information
Patent Citations
Efficient parallel handling of floating-point exceptions in one processor
DE102009030525A1
DENORMALIZATION IN A MULTIGERACY SLIDING-POINT Arithmetic Circuit Arrangement
DE112017004287T5
Multifunctional two-part reference table
DE69801678T2
METHOD, DEVICE AND CALCULATOR PROGRAM PRODUCTS FOR ACCUMULATION OF LOGARITHMIC VALUES
DE69832519T2
Reuse of rounder for fixed conversion of log instructions
US20100174764A1