Data processing system, circuit and method
By introducing a dual-mode adder, maximum value detector circuit, area detector circuit and alignment circuit into the data processing circuit, the problem of exponential and mantissa adjustment in floating-point number operation is solved, and efficient and accurate floating-point number operation is achieved.
Patent Information
- Application Number
- CN202510059743.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
In floating-point number addition or subtraction operations, it is difficult for the prior art to efficiently adjust the exponent and mantissa of floating-point numbers to achieve fast and accurate floating-point number operations.
Provided is a data processing circuit, including a dual-mode adder, a maximum value detector circuit, a region detector circuit and an alignment circuit. The circuit detects the maximum value by the first part of the product, classifies it to the region, and aligns the mantissa of the floating point number according to the region and the second part of the product to generate the aligned mantissa for floating point number operation.
Through the design of this circuit, the exponent and mantissa of floating point numbers can be effectively adjusted, the efficiency and accuracy of floating point number operations can be improved, and energy consumption and area use can be reduced.
Smart Images

Figure CN120066451A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data processing system, and more particularly to a data processing system for a neural network. Background Art
[0002] A floating point number is typically represented by three components: a sign, an exponent, and a mantissa. When performing addition or subtraction of two floating point numbers, the exponents and mantissas need to be aligned. Specifically, before adding or subtracting the significands of two floating point numbers, the exponents of the two floating point numbers must be adjusted to be equal to each other, and the mantissas of the two floating point numbers are shifted according to the adjustment. Summary of the Invention
[0003] In some embodiments, a data processing circuit is provided. The circuit includes a dual-mode adder, a maximum detector circuit, a region detector circuit, and an alignment circuit. The dual-mode adder generates a product between a first exponent of a first floating point number and a second exponent of a second floating point number. The maximum detector circuit detects a maximum value in a first part of the product. The region detector circuit classifies the first part into a region by comparing the first part with the maximum value. The alignment circuit aligns a first mantissa of the first floating point number according to the region and a second part of the product to generate an aligned mantissa for floating point arithmetic.
[0004] In some embodiments, a data processing system is provided. The system includes a data processing circuit and a memory device. The data processing circuit includes a first comparator circuit, a second comparator circuit, and an alignment circuit. The first comparator circuit detects a maximum value in a first part of a product between a first activation and a weight, where the first activation and the weight are floating point numbers. The second comparator circuit compares the first part of the product with a reference to generate a region flag. The reference is based on the maximum value. The alignment circuit aligns each first mantissa according to a corresponding region flag of the region flag to generate an aligned mantissa. The memory device includes a compute-in-memory (CIM) array. The CIM array performs a multiply-accumulate operation between the aligned mantissa and a second mantissa of the weight.
[0005] In some embodiments, a data processing method is provided. The method includes the steps of: adding a first exponent of a first floating point number and a second exponent of a second floating point number to generate a product; identifying a maximum value in a first part of the product; generating a region flag corresponding to the product according to the maximum value, where each region flag represents a region to which the corresponding product is classified; and aligning a first mantissa of the first floating point number according to the region flag and a second part of the product to generate an aligned mantissa. Brief Description of the Drawings
[0006] Aspects of the present disclosure may be best understood from the following detailed description when taken in conjunction with the accompanying drawings. Note that, in accordance with standard practice in the industry, the various features are not drawn to scale. In fact, for clarity of discussion, the dimensions of the various features may be arbitrarily increased or decreased.
[0007] Figure 1 Schematic diagram of a system according to various embodiments of the present disclosure;
[0008] Figure 2 For various embodiments of the present disclosure, as Figure 1 Schematic diagram of a data processing circuit shown;
[0009] Figure 3 For some embodiments of the present disclosure, as Figure 2 Flowchart of an operation method of a data processing circuit shown;
[0010] Figure 4A For some embodiments of the present disclosure, corresponding to as Figure 1 And Figure 2 Schematic diagram of an alignment operation of a data processing circuit shown and a method as Figure 3 Shown;
[0011] Figure 4B For some embodiments of the present disclosure, corresponding to as Figure 1 And Figure 2 Schematic diagram of a data processing circuit shown and a method as Figure 3 Shown;
[0012] Figure 5 For various embodiments of the present disclosure, relative to as Figures 1 to 3 , Figure 4A And Figure 4B Schematic diagram of a data processing circuit of a data processing circuit configuration shown;
[0013] Figure 6 And Figure 7 For various embodiments of the present disclosure, corresponding to as Figure 1 , Figure 2 And Figure 5 Schematic diagram of an alignment operation table of a data processing circuit shown and a method as Figure 3 Shown;
[0014] Figure 8 For various embodiments of the present disclosure, corresponding to as Figure 1 , Figure 2 And Figure 5 Schematic diagram of a data processing circuit shown and a method as Figure 3 Shown;
[0015] Figure 9A and Figure 9B is a schematic diagram of the floating-point operation of the system as shown in Figures 1 to 8 accordance with various embodiments of the present disclosure;
[0016] Figure 10 is a schematic diagram of the region detector circuit and the arithmetic circuit as shown in Figures 1 to 8 and Figures 9A to 9B accordance with various embodiments of the present disclosure;
[0017] Figure 11A and Figure 11B is a schematic diagram of the operation of the arithmetic circuit as shown in Figures 1 to 3 and Figure 4A and Figure 4B and Figures 5 to 8 and Figures 9A to 9B accordance with various embodiments of the present disclosure;
[0018] Figure 12A is a schematic diagram of an example of the floating-point mode of the system corresponding to as shown in Figures 1 to 3 and Figure 4A and Figure 4B and Figures 5 to 8 and Figure 9A and Figure 9B and Figure 10 and Figure 11A and Figure 11B accordance with various embodiments of the present disclosure;
[0019] Figure 12B is a schematic diagram of an example of the floating-point mode of the system corresponding to as shown in Figures 1 to 3 and Figure 4A and Figure 4B and Figures 5 to 8 and Figure 9A and Figure 9B and Figure 10 and Figure 11A and Figure 11B accordance with various embodiments of the present disclosure;
[0020] Figure 13 is a schematic diagram of the system with respect to the system configuration as shown in Figures 1 to 3 and Figure 4A and Figure 4B and Figures 5 to 8 and Figure 9A and Figure 9B and Figure 10 and Figure 11A and Figure 11B and Figure 12A and Figure 12B accordance with various embodiments of the present disclosure;
[0021] Figure 14Schematic diagram of a system as shown in Figure 13 for various embodiments according to the present disclosure;
[0022] Figure 15 For various embodiments according to the present disclosure, such as Figure 13 and Figure 14 Flowchart of an operation method of the system shown.
[0023]
Symbol Description
[0024] 10, 20: System
[0025] 100: Data processing circuit
[0026] 101: Dual-mode adder
[0027] 102: Maximum detector circuit
[0028] 103: Register
[0029] 104: Region detector circuit
[0030] 104a: Region bias circuit
[0031] 104b: Region detector
[0032] 105: Alignment circuit
[0033] 105a: Input processing circuit
[0034] 111: Comparator circuit
[0035] 112: Subtractor
[0036] 113: Inverter
[0037] 200: Arithmetic circuit
[0038] 201: Memory writer
[0039] 202: CIM array
[0040] 203: ADC circuit
[0041] 204: Accumulator circuit
[0042] 300, 1500: Method
[0043] 400, 600, 700, 800: Table
[0044] 500: Data processing circuit
[0045] 1301: Calculation circuit
[0046] 1401: Multiplier
[0047] 1402: Adder tree
[0048] 1501 to 1505, s1 to s4: Steps
[0049] t1 to t5: Time Detailed implementation manners
[0050] The following disclosure provides different embodiments or examples for implementing the features of the provided subject matter. The following describes specific examples of components, materials, values, steps, arrangements, etc. to simplify the content of the embodiments of this disclosure. Of course, these are only examples and are not intended to be limiting. Other components, materials, values, steps, arrangements, etc. can be expected. For example, in the following description, forming the first feature above or on top of the second feature may include embodiments in which the first and second features are formed in direct contact, and may also include embodiments in which additional features are formed between the first and second features so that the first and second features may not be in direct contact. In addition, the present disclosure may repeat element symbols and / or letters in various examples. This repetition is for the purpose of simplicity and clarity and does not itself specify the relationship between the various embodiments or configurations discussed.
[0051] The terms used in the following specification and the scope of the invention patent application generally have the ordinary meanings clearly established in the art or in the specific context where each term is used. Those of ordinary skill in the art will understand that a component or process may be represented by different names. The numerous different embodiments detailed in this specification are for illustrative purposes only and do not limit the scope and spirit of the disclosure or any of the exemplary terms.
[0052] It should be noted that terms such as "first" and "second" are used herein to describe various elements or processes and are intended to distinguish the elements or processes. However, the elements, processes, and their order are not limited by these terms. For example, without departing from the scope of the embodiments of this disclosure, the first element may be referred to as the second element, and the second element may be similarly referred to as the first element.
[0053] In the following discussion and the scope of the invention patent application, the terms "comprising", "including", "containing", "having", "involving", etc. should be understood as open-ended, that is, interpreted as including but not limited to. As used herein, the term "and / or" is not mutually exclusive and includes any relevant listed items and all combinations of one or more of the relevant listed items.
[0054] As used herein, "about", "approximately", "near" or "substantially" shall generally refer to any approximation of a given value or range, which varies depending on the various fields involved and the scope of which shall be consistent with the broadest interpretation understood by those skilled in the art to cover all such modifications and similar constructions. In some embodiments, it shall generally refer to within twenty percent of a given value or range, preferably within ten percent, more preferably within five percent. The numerical values given herein are approximate, meaning that the terms "about", "approximately", "near" or "substantially" can be inferred if not explicitly stated, or other approximations are meant.
[0055] This application relates to floating-point data processing. In high-performance computing applications, a large number of floating-point operations are typically required. To achieve high speed and high performance, some systems, such as compute-in-memory (CIM) systems and near-memory-computing (NMC) systems, may include floating-point arithmetic circuits for performing floating-point operations and floating-point processing circuits for generating preprocessed floating-point data as inputs to the floating-point arithmetic circuits. For example, the floating-point processing circuit may generate preprocessed floating-point data with the mantissas aligned as inputs to the floating-point arithmetic circuit, which performs addition and subtraction of floating-point numbers.
[0056] Now refer to Figure 1 . Figure 1 FIG. 10 is a schematic diagram of a system 10 according to various embodiments of the present disclosure. In some embodiments, the system 10 is a CIM system. The system 10 includes a data processing circuit 100 and an arithmetic circuit 200. In some embodiments, the data processing circuit 100 is a floating-point processing circuit, and the arithmetic circuit 200 is a floating-point arithmetic circuit. In some embodiments, the arithmetic circuit 200 is implemented by a memory device that performs CIM operations. The arithmetic circuit 200 performs floating-point operations.
[0057] In some embodiments, the data processing circuit 100 aligns the number "n" of floating-point numbers IN[1] to IN[n] to generate aligned mantissas IN MA [1] to IN MA [n] as inputs to the arithmetic circuit 200. In some embodiments, the arithmetic circuit 200 performs multiply-and-accumulate (MAC) operations (dot product operations) on the floating-point numbers IN[1] to IN[n] and the floating-point numbers W[1] to W[n] to generate a MAC result as the output OUT. In other embodiments, the arithmetic circuit 200 performs operations on the aligned mantissas IN MA [1] to IN MA [n] and the mantissas W of the floating-point numbers W[1] to W[n]M [1] to W M [n] perform a MAC operation to generate a MAC result as output OUT. In some embodiments, the floating-point numbers IN[1] to IN[n] are activations of a neural network, and the floating-point numbers W[1] to W[n] are weights of the neural network. In some embodiments, the floating-point numbers IN[1] to IN[n] correspond to media information (e.g., images, audio, video data) or electrical signals for detection, classification, recognition, adjustment, conversion, or any suitable application. The output OUT corresponds to the inference result of the neural network.
[0058] For practical applications, the arithmetic circuit 200 and the neural network in the disclosure can be used in various fields, such as machine vision, image classification, or data classification. For example, the arithmetic circuit 200 and the neural network can be used to classify medical images. For example, they can be used to classify X-ray images under normal conditions, including pneumonia, bronchitis, or heart disease. These methods can also be used to classify ultrasound images of normal or abnormal fetal positions. On the other hand, the arithmetic circuit 200 and the neural network can also be used to classify images collected during autonomous driving, such as differentiating images of normal roads, roads with obstacles, and road conditions of other vehicles. In addition, the arithmetic circuit 200 and the neural network can be used in other similar fields, such as music spectrum recognition, spectrum recognition, big data analysis, data feature recognition, and other related machine learning fields.
[0059] As Figure 1 shown, in some embodiments, the arithmetic circuit 200 includes a memory writer 201, a compute-in-memory (CIM) array 202, an analog-to-digital converter (ADC) circuit 203, and an accumulator circuit 204. The memory writer 201 is coupled to the CIM array 202. The CIM array 202 is coupled to the ADC circuit 203. The ADC circuit 203 is coupled to the accumulator circuit 204. It should be understood that within the scope of the present disclosure embodiments, the description of "coupled" herein includes any direct and indirect connection manners. Therefore, if the first element is described as being coupled to the second element, it means that the first element can be directly connected to the second element via an electrical connection or via other elements or connections.
[0060] In some embodiments, the CIM array 202 includes memory cells that store weights of a neural network and receive activations of the neural network as inputs. In some embodiments, the memory cells store activations of the neural network and receive weights of the neural network as inputs. The memory cells include non-volatile RAM (RRAM, SRAM) and volatile RAM (DRAM) or other suitable memories.
[0061] The memory writer 201 writes the floating-point numbers W[1] to W[n] or the mantissas W M [1] to W M [n] into the CIM array 202. The CIM array 202 receives the aligned mantissas IN MA [1] to IN MA [n] as inputs, and performs a multiply-accumulate operation on the aligned mantissas IN MA [1] to IN MA [n] and the mantissas W M [1] to W M [n] stored in the memory cells of the CIM array 202. The ADC circuit 203 converts the result of the CIM array 202 from analog to digital. The accumulator circuit 204 accumulates the corresponding results from the ADC circuit 203 to generate the MAC results of the floating-point numbers IN[1] to IN[n] and the floating-point numbers W[1] to W[n], or the MAC results of the aligned mantissas IN MA [1] to IN MA [n] and the mantissas W M [1] to W M [n].
[0062] Figure 1 The configurations shown are for illustrative purposes only. Various tools are within the contemplated scope of the embodiments of the present disclosure. For example, in some embodiments, the system 10 further includes a floating-point compression circuit coupled between the data processing circuit 100 and the CIM array 202. In some embodiments, the arithmetic circuit 200 further includes a multiplexer (MUX) circuit coupled between the CIM array 202 and the ADC circuit 203.
[0063] Now refer to Figure 2 . Figure 2 FIG. is a schematic diagram of the data processing circuit 100 according to various embodiments of the present disclosure. For embodiments relative to Figure 1 shown, for ease of understanding, Figure 1 similar elements in Figure 2 are designated with the same element symbols. For brevity, the specific operations of similar elements that have been discussed in detail in the above paragraphs are omitted herein.
[0064] In some embodiments, before performing a MAC operation between floating-point numbers IN[1] to IN[n] and floating-point numbers W[1] to W[n], data processing circuit 100 performs a mantissa alignment operation on the floating-point numbers IN[1] to IN[n] according to the exponents IN E [1] to IN E [n] and the exponents W E [1] to W E [n] of the floating-point numbers W[1] to W[n]. In the mantissa alignment operation, data processing circuit 100 aligns the mantissas IN M [1] to IN M [n] to respectively generate aligned mantissas IN MA [1] to IN MA [n].
[0065] For illustration, data processing circuit 100 includes a dual-mode adder 101, a maximum detector circuit 102, a register 103, a region detector circuit 104, and an alignment circuit 105. As Figure 2 shown, the dual-mode adder 101 is coupled to the maximum detector circuit 102 and the alignment circuit 105. The maximum detector circuit 102 is coupled to the register 103 and the region detector circuit 104. The region detector circuit 104 is coupled to the alignment circuit 105. The operation of data processing circuit 100 will be described in the following paragraphs with reference to Figure 3 the following.
[0066] Now refer to Figure 2 and Figure 3 . Figure 3 FIG. 300 is a flowchart of an operation method 300 of the data processing circuit 100 according to some embodiments of the present disclosure. It should be understood that additional steps may be provided before, during, and after the steps shown in Figure 2 FIG. 300, and for additional embodiments of method 300, some of the steps described below may be replaced or eliminated. The order of the steps may be interchanged. In various views and illustrative embodiments, like element symbols are used to designate like elements. Method 300 includes steps s1 to s4 described below with reference to the data processing circuit 100 shown in Figure 3 FIG. 300. Figure 2 FIG. 300.
[0067] In step s1, a product PD E [i] of the exponent IN E [i] and the exponent W E [i] is generated, where the number "i" is between the numbers "1" and "n". Specifically, the dual-mode adder 101 receives the exponents IN E [1] to IN E [n] and the exponents W E[1] to W E [n], and add the exponent IN E [1] to IN E [n] to the exponent W E [1] to W E [n] respectively to generate the exponent IN E [1] to IN E [n] and the exponent W E [1] to W E [n] to obtain the product PD E [1] to PD E [n].
[0068] In step s2, perform a maximum product exponent search to identify the higher-order "M" bits PD E [1][M + N - 1:N] to PD E [n][M + N - 1:N] of the maximum value PD E-MAX [M + N - 1:N]. Specifically, the maximum value detector circuit 102 uses the product PD E [1] to PD E [n] and performs a maximum product exponent search on the higher-order "M" bits in the total "M + N" bits of each of them. M and N are positive integers. For example, the maximum value detector circuit 102 receives the higher-order "M" bits PD E [1] to PD E [n] which are PD E [1][M + N - 1:N] to PD E [n][M + N - 1:N], and detects the maximum value PD E [1][M + N - 1:N] to PD E [n][M + N - 1:N] which is PD E-MAX [M + N - 1:N]. In some embodiments, the number "N" is used as the logarithm of the base-2 mantissa digit width number of the mantissa IN M [1] to IN M [n].
[0069] In some embodiments, the maximum value detector circuit 102 includes a comparator circuit 111 that compares two higher-order "M" bits PD E [1][M + N - 1:N] to PD E [n][M + N - 1:N] (for example, PD E [1][M + N - 1:N] and PD E[2][M+N-1:N]), and determines the larger one as the temporary maximum value stored in register 103. Then, comparator circuit 111 compares the temporary maximum value with the remaining higher-order "M" bits of PD E [1][M+N-1:N] to PD E [n][M+N-1:N] to determine the larger one as the updated temporary maximum value stored in register 103. Comparator circuit 111 repeats this comparison to detect the maximum value PD E [1][M+N-1:N] to PD E [n][M+N-1:N] among the higher-order "M" bits of PD E-MAX [M+N-1:N].
[0070] In one example, the total "M+N" bits of each of the products PD E [1] to PD E [n] are 9 bits. The higher-order "M" bits are 6 bits. The lower-order "N" bits are 3 bits.
[0071] Compared with some methods that perform the maximum product exponent search using all the "M+N" bits of each of the products PD E [1] to PD E [n], the maximum value detector circuit 102 only uses the higher-order "M" bits to perform the maximum product exponent search, simplifying the hardware design to reduce area and power consumption.
[0072] In step s3, the region detector circuit 104 generates region flags ZFG[1] to ZFG[n] corresponding to the products PD E-MAX [M+N-1:N] for the products PD E [1] to PD E [n]. Each of the region flags ZFG[1] to ZFG[n] indicates the region to which the corresponding one of the products PD E [1] to PD E [n] is classified. The following paragraphs describe more details of step s3.
[0073] In some embodiments, the region detector circuit 104 classifies each of the products PD E [1] to PD E [n] into one of "K" regions. Specifically, the region detector circuit 104 receives the higher-order "M" bits of PD E [1][M+N-1:N] to PD E [n][M+N-1:N] and the maximum value PD E-MAX [M+N-1:N], and based on the higher-order "M" bits of PD E[1][M+N-1:N] to PD E [n][M+N-1:N] and the maximum value PD E-MAX [M+N-1:N] will multiply PD E [1] to PD E [n] Each of them is classified into one of the "K" regions. In some embodiments, the region detector circuit 104 is based on the higher-order "M" bits PD E [i][M+N-1:N] and the reference PD E-REF1 to PD E-REFK Between the comparisons, multiply PD E [i] is classified into one of the "K" regions.
[0074] In some embodiments, the reference PD E-REF1 Is a binary number of "M+N" bits, where the "M" higher-order bits are the maximum value PD E-MAX [M+N-1:N], and the "N" lower-order bits are the "N" bits of the binary bits. The reference PD E-REF2 Is used as the reference PD E-REF1 Subtract the decimal number "2 N ", the reference PD E-REF3 Is used as the reference PD E-REF1 Subtract the decimal number "2×2 N "... and the reference PD E-REFK Is used as the reference PD E-REF1 Subtract the decimal number "(K-1)×2 N ".
[0075] In some embodiments, the region detector circuit 104 includes generating a reference PD by performing subtraction on the maximum value PD E-MAX [M+N-1:N] to generate the reference PD E-REF1 to PD E-REFK Of the M-bit subtractor 112. For example, the subtractor 112 subtracts the maximum value PD from the decimal number 1 E-MAX [M+N-1:N], to generate the "M" higher-order bits of the reference PD E-REF2 The subtractor 112 subtracts the maximum value PD from the decimal number 2 E-MAX [M+N-1:N], to generate the "M" higher-order bits of the reference PD E-REF3 "... and the subtractor 112 subtracts the maximum value PD from the decimal number K-1 E-MAX [M+N-1:N], to generate the "M" higher-order bits of the reference PD E-REFK Then, the region detector circuit 104 fills each "M" higher-order bit of the reference PD E-REF2 to PD E-REFK To generate the reference PDE-REF2 to PD E-REFK .
[0076] In some embodiments, when the product PD E [i] is less than or equal to the reference PD E-REF1 and greater than the reference PD E-REF2 , the region detector circuit 104 classifies the product PD E [i] into region 1 in the "K" region. When the product PD E [i] is less than or equal to the reference PD E-REF2 and greater than the reference PD E-REF3 , the product PD E [i] is classified into region 2 in the "K" region,..., when the product PD E [i] is less than or equal to the reference PD E-REFK-1 and greater than the reference PD E-REFK , the product PD E [i] is classified into region K - 1 in the "K" region, and when the product PD E [i] is less than or equal to the reference PD E-REFK , the product PD E [i] is classified into region K in the "K" region.
[0077] In some embodiments, the region detector circuit 104 includes an "M" - bit comparator. When the higher - order "M" - bit PD E [i][M + N - 1:N] is less than or equal to the maximum value PD E-MAX [M + N - 1:N] and greater than the maximum value PD E-MAX [M + N - 1:N] minus the decimal number 1, the region detector circuit 104 classifies the product PD E [i] into region 1. When the higher - order "M" - bit PD E [i][M + N - 1:N] is less than or equal to the maximum value PD E-MAX [M + N - 1:N] minus the decimal number 1 and greater than the maximum value PD E-MAX [M + N - 1:N] minus the decimal number 2, the product PD E [i] is classified into region 2,..., when the higher - order "M" - bit PD E [i][M + N - 1:N] is less than or equal to the maximum value PD E-MAX [M + N - 1:N] minus the decimal number "K - 1" and greater than the maximum value PD E-MAX [M + N - 1:N] minus the decimal number "K", the product PD E [i] is classified into region K - 1, and when the product PD E [i] is less than or equal to the decimal number "K", the product PDE Classify into region “K” within the “K” region.
[0078] In some embodiments, determine the number of regions “K” according to the number of bits (bit width) “B” of each of IN MA [1] to IN MA [n]. In some embodiments, the number “K” is equal to the number of bits of each of IN MA [1] to IN MA [n] divided by “2 N ”, and then adding 1 (K = (B / 2 N ).
[0079] In some embodiments, the region detector circuit 104 generates region flags (region bits) ZFG[1] to ZFG[n], respectively indicating the regions corresponding to the products PD E [1] to PD E [n]. In some embodiments, each of the region flags ZFG[1] to ZFG[n] includes one of the numbers “1” to “K” in binary form to respectively represent one of regions 1 to region K. For example, when the number of regions is equal to 3, each of the region flags ZFG[1] to ZFG[n] includes one of the binary numbers “01”, “10”, and “11” to respectively represent regions 1 to 3. When the product PD E [1] is classified into region 1, the region detector circuit 104 generates a region flag ZFG[1] including the binary number “11”.
[0080] In step s4, the alignment circuit 105 aligns the mantissa IN E [1][N - 1:0] to PD E [N][N - 1:0] according to the region flags ZFG[1] to ZFG[N] and the “N” lower order bits of the mantissa IN M [1] to IN M [N] to generate the mantissa IN MA [1] to IN MA [n].
[0081] Perform an alignment operation to shift the mantissa IN M [i]. The mantissa IN M [i] is shifted according to the “relative exponent value”, which is defined as the difference between the product PD E [i] and the upper boundary of the region of the product PD E [i]. For example, when the product PD E [1] is located in region 1, the relative exponent value is the product PD E [1] and the reference PD E-REF1The difference between, where the reference PD E-REF1 is the upper boundary of Region 1. According to an embodiment of the present disclosure, corresponding to the mantissa IN M [i], the relative exponent value is equal to the inverse of the lower N bits PD E [i][N-1:0]. Further reference is made below to Figure 4A and Figure 4B for an example of the alignment operation.
[0082] Now refer to Figure 2 , Figure 3 , Figure 4A and Figure 4B . Figure 4A is a schematic diagram of the alignment operation corresponding to the data processing circuit 100 as shown in Figure 1 and Figure 2 and the method 300 as shown in Figure 3 according to some embodiments of the present disclosure. Figure 4B is a table 400 of the alignment operation corresponding to the data processing circuit 100 as shown in Figure 1 and Figure 2 and the method 300 as shown in Figure 3 according to some embodiments of the present disclosure. For the embodiment relative to Figures 1 to 3 , Figures 4A to 4B the similar elements in
[0083] are designated with the same element symbols for ease of understanding. Figure 4A In the example shown in E the bit width of each of the products PD E [1] to PD E [n] is 9. The number “M” is 6, and the number “N” is 3. The products PD E [1] to PD E-MAX [n] are classified into three regions: Region 1 to Region 3. PD E-REF1 [M+N-1:N] is the 6-bit “011111”. Therefore, the reference PD E-REF2 is “011111111”, the reference PD E-REF3 is “011110111”, and the reference PD E [0] is “0111111101”. The product PD E [1] is “0111110011”. The product PD E [n] is “0111101100”.
[0084] The product PD E-REF1 located between the reference PD E-REF2 and PD E[0] Classify it into Region 1 in Step S3. Then, in Step S4, based on the lower three bits of the product PD E [0] determine the difference between the reference PD E-REF1 (the upper boundary of Region 1) and the product PD E [0]. As Figure 4A shown, the binary number "010" corresponding to the difference 2 is the inverse of the lower three bits "101". Based on the difference between the reference PD E-REF1 and the product PD E [0] and the product PD E [0] being classified into Region 1, determine the shift bit of the aligned mantissa IN M [0] as 2.
[0085] Similarly, the product PD E-REF2 located between the reference PD E-REF3 and PD E [1] is classified into Region 1 in Step S3. In Step S4, based on the lower three bits of the product PD E [1] determine the difference between the reference PD E-REF2 (the upper boundary of Region 2) and the product PD E [1]. As Figure 4A shown, the binary number "100" corresponding to the difference 4 is the inverse of the lower three bits "011". Based on the difference between the reference PD E-REF1 and the product PD E [0] and the product PD E [0] being classified into Region 2, determine the shift bit of the aligned mantissa IN M [1] as 4 plus 8.
[0086] The product PD E-REF3 less than the reference PD E [n] is classified into Region 3 in Step S3. Then, in Step S4, based on the lower three bits of the product PD E [n] determine the difference between the reference PD E-REF3 (the upper boundary of Region 3) and the product PD E [n]. As Figure 4A shown, the binary number "011" corresponding to the difference 3 is the inverse of the lower three bits "100". Based on the difference between the reference PD E-REF3 and the product PD E [0] and the product PD E [0] being classified into Region 3, determine the shift bit of the aligned mantissa IN M [n] as 3 plus 16.
[0087] In the alignment operation of Step S4, refer to Figure 2, the alignment circuit 105 receives the region flags ZFG[1] to ZFG[n] from the region detector circuit 104, and further receives the lower "N" bits PD E [1][N-1:0] to PD E [n][N-1:0]. The region detector circuit 104 aligns the mantissa IN E [1][N-1:0] to PD E [n][N-1:0] further according to the region flags ZFG[1] to ZFG[n] and the lower "N" bits PD M [1] to IN M [n], to generate the aligned mantissa IN MA [1] to IN MA [n].
[0088] In some embodiments, the alignment circuit 105 pads the mantissa IN with binary zeros on the most significant bit (MSB) side by according to the region flags ZFG[1] to ZFG[n] and the lower "N" bits PD E [1][N-1:0] to D E [n][N-1:0], to generate the aligned mantissa IN M [1] to IN M [n] MA [1] to IN MA [n].
[0089] In some embodiments, the alignment circuit 105 uses the aligned mantissa IN MA [1] to IN MA [n] to generate the corresponding floating point number IN A [1] to IN A [n]. Each of the floating point numbers IN A [1] to IN A [n] has an exponent of the reference PD E-REF1 . In some embodiments, the arithmetic circuit 200 performs floating point operations on the floating point numbers IN A [1] to IN A [n].
[0090] In some embodiments, referring to Figure 4B , when the product PD E [i] is in region 1, the mantissa IN M [i] is padded with bits of binary zeros on the MSB side. The number of padded bits of binary zeros is based on the lower "N" bits PD EThe bitwise inverse of [i][N-1:0]. In some embodiments, the number of padding bits of binary zero is equal to the lower-order "N" bits of PD in decimal form E The bitwise inverse of [i][N-1:0]. For example, when the number "N" is 3 and the lower-order "N" bits of PD E [i][N-1:0] is the three-bit "101" in some embodiments, the number of padding bits of binary zero is equal to 2. The number 2 corresponds to the bits "010" in decimal form, where the bits "010" are the bitwise inverse of the bits "101". In this embodiment, the mantissa IN M [i] is padded with two bits from the MSB side.
[0091] When the product PD E [i] is in region 2, the mantissa IN M [i] is padded with bits of binary zero on the MSB side. However, the number of padding bits of binary zero is equal to "(2 - 1) × 2 N " plus the bitwise inverse of the lower-order "N" bits of PD E [i][N-1:0]. For example, when the number "N" is 3 and the lower-order "N" bits of PD E [i][N-1:0] is the three-bit "101" in some embodiments, the number of padding bits of binary zero is equal to 8 plus 2. The number 8 is equal to 2 3 , and the number 2 corresponds to the bits "010" in decimal form, where the bits "010" are the bitwise inverse of the bits "101". In this embodiment, the mantissa IN M [i] is padded with 8 plus 2 bits from the MSB side.
[0092] Similarly, when the product PD E [i] is in region 3, the number of padding bits of binary zero is equal to "(3 - 1) × 2 N " plus the bitwise inverse of the lower-order "N" bits of PD E [i][N-1:0].
[0093] For the product PD E [i] in the case of regions 4 to region K, the number of padding bits of binary zero is configured in a similar manner as described above. For example, when the product PD E [i] is in region K, the number of padding bits of binary zero is equal to "(K - 1) × 2 N " plus the bitwise inverse of the lower-order "N" bits of PD E [i][N-1:0].
[0094] As described above, the relative exponent value (i.e., generating an aligned mantissa IN MAThe mantissa of [i], IN M The bit shift amount of [i] is based on the lower-order "N" bits of PD E [1][N-1:0] to PD E [n][N-1:0] for inverse determination. Therefore, a subtractor is not required in the alignment circuit 105 to calculate the relative exponent value. In some embodiments, the alignment circuit 105 includes an inverter 113 that receives the lower-order "N" bits of PD E [1][N-1:0] to PD E [n][N-1:0] to generate the lower-order "N" bits of PD E [1][N-1:0] to PD E [n][N-1:0] for the inverse to determine the number of padded bits of binary zero as described above.
[0095] In some embodiments, when the number of padded bits of binary zero is equal to or greater than the total number of bits of the mantissa IN M [i], align the mantissa IN MA [i] is set to zero.
[0096] In some embodiments, when the product PD E [i] is in region K, no alignment is performed, and the aligned mantissa IN MA [i] is directly set to zero.
[0097] Now refer to Figure 5 . Figure 5 For various embodiments according to the present disclosure, with respect to Figures 1 to 3 , Figure 4A and Figure 4B as shown, a schematic diagram of the data processing circuit 500 with respect to the configuration of the data processing circuit 100. With respect to Figures 1 to 3 , Figure 4A and Figure 4B in the embodiments, Figure 5 similar elements are designated with the same element symbols for ease of understanding.
[0098] As Figure 5 shown, the dual-mode adder 101 receives the M+N bits of each of the exponents IN E [1] to IN E [n] via n×(M+N) data lines, and receives the M+N bits of each of the exponents W E [1] to W E [n] via other n×(M+N) data lines.
[0099] In some embodiments, the maximum detector circuit 102 is coupled to the dual-mode adder 101 via n×M data lines to receive the higher-order "M" bits of PDE [1][M+N-1:N] to PD E [n][M+N-1:N] to perform a maximum product index search.
[0100] In some embodiments, the region detector circuit 104 is coupled to the maximum value detector circuit 102 via (n+1)×M data lines to receive the higher order “M” bits PD E [1][M+N-1:N] to PD E [n][M+N-1:N] and maximum value PD E-MAX [M+N-1:N] performs region classification. The data processing circuit 100 provides less data transmission instead of using the product PD E [i] (i.e., using the “M+N” data lines) to perform a maximum product index search and detect the maximum product value and product PD E [i], thereby providing lower area and energy consumption.
[0101] In some embodiments, the data line of the data processing circuit 500 is a metal line in a metal layer of a semiconductor device. For example, the data line is located in a single metal layer.
[0102] Figure 2 , Figure 3 , Figure 4A , Figure 4B and Figure 5 The configuration is for illustration purposes only. Various tools are within the contemplated scope of the present disclosure. For example, in some embodiments, a portion of the dual-modulus adder 101 is included in the arithmetic circuit 200.
[0103] See now Figure 6 and Figure 7 . Figure 6 and Figure 7 According to various embodiments of the present disclosure, Figure 1 , Figure 2 and Figure 5 The data processing circuits 100 and 500 shown in FIG. Figure 3 Schematic diagram of alignment operation tables 600 and 700 of method 300 shown in FIG. Figures 1 to 5 An embodiment of Figure 6 and Figure 7 Similar components in the drawings are designated with the same reference numerals to facilitate understanding.
[0104] In some embodiments, corresponding to Figure 1 , Figure 2 and Figure 5 The data processing circuits 100 and 500 shown in FIG. Figure 3The alignment operation of the method 300 shown is for single-phase mode or two-phase mode. In single-phase mode, the aligned mantissas IN MA [1] to IN MA [n] have the number of bits of the mantissas IN M [1] to IN M [n]. Specifically, each of the aligned mantissas IN MA [1] to IN MA [n] has the same number of bits as each of the mantissas IN M [1] to IN M [n].
[0105] An example of the alignment operation in single-phase mode is described below with reference to Table 600. In this example, the products PD E [1] to PD E [n] are classified into three regions. The number of padding bits of binary zeros for mantissa alignment is based on the bitwise inverse of the lower three bits PD E [1][2:0] to PD E [n][2:0]. Each of the mantissas IN M [1] to IN M [n] has 8 bits. Each of the aligned mantissas IN MA [1] to IN MA [n] has 8 bits. The mantissa IN M [i] includes 8 bits "1, IN 6 , IN 5 , IN 4 , IN 3 , IN 2 , IN 1 , IN 0 ".
[0106] In Table 600, the first row corresponds to the region of the product PD E [i]. The second row corresponds to the lower three bits PD E [i][2:0]. The third row corresponds to the aligned mantissa IN MA [i] and the lower three bits PD E [i][2:0] generated by the single-phase mode alignment operation corresponding to the region. For example, as shown in Table 600, when the product PD E [i] is classified into region 1 and the lower three bits PD E [i][2:0] is equal to "101", the aligned mantissa IN MA [i] is equal to the mantissa IN M [i] filled (shifted) with two bits of zeros from the MSB side, including 8 bits "0, 0, 1, IN 6 , IN5 , IN 4 , IN 3 , IN 2 ". "XXX" in Table 600 represents any value of the lower three bits PD E [i][2:0]. When the product PD E [i] is classified into Region 1, the alignment mantissa IN MA [i] is directly set to zero.
[0107] Another example of the alignment operation in single-phase mode is described below with reference to Table 700. In this example, the products PD E [1] to PD E [n] are classified into three regions. The number of padding bits of binary zeros for mantissa alignment is based on the bitwise inverse of the lower two bits PD E [1][1:0] to PD E [n][1:0]. Each of the mantissas IN M [1] to IN M [n] has 8 bits. The alignment mantissas IN MA [1] to IN MA [n] each have 8 bits. The mantissa IN M [i] includes 8 bits "1, IN 6 , IN 5 , IN 4 , IN 3 , IN 2 , IN 1 , IN 0 ".
[0108] In Table 700, the first row corresponds to the region of the product PD E [i]. The second row corresponds to the lower three bits PD E [i][1:0]. The third column corresponds to the alignment mantissa IN MA [i] generated by the single-phase mode alignment operation corresponding to the region and the lower two bits PD E [i][1:0]. For example, as shown in Table 400, when the product PD E [i] is classified into Region 1 and the lower two bits PD E [i][1:0] is equal to "10", the alignment mantissa IN MA [i] is equal to the mantissa IN M [i] filled (shifted) with one bit zero from the MSB side, including 8 bits "0, 1, IN 6 , IN 5 , IN 4 , IN 3 , IN 2, IN 1 ”. When the product PD E [i] is classified into region 2 and the lower three bits PD E [i][1:0] is equal to “10”, align the mantissa IN MA [i] is equal to the mantissa IN filled (shifted) with four bits plus one bit zero from the MSB side M [i], including 8 bits “0, 0, 0, 0, 0, 1, IN 6 , IN 5 ”.
[0109] Now refer to Figure 8 . Figure 8 For the corresponding alignment operation table 800 of the data processing circuits 100 and 500 as shown in Figure 1 , Figure 2 and Figure 5 and the method 300 as shown in Figure 3 according to various embodiments of the present disclosure. With respect to the embodiment of Figures 1 to 7 , the similar elements in Figure 8 are designated with the same element symbols for easy understanding.
[0110] Table 800 corresponds to the alignment operation in the two-phase mode. In the two-phase mode, each of the aligned mantissas IN MA [1] to IN MA [n] includes phase Ph1 and phase Ph2. Phase Ph1 includes the higher-order bits, and the bit width of these higher-order bits is equal to the bit width of each of the mantissas IN M [1] to IN M [n]. Phase Ph2 includes the extended bits for mantissa extension.
[0111] An example of the alignment operation in the two-phase mode is described below with reference to Table 800. In this example, the products PD E [1] to PD E [n] are classified into three regions. The number of filled bits of binary zeros for mantissa alignment is based on the bitwise inversion of the lower three bits PD E [1][2:0] to PD E [n][2:0]. The number of bits of each of the mantissas IN M [1] to IN M [n] is 8. The number of bits of each of the aligned mantissas IN MA [1] to IN MA [n] is 16. The mantissa IN M [i] includes 8 bits “1, IN 6 , IN 5 , IN 4 , IN3 , IN 2 , IN 1 , IN 0 ”.
[0112] In Table 800, the first row corresponds to the region of the product PD E [i]. The second row corresponds to the lower-order three-bit PD E [i][2:0]. The third and fourth rows correspond to the alignment mantissa IN MA [i] generated by the two-phase mode alignment operation corresponding to the region, the phases Ph1 and Ph2 of IN E [i], and the lower-order three-bit PD
[0113] For illustration, as shown in Table 800, when the product PD E [i] is classified into Region 1 and the lower-order three-bit PD E [i][2:0] is equal to "101", the alignment mantissa IN MA [i] is equal to the mantissa IN M [i] filled (shifted) with two-bit zeros from the MSB side. The phase Ph1 includes 8 bits "0, 0, 1, IN 6 , IN 5 , IN 4 , IN 3 , IN 2 ". The phase Ph2 includes 8 bits "IN 1 , IN 0 , 0, 0, 0, 0, 0, 0".
[0114] According to various embodiments, the single-phase mode provides higher energy efficiency and no mantissa expansion, while the two-phase mode provides better computational accuracy of the arithmetic circuit 200 with mantissa expansion.
[0115] Now refer to Figure 9A and Figure 9B , Figure 9A and Figure 9B are schematic diagrams of the floating-point operations of the system 10 according to various embodiments of the present disclosure, such as Figures 1 to 3 , Figure 4A , Figure 4B and Figures 5 to 8 shown. In some embodiments, the floating-point operations on the phases Ph1 and Ph2 (e.g., corresponding to Figure 8 ) are performed separately. In some embodiments, the floating-point operations on the phases Ph1 and Ph2 are performed in parallel with pipeline execution.
[0116] Figure 9ADepicts an example of floating-point operations performed without pipelines in a two-phase mode. For illustration, the clock CLK is the clock signal for system 10. In task 1, system 10 performs MAC operations on floating-point numbers IN[1] to IN[n] and floating-point numbers W[1] to W[n]. "EXP" represents an exponential operation, which includes the following steps: detecting the maximum exponent corresponding to the maximum exponent detector circuit 102 in Figure 2 ; the input mantissa shift bit amount; and shifting the input mantissa corresponding to the alignment circuit 105 in Figure 2 . "MAN" represents the mantissa product-sum operation corresponding to the CIM array 202 in Figure 1 . "ADD" represents the mantissa accumulation operation corresponding to the accumulator circuit 204 in Figure 1 .
[0117] As shown in Figure 9A , during the time period from time t1 to time t3, system 10 performs the exponential operation EXP on floating-point numbers IN[1] to IN[n], performs the mantissa product-sum operation MAN on the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] in phase Ph1, and performs the mantissa accumulation operation ADD on the result of the mantissa product-sum operation MAN and the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] in phase Ph1. System 10 performs operations from time t1 to time t3 to generate the MAC result for phase Ph1.
[0118] Then, during the time period from time t3 to time t5, system 10 performs the mantissa product-sum operation on the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] in phase Ph2, and performs the mantissa accumulation operation on the result of the mantissa product-sum operation MAN and the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] in phase Ph2. System 10 performs operations from time t3 to time t5 to generate the MAC result for phase Ph2. System 10 further generates the complete MAC result for task 1 based on the MAC results of phases Ph1 and Ph2.
[0119] Figure 9BAn example of floating-point operation and pipeline execution (simultaneous execution) in a two-phase mode is depicted. For example, in Task 1, System 10 performs MAC operations on floating-point numbers IN[1] to IN[n] and floating-point numbers W[1] to W[n]. In Task 2, System 10 performs MAC operations on floating-point numbers IN[n + 1] to IN[2n] and floating-point numbers W[n + 1] to W[2n].
[0120] As Figure 9B shown, within the time period from time t1 to time t2, System 10 performs an exponential operation EXP on floating-point numbers IN[1] to IN[n], and performs a mantissa product-sum operation on the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] of phase Ph1. From time t2 to time t3, System 10 performs a mantissa accumulation operation ADD on the result of the mantissa product-sum operation MAN and the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] of phase Ph1. Within the time period from time t2 to time t3, System 10 further performs a mantissa product-sum operation MAN on the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] of phase Ph2, while performing a mantissa accumulation operation ADD on the mantissa product-sum operation and the result of phase Ph1.
[0121] From time t3 to time t4, System 10 performs an exponential operation EXP on floating-point numbers IN[n + 1] to IN[2n], and performs a mantissa product-sum operation MAN on the aligned mantissas IN MA [n + 1] to IN MA [2n] and mantissas W M [n + 1] to W M [2n] of phase Ph1. Within the time period from time t3 to time t4, System 10 further performs a mantissa accumulation operation ADD on the result of the mantissa product-sum operation MAN and the aligned mantissas IN MA [1] to IN MA [n] and mantissas W M [1] to W M [n] of phase Ph2, while performing a mantissa accumulation operation ADD on floating-point numbers IN[n + 1] to IN[2n], the mantissa product-sum operation MAN, and the aligned mantissas IN MA [n + 1] to IN MA [2n] and mantissas W M[n + 1] to W M Perform exponentiation EXP on the Ph1 phase of [2n]. For the mantissa product sum operation MAN and the aligned mantissa IN MA [1] to IN MA [n] and the mantissa W M [1] to W M After performing the mantissa accumulation operation ADD on the result of the phase Ph2 of [n], the system 10 further generates the complete MAC result of task 1 based on the MAC results of phases Ph1 and Ph2.
[0122] As Figure 9A and Figure 9B shown, in some embodiments, compared to a run without pipeline execution, a run with pipeline execution provides less latency. For illustration, Figure 9B The pipelined task 1 shown is completed earlier than Figure 9A the non - pipelined task 1 shown.
[0123] Figures 6 to 8 , Figure 9A and Figure 9B The configurations of... are for illustrative purposes only. Various tools are within the scope contemplated by the present disclosure embodiments. For example, in some embodiments, it takes two clock cycles to complete the mantissa accumulation operation ADD, as Figure 9A and Figure 9B shown.
[0124] Now refer to Figure 10 , Figure 11A and Figure 11B . Figure 10 For various embodiments according to the present disclosure, such as Figures 1 to 8 , Figure 9A and Figure 9B shown, is a schematic diagram of the region detector circuit 104 and the arithmetic circuit 200. Figure 11A and Figure 11B For various embodiments according to the present disclosure, such as Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , Figure 9A and Figure 9B shown, is a schematic diagram of the operation of the arithmetic circuit 200. Relative to Figures 1 to 8 , Figure 9A and Figure 9B embodiments, for ease of understanding, Figure 10 the similar elements in... are designated with the same element symbols.
[0125] For illustration, in some embodiments, the region detector circuit 104 further receives a signal MODE indicating a floating-point mode or an integer (INT) mode. In the floating-point mode, the region detector circuit 104 performs exponent product region classification of floating-point numbers as described in the previous paragraphs.
[0126] In the INT mode, the region detector circuit 104 further receives an integer IN. In some embodiments, the arithmetic circuit 200 performs MAC operations on the integer IN and the weights stored in the CIM array 202. In some embodiments, the integer IN is processed, and the arithmetic circuit 200 performs MAC operations on the processed integer IN and the weights stored in the CIM array 202.
[0127] In the INT mode, the region detector circuit 104 performs sparsity detection to compare the integer IN with zero. When the integer IN is less than or equal to zero, the region detector circuit 104 outputs a signal SPAR having a first value (e.g., a binary value) to indicate sparse-aware activation. When the integer IN is greater than zero, the region detector circuit 104 outputs a signal SPAR having a second value different from the first value (e.g., binary zero) to indicate sparse-aware deactivation. In some embodiments, the arithmetic circuit 200 operates according to the signal SPAR. In some embodiments, the arithmetic circuit 200 disables some parts of the arithmetic circuit 200 in response to sparse-aware activation. In some embodiments, the region detector circuit 104 generates an output with a number of zeros of the arithmetic circuit 200 as an input to integer operations according to sparse-aware activation.
[0128] See Figure 11A and Figure 11B , in the INT mode, the arithmetic circuit 200 performs integer operations that include MAC operations and an accumulation operation ADD on the results of the MAC operations. As Figure 11A shown, in some embodiments, the integer operations for tasks 1 and 2 corresponding to different integer inputs are performed sequentially. As Figure 11B shown, in some embodiments, the integer operations for tasks 1 and 2 corresponding to different integer inputs are pipelined (performed simultaneously).
[0129] Now see Figure 12A and Figure 12B . Figure 12A For various embodiments according to the present disclosure corresponding to as Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , Figure 9A , Figure 9B , Figure 10 , Figure 11A and Figure 11BSchematic diagram of an example of the floating mode of the system 10 shown. Figure 12B For corresponding to various embodiments according to the present disclosure, as Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , Figure 9A , Figure 9B , Figure 10 , Figure 11A and Figure 11B Schematic diagram of an example of the floating-point mode of the system 10 shown. Relative to Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , Figure 9A , Figure 9B , Figure 10 , Figure 11A and Figure 11B embodiments, Figure 12A and Figure 12B Similar elements in are designated with the same element symbols for ease of understanding.
[0130] As Figure 12A shown, in the floating point (FP) mode, the system 10 first performs an exponentiation operation EXP to calculate the product of exponents and / or detect the maximum exponent. Then the system 10 performs a mantissa alignment operation ALIGN to generate an aligned mantissa.
[0131] In the two-phase mode, each aligned mantissa includes a phase Ph1 and a phase Ph2. Taking the aligned mantissa IN MA with a bit width of 16 as an example, the higher-order 8 bits IN MA [15:8] correspond to the phase Ph1, and the lower-order 8 bits IN MA [7:0] correspond to the phase Ph1.
[0132] In the two-phase mode, the system 10 performs a multiplication operation MULT and an addition operation ADD on the phases Ph1 and Ph2 respectively to generate a partial MAC result pMAC Ph1 and a partial MAC result pMAC Ph2 . In some embodiments, the partial MAC result pMAC Ph1 and the partial MAC result pMAC Ph2 are combined to generate a more accurate MAC result.
[0133] As Figure 12BAs shown, in integer (INT) mode, the system 10 first performs a preprocessing operation PRE on the integer data. Then the system 10 performs sparsity detection DET. When sparsity sensing is in an inactive state, the system 10 performs a multiplication operation MULT and an addition operation ADD on the input integer.
[0134] In some embodiments, hardware for floating point mode and integer mode is reused. For example, the circuit for exponential operation EXP in floating point mode is reused in preprocessing operation PRE in integer mode. The circuit for alignment operation ALIGN in floating point mode is reused in sparsity detection DET in integer mode.
[0135] Figure 10 , Figure 11A , Figure 11B , Figure 12A and Figure 12B The configuration is for illustration purposes only. Various tools are contemplated within the scope of the present disclosure. For example, in some embodiments, the region detector circuit 104 outputs the signal SPAR to an input processing circuit coupled between the region detector circuit 104 and the operation circuit 200.
[0136] See now Figure 13 . Figure 13 For various embodiments according to the present disclosure, Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , Figure 9A , Figure 9B , Figure 10 , Figure 11A , Figure 11B , Figure 12A and Figure 12B Schematic diagram of system 20 configured in system 10 shown. Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , Figure 9A , Figure 9B , Figure 10 , Figure 11A , Figure 11B , Figure 12A and Figure 12B An embodiment of Figure 13 Similar components in the drawings are designated with the same reference numerals for easier understanding.
[0137] Compared with the region detector circuit 104 of system 10, the region detector circuit 104 of system 20 further includes a region bias circuit 104a and a region detector 104b. Compared with the alignment circuit 105 of system 10, the alignment circuit 105 of system 20 further includes a plurality of input processing circuits 105a. Compared with the arithmetic circuit 200 of system 10, the arithmetic circuit 200 of system 20 further includes a plurality of calculation circuits 1301.
[0138] As Figure 13 shown, the region bias circuit 104a is coupled to the maximum detector circuit 102 and the region detector 104b. The region detector 104b is coupled to each input processing circuit 105a and each calculation circuit 1301. Each input processing circuit 105a is coupled to a corresponding calculation circuit 1301.
[0139] In some embodiments, the maximum detector circuit 102 outputs PD E-MAX [M+N-1:N] to the region bias circuit 104a. The region bias circuit 104a generates a reference PD E-REF1 to PD E-REFK .
[0140] Each calculation circuit 1301 outputs one of the higher-order "M" bit PD E [1][M+N-1:N] to PD E [N][M+N-1:N] to the region detector 104b. The region detector 104b performs the region detection described in the paragraph corresponding to step s3 of method 300 to generate a region flag ZFG[i] by comparing the higher-order "M" bit PD E [i][M+N-1:N] with PD E-REF1 to PD E-REFK .
[0141] The input processing circuit 105a receives the region flag ZFG[i], the mantissa IN M [i], the exponent IN E [i], and the lower-order bit PD E [i][N-1:0]. The input processing circuit 105a performs the alignment of the mantissa IN M [i] as described in the paragraph corresponding to step s4 of method 300 to generate an aligned mantissa IN E [i] according to the region flag ZFG[i] and the lower-order bit PD MA [i][N-1:0]. The input processing circuit 105a outputs the aligned mantissa IN MA [i] and the exponent IN E [i] to the corresponding calculation circuit 1301. In some embodiments, the calculation circuit 1301 performs the mantissa INMA [i] and the mantissa W of the weight W M of [i].
[0142] In some embodiments, the region bias circuit 104a receives the signal MODE. When the signal MODE indicates the INT mode, the region bias circuit 104a sets the reference to zero. The region detector 104b compares the zero reference with the input integer to output the signal SPAR.
[0143] Now refer to Figure 14 . Figure 14 is a schematic diagram of the system 20 according to various embodiments of the present disclosure. Relative to Figure 13 , Figures 1 to 3 , Figure 4A , Figure 4B , Figures 5 to 8 , [[ID= , , , , , , and embodiments, similar elements in
[0144] are designated with the same element symbols as those in for ease of understanding.
[0145] Compared with the system 10 shown in
[0146] as shown, the dual-mode adder 101 of the system 20 is included in the computing circuit 1301 instead of being included in the data processing circuit 100. In some embodiments, each computing circuit 1301 includes a dual-mode adder 101. In some embodiments, each computing circuit 1301 further includes a plurality of multipliers 1401.
[0147] Now further refer to . is a schematic diagram of the system 20 according to various embodiments of the present disclosure as shown in and Flowchart of the operation method 1500 of the system 20 shown. It can be understood that additional steps may be provided before, during, and after the steps shown, and for additional embodiments of the method 1500, some of the steps described below may be replaced or eliminated. The order of the steps may be interchanged. In various views and illustrative embodiments, like element symbols are used to designate like elements. The method 1500 includes the following steps 1501 to 1505 described with reference to the system 20 as shown in and shown.
[0148] In step 1501, the dual-mode adder 101 performs an addition between the exponent IN E [i] of the input IN[i] and the exponent W E [i] of the weight W[i], as described in the paragraph corresponding to step s1 of the method 300, to generate the product PD E [i]. The computing circuit 1301 outputs the product PD E [i] to the data processing circuit 100.
[0149] In step 1502, the data processing circuit 100 aligns the mantissa IN E [i] based on the product PD M [i]. In some embodiments, the data processing circuit 100 outputs the aligned mantissa IN MA [i] to the multiplier 1401.
[0150] In step 1503, the computing circuit 1301 calculates the product PD MA [i] between the aligned mantissa IN M [i] and the mantissa W M [i] of the weight W[i]. The multiplier 1401 performs a multiplication between the aligned mantissa IN MA [i] and the mantissa W M [i] to generate the product PD M [i].
[0151] In step 1504, in some embodiments, the computing circuit 1301 outputs the products PD M [1] to PD M [L] to the adder tree 1402. The adder tree accumulates the products PD M [1] to PD M [L] to generate the result MAC M of the MAC operation of the mantissa of the aligned mantissa and the mantissa of the weight.
[0152] In step 1505, the maximum value PD E-MAX and the result MAC MCombines to produce a MAC result of inputs IN[1] to IN[L] and weights W[1] to W[L].
[0153] The configurations of to are for illustrative purposes only. Various tools are within the scope contemplated by this disclosure. For example, in some embodiments, each input processing circuit 105a includes an inverter 113.
[0154] As described above, this disclosure provides a system, circuit, and method for performing a mantissa alignment operation on floating-point numbers. The proposed maximum product exponent search for mantissa alignment uses a portion rather than all bits of the exponent product. Using fewer bits simplifies routing and power consumption. In addition, using the proposed region detector circuit, mantissa alignment is performed based on relative exponent values, and the mantissa bits are shifted via the direct inverse of a portion of the product exponent. Compared with some methods, the proposed circuit and method reduce data transmission by approximately 37%. Compared with some methods, the proposed circuit and method reduce the mantissa alignment power consumption by approximately 39% and the area usage by approximately 14%.
[0155] In some embodiments, a data processing circuit is provided. The circuit includes a dual-mode adder, a maximum value detector circuit, a region detector circuit, and an alignment circuit. The dual-mode adder produces a product between a first exponent of a first floating-point number and a second exponent of a second floating-point number. The maximum value detector circuit detects a maximum value in a first portion of the product. The region detector circuit classifies the first portion into regions by comparing the first portion with the maximum value. The alignment circuit aligns a first mantissa of the first floating-point number based on the regions and a second portion of the product to produce an aligned mantissa for floating-point operations.
[0156] In some embodiments, the maximum value detector circuit further includes a register and a comparator circuit. The register stores a temporary maximum value. The comparator circuit compares the first portion with the temporary maximum value and updates the temporary maximum value based on these comparisons, where the temporary maximum value is used as the maximum value after the comparisons are completed.
[0157] In some embodiments, the comparator circuit is an M-bit comparator, where the number M is equal to the number of bits of each first portion, and the number M is less than the number of bits of each product.
[0158] In some embodiments, the region detector circuit further includes a subtractor circuit. The subtractor circuit performs a subtraction between the first portion and the maximum value to produce a reference, and classifies the first portion by comparing the first portion with the maximum value and the reference.
[0159] In some embodiments, the circuit further includes metal lines having a number M+1, which are coupled between the maximum detector circuit and the region detector circuit to transmit the first part and the maximum value, wherein the number M+1 of the metal lines is less than the number of bits of each product.
[0160] In some embodiments, the alignment circuit further includes an inverter circuit. The inverter circuit inverts the second part to generate a number of shift bits, and generates an aligned mantissa according to the number of shift bits.
[0161] In some embodiments, each inversion includes a binary number, which corresponds to a respective one of the number of shift bits.
[0162] In some embodiments, each aligned mantissa includes a first bit, the first bit width of these first bits being equal to the bit width of each first mantissa, and includes a second bit as an extended bit.
[0163] In some embodiments, the region detector circuit further receives an integer, and generates a sparse awareness signal according to a comparison between the integer and a zero number.
[0164] In some embodiments, when the integer is less than zero, the region detector circuit is further used to generate a sparse signal to disable a part of the memory device that performs integer operations on the integer.
[0165] In some embodiments, a data processing system is provided. The system includes a data processing circuit and a memory device. The data processing circuit includes a first comparator circuit, a second comparator circuit, and an alignment circuit. The first comparator circuit detects a maximum value in a first part of a product between a first activation and a weight, wherein the first activation and the weight are floating-point numbers. The second comparator circuit compares the first part of the product with a reference to generate a region flag. The reference is based on the maximum value. The alignment circuit aligns each first mantissa according to the corresponding region flag of the region flag to generate an aligned mantissa. The memory device includes a compute-in-memory (CIM) array. The CIM array performs a multiply-accumulate operation between the aligned mantissa and a second mantissa of the weight.
[0166] In some embodiments, the second comparator circuit compares a zero point with a second activation to determine whether the integer is less than zero, wherein these second activation points are integers, and when the integer is less than zero, the second comparator circuit generates a signal to disable a part of the memory device.
[0167] In some embodiments, the system further includes a register. The register stores a temporary maximum value, wherein the first comparator circuit compares the first part with the temporary maximum value, and updates the temporary maximum value according to these comparisons, wherein the temporary maximum value is used as the maximum value after the comparison is completed.
[0168] In some embodiments, the alignment circuit aligns each first mantissa according to the corresponding region flag and the corresponding second part of the product.
[0169] In some embodiments, each aligned mantissa includes a first bit, the width of the first bit being equal to the bit width of each first mantissa, and includes a second bit as an extended bit, wherein the memory device further performs a multiply-accumulate operation on the first bit and the second bit in parallel.
[0170] In some embodiments, a data processing method is provided. The method includes the steps of: adding a first exponent of a first floating-point number and a second exponent of a second floating-point number to generate a product; identifying a maximum value in a first part of the product; generating a region flag corresponding to the product according to the maximum value, wherein each region flag indicates a region to which the corresponding product is classified; and aligning a first mantissa of the first floating-point number according to the region flag and a second part of the product to generate an aligned mantissa.
[0171] In some embodiments, the step of generating the region flag includes the steps of: determining a reference according to the maximum value and the bit width of each second part; and comparing the reference with the product to generate the region flag.
[0172] In some embodiments, the reference has a tolerance of 2 N , where N corresponds to the bit width of each second part.
[0173] In some embodiments, the step of aligning the first mantissa includes the steps of: determining the inverse of the second part; and aligning the first mantissa according to the inverse and the region flag.
[0174] In some embodiments, the step of determining the maximum value includes the steps of: sequentially comparing one of the first parts with a temporary maximum value; and updating the temporary maximum value according to the comparison, wherein the temporary maximum value is used as the maximum value after comparing all the first parts.
[0175] The features of several embodiments are outlined above so that those skilled in the art can better understand various aspects of the present disclosure. Those skilled in the art should understand that those skilled in the art can easily use the present disclosure as a basis for designing or modifying other processes and structures to achieve the same purposes and / or achieve the same advantages as the embodiments described herein. Those skilled in the art should also recognize that these equivalent structures do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and alterations can be made to these equivalent structures without departing from the spirit and scope of the present disclosure.
Claims
1. A data processing circuit, characterized in that: Include: a dual-modulus adder for generating a plurality of products between a plurality of first exponents of a plurality of first floating-point numbers and a plurality of second exponents of a plurality of second floating-point numbers; a maximum value detector circuit for detecting a maximum value in the plurality of first portions of the plurality of products; a region detector circuit for classifying the plurality of first portions into a plurality of regions by comparing the plurality of first portions with the maximum value; and An alignment circuit is used for aligning a plurality of first mantissas of the plurality of first floating point numbers according to the plurality of regions and a plurality of second parts of the plurality of products to generate a plurality of aligned mantissas for a floating point operation.
2. The circuit according to claim 1, characterized in that Wherein the maximum value detector circuit further comprises: A register for storing a temporary maximum value; and a comparator circuit for performing a plurality of comparisons between the plurality of first portions and the temporary maximum value, and updating the temporary maximum value according to the plurality of comparisons, wherein the temporary maximum value is used as the maximum value after the plurality of comparisons are completed, The comparator circuit is an M-bit comparator, wherein the number M is equal to the number of bits of each of the plurality of first parts, and wherein the number M is less than the number of bits of each of the plurality of products.
3. The circuit according to claim 1, characterized in that The region detector circuit further comprises: a subtractor circuit for performing a plurality of subtractions between the plurality of first portions and the maximum value to generate a plurality of references, The region detector circuit classifies the first portions by comparing the first portions with the maximum value and the references.
4. The circuit according to claim 1, characterized in that Further including: A plurality of metal lines having a number (M+1) and coupled between the maximum value detector circuit and the region detector circuit to transmit the plurality of first portions and the maximum value, wherein the number (M+1) of the plurality of metal lines is less than the number of bits of each of the plurality of products.
5. The circuit according to claim 1, characterized in that Wherein the alignment circuit further comprises: an inverter circuit for performing a plurality of inversions on the plurality of second parts to generate a plurality of shift bit numbers, and generating the plurality of aligned mantissas according to the plurality of shift bit numbers, Each of the plurality of inversions comprises a binary number corresponding to a corresponding one of the plurality of shift bit numbers.
6. The circuit according to claim 1, characterized in that, The region detector circuit is further configured to receive an integer and generate a sparse sensing signal according to a comparison between the integer and zero. When the integer is less than zero, the region detector circuit is further configured to generate the sparse sensing signal to enable a portion of a memory device to perform an integer operation on the integer.
7. A data processing system, characterized in that: Include: A data processing circuit comprising: a first comparator circuit for detecting a maximum value among first portions of products of first activations and weights, wherein the first activations and weights are floating point numbers; a second comparator circuit for comparing the first portions of the products with a plurality of references to generate a plurality of region flags, wherein the plurality of references are based on the maximum value; and an alignment circuit for aligning each of the plurality of first activated first mantissas according to a corresponding region mark of the plurality of region marks to generate a plurality of aligned mantissas; and A memory device comprising: An in-memory computation array is used to perform a product-sum operation between the plurality of aligned mantissas and the plurality of second mantissas of the plurality of weights.
8. The system according to claim 7, characterized in that wherein each of the plurality of alignment mantissas comprises a plurality of first bits, a first bit width of the plurality of first bits is equal to a bit width of each of the plurality of first mantissas, and comprises a plurality of second bits as a plurality of extension bits, The memory device is further used to perform the product-sum operation on the first bits and the second bits simultaneously.
9. A method of data processing, characterized in that: The following steps are involved: adding a plurality of first exponents of the plurality of first floating point numbers and a plurality of second exponents of the plurality of second floating point numbers to generate a plurality of products; identifying a maximum value among the plurality of first portions of the plurality of products; generating a plurality of region flags corresponding to the plurality of products according to the maximum value, wherein each of the plurality of region flags indicates a region to which a corresponding one of the plurality of products is classified; and A plurality of first mantissas of the plurality of first floating point numbers are aligned according to the plurality of region flags and a plurality of second parts of the plurality of products to generate a plurality of aligned mantissas.
10. The method according to claim 9, characterized in that The step of generating the plurality of area marks comprises the following steps: determining a plurality of references according to the maximum value and a bit width of each of the plurality of second parts; and comparing the plurality of references with the plurality of products to generate the plurality of region flags, Wherein the plurality of references have a tolerance 2 N , where N corresponds to the bit width of each of the plurality of second portions.