Processors and methods of operating processors

CN114546333BActive Publication Date: 2026-09-25SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111325168.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-23
Filing Date
2021-11-10
Publication Date
2026-09-25
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

[0003]用于神经网络的处理器可执行大量的乘法运算和加法运算,因为正在处理的数字的很大一部分可能相对较小,只有一小部分离群值(outlier value)可能相对较大,所以乘法运算和加法运算中的一些可能是对处理资源的不良使用

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114546333B_ABST
    Figure CN114546333B_ABST
Patent Text Reader

Abstract

Processors and methods of operating processors are disclosed. In some embodiments, the method includes forming a first set of products and forming a second set of products. Forming the first set of products can include multiplying a first activation value by a first least significant subword, a second least significant subword, and a most significant subword in a first multiplier, a second multiplier, and a third multiplier, and adding resulting first partial products and resulting second partial products. Forming the second set of products can include forming a first floating point product, forming the first floating point product including multiplying a first subword of a mantissa of the activation value by a first subword of a mantissa of the weight in the first multiplier to form a third partial product.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority and benefit to U.S. Provisional Application No. 63 / 112,299, filed November 11, 2020, entitled “System and method for improving area and power efficiency by redistributing weight nibbles and supporting FP16,” and U.S. Application No. 17 / 133,288, filed December 23, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0002] One or more aspects of embodiments of this disclosure relate to processing circuitry, and more specifically, to a processor and a method of operating the processor. Background Technology

[0003] Processors used for neural networks can perform a large number of multiplication and addition operations. Because a large portion of the numbers being processed may be relatively small, and only a small subset of outlier values ​​may be relatively large, some of these multiplication and addition operations could be an inefficient use of processing resources. Furthermore, some operations in such systems can be integer operations, and some can be floating-point operations. If these operations were performed on separate dedicated hardware sets, they would consume significant amounts of chip area and power.

[0004] Therefore, there is a need for a system and method for performing multiple sets of multiplications in a manner that accommodates outliers and is capable of performing both integer and floating-point operations. Summary of the Invention

[0005] According to an embodiment of the present invention, a method is provided, the method comprising: forming a first set of products, each product in the first set of products being an integer product of a first activation value and a corresponding weight in a first plurality of weights; and forming a second set of products, each product in the second set of products being a floating-point product of a second activation value and a corresponding weight in a second plurality of weights, each weight in the first plurality of weights including a least significant subword and a most significant subword, the most significant subword of the first weight in the first plurality of weights being non-zero, and the most significant subword of the second weight in the first plurality of weights being zero, the step of forming the first set of products comprising: in a first multiplier The steps of multiplying the first activation value with the least significant subword of the first weight to form a first partial product, multiplying the first activation value with the least significant subword of the second weight in a second multiplier, multiplying the first activation value with the most significant subword of the first weight in a third multiplier to form a second partial product, and adding the first partial product and the second partial product to form a second set of products include forming a first floating-point product. The steps of forming the first floating-point product include multiplying the first subword of the mantissa of the second activation value with the first subword of the mantissa of the first weight in the second plurality of weights in a first multiplier to form a third partial product.

[0006] In some embodiments, the first multiplier is configured to receive a first independent variable and a second independent variable, the first independent variable having a first independent variable size, the second independent variable having a second independent variable size, and the first independent variable size being greater than the second independent variable size.

[0007] In some embodiments, the step of forming a first floating-point product includes: receiving a first independent variable by a first multiplier; receiving a second independent variable by a first multiplier; and multiplying the first independent variable by the second independent variable, wherein the first independent variable includes a first subword of the mantissa of a second activation value and a second subword of the mantissa of the second activation value, and the second independent variable includes a first subword of the mantissa of a first weight among a second plurality of weights.

[0008] In some embodiments, the step of forming a first floating-point product includes: receiving a first independent variable by a first multiplier; receiving a second independent variable by a first multiplier; and multiplying the first independent variable by the second independent variable, wherein the first independent variable includes a first subword of the mantissa of a first weight in a second plurality of weights and a second subword of the mantissa of the first weight in a second plurality of weights, and the second independent variable includes a first subword of the mantissa of a second activation value.

[0009] In some embodiments, the second multiplier is configured to receive a first argument having a first size and a second argument having a second size, wherein the first size is greater than the second size.

[0010] In some embodiments, the step of forming a second set of products further includes forming a second floating-point product, the step of forming a second floating-point product including: receiving a first independent variable by a second multiplier; receiving a second independent variable by a second multiplier; and multiplying the first independent variable received by the second multiplier with the second independent variable received by the second multiplier, wherein the first independent variable received by the second multiplier includes a first subword of the mantissa of a first activation value and a subword consisting of zeros, and the second independent variable received by the second multiplier includes a third subword of the mantissa of a first weight in a second plurality of weights.

[0011] In some embodiments, the step of adding the first partial product and the second partial product includes performing an offset addition in a first offset adder.

[0012] In some embodiments, the step of forming the second set of products further includes forming a second floating-point product, the step of forming the second floating-point product including: multiplying the first subword of the second activated mantissa with the second subword of the mantissa of the first weight in the second plurality of weights in a third multiplier.

[0013] In some embodiments, the method further includes adding the first floating-point product and the second floating-point product.

[0014] In some embodiments, the step of adding the first floating-point product and the second floating-point product includes performing an offset addition in a first offset adder.

[0015] According to an embodiment of the present invention, a system is provided, the system comprising: a processing circuit including a first multiplier, a second multiplier, and a third multiplier, the processing circuit being configured to: form a first set of products and form a second set of products, each product in the first set of products being an integer product of a first activation value and a corresponding weight in a first plurality of weights, each product in the second set of products being a floating-point product of a second activation value and a corresponding weight in a second plurality of weights, each weight in the first plurality of weights including a least significant word and a most significant word, the most significant word of the first weight in the first plurality of weights being non-zero, and the most significant word of the second weight in the first plurality of weights being zero. The steps for forming the first set of products include: multiplying the first activation value with the least significant subword of the first weight in a first multiplier to form a first partial product; multiplying the first activation value with the least significant subword of the second weight in a second multiplier; multiplying the first activation value with the most significant subword of the first weight in a third multiplier to form a second partial product; and adding the first partial product and the second partial product together. The steps for forming the second set of products include forming a first floating-point product. The steps for forming the first floating-point product include: multiplying the first subword of the mantissa of the second activation value with the first subword of the mantissa of the first weight in the second plurality of weights in a first multiplier to form a third partial product.

[0016] In some embodiments, the first multiplier is configured to receive a first independent variable and a second independent variable, the first independent variable having a first independent variable size, the second independent variable having a second independent variable size, and the first independent variable size being greater than the second independent variable size.

[0017] In some embodiments, the step of forming a first floating-point product includes: receiving a first independent variable by a first multiplier; receiving a second independent variable by a first multiplier; and multiplying the first independent variable by the second independent variable; wherein the first independent variable includes a first subword of the mantissa of a second activation value and a second subword of the mantissa of the second activation value, and the second independent variable includes a first subword of the mantissa of a first weight in a second plurality of weights.

[0018] In some embodiments, the step of forming a first floating-point product includes: receiving a first independent variable by a first multiplier; receiving a second independent variable by a first multiplier; and multiplying the first independent variable by the second independent variable; wherein the first independent variable includes: a first subword of the mantissa of a first weight in a second plurality of weights and a second subword of the mantissa of the first weight in a second plurality of weights, and the second independent variable includes: a first subword of the mantissa of a second activation value.

[0019] In some embodiments, the second multiplier is configured to receive a first argument having a first size and a second argument having a second size, wherein the first size is greater than the second size.

[0020] In some embodiments, the step of forming a second set of products further includes forming a second floating-point product, the step of forming a second floating-point product including: receiving a first independent variable by a second multiplier; receiving a second independent variable by a second multiplier; and multiplying the first independent variable received by the second multiplier with the second independent variable received by the second multiplier; the first independent variable received by the second multiplier includes: a first subword of the mantissa of a first activation value and a subword consisting of zeros; the second independent variable received by the second multiplier includes: a third subword of the mantissa of a first weight in a second plurality of weights.

[0021] In some embodiments, the step of adding the first partial product and the second partial product includes performing an offset addition in a first offset adder.

[0022] In some embodiments, the step of forming the second set of products further includes forming a second floating-point product, the step of forming the second floating-point product including: multiplying the first subword of the second activated mantissa with the second subword of the mantissa of the first weight in the second plurality of weights in a third multiplier.

[0023] In some embodiments, the processing circuitry is further configured to add the first floating-point product and the second floating-point product.

[0024] According to an embodiment of the present invention, a system is provided, the system comprising: a processing means including a first multiplier, a second multiplier, and a third multiplier, the processing means being configured to: form a first set of products and form a second set of products, each product in the first set of products being an integer product of a first activation value and a corresponding weight in a first plurality of weights, each product in the second set of products being a floating-point product of a second activation value and a corresponding weight in a second plurality of weights, each weight in the first plurality of weights including a least significant subword and a most significant subword, the most significant subword of the first weight in the first plurality of weights being non-zero, and the most significant subword of the second weight in the first plurality of weights being non-zero. The steps for forming the first product when the subword is zero include: multiplying the first activation value with the least significant subword of the first weight in the first multiplier to form a first partial product; multiplying the first activation value with the least significant subword of the second weight in the second multiplier; multiplying the first activation value with the most significant subword of the first weight in the third multiplier to form a second partial product; and adding the first partial product and the second partial product to form the second product. The steps for forming the first floating-point product include: multiplying the first subword of the mantissa of the second activation value with the first subword of the mantissa of the first weight in the second plurality of weights in the first multiplier to form a third partial product. Attached Figure Description

[0025] These and other features and advantages of this disclosure will be appreciated and understood by referring to the specification, claims and drawings, in which:

[0026] Figure 1 This is a block diagram of a portion of a neural network processor according to embodiments of the present disclosure;

[0027] Figure 2A This is a block diagram of a portion of a hybrid processing circuit according to an embodiment of the present disclosure;

[0028] Figure 2B This is a data mapping diagram according to embodiments of the present disclosure;

[0029] Figure 3A It is a table of floating-point mantissa nibbles according to embodiments of the present disclosure; and

[0030] Figure 3B This is a block diagram of a portion of a hybrid processing circuit according to an embodiment of the present disclosure. Detailed Implementation

[0031] The specific embodiments described below with reference to the accompanying drawings are intended as exemplary descriptions of a processor for fine-grain sparse integer and floating-point arithmetic provided in this disclosure, and are not intended to represent the only form in which this disclosure may be constructed or utilized. The description, in conjunction with the illustrated embodiments, illustrates the features of this disclosure. However, it should be understood that the same or equivalent functionality and structure may be implemented by different embodiments, which are also intended to be included within the scope of the disclosure. As indicated elsewhere herein, the same element reference numerals are intended to indicate the same elements or features.

[0032] Neural networks (e.g., when performing inference) perform a large amount of computation in which activations (or "activation values") (elements of the input feature map (IFM)) are multiplied by weights. The product of activations and weights can form a multidimensional array, which can be summed along one or more axes to form an array, or "tensor," that may be called an output feature map (OFM). (See reference...) Figure 1 Specialized hardware can be used to perform such calculations. Activations can be stored in static random access memory (SRAM) 105 and fed into a multiplier-accumulator (MAC) array, which may include (i) multiple blocks (which may be referred to as “bricks”) 110 (each block may include multiple multipliers (or cells) for multiplying activations by weights), (ii) one or more adder trees for summing the products generated by the bricks, and (iii) one or more accumulators for accumulating the sums generated by the adder trees. Each activation value can be broadcast to... Figure 1 The representation shows multiple multipliers conceptually arranged in rows. Multiple adder trees 115 can be used to form sums.

[0033] In computation, weights may fall within a numerical range, and the distribution of weight values ​​is such that relatively small weights are significantly more common than relatively large weights. For example, if each weight is represented as an 8-bit number, many weights (e.g., most weights, or more than 3 / 4 of the weights) may have values ​​less than 16 (i.e., the most significant nibble is zero); weights with a non-zero most significant nibble can then be referred to as "outliers." In some embodiments, appropriately constructed hardware can leverage these characteristics of weights to achieve improved speed and power consumption.

[0034] Figure 2A A portion of the mixed-signal processing circuitry is shown (it's called "mixed" because it handles both integer and floating-point operations). See reference. Figure 2AIn some embodiments, multiple multipliers 205, 210 are used to multiply weights by activations (e.g., one "half-byte" at a time). Each multiplier may be an 8×4 (i.e., 8-bit × 4-bit) multiplier, having a first input configured to receive an activation byte (which may be broadcast to all multipliers) and a second input configured to receive a corresponding weight half-byte. Figure 2A An embodiment includes four multipliers (which may be referred to as standard multipliers 205) and one reserve multiplier 210 (in some embodiments, there are more or fewer standard multipliers 205, or more reserve multipliers), two offset adders 215, four shifters 220, and four output multiplexers 225. Each of the standard multipliers 205 and the reserve multipliers 210 may be an 8×4 (i.e., 8-bit × 4-bit) multiplier. The standard multipliers 205 and the reserve multipliers 210 may be the same circuitry, differing only in how they are connected (e.g., Figure 2A (as shown in the diagram) and how they are used in the operation (as discussed in further detail below). Thus, each of these multipliers 205, 210 can receive a first argument (e.g., an activation byte) and a second argument (e.g., a weight nibble), the first argument having a first argument size (e.g., 8 bits (8b), the size of the activation byte) and the second argument having a second argument size (e.g., 4 bits, the size of the weight nibble).

[0035] Figure 2A The embodiment can be used to perform integer multiplication as follows. Weights are fed from weight buffer 230 (only its output line is shown) to the multipliers one "nibble" at a time, and an 8-bit activation value is broadcast to all multipliers 205, 210. Weights with a non-zero most significant nibble can be treated differently from weights with a zero most significant nibble (as used herein, a "zero nibble" (e.g., a "zero most significant nibble") is a nibble with a zero value). Figure 2A In the example, the leftmost multiplier forms (i) the product of an 8-bit activation value and (ii) a first weight (which has a zero most significant nibble). In the leftmost standard multiplier 205, the least significant nibble L0 of the first weight is multiplied by the 8-bit activation value to produce a 12-bit (12b) product. This 12-bit product is converted to a 16-bit (16b) number by the leftmost shifter 220 (so that it can be added through the adder tree) and fed to the leftmost output multiplexer 225.

[0036] The second standard multiplier 205 from the left similarly forms the product of an 8-bit activation value and a second weight (with zero most significant nibble and non-zero least significant nibble L1). Figure 2AThe third weight in the example has a non-zero most significant nibble M2 and least significant nibble L2. This weight is multiplied by an 8-bit activation value in the third and fourth standard multipliers 205 from the left. The third standard multiplier 205 from the left forms the first partial product by multiplying the 8-bit activation value by the least significant nibble L2, and the fourth standard multiplier 205 from the left forms the second partial product by multiplying the 8-bit activation value by the most significant nibble M2. The offset sum of the two partial products (the latter of the two partial products has 4 more significant bits than the former) is then formed in the offset adder 215 connected to the two multipliers. The offset of the offset adder 215 ensures that the bits of the two partial products are correctly aligned.

[0037] As used herein, the “offset sum” of two values ​​is the result of an “offset addition,” which is the sum of (i) the first of the two values ​​and (ii) the second of the two values ​​(shifted left by several bits, e.g., four bits), and an “offset adder” is an adder that performs the addition of two numbers with an offset between the positions of their least significant bits. As used herein, the “significant bits” of a nibble (or more generally, a sub-word (discussed in further detail below)) are the positions it occupies within a word (a nibble is a part of a word) (e.g., whether the nibble is the most significant nibble or the least significant nibble of an 8-bit word). Therefore, the most significant nibble of an 8-bit word has four more significant bits than the least significant nibble.

[0038] Then, through Figure 2A The circuit produces a product (i.e., the offset sum of the two partial products) at the output of the third output multiplexer 225 from the left. If all four weights have zero most significant nibble, then four products of (i) the activation value and (ii) the four least significant nibs L0, L1, L2, and L3 can be formed in the four standard multipliers 205, and the result is routed to the output of the four output multiplexers 225 via four shifters 220. However, in Figure 2A In the example, the fourth standard multiplier 205 is used to form the second part of the product of the activation value and the third weight (consisting of L2 and M2). In this example, the product of the fourth weight (which has the least significant nibble L3 and zero most significant nibble) is thus formed in the spare multiplier 210. If any of the other three weights besides the third weight has a non-zero most significant nibble, a similar configuration can be used, where one of the other weights (with zero most significant nibble) is processed in the spare multiplier 210 to free up the multiplier used to form the second part of the product of the weight with a non-zero most significant nibble. Therefore, if at most one of the four weights has a non-zero most significant nibble, then Figure 2AThe circuitry can calculate the product of the activation value and any four weights within one clock cycle. In another embodiment, similar to... Figure 2A However, with two spare multipliers 210, if at most two of the four weights have a non-zero highest significant nibble, the product of the activation value and any four weights can be calculated in a similar manner in one clock cycle.

[0039] like Figure 2B As shown, the arrangement of weight nibbles in weight buffer 230 can be the result of preprocessing. The original weight array 240 may include, as shown, a first row of least significant nibbles (e.g., L0, L1, L2, and L3) and a second row of most significant nibbles (in...). Figure 2B In the example, the second line contains only a non-zero most significant nibble (M2). Figure 2B As shown in the blank cells, the remaining most significant nibble can be zero. Preprocessing can rearrange these nibbles when filling the weight buffer (e.g., as shown in the blank cells). Figure 2B (As indicated by the arrow in the diagram), this causes the weight buffer to contain a smaller proportion of zero-value nibbles than the original weight array 240. Figure 2B In the example, the four weights (each composed of the least significant nibble and the most significant nibble) are rearranged such that zero-value nibbles are discarded and non-zero nibbles are placed in the five positions of a row of the weight buffer, so that the row of the weight buffer can be accessed by five multipliers 205, 210, and 210. Figure 2A The data is processed once (e.g., in one clock cycle). The preprocessing operation may also generate an array of control signals for controlling multiplexers (e.g., output multiplexer 225) in the mixed processing circuitry, which perform data routing according to the above rearrangement.

[0040] Figure 3A and Figure 3B Show how they can be combined Figure 2A Two copies of the circuit are used to form a first floating-point processing circuit 305 and a second floating-point processing circuit 310, each floating-point processing circuit being adapted to form a floating-point product of FP16 (half-precision floating-point) activation and FP16 weights. The multipliers A and B of the first floating-point processing circuit 305 are... Figure 2A The first copy of the circuit 315 has two leftmost standard multipliers 205, and the multipliers C and D of the first floating-point processing circuit 305 are... Figure 2A The second copy of the circuit 320 has two leftmost standard multipliers 205, and the multiplier E of the first floating-point processing circuit 305 is... Figure 2A The first copy of the circuit 315 includes a spare multiplier 210. Similarly, the multipliers A, B, C, D, and E of the second floating-point processing circuit 310 include... Figure 2AThe first copy of the circuit 315 has two standard multipliers 205. Figure 2A The second copy of the circuit 320 has two standard multipliers 205, and Figure 2A The second copy of the circuit 320 is a spare multiplier 210. Figure 3B For ease of illustration, each spare multiplier 210 is shown in the middle of a set of standard multipliers 205 (instead of...). Figure 2A (As shown on the right).

[0041] Each floating-point number can be an FP16 floating-point number with a sign bit, an 11-bit mantissa (or “significant digits”) (represented by a 10-bit implicit lead bit or “hidden bit”), and a five-bit exponent (using, for example, a format according to the IEEE 754-2008 standard). The 11-bit mantissa can be padded with a zero and split into three “nibbles”: a “high” (most significant) nibble, a “low” (least significant) nibble, and a “middle” nibble (so that concatenating the high, middle, and low nibbles in sequence produces a 12-bit (padded) mantissa). For brevity, the qualifier “mantissa” may be omitted when describing these nibbles in this disclosure.

[0042] Figure 3A The nine cells of the 3×3 table show the mapping of the product of the three "half-bytes" of the activation value (corresponding to three rows, labeled H_A (for the high half-byte of the activation value), M_A (for the middle half-byte of the activation value), and L_A (for the low half-byte of the activation value)) with the three "half-bytes" of the weight (corresponding to three columns, labeled H_W (for the high half-byte of the weight), M_W (for the middle half-byte of the weight), and L_W (for the low half-byte of the weight)) to the corresponding multiplier in each of the first floating-point processing circuit 305 and the second floating-point processing circuit 310.

[0043] like Figure 3A As indicated by the rectangle marked A, in the first floating-point processing circuit 305, the standard multiplier marked A can multiply (i) the high nibble H_A and the middle nibble M_A of the activation value received at the first (8-bit) input of the standard multiplier A with (ii) the high nibble H_W of the weight. Because the standard multiplier A has an 8-bit wide input and a 4-bit wide (nibble-byte wide) input, it is possible to multiply (i) the high nibble H_A of the activation value with the high nibble H_W of the weight, and (ii) multiply the middle nibble M_A of the activation value with the high nibble H_W of the weight in one operation.

[0044] In this way, five corresponding partial products can be formed (which may be referred to as partial products). and ). and partial product Compared to the effective bits, partial product It has four significant bits, and the two partial products are added together in an offset adder 215 connected to standard multiplier A and standard multiplier B. Similarly, with the partial products... Compared to the effective bits, partial product Having four significant bits, the product of these two parts is added together in offset adders 215 connected to multipliers C and D. Alternate multiplier E can multiply the active high nibble H_A with the weighted low nibble L_W (the unused 4 bits of the first input of alternate multiplier E can be set to zero). The sum produced by the two offset adders 215 and the output of alternate multiplier E can then be added together in the adder tree (which is connected to the output of output multiplexer 225).

[0045] Although some examples are shown herein with respect to an embodiment having 8-bit weights, 8-bit activation values, a weight buffer five weight widths, and a weight and activation that can process one "nibble" at a time, it will be understood that these and other similar parameters in this disclosure are used only as specific, concrete examples for ease of explanation, and any of these parameters may be changed. Thus, for example, the size of a weight can be a "word," and the size of a portion of a weight can be a "subword," wherein, in Figure 2A and Figure 2B In one embodiment, the word size is one byte, and the subword size is one "half-byte". In other embodiments, for example, the word can be 12 bits and the subword can be 6 bits, or the word can be 16 bits and the subword can be one byte.

[0046] As used herein, “part” means “at least some” of the thing, and therefore can mean less than or all of the thing. Thus, “part” of the thing includes the whole thing as a special case (i.e., the whole thing is an example of a part of the thing). As used herein, the term “or” should be interpreted as “and / or”, such that, for example, “A or B” means either “A”, or “B”, or “A and B”.

[0047] Each of the terms “processing circuit” and “means for processing” is used herein to refer to any combination of hardware, firmware, and software for processing data or digital signals. Processing circuit hardware may include, for example, application-specific integrated circuits (ASICs), general-purpose or special-purpose central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), and programmable logic devices (such as field-programmable gate arrays (FPGAs)). As used herein, in processing circuitry, each function is performed by hardware configured (i.e., hardwired) to perform that function, or by more general-purpose hardware (such as a CPU) configured to execute instructions stored in a non-transitory storage medium. Processing circuitry may be fabricated on a single printed circuit board (PCB) or distributed across several interconnected PCBs. Processing circuitry may include other processing circuitry; for example, processing circuitry may include two processing circuits, an FPGA and a CPU, interconnected on a PCB.

[0048] As used herein, when a method (e.g., adjustment) or a first quantity (e.g., a first variable) is referred to as “based on” a second quantity (e.g., a second variable), this means that the second quantity is an input to the method or affects the first quantity (e.g., the second quantity may be an input to a function that calculates the first quantity (e.g., a unique input or one of several inputs), or the first quantity may be equal to the second quantity, or the first quantity may be the same as the second quantity (e.g., the first quantity and the second quantity are stored in the same one or more locations in memory)).

[0049] It will be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various elements, components, regions, layers, and / or portions, these elements, components, regions, layers, and / or portions should not be limited by these terms. These terms are used only to distinguish one element, component, region, layer, or portion from another. Therefore, without departing from the spirit and scope of the inventive concept, the first element, first component, first region, first layer, or first portion discussed herein may be referred to as a second element, second component, second region, second layer, or second portion.

[0050] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the inventive concept. As used herein, the terms “substantially,” “about,” and similar terms are used as approximate terms rather than terms of degree and are intended to take into account the inherent biases of measurements or calculations that will be recognized by those skilled in the art. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. It will also be understood that the terms “comprising” and / or “including,” when used in this specification, indicate the presence of the stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items. Expressions such as “at least one of…” modify the entire list of elements when following a list of elements, without modifying any individual element in the list. Furthermore, the word “may” is used when describing embodiments of the inventive concept to mean “one or more embodiments of this disclosure.” Additionally, the term “exemplary” is intended to indicate an example or illustration. As used herein, the term “use” may be considered synonymous with the term “utilize”.

[0051] It will be understood that when an element or layer is referred to as being "on" another element or layer, "connected to", "bonded to", or "adjacent to" another element or layer, it may be directly on, directly connected to, directly bonded to, or immediately adjacent to the other element or layer, or one or more intermediate elements or layers may be present. In contrast, when an element or layer is referred to as being "directly on" another element or layer, "directly connected to", "directly bonded to", or "immediately adjacent to" another element or layer, no intermediate elements or layers are present.

[0052] Any numerical range described herein is intended to include all subranges containing the same numerical precision within the described range. For example, the range “1.0 to 10.0” or “between 1.0 and 10.0” is intended to include all subranges between the described minimum value 1.0 and the described maximum value 10.0 (and includes both the described minimum value 1.0 and the described maximum value 10.0) (i.e., all subranges having a minimum value equal to or greater than 1.0 and a maximum value equal to or less than 10.0 (e.g., 2.4 to 7.6)). Any maximum numerical limit described herein is intended to include all lower numerical limits contained therein, and any minimum numerical limit described in this specification is intended to include all higher numerical limits contained therein.

[0053] Although exemplary embodiments of processors for fine-grained sparse integer and floating-point operations have been specifically described and illustrated herein, many modifications and variations will be apparent to those skilled in the art. Therefore, it should be understood that processors for fine-grained sparse integer and floating-point operations constructed according to the principles of this disclosure can be implemented in ways other than those specifically described herein. The invention is also defined in the appended claims and their equivalents.

Claims

1. A method of operating a processor, the processor comprising processing circuitry including a first multiplier, a second multiplier, and a third multiplier, the method comprising: The processing circuit forms a first set of products, each product in the first set of products being an integer product of a first activation value and the corresponding weight in a first plurality of weights; as well as The processing circuitry forms a second set of products, each of which is a floating-point product of the second activation value and the corresponding weight from a second set of weights. Each of the multiple weights includes the least significant subword and the most significant subword. The highest valid subword of the first weight among multiple weights is non-zero. In the first set of multiple weights, the highest valid word of the second weight is zero, and the lowest valid word of the second weight is non-zero. The steps to form the first product include: In the first multiplier, the first activation value is multiplied by the least significant subword of the first weight to form the first partial product. In the second multiplier, the first activation value is multiplied by the least significant subword of the second weight. In the third multiplier, the first activation value is multiplied by the highest effective subword of the first weight to form the second part of the product, and Add the first part of the product and the second part of the product together. The steps to form the second product include forming the first floating-point product. The steps to form the first floating-point product include: multiplying the first subword of the mantissa of the second activation value with the first subword of the mantissa of the first weight in the second plurality of weights in the first multiplier to form the third part of the product.

2. The method according to claim 1, wherein, The first multiplier is configured to receive a first argument and a second argument. The first independent variable has the magnitude of the first independent variable. The second independent variable has a second independent variable magnitude, and The size of the first independent variable is greater than the size of the second independent variable.

3. The method according to claim 2, wherein, The steps to form the first floating-point product include: The first independent variable is received by the first multiplier. The second independent variable is received by the first multiplier, and Multiply the first independent variable by the second independent variable; The first independent variable includes: the first subword of the second activation value and the second subword of the second activation value; The second independent variable includes: the first subword of the last digit of the first weight in the second set of weights.

4. The method according to claim 2, wherein, The steps to form the first floating-point product include: The first independent variable is received by the first multiplier. The second independent variable is received by the first multiplier, and Multiply the first independent variable by the second independent variable. The first independent variable includes: the first sub-word of the last digit of the first weight in the second plurality of weights and the second sub-word of the last digit of the first weight in the second plurality of weights; The second independent variable includes the first subword of the second activation value's mantissa.

5. The method according to claim 1, wherein, The second multiplier is configured to receive a first argument of a first size and a second argument of a second size, wherein the first size is greater than the second size.

6. The method according to claim 5, wherein, The steps to form the second product also include forming the second floating-point product; The steps to form the second floating-point product include: The first independent variable is received by the second multiplier. The second independent variable is received by the second multiplier, and Multiply the first independent variable received by the second multiplier by the second independent variable received by the second multiplier; The first argument received by the second multiplier includes: the first subword of the mantissa of the second activation value and a subword consisting of zeros; The second independent variable received by the second multiplier includes the third subword of the mantissa of the first weight in the second plurality of weights.

7. The method according to any one of claims 1 to 6, wherein, The step of adding the first part product and the second part product includes performing offset addition in the first offset adder.

8. The method according to claim 1, wherein, The step of forming the second product also includes forming a second floating-point product, which includes multiplying the first subword of the mantissa of the second activation value with the second subword of the mantissa of the first weight in the second plurality of weights in a third multiplier.

9. The method according to claim 8, further comprising: Add the first floating-point product and the second floating-point product together.

10. The method according to claim 9, wherein, The step of adding the first floating-point product and the second floating-point product includes performing offset addition in the first offset adder.

11. A processor, comprising: Processing circuitry, including: First multiplication device The second multiplication instrument, and The third vehicle The processing circuit is configured as follows: This forms a first set of products, where each product in the first set is an integer product of a first activation value and the corresponding weight from a first plurality of weights. This forms a second set of products, where each product is a floating-point product of the second activation value and the corresponding weight from the second set of weights. Each of the multiple weights includes the least significant subword and the most significant subword. The highest valid subword of the first weight among multiple weights is non-zero. In the first set of multiple weights, the highest valid word of the second weight is zero, and the lowest valid word of the second weight is non-zero. The steps to form the first product include: In the first multiplier, the first activation value is multiplied by the least significant subword of the first weight to form the first partial product. In the second multiplier, the first activation value is multiplied by the least significant subword of the second weight. In the third multiplier, the first activation value is multiplied by the highest effective subword of the first weight to form the second part of the product, and Add the first part of the product and the second part of the product together. The steps to form the second product include forming the first floating-point product. The steps to form the first floating-point product include: multiplying the first subword of the mantissa of the second activation value with the first subword of the mantissa of the first weight in the second plurality of weights in the first multiplier to form the third part of the product.

12. The processor according to claim 11, wherein, The first multiplier is configured to receive a first argument and a second argument. The first independent variable has the magnitude of the first independent variable. The second independent variable has a second independent variable magnitude, and The size of the first independent variable is greater than the size of the second independent variable.

13. The processor according to claim 12, wherein, The steps to form the first floating-point product include: The first independent variable is received by the first multiplier. The second independent variable is received by the first multiplier, and Multiply the first independent variable by the second independent variable; The first independent variable includes: the first subword of the second activation value and the second subword of the second activation value; The second independent variable includes: the first subword of the last digit of the first weight in the second set of weights.

14. The processor according to claim 12, wherein, The steps to form the first floating-point product include: The first independent variable is received by the first multiplier. The second independent variable is received by the first multiplier, and Multiply the first independent variable by the second independent variable; The first independent variable includes: the first sub-word of the last digit of the first weight in the second plurality of weights and the second sub-word of the last digit of the first weight in the second plurality of weights; The second independent variable includes the first subword of the second activation value's mantissa.

15. The processor according to claim 11, wherein, The second multiplier is configured to receive a first argument of a first size and a second argument of a second size, wherein the first size is greater than the second size.

16. The processor of claim 15, wherein, The steps to form the second product also include forming the second floating-point product; The steps to form the second floating-point product include: The first independent variable is received by the second multiplier. The second independent variable is received by the second multiplier, and Multiply the first independent variable received by the second multiplier by the second independent variable received by the second multiplier; The first argument received by the second multiplier includes: the first subword of the mantissa of the second activation value and a subword consisting of zeros; The second independent variable received by the second multiplier includes the third subword of the mantissa of the first weight in the second plurality of weights.

17. The processor according to any one of claims 11 to 16, wherein, The step of adding the first part product and the second part product includes performing offset addition in the first offset adder.

18. The processor of claim 11, wherein, The step of forming the second product also includes forming a second floating-point product, which includes multiplying the first subword of the mantissa of the second activation value with the second subword of the mantissa of the first weight in the second plurality of weights in a third multiplier.

19. The processor of claim 18, wherein, The processing circuit is also configured to add the first floating-point product and the second floating-point product.

20. A processor, comprising: A device for processing, the device for processing includes: First multiplication device The second multiplication instrument, and The third vehicle The device for processing is configured as follows: This forms a first set of products, where each product in the first set is an integer product of a first activation value and the corresponding weight from a first plurality of weights. This forms a second set of products, where each product is a floating-point product of the second activation value and the corresponding weight from the second set of weights. Each of the multiple weights includes the least significant subword and the most significant subword. The highest valid subword of the first weight among multiple weights is non-zero. In the first set of multiple weights, the highest valid word of the second weight is zero, and the lowest valid word of the second weight is non-zero. The steps to form the first product include: In the first multiplier, the first activation value is multiplied by the least significant subword of the first weight to form the first partial product. In the second multiplier, the first activation value is multiplied by the least significant subword of the second weight. In the third multiplier, the first activation value is multiplied by the highest effective subword of the first weight to form the second part of the product, and Add the first part of the product and the second part of the product together. The steps to form the second product include forming the first floating-point product. The steps for forming the first floating-point product include: multiplying the first subword of the mantissa of the second activation value with the first subword of the mantissa of the first weight among the second plurality of weights in the first multiplier to form the third part of the product. The step of forming the second product further includes forming a second floating-point product, which involves multiplying the first subword of the mantissa of the second activation value with the second subword of the mantissa of the first weight among the second plurality of weights in the third multiplier. The processing device is further configured to add the first floating-point product and the second floating-point product. The step of adding the first floating-point product and the second floating-point product includes performing offset addition in the first offset adder.

Citation Information

Patent Citations

  • Mixed-precision neural-processing unit tile

    US20200349106A1