Multiplication circuit, device, system, chip-containing product, method and computer readable medium

By sharing Booth coding circuits among adder arrays and independently controlling signals to enable/disable subsets, the adder arrays support multiplication operations with different data element sizes, solving the problem of low efficiency of multiplication circuits in the prior art and achieving more efficient multiplication operations.

CN121175652APending Publication Date: 2025-12-19ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480021932.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-03
Filing Date
2024-01-23
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

When performing multiplication operations, existing processors require hardware circuit design optimization to improve performance and energy efficiency, especially when handling multiplication operations with different data element sizes, where existing technologies struggle to efficiently utilize circuit resources.

Method used

By sharing a Booth coding circuit among multiple adder arrays, and enabling or disabling subsets of the adder arrays through independent control signals, multiplication operations on data elements of different sizes can be achieved. Furthermore, a cooperative mode can be used to support larger multiplication operations, thereby reducing circuit area and power consumption.

Benefits of technology

It improves the energy efficiency and performance of the multiplication circuit, supports parallel multiplication operations with various data element sizes, reduces circuit timing pressure, and is suitable for multiplication operations within the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121175652A_ABST
    Figure CN121175652A_ABST
Patent Text Reader

Abstract

A multiplication circuit includes at least two adder arrays each for adding a respective set of partial products to generate a respective product representation representing a multiplication result of a respective pair of bit portions selected from a first operand and a second operand. The adder array includes separate instances of hardware circuitry, the adder array having at least two separate enable control signals for independently controlling whether at least two subsets of the adder array are enabled or disabled. A Booth encoding circuit is shared between the adder arrays to Booth encode the first operand to generate partial product selection indicators, each corresponding to a Booth encoding of a respective Booth number of the first operand. At least two adder arrays operate respective partial products selected by a partial product selection circuit based on a same partial product selection indicator generated by the shared Booth coding circuit based on a same Booth number of the first operand.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND TECHNICAL FIELD

[0001] The present technology relates to the field of data processing. BACKGROUND

[0002] Processors can have logic circuitry for implementing various arithmetic or logical operations. One arithmetic operation to be supported by a processor can be a multiplication operation. While the arithmetic operation for a multiplication operation is well known, there are design choices to be made in how to implement the hardware circuitry for performing the operation within a processor. Design decisions made by circuit designers can have an impact on processing performance and / or energy efficiency. SUMMARY

[0003] At least some examples of the present technology provide a multiplication circuit comprising:

[0004] a plurality of adder arrays each for adding a respective set of partial products to generate a respective product representation value representing a multiplication result of a respective pair of bit portions selected from a first operand and a second operand, the plurality of adder arrays comprising separate instances of hardware circuitry, the plurality of adder arrays having at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled;

[0005] a shared Booth encoding circuit shared between the plurality of adder arrays for Booth encoding the first operand to generate a plurality of partial product selection indicators, each partial product selection indicator corresponding to a Booth encoding of a respective Booth number of the first operand; and

[0006] a partial product selection circuit for selecting the partial products to be added by the plurality of adder arrays based on the second operand and the plurality of partial product selection indicators; wherein:

[0007] at least two of the adder arrays are configured to operate on respective partial products selected by the partial product selection circuit based on a same partial product selection indicator generated by the shared Booth encoding circuit based on the Booth encoding of a same Booth number of the first operand.

[0008] At least some examples of the present technology provide an apparatus comprising processing circuitry for performing data processing in response to instructions, the processing circuitry comprising the multiplication circuit described above.

[0009] At least some examples of the present technology provide a method comprising:

[0010] Booth encoding the first operand using a shared Booth encoding circuit shared between the plurality of adder arrays to generate a plurality of sets of partial product selection indicators, each partial product selection indicator corresponding to a Booth encoding of a respective Booth number of the first operand, wherein the plurality of adder arrays comprise separate instances of hardware circuitry, the plurality of adder arrays having at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled;

[0011] selecting the partial products to be added by the plurality of adder arrays based on the second operand and the plurality of partial product selection indicators; and

[0012] adding the partial products using the enabled adder arrays of the plurality of adder arrays to generate respective product representation values each representing a multiplication result of a respective pair of bit portions selected from the first operand and the second operand,

[0013] wherein at least two of the adder arrays are configured to operate on respective partial products selected based on a same partial product selection indicator generated by the shared Booth encoding circuit based on the Booth encoding of a same Booth number of the first operand.

[0014] At least some examples of the present technology provide a non-transitory computer readable medium for storing computer readable code for manufacturing a multiplication circuit, the multiplication circuit comprising:

[0015] a plurality of adder arrays each for adding a respective set of partial products to generate a respective product representation value representing a multiplication result of a respective pair of bit portions selected from a first operand and a second operand, the plurality of adder arrays comprising separate instances of hardware circuitry, the plurality of adder arrays having at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled;

[0016] a shared Booth encoding circuit shared between the plurality of adder arrays for Booth encoding the first operand to generate a plurality of partial product selection indicators, each partial product selection indicator corresponding to a Booth encoding of a respective Booth number of the first operand; and

[0017] a partial product selection circuit for selecting the partial products to be added by the plurality of adder arrays based on the second operand and the plurality of partial product selection indicators; wherein:

[0018] At least two of the adder arrays are configured to operate on respective partial products selected by the partial product selection circuit based on a same partial product selection indicator generated by the shared Booth encoding circuit based on the Booth encoding of a same Booth number of the first operand.

[0019] At least some examples provide a system comprising the multiplication circuit described above implemented in at least one package chip; at least one system component; and a board, wherein the at least one package chip and the at least one system component are assembled on the board.

[0020] At least some examples provide a chip-in product comprising the system described above assembled with at least one other product component on a further board.

[0021] Further aspects, features, and advantages of the technology will be apparent from the following description of examples, read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 Examples of data processing apparatus are illustrated;

[0023] Figure 2 Examples of multiplication circuit are illustrated;

[0024] Figure 3 Methods are illustrated that each adder array used for comparison has its own private Booth encoding circuit;

[0025] Figure 4 Examples are illustrated of shared Booth encoding circuit shared between adder arrays;

[0026] Figure 5 Adder arrays are illustrated in more detail;

[0027] Figures 6A to 6D Mapping of adder arrays onto partials of multiplication operations performed in different data element sizes is illustrated;

[0028] Figure 7 Bit selection of Booth numbers of the first operand by each Booth encoder of the shared Booth encoding circuit is illustrated;

[0029] Figure 8 Which partial (and which Booth encoder) of the second operand to use to select each partial product to be added by each adder array for each adder array is illustrated;

[0030] Figure 9 Addition of partial products for an 8-bit adder array is illustrated;

[0031] Figure 10A and Figure 10B Addition of partial products of a 16-bit adder array is illustrated;

[0032] Figure 11A and Figure 11B Addition of partial products of a 32-bit adder array is illustrated;

[0033] Figure 12 Addition of partial products of an adder array used in 64-bit multiplication to provide extra bits not needed for 8-bit, 16-bit, or 32-bit multiplication is illustrated;

[0034] Figure 13 Relative alignment for addition of product representation values in a cooperative mode is shown;

[0035] Figure 14 is a flowchart illustrating a method of performing multiplication using a multiplication circuit; and

[0036] Figure 15 Examples of systems and chip products are illustrated. DETAILED DESCRIPTION

[0037] A multiplication circuit includes two or more adder arrays each to add a respective set of partial products to generate a respective product representation value representing a multiplication result of a respective pair of bit portions selected from a first operand and a second operand. Each adder array includes a separate instance of hardware circuitry. The adder arrays have at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled. For example, respective subsets of the adder arrays each having independent enable / disable control can be used to support multiplication operations on portions of the first and second operands corresponding to different data types, with each adder array capable of providing a respective product representation value representing a multiplication result of a respective pair of bit portions from the first and second operands.

[0038] The partial products summed by each adder array are selected based on the second operand and based on partial product selection indicators generated by a Booth encoding circuit. Booth encoding is a technique used in multiplication circuits to reduce the number of partial products that need to be added to produce a multiplication result. Booth encoding is based on treating the first operand as a series of overlapping Booth numbers (each Booth number comprising a certain number of bits of the first operand), and analyzing the bit pattern of each Booth number to identify where strings of consecutive ones begin / end in the first operand. For a given Booth number of the first operand, a corresponding partial product selection indicator is generated that represents the Booth encoding of the Booth number. The partial product selection indicator controls a partial product selection circuit to select a corresponding partial product to be summed by one of the adder arrays. For example, a given partial product can be a selected multiple of a portion of the second operand, where the corresponding partial product selection indicator is used to select between different multiples of that portion of the second operand.

[0039] In cases where several independent subsets of the adder arrays are provided for generating respective product representation values, the general approach would be to provide an instance of the Booth encoding circuit that is private to each adder array rather than shared with the Booth encoding circuits used by the other arrays.

[0040] However, the inventors recognized that there is an opportunity to save circuit area and power consumption based on the recognition that the same operand or the same sub-segment of that operand can be Booth encoded for different adder arrays, such that the Booth encoding circuit can be shared among the adder arrays. Accordingly, a shared Booth encoding circuit is provided that is shared among multiple adder arrays to Booth encode the first operand to generate partial product selection indicators, each corresponding to a Booth encoding of a respective Booth number of the first operand. At least two of the adder arrays are configured to operate on respective partial products selected by the partial product selection circuit based on the same partial product selection indicator that is generated by the shared Booth encoding circuit based on the Booth encoding of the same Booth number of the first operand. By providing a shared Booth encoding circuit, the overall circuit area and power consumption of the multiplication circuit can be reduced. This technique can also help to reduce pressure on conforming to circuit timing aspects compared to implementations where each adder array has a private, non-shared Booth encoder, as there is no need to reserve many bits for the partial product selection indicators.

[0041] Each adder array generates a respective product representation value indicative of a numerical result of a respective pair of bit portions selected from the first operand and the second operand. Thus, one adder array can generate a product representation value representing a product of a first portion of the first operand and a corresponding first portion of the second operand. Another adder array can generate a product representation value representing a product of a second portion of the first operand (which can or can not overlap the first portion) and a corresponding second portion of the second operand. Each product representation value can represent a numerical result of A-bit * B-bit multiplication of respective bit portions taken from the first operand and the second operand, respectively (where A can equal B, or A can be greater than or less than B, and the values of A and B can be the same for two or more adder arrays or can differ from one adder array to another). Each product representation value represents a numerical result of A-bit * B-bit multiplication (so not just an intermediate step representing a partial sum of only some of the partial products corresponding to A-bit * B-bit multiplication). However, the numerical result can be represented in different forms. For example, the product representation value can represent the numerical result using a carry-save format of sum terms and carry terms. For example, the adder array can include a carry-save-adder tree that does not pass carries from one bit lane to another, but rather performs a series of 3:2 carry-save addition reductions to reduce the set of partial products into sum terms and carry terms that, when added together by a further carry-propagating adder (not part of the adder array), can yield a number representing the numerical result of A-bit * B-bit multiplication in binary representation. Thus, it should be understood that the output of each adder array can represent the product representation value in carry-save form.

[0042] The multiplication circuit is configured to support at least two data element size configurations for multiplication of one or more pairs of respective data elements selected from the first operand and the second operand, each data element size configuration corresponding to a different combination of data element sizes of data elements selected from the first operand and the second operand. For example, the data element size configurations can correspond to different SIMD element sizes of a SIMD (single instruction multiple data) multiplication operation. For example, the first operand and the second operand can be vector operands, and each data element can be a respective vector element. In another example, the first operand and the second operand can be matrix operands, and each data element can be a respective matrix element.

[0043] For example, the size of each individual adder array can be designed to be suitable for a particular data element size configuration, and the selection of which adder arrays to enable / disable can depend on the current element size configuration in use. Thus, the enable control circuitry can select which of the adder arrays are to be enabled and which of the adder arrays are to be disabled based on the current data element size configuration to be used for the multiplication operation. Using at least two subsets of the adder arrays with independent enable / disable control can provide greater energy efficiency for implementing multiplications with different data element size configurations compared to a particular implementation that provides a single large adder array and relies on setting 0s in some partial product bits when performing multiplications with data element sizes smaller than the maximum supported size.

[0044] The shared Booth encoding circuitry can generate the Booth encoding of the given Booth number based on the least significant bit of the given Booth number being set to 0 for at least one of the data element size configurations for which the given Booth number of the first operand straddles a data element boundary; and generate the Booth encoding of the given Booth number based on the least significant bit of the given Booth number being set to the value of the corresponding bit of the first operand for at least one other of the data element size configurations for which the given Booth number of the first operand does not straddle a data element boundary. Thus, by providing selection circuitry in the shared Booth encoding circuitry to select whether to treat the least significant bit of the given Booth number as 0 or the corresponding bit of the first operand, this allows the same Booth encoding circuitry to be shared between different adder arrays even when those adder arrays are based on data element size configurations that mean the given Booth number of the first operand will straddle an element boundary for one configuration but not for another.

[0045] The adder array can comprise at least:

[0046] • a first subset of the adder array, each respective product representation value of the first subset representing the multiplication result of a respective pair of data elements selected from the first operand and the second operand according to a first data element size configuration; and

[0047] • a second subset of the adder array, each respective product representation value of the second subset representing the multiplication result of a respective pair of data elements selected from the first operand and the second operand according to a second data element size configuration.

[0048] With respect to sharing of the Booth encoder between the adder arrays, at least two of the adder arrays described above (those whose respective partial products are selected based on the same partial product selection indicator) can comprise the adder arrays of the first subset and the adder arrays of the second subset. This can reflect that, in the first data element size configuration / second data element size configuration, the first subset of adder arrays / second subset of adder arrays operate on the same sub-sections of the first operand (although those sub-sections can be logically divided into different data element configurations), so the same hardware circuitry logic can be used to analyse the bit pattern of a given Booth number of the first operand and provide a corresponding partial product selection indicator used by the partial product selection circuit to select the respective partial product for both the adder arrays of the first subset and the adder arrays of the second subset, regardless of which data element configuration is active.

[0049] While the same partial product selection indicator is used for both the adder arrays of the first subset and the adder arrays of the second subset, the partial products selected for the adder arrays of the first subset based on the same partial product selection indicator can depend on a first subset of bits of the second operand when the multiplication is to be performed in accordance with the first data element size configuration, and the partial products selected for the adder arrays of the second subset based on the same partial product selection indicator can depend on a second subset of bits of the second operand when the multiplication is to be performed in accordance with the second data element size configuration. Thus, while the same partial product selection indicator is used for the respective adder arrays in both the first subset and the second subset, the values of the corresponding partial products used by those respective adder arrays in the first subset / second subset can still differ as those partial products can be based on different portions of the second operand.

[0050] The enable control circuitry can be configured to set one or more enable control signals for the first subset of adder arrays to disable the first subset of adder arrays in response to the current data element size configuration information specifying that the first data element size configuration is not required for a given multiplication operation performed on the first operand and the second operand, and to set one or more enable control signals for the second subset of adder arrays to disable the second subset of adder arrays in response to the current data element size configuration information specifying that the second data element size configuration is not required for the given multiplication operation. Thus, the first subset of adder arrays can be enabled when the first data element size configuration is required, and the second subset of adder arrays can be enabled when the second data element size configuration is required. If one of the first data element size configuration / second data element size configuration is not required, the corresponding one of the first subset of adder arrays / second subset of adder arrays can be disabled to save power.

[0051] The enable control circuit can support performing multiplication operations on the first and second operands in parallel using both the first data element size configuration and the second data element size configuration by enabling both the first subset of adder arrays and the second subset of adder arrays using the enable control signal. While only one of the data element size configurations can typically be needed at a given time, there can be some use cases in which it is useful to process the same first and second operands using more than one different data element size configuration to generate a corresponding set of product results in two or more data element size configurations. This can be useful, for example, in some machine learning or graphics processing applications. For example, two vector operands J and K can be processed to generate a first vector result L that provides a set of product representation values each corresponding to the result of multiplying a corresponding pair of 32-bit elements of J and K, and a second vector result M that provides a set of product representation values each corresponding to the result of multiplying a corresponding pair of 16-bit elements of J and K. This can support such parallel multiplication operations performed on the same operands according to more than one data element size configuration by implementing the multiplication circuit using several subsets of adder arrays each corresponding to a given data element size configuration and each having independent enable / disable control.

[0052] While the examples discussed refer to a first subset of adder arrays and a second subset of adder arrays, it should be understood that the same techniques can extend to three or more subsets of adder arrays. For example, the multiple adder arrays can also include a third subset of adder arrays, each respective product representation value of the third subset of adder arrays representing the multiplication result of a respective pair of data elements selected from the first and second operands according to a third data element size configuration. For example, the first subset of adder arrays, the second subset of adder arrays, and the third subset of adder arrays can be used for 8-bit, 16-bit, 32-bit multiplication, respectively.

[0053] The adder array can also be used in a cooperative mode in which it cooperates to perform respective portions of a larger multiplication. The product addition circuitry can operate multiplication operations for which two or more of the adder arrays operate in the cooperative mode, add respective product representation values generated by the two or more of the adder arrays to generate a product result value representing a multiplication of the first operand and the second operand that is wider than the bit portions used to generate the respective product representation values for any of the two or more of the adder arrays. For example, where respective adder arrays are provided that support at least two of 8-bit, 16-bit, and 32-bit data element configurations respectively, those adder arrays can also be mapped to respective portions of the partial product addition needed for a larger multiplication operation (e.g., with 64-bit elements). Using several independently enabled / disabled smaller adder arrays in this way to implement larger multiplication operations can provide a more power efficient way to implement options for both smaller and larger size multiplication operations compared to alternative approaches such as putting zeros into a single monolithic large multiplier as discussed above.

[0054] For at least one of the adder arrays, the portion of the second operand used to form the partial product for that adder array can vary depending on whether the multiplication operation is to be performed in a cooperative mode. For example, the partial product selection circuitry can select the portion of the second operand used to form the partial product for that adder array as a first portion of the second operand when the multiplication operation is to be performed in a non-cooperative mode in which each adder array independently operates to produce an independent product representation value, and as a second portion of the second operand when the multiplication is to be performed in a cooperative mode in which the adder arrays generate respective product representation values that can be further added together by the product addition circuitry to generate a product result value representing a result of a wider multiplication.

[0055] In some examples, the wider bit portion of the first operand and the second operand includes all magnitude-indicating bits of the first operand and the second operand. For example, where the first operand / second operand is a vector operand based on a signed / unsigned integer data type, the cooperative mode can provide a full operand width multiplication of the quantity represented by all of the bits of the first operand and the quantity represented by all of the bits of the second operand. Further, for examples in which the first operand / second operand is a floating point operand including bits representing a sign (positive or negative), an exponent, and a fraction, the magnitude-indicating bits can be the bits of the significand represented by the fraction (so the bits of the first operand / second operand corresponding to the sign and the exponent can not be considered magnitude-indicating bits of the first operand / second operand).

[0056] In other examples, the wider bit portions of the first and second operands need not include all of the magnitude indication bits of the first and second operands. For example, a cooperation mode can be used to cause the adder arrays to work together to compute a multiplication result for a data type that does not correspond to the largest supported size that can be represented using the first / second operands, but for which the adder array does not natively support. In this case, only some of the adder arrays in the adder array can need to cooperate to provide the multiplication result for the data type, so even in the cooperation mode, some of the adder arrays can not be needed (and thus can be disabled using the corresponding enable control signal) according to the configuration option selected for the particular multiplication operation.

[0057] In one example, the adder array comprises:

[0058] a plurality of subsets of the adder array corresponding to different data element size configurations, each subset of the adder array comprising two or more adder arrays for generating, in a non-cooperation mode, a respective product representation value representing a multiplication of a respective pair of data elements selected from the first and second operands according to the data element size configuration corresponding to the subset of the adder array, wherein in a cooperation mode, each of the plurality of subsets of the adder array is assigned to a portion of the multiplication of the wider portion; and

[0059] a further adder array for generating a further product representation value representing a result of a remaining portion of the multiplication of the wider portion other than the portion assigned to the plurality of subsets of the adder array; and

[0060] the multiplication circuit comprises an enable control circuit for disabling the further adder array in the non-cooperation mode.

[0061] The plurality of subsets of the adder array can be the same subsets as the at least two subsets mentioned above that are provided with separate enable control signals.

[0062] Thus, in some examples, supporting larger multiplications in the cooperation mode can require adding some additional partial products that are not needed by any of the data element size configurations natively supported by the plurality of subsets of the adder array. Thus, a further adder array (which is disabled in the non-cooperation mode) can be provided to compute a further product representation value representing a result of a remaining portion of the multiplication of the wider portion that is multiplied in the cooperation mode. The further adder array can also share the same Booth encoding circuitry of the other subsets of the adder array used in the non-cooperation mode.

[0063] In some examples, separate enable control for the adder array can be provided on a per-subset basis, such that each subset of the adder array corresponding to a given data element size configuration has its own independent enable / disable control, but it is not necessary to provide finer-grained control to independently enable / disable each individual adder array within a subset. For example, a first subset of the adder array designed for multiplication based on a first data element size configuration (e.g., 8-bit * 8-bit multiplication) can be collectively enabled / disabled based on a single shared enable control signal, where this enable control signal is different from the enable control signal used to enable / disable a second subset of the adder array designed for multiplication based on a second data element size configuration (e.g., 16-bit * 16-bit multiplication).

[0064] However, in other examples, each adder array can have its own independent enable control signal, allowing each adder array to be independently enabled / disabled by this enable control signal. For example, if a given data element lane of the first operand / second operand is asserted to disable action in that lane, this can be used, for example, to allow some adder arrays to be disabled even when the current selected data element size configuration is a data element size configuration associated with that adder array. Furthermore, providing independent enable / disable control for respective adder arrays within a subset of the adder array designed for a given data element size configuration can also help allow a set of adder arrays designed to support a certain maximum operand length of the first operand / second operand to also support multiplication performed on operands of smaller length. For example, while a set of adder arrays can be provided to support 64-bit operands, when multiplication is to be applied to 32-bit operands, half of the adder arrays for a given data element size configuration can be disabled.

[0065] The separate enable control signal for each adder array can be implemented in various ways. For example, power gating can be used to provide separate enable / disable control by providing the ability to selectively isolate each adder array from a power supply node.

[0066] However, in one example, the separate enable control signal comprises a separate clock signal. For example, when a selected adder array is to be disabled, the enable control circuit described above can clamp the clock signal to the selected adder array to a fixed value. Keeping the clock signal to the selected adder array static can prevent that selected adder array from acting, and can save dynamic power. Controlling the enable / disable of the adder array based on whether the corresponding clock signal is allowed to toggle bi-stably can be a more area-efficient way of providing independent enable / disable control for each adder array, as it can require less additional logic gates than alternative solutions such as power gating.

[0067] An apparatus can comprise processing circuitry that performs data processing in response to instructions; and the processing circuitry comprises the multiplication circuitry described above. For example, the processing circuitry can be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or other processing unit within a data processing system (e.g., a neural processing unit provided for performing neural network processing or other machine learning operations).

[0068] Figure 1 An example of a data processing apparatus 2 is schematically illustrated. The data processing apparatus has a processing pipeline 4 (an example of processing circuitry) comprising a number of pipeline stages. In this example, the pipeline stages comprise: a fetch stage 6 for fetching instructions from an instruction cache 8; a decode stage 10 for decoding fetched program instructions to generate micro-operations to be processed by the remaining stages of the pipeline; an issue stage 12 for checking that operands required by a micro-operation are available in a register file 14 and issuing the micro-operation for execution once the required operands for a given micro-operation are available; an execution stage 16 for performing data processing operations corresponding to micro-operations by processing operands read from the register file 14 to generate result values; and a writeback stage 18 for writing back results of processing to the register file 14. It will be appreciated that this is just one example of a possible pipeline architecture, and other systems can have additional stages or a different configuration of stages. For example, in an out-of-order processor, an additional register renaming stage can be included for mapping architectural registers specified by program instructions or micro-operations to physical register specifiers identifying physical registers in the register file 14.

[0069] The execution stage 16 includes several processing units for performing different categories of processing operations. For example, the execution units can include an integer or fixed point arithmetic / logical unit 20 for performing arithmetic or logical operations on scalar or vector operands read from the register file 14, a floating point unit 22 for performing operations on floating point values, a branch unit 24 for evaluating the results of branch operations and adjusting the program counter representing the current point of execution accordingly, and a load / store unit 28 for performing load / store operations to access data in the memory system 8, 30, 32, 34. In this example, the memory system includes a level one data cache 30, a level one instruction cache 8, a shared level two cache 32, and a main system memory 34. It will be appreciated that this is just one example of a possible memory hierarchy, and other configurations of caches can be provided. The specific types of processing units 20-28 shown in the execution stage 16 are just one example, and other implementations can have a different set of processing units or can include multiple instances of the same type of processing unit so that multiple micro-operations of the same type can be handled in parallel. It will be appreciated that Figure 1 is a simplified representation of just some components of a possible processor pipeline architecture, and that processors can include many other elements not illustrated for the sake of brevity, such as branch prediction mechanisms or address translation or memory management mechanisms.

[0070] One example of an operation that can be supported by processing circuitry 4 (e.g., within ALU 20 or floating point unit 22) is a multiplication operation. In some systems, a special purpose execution unit known as a multiply-accumulate (MAC) unit can be provided to handle multiplication due to multiply-accumulate operations, in which two operands are multiplied and the result is added to an accumulator value. For example, multiply-accumulate operations can be used frequently in digital signal processing algorithms, so any techniques for improving energy efficiency and reducing pressure on circuit timing can be highly helpful.

[0071] While the following examples discuss multiplication operations, this is intended to encompass multiply-add or multiply-accumulate operations, so even if a subsequent adder for adding the multiplication result to a third operand is not shown, such a subsequent adder can be provided. It is also possible to provide independent multiplication operations that produce a multiplication result without also adding that multiplication result to a third operand.

[0072] The multiplication circuit 50 described below is based on a technique known as Booth multiplication, which is based on the principle that when multiplying a first value (multiplicand M) by a second value (multiplier R) to obtain a multiplication result M*R, within the multiplicand M, a string of consecutive binary Is can effectively be replaced with +1 at bit positions one higher than the high end of the string and -1 at bit positions corresponding to the low end of the string, which can help reduce many of the partial products in the partial products to zero and thus make the processor logic implementation more straightforward. This is analogous to the way that in decimal, 999 is equivalent to 1000 - 1. Thus, if one considers the multiplication of 999 * R, the "textbook" long multiplication method would perform a series of additions of partial products 900 * R + 90 * R + 9 * R. With the Booth method, this can be reduced to 1000 * R - 1 * R. The corresponding overlapping groups of bits of the multiplicand M (referred to below as "Booth digits") can be analyzed to find patterns that represent the start / end of strings of consecutive Is, and this can be used to infer each multiple of R that is to be selected as a corresponding partial product to be added to form the multiplication result.

[0073] Booth multiplication involves three stages:

[0074] 1. Booth encoding of multiplicand M .

[0075] The multiplicand M is logically divided into a series of overlapping Booth digits corresponding to subsets of bits of the multiplicand M. For each Booth digit, a Booth encoder analyzes the pattern of bits in that Booth digit and outputs a partial product selection indicator as the Booth encoding of that Booth digit, which indicates which of several different multiples of the multiplier R should be selected as the corresponding partial product to be included in the set of partial products that are added to produce the multiplication result M * R. Different "radix" versions of the Booth encoding scheme can be provided, where the radix indicates how many bits of the multiplicand M are considered to be in each Booth digit. Adjacent Booth digits overlap by 1 bit. The least significant Booth digit is padded with a fixed bit of 0b0 at the bottom. The most significant Booth digit is padded with at least one bit above the most significant bit of M. The padding bits correspond to the sign extension of the multiplicand.

[0076] For example, for radix-4 Booth multiplication, each Booth digit includes 3 bits, and adjacent Booth digits overlap by 1 bit. For example, for an 8-bit multiplicand M having bits M[7:0], the Booth digits can include:

[0077] - M[1:0]: 0b0 (the two lower bits of M concatenated with a fixed value of 0 at the bottom).

[0078] - M[3:1]

[0079] - M[5:3]

[0080] - M[7:5]

[0081] - a 3-bit sign extension including M[7] (e.g., for unsigned values, this would be 0b00:M[7], and for signed values represented in two’s complement, all three bits are set equal to M[7]).

[0082] The Booth encoder implements hardware circuit logic that determines which multiple of the multiplier R should be selected for a corresponding partial product based on the pattern of bits in a given Booth number. The rule of which multiple to select for a given pattern of bits in a Booth number is based on whether a string of consecutive 1s begins or ends within that Booth number. For example, if all bits within a Booth number are either 0 or all are 1, then the multiple to select is 0*R since no string of 1s begins or ends within the Booth number. For Booth numbers involving a mix of 0s and 1s, the multiple to select depends on where any transitions from 0 to 1 or from 1 to 0 occur, such that the multiple implements a combination effect of (i) adding a +R multiple at a bit position that is one higher than the top bit of any string of consecutive 1s occurring within the Booth number, and (ii) adding a -R multiple at a bit position corresponding to the bottom bit of any string of consecutive 1s occurring within the Booth number. However, since multiple bits are evaluated at a time for a Booth number, for Booth encoding of base -4 or higher bases, higher multiples of R such as ±2*R are considered to account for the fact that a +R or -R multiple can be placed at different bit positions within a Booth number. Some examples of operation for base -4 and base -8 are discussed below, but it should be understood that other base values can be used. For each Booth number of the multiplicand M, the Booth encoder outputs a partial product selection indicator indicating which of a set of candidate multiples of the multiplier R should be selected for the corresponding partial product.

[0083] 2. Selection of partial products

[0084] The required multiples of the multiplier R are prepared (this can be done in parallel with the Booth encoding). For example, for base -4 operations, the multiples of R that can be selected for a given Booth number are: +2*R, +R, 0, -R, -2*R. For base -8, the multiples extend from +4*R to -4*R. Thus, the partial product selection circuit can form multiple values. Forming multiple values can include taking the negative of R for forming negative multiples, left shifting R to form multiples of powers of 2 such as ±2*R and ±4*R, and adding to other multiples if the base requires multiples of non-powers of 2 such as ±3*R (e.g., adding ±2*R and ±R to form ±3*R multiples).

[0085] The partial product selection circuit selects one of those multiples for a given Booth number from the candidate multiples based on the partial product selection indicator provided by the Booth encoder for that given Booth number.

[0086] 3. Addition of partial products

[0087] The selected partial products for each Booth number are added together (with proper alignment between partial products of adjacent Booth numbers based on considering the relative magnitudes of the partial products based on the position within the multiplicand M where the Booth number was found).

[0088] To illustrate Booth multiplication, consider the multiplication M*R, where M and R are each 8-bit values, and the decimal values corresponding to M and R are M=56 and R=47:

[0089] M = 00111000

[0090] R = 00101111

[0091] Using the conventional "textbook" long multiplication method, this can be converted into the following series of partial products (PPi is +R if the corresponding bit i of M is 1, and 0 if the corresponding bit i of M is 0):

[0092] Adding the partial products (each time shifting by 1):

[0093]

[0094] 9 partial products for 8*8 bit multiplication

[0095] Using the textbook method, the string of consecutive 1's in M causes three partial products, PP3, PP4, PP5, to include +R multiples. Using the Booth method, the same result could have been achieved by adding +R for PP6 and -R for PP3, but using base -2 (which would correspond to Booth numbers each including 2 bits) this would not reduce the number of partial products. Thus, most practical Booth implementations use base -4 or higher.

[0096] For base -4, as explained previously, the Booth numbers are selected based on the bits of M, and the encoding rules are as follows:

[0097] Selected bus number M i+1 , M i , M i-1 ]]> Multiples of R selected as partial products 000 0 001 +1 010 +1 011 +2 100 -2 101 -1 110 -1 111 0

[0098] Using the same example of M=56 and R=47, the multiple values that can be used for each partial product selection are:

[0099] +2R = 01011110

[0100] +R = 00101111

[0101] 0 = 00000000

[0102] -R = 11010001

[0103] -2R = 10100010

[0104] And the multiplicand is:

[0105] M = 00111000

[0106] Thus, the Booth numbers selected from the multiplicand M and the corresponding partial products selected for each Booth number according to the encoding rules shown above are as follows:

[0107] BD0 = M1,M0,M -1 = 00 (0) -> PP0 = 0

[0108] BD1 = M3,M2,M1 = 100 -> PP1 = -2R = 10100010

[0109] BD2 = M5,M4,M3 = 111 -> PP2 = 0

[0110] BD3 = M7,M6,M5 = 001 -> PP3 = +R = 00101111

[0111] BD4 = SE,SE,M7 = 000 -> PP4 = 0

[0112] Adding the partial products (each shifted by 2 to account for the relative alignment of each Booth number) gives:

[0113]

[0114] Thus, now only 5 rather than 9 partial products can be added to achieve the same numerical result.

[0115] Similarly, for base -8, the Booth encoding rules are as follows:

[0116] Selected bus number M i+2 : M i-1 ]]> Multiples of R selected as partial products 0000 0 0001 +1 0010 +1 0011 +2 0100 +2 0101 +3 0110 +3 0111 +4 1000 -4 1001 -3 1010 -3 1011 -2 1100 -2 1101 -1 1110 -1 1111 0

[0117] Applying this to the same example of M = 56 and R = 47 gives:

[0118] M = 00111000

[0119] +4R = 10111100

[0120] +3R = 01110011

[0121] +2R = 01011110

[0122] +R = 00101111

[0123] 0 = 00000000

[0124] -R = 11010001

[0125] -2R = 10100010

[0126] -3R = (1)01110011

[0127] -4R = (1)01000100

[0128] BD0 = M2:M -1 = 000(0) -> PP0 = 0

[0129] BD1 = M5:M2 = 1110 -> PP1 = -R = 11010001

[0130] BD2 = M8:M5 = (0)0011 (bit 8 sign-extended from bit 7)

[0131] -> PP2 = +R = 00101111

[0132] Add the partial products together (shift by 3 each time):

[0133]

[0134] The result can now be achieved in 3 partial products.

[0135] Thus, the approach using a higher base number can treat the multiplicand M as including fewer Booth numbers and thus requiring fewer partial products to be added, but this comes at the cost of increased complexity in having more options for multiple selection (which will increase the circuit complexity of the multiplier for generating multiple values and the multiplexer for selecting multiples, as well as the complexity of the Booth encoder).

[0136] Figure 2An example of a multiplication circuit 50 that can be included within the execution stage 16 of the processor 2 (within the ALU 20, or within the floating point unit 22, or in another execution unit such as a vector ALU or matrix processing unit separate from the ALU for scalar operations) is illustrated. The following example assumes that the multiplication operation is performed on data elements having a number of bits equal to a power of 2, which can be common if the operation is performed on integer operands. However, it is not necessary that the data elements have a number of bits corresponding to a power of 2. For example, for multiplication performed on the significands of floating point operands, since some of the bits of a 2's power size floating point representation are used for the sign and exponent, the significands can have a non-2's power number of bits. Thus, it should be understood that the following example can be adjusted for handling other data element sizes. Furthermore, while the following example assumes that the pair of portions of the two operands that are multiplied together have an equal number of bits (e.g., 16-bit * 16-bit multiplication, or 32-bit * 32-bit multiplication), this is not necessary, and it is also possible to provide data element size configurations that multiply portions having asymmetric sizes (e.g., 16-bit * 8-bit multiplication). For example, a multiplication circuit having asymmetric operand sizes can be used for machine learning processing, where for example kernel weights for a neural network can have fewer bits than the activation values that are multiplied by the kernel weights.

[0137] As Figure 2 shown, the multiplication circuit 50 includes a shared Booth encoding circuit 52, a partial product selection circuit 54, a set of adder arrays 56, a cooperative mode adder (also referred to as a "product addition circuit") 58, and an enable control circuit 60. The multiplication circuit 50 receives a first operand src_a, a second operand src_b, and data element size configuration information indicating a data element size configuration to be used in processing the first operand src_a and the second operand src_b. The first and second operands can be SIMD (single instruction multiple data) operands having a number of independent data elements each representing separate data values within the same operand. For example, the first and second operands can be vector operands representing a one-dimensional array of independent vector elements, or matrix operands representing a two-dimensional array of independent matrix elements. The first and second operands can be obtained from source registers specified by a multiplication instruction executed by the processing circuit, and / or can be forwarded from results generated by earlier instructions in the processing pipeline 4.

[0138] For one data element size configuration, the first and second operands can be treated as a single data element to be multiplied together, but for other data element size configurations, each of the first and second operands can be logically divided into multiple independent data elements, and the product result to be generated can be a vector or matrix of several result data elements each corresponding to a product of a pair of corresponding data elements of the first and second operands. For a given multiplication operation, the data element size configuration information can depend on an immediate or register operand of the instruction executed by the processing circuitry 4, and / or be based on element size mode information stored within a system register of the processing circuitry 4. The data element size configuration information can vary from multiplication operation to multiplication operation.

[0139] The shared Booth encoding circuitry 52 Booth encodes the first operand src_a to generate a set of partial product selection indicators 62 each corresponding to a Booth encoding of a respective Booth number of the first operand. The Booth encoding is generated based on the bit pattern of the corresponding Booth number according to the encoding rules shown above for base-4 or base-8 (or alternatively, similar rules for a higher base if a higher base is used for Booth encoding).

[0140] The partial product selection circuit 54 selects groups of partial products to be added by each of the adder arrays 56 based on the second operand src b and the partial product selection indicator 62. For example, a first group of partial products "pps 0" is selected for adder array 56-0, a second group of partial products "pps 1" is selected for adder array 56-1, and so on. The partial product selection can also depend on data element size configuration information (e.g., depending on whether the data element size configuration information indicates to use a cooperative mode or a non-cooperative mode described further below, different portions of the second operand src b can be used to select the partial products). For each partial product, the partial product is a selected multiple of a corresponding portion of the second operand src b, where the multiple ranges from +2*R to -2*R for a base of -4, and from +4*R to -4*R for a base of -8 (where R is a value corresponding to the selected bit portion of the second operand src b that is relevant to the given adder array 56). Different adder arrays 56 can select their partial products based on different portions of the second operand src b. As well as selecting the partial products, the partial product selection circuit 54 can also include circuitry for generating the multiple values that can be used to select the partial products, for example including offset circuitry, negative number circuitry, and / or addition circuitry to generate the required 2*R to -2*R or 4*R to -4*R multiples for each portion of the second operand src b. The circuitry for generating the multiple values based on src b can operate in parallel with the Booth encoding circuitry 52 that generates the partial product selection indicator 62 based on the first operand src a.

[0141] Each adder array 56 receives its partial product set and, when enabled based on a corresponding enable control signal 64 provided by enable control circuit 60, adds its partial products to generate a corresponding product representation value 66 representing the product of a respective pair of bit portions selected from the first operand src a and the second operand src b. A given product representation value 66 represents the numerical result M*R of the product of M (the quantity represented by the selected bit portions of the first operand src a) and R (the quantity represented by the selected portions of the second operand src b). The portions from which M and R are selected can differ for each adder array 56. To speed the addition of the partial products, each adder array 56 can be implemented as a carry-save-adder tree that performs a series of carry-save additions (rather than carry-propagate additions), which reduces processing time by allowing the additions on different bit channels to be processed in parallel because the additions on one bit channel are not dependent on the carry generated on a lower bit channel. Thus, a given adder's product representation value 66 can be represented using a carry-save representation of summand and carry terms. To generate a binary result of M*R represented in two's complement, this can require further addition of the summand and carry terms using a carry-propagate adder (not shown in FIG. 1). Figure 2

[0142] The size of the respective adder arrays 56 is designed to handle different data element configurations within the first and second operands. For example, one subset of adder arrays 56 can implement the addition of partial products for respective pair-wise multiplications of 8-bit data element pairs within the first and second operands. A second subset of adder arrays 56 can implement partial product addition for respective pair-wise multiplications of 16-bit data element pairs within the first and second operands. A third subset of adder arrays 56 can implement partial product addition for respective pair-wise multiplications of 32-bit data element pairs within the first and second operands. It should be understood that this is just one example of different data element configurations that can be implemented. However, it can be useful to provide separate distinct adder arrays sized to accommodate each data element size configuration rather than implementing all data element sizes using a single larger adder array, as this can be more energy efficient due to its allowing the deactivation of adder arrays 56 corresponding to data element size configurations not needed for a given multiplication operation to save power.

[0143] ​In this example, each adder array 56 has its own independent enable control signal 64 that is independently set by the enable control circuit 60 to independently control whether each adder array is currently enabled or disabled. For example, each enable control signal 64 can be a clock signal that clocks components of the adder array 56, so the enable control circuit 60 can disable a given adder array by clamping the corresponding clock signal to a fixed value. By preventing the clock signal from toggling, the adder array can be disabled and save dynamic power. Other implementations can use different forms of enable control, such as power gating that controls the enable / disable of the adder array 56 by turning the adder array 56 on / off from a power supply node that it is coupled to or isolated from.

[0144] Other examples can control the independent enable / disable of adder arrays at a coarser granularity on a per-subset basis. For example, the first subset, second subset, and third subset of adder arrays 56 described above can each be provided with independent enable control signals 64, but the adder arrays within the same subset can be collectively enabled / disabled based on the same enable control signal. In practice, however, as shown in the example of FIG. 5, the first subset of adder arrays 56 can be collectively enabled / disabled by the enable control signal 64 for the first subset, the second subset of adder arrays 56 can be collectively enabled / disabled by the enable control signal 64 for the second subset, and the third subset of adder arrays 56 can be collectively enabled / disabled by the enable control signal 64 for the third subset. Figure 2 The example shown in FIG. 5 with independent enable / disable control for each adder array 56 can provide greater power saving opportunities, for example, by allowing the adder arrays 56 corresponding to masked data elements of the first operand / second operand (elements masked by the prediction) to be disabled to save power, or by allowing the adder arrays 56 within a subset to be disabled when a multiplication operation applied to operands of a length shorter than the maximum supported size does not require them.

[0145] The respective adder arrays 56 can also cooperate to implement larger multiplications, such as multiplications of src_a and src_b that include wider bit portions of the first operand and second operand that include all of the magnitude indicator bits. When the data element size configuration information indicates that the cooperative mode is to be used, the respective product representation values 66 generated by at least one subset of adder arrays (e.g., all of the adder arrays) are added together by the cooperative mode adder 58 to produce a product result that indicates a product corresponding to a wider portion of src_a and src_b than would be considered by any individual adder array. For example, in implementations discussed in more detail below, the cooperative mode implements 64-bit * 64-bit multiplications by adding product representation values 66 generated by various adder arrays 56 that would handle smaller 8-bit, 16-bit, or 32-bit multiplications in a non-cooperative mode, and additional product representation values generated by additional adder arrays 56 that handle the spare bits of the 64-bit multiplication that are not covered by the other adder arrays 56.

[0146] Thus, the larger adder array for the multiplier is constructed from several smaller (sub) arrays 56. Some of the smaller sub-arrays are sized (in terms of power and area) to be used in non-cooperative mode to multiply with smaller data types. The sub-arrays 56 can also be cooperatively used to construct larger logical arrays (e.g., for the largest data type).

[0147] In the example discussed below, it should be recognized that there is an opportunity to save area by implementing the Booth encoding of the sub-arrays 56 from the same operand src_a or the same sub-section of that operand src_a. They do so whether they are running native size multiplication in non-cooperative mode or working together in cooperative mode to compute larger data types. The area and timing pressure reduction solution is to pull the Booth encoders out of each sub-array and implement a set of Booth encoders (collectively referred to as shared Booth encoding circuitry 52) shared among all of the adder arrays 56 such that at least two of the adder arrays each receive a partial product selected based on the same partial product selection indicator 62 corresponding to the Booth encoding of the same Booth digit generated by the same set of hardware circuitry logic in the shared Booth encoding circuitry.

[0148] Figure 3 An alternative approach for comparison is shown that is based on private Booth encoding circuitry specific to each adder array 56. In this example, there are 2 32-bit adder arrays 56, 4 16-bit adder arrays 56, and 8 8-bit adder arrays 56 that are each dedicated to handling partial product addition for 32*32-bit, 16*16-bit, and 8*8-bit multiplications, respectively. There is also an "extra bit" adder array 56 that adds only the portions of src_a and src_b needed for cooperative 64-bit multiplication. Each adder array 56 has a corresponding portion of the partial product selection circuitry 54 including multiplexers for selecting each partial product to be added by that adder array. Each adder array 56 also has its own private instance of Booth encoding circuitry 70 for Booth encoding the first operand src_a to generate partial product selection indicators used by the partial product selection circuitry 54 to select partial products for that adder array 56. There is no sharing of Booth encoders among the adder arrays, each adder array 56 receives partial products selected based on Booth encoding generated by a separate segment of hardware circuitry logic not used by any other adder array 56.

[0149] When the data element size configuration to be used is 8-bit, 16-bit, or 32-bit, a corresponding subset of the adder arrays is enabled, and each adder array within that subset receives partial products selected from a pair of corresponding 8-bit / 16-bit / 32-bit data elements within the first operand src a and the second operand src b. The results of each adder array can be assembled into a vector result (e.g., by concatenating the results of the adder arrays in the order in which the adder arrays are arranged, or by concatenating the results of the adder arrays in the order in which the adder arrays are arranged and then reversing the order of the concatenated results, or by concatenating the results of the adder arrays in the order in which the adder arrays are arranged and then reversing the order of the concatenated results and then shifting the results of the adder arrays within the reversed order by a number of positions equal to the number of adder arrays in the subset, or by some other suitable order). Figure 4 Figure 2 In the non-cooperative mode, the adder arrays corresponding to data element size configurations that are not in use can be disabled by the enable control circuit 60 for power conservation. Optionally, the adder arrays corresponding to the data element size configuration in use can also be disabled, e.g., if the corresponding data elements on which they act are masked by the predicate, or because they are not needed for an operation on operands of a shorter length than the maximum operand length supported.

[0150] The adder arrays can be operable such that a subset of the adder arrays for more than one of the data element size configurations is enabled in parallel to produce a first vector result corresponding to one size configuration (e.g., 32-bit) and a second vector result corresponding to a second size configuration (e.g., 16-bit) (based on the same source operands src a, src b). This can require duplication of the result assembly circuitry to allow multiple independent results to be output in the same cycle.

[0151] In the 64-bit cooperative configuration, all of the adder arrays are enabled, and the respective product representation values 66 generated by the adder arrays are further added by the cooperative mode adder 58 to produce a 64-bit multiplication result. Optionally, for a multiply-accumulate operation, an additional adder 72 can add the multiplication result (or the multiplication result for each data element lane) to a corresponding element of a third operand (still in carry-save form to speed up the additional adder 72 compared to a carry-propagate adder).

[0152] Optionally, the carry-in and save-out terms output by the result assembly circuitry, the cooperative mode adder 58, or the additional adder 72 can be added by a carry-propagate adder to produce a result in two's complement representation, but this is not necessary since the multiplication operation is typically one of a series of multiply-accumulate operations, and thus it can be more efficient to hold the result in carry-save form to allow the additional adder 72 to perform a faster addition on a previous accumulated result also in carry-save form (where the carry-propagate operation to convert to two's complement is delayed until after the final accumulation is performed).

[0153] In contrast, Figure 4 ​An implementation is illustrated in which the Booth encoders for each of the adder arrays 56 are pulled out of the adder arrays 56 and shared across all of the adder arrays 56. In addition to the sharing of the Booth encoders, the partial product selection circuit 54 and adder tree 56, result assembly circuit / cooperative mode adder 58, and vector adder 72 are the same as in Figure 3 and thus Figure 4 these features of Figure 3 are discussed above with respect to

[0154] Since the Booth encoders are shared among the adder arrays, this means that more than one adder array 56 receives partial products that are selected using the same partial product selection indicator generated by the same Booth encoder (the same piece of hardware circuit logic). It should be recognized that all subarrays 56 should receive partial products based on the same section or sub-section of Booth encoding of one input operand src a, and thus there can be a set of global Booth encoders across a 64-bit input operand. These shared global Booth encoders supply each of the 15 subarray multipliers. When the Booth encoders are shared among adder arrays 56 that support different data element size configurations, the Booth digit that is across a data element boundary in one configuration but at a mid-point within a data element for another configuration can require different input bits for different element size configurations. Thus, for example, for an 8-bit boundary within the first operand, there is a special Booth encoder with selection logic to select whether to treat the low bits of the Booth digit as 0 or to select from src a. This is described in more detail below.

[0155] Sharing the Booth encoders among the adder arrays 56 enables a substantial savings in circuit area and power. In Figure 3 the approach shown, to support signed multiplication operations, eight 8-bit multipliers each have 5 private Booth encoders; four 16-bit multipliers each have 9 private Booth encoders; two 32-bit multipliers each have 17 private Booth encoders; and the extra bit multiplier has 33 private Booth encoders. Thus, there are a total of 143 private Booth encoders.

[0156] In contrast, in the Figure 4 with shared global Booth encoders supplying all 15 adder arrays, there are now 32 Booth encoders to support signed multiplication (plus the 8 smaller Booth encoders for unsigned multiplication described further below). This represents a substantial reduction in area compared to Figure 3 .

[0157] The significant reduction in encoder size also allows for the economical splitting of the flip-flop stage between Booth coding 52, Booth multiplexing (partial product selection 54), and partial product compression (partial product addition 56). Previously, 143 sets of Booth-coded values ​​would have needed to be modified for this split; now only 32 sets need modification (plus 8 single-bit flip-flops for unsigned multiplication). This is similar in size to the modified 64-bit operands, which no longer need to be encoded in a second flip-flop stage due to the large number of flip-flops required for this operation. This is the timing pressure reduction achieved by this technique.

[0158] Figures 5 to 12 A multiplication circuit 50 is illustrated in more detail as an example of implementing base-4 Booth multiplication (it should be understood that other examples may use higher bases). Figure 5 Examples of the same Figure 4 The same adder array 56 and shared Booth encoding circuit 52 are used, but each subset of the adder array (32-bit subset, 16-bit subset, and 8-bit subset) is expanded to show the individual adder array 56 in each subset. Therefore, there are two 32-bit adder arrays 56, four 16-bit adder arrays, eight 8-bit adder arrays, and an "extra bit" adder array 56 used in cooperative 64-bit mode. The enable control input for each adder array 56 is... Figure 5 The markers indicate the data element size configuration for which the adder array is enabled. Thus, an 8-bit adder array is enabled for 8-bit and 64-bit configurations but can be disabled for 16-bit and 32-bit configurations; a 16-bit adder array is enabled for 16-bit and 64-bit configurations but can be disabled for 8-bit and 32-bit configurations; and a 32-bit adder array is enabled for 32-bit and 64-bit configurations but can be disabled for 8-bit and 16-bit configurations. For example, if an individual adder array is not needed because the corresponding data element of the first operand / second operand is masked by prediction, then the adder array can be disabled even if the currently selected configuration is one associated with that adder array.

[0159] Figures 6A to 6DEach diagram is illustrated as a diamond, where the top edge of the diamond represents bits of the first operand src a and the right side edge of the diamond represents bits of the second operand src b, where each operand src a, src b has 64 bits. The diamond shape represents the partial product additions that would be performed (adding 64 partial products, with a 1-bit offset between each row, corresponding to src a if the corresponding bit of src b is 1, and to 0 if the corresponding bit of src b is 0) in the case of performing a full-width 64-bit * 64-bit multiplication using the traditional "textbook" long multiplication method. It should be understood that, since the Booth multiplication scheme is used, the adder arrays do not actually add 64 partial products in this way, but it is useful for a theoretical example to consider adding 64 partial products to explore how the portions of src a and src b are mapped to the adder arrays.

[0160] Figure 6A The mapping of the adder arrays in a configuration with a 32-bit data element size is illustrated. In this case, the 64-bit operands src a and src b are treated as a pair of 32-bit data elements, and two 32-bit adder arrays (32-bit mul 0 and 32-bit mul 1) are assigned to add the partial products corresponding to respective full word (32-bit) multiplications FW 0 and FW 1, where FW 0 corresponds to the multiplication of the lower elements src a [31 :0] and src b [31 :0], and FW 1 corresponds to the multiplication of the upper elements src a [63:32] and src b [63:32]. As explained further below, Figure 6A The area denoted with "0" in the middle is not assigned to any of the adder arrays 56.

[0161] Figure 6BThe mapping of the adder array in a 16-bit data element size configuration is illustrated. In this case, the 64-bit operands src_a and src_b are each treated as four 16-bit data elements, and four 16-bit adder arrays (16-bit mul 0 through 16-bit mul 3) are assigned to add the partial products corresponding to the four respective halfword (16-bit) multiplication HW_0 through HW_3, respectively, where HW_0 corresponds to the multiplication of the lower elements src_a[15:0] and src_b[15:0], HW_1 corresponds to the multiplication of the next lower elements src_a[31:16] and src_b[31:16], HW_2 corresponds to the multiplication of the second higher elements src_a[47:32] and src_b[47:32], and HW_3 corresponds to the multiplication of the upper elements src_a[63:48] and src_b[63:48].

[0162] Figure 6C The mapping of the adder array in an 8-bit data element size configuration is illustrated. In this case, the 64-bit operands src_a and src_b are each treated as eight 8-bit data elements, and eight 8-bit adder arrays (8-bit mul 0 through 8-bit mul 7) are assigned to add the partial products corresponding to the eight respective byte multiplications Byte0 through Byte7, respectively (each operating on a respective pair of 8-bit elements within src_a and src_b).

[0163] The 8-bit, 16-bit, and 32-bit configurations are examples of non-cooperative modes, in that each adder array produces an independent partial product representation value that does not require further addition with other partial product representation values to produce the multiplication result (which is subject to assembly into a vector result value, and possibly conversion from a carry-save form to two's complement).

[0164] In contrast, Figure 6D The adder arrays 56 are shown together in a cooperative mode for producing respective partial product representation values that represent sub-portions of a wider 64-bit * 64-bit multiplication. The respective partial product representation values from each adder array can be added together (with appropriate bit channel offsets depending on the relative positions of the portions of the multiplication corresponding to the adder array, within the 64-bit multiplication "diamond" shown), to generate a result value representing the product of the 64-bit portions of src_a and src_b. The relative alignment of the addition of the partial product representation values is discussed in more detail below with respect to Figure 6D The relative alignment of the addition of the partial product representation values is discussed in more detail below with respect to Figure 13

[0165] Figure 6D ​The way in which each of the adder arrays 56 is mapped onto a portion of a 64-bit multiplication is shown. Two full-word (32-bit) adder arrays each implement the portion of the multiplication corresponding to src_a[31 :0]*src_b[31 :0] and src_a[63:32]*src_b[31 :0], respectively. Four half-word (16-bit) adder arrays each implement the portion of the multiplication corresponding to src_a[15:0]*src_b[55:40], src_a[31 :16]*src_b[55:40], src_a[47:32]*src_b[55:40], and src_a[63:48]*src_b[55:40], respectively. Eight byte (8-bit) adder arrays each implement the portion of the 64-bit multiplication corresponding to src_a[7:0]*src_b[63:55], src_a[15:8]*src_b[63:55], src_a[23:16]*src_b[63:55], src_a[31 :24]*src_b[63:55], src_a[39:32]*src_b[63:55], src_a[47:40]*src_b[63:55], src_a[55:48]*src_b[63:55], and src_a[63:56]*src_b[63:55], respectively. This leaves a portion of the extra bits not covered by the 8-bit, 16-bit, and 32-bit adder arrays, and so an extra-bit adder array performs the portion of the partial product addition corresponding to the remaining portion src_a[63:0]*src_b[39:32] not covered elsewhere.

[0166] It should be understood Figure 6D Only one way in which the required partial product addition of a 64-bit multiplication can be split among the respective adder arrays is shown. Other examples can map the individual 8-bit, 16-bit, or 32-bit adder arrays in different ways, such that the "extra-bit" reserve is available Figure 6D for the region of src_a[63:0]*src_b[39:32] shown.

[0167] Figures 6A to 6D The advantages of the approach discussed above of using separate, independently enabled / disabled subsets of adder arrays to target specific data element size configurations are also helped to be illustrated. An alternative would be to provide a single large monolithic adder array having enough adder elements and sized to be able to handle operand widths of 64-bit*64-bit multiplications. Such a large 64-bit adder array could then be mapped to handle Figure 6A and Figure 6B corresponding to the src_a[63:0]*src_b[39:32] region marked as 0 (and with Figure 6Call partial product bits corresponding to regions of the multiplication that do not correspond to regions of the multiplication that correspond to non-shaded regions in FIG. 6) are set to 0 to re-purpose the smaller 8-bit * 8-bit, 16-bit * 16-bit, or 32-bit * 32-bit multiplication to be performed. In the carry-save adder tree, there is no carry propagation between the lanes, and thus each lane of the addition remains independent within the final 64-bit result output by the adder tree. This means that if the non-zero partial product bits are only in the regions labeled FW_0 and FW_1 (for Figure 6A ), HW_0 through HW_3 (for Figure 6B ), and Byte0 through Byte7 (in Figure 6C ), then the result is effectively a vector of multiplication results. However, a disadvantage of using this approach is that even if some of the partial product lanes have been set to 0, the adder elements in some lanes of the multiplication that are labeled "0" still need to propagate through the bits added in the earlier rows of the adder array where there are non-zero bits. For example, the region labeled 0 corresponding to src_a[31 :0] * src_b[63 :32] in FIG. 6 still needs to propagate through the addition results in the region labeled FW_0. Thus, even in the possible "unused" portions of the large single-block adder array, the adder elements of the adder array can still need to be enabled (powered and clocked) to allow the final multiplication result to correctly represent the element-wise vector multiplication result. Figure 6A

[0168] In contrast, with the approach shown in FIGS. 6, 7, and 8, where the adder array is sized according to the particular data element size configuration and is independently enabled / disabled by the enable control circuit 60 that is separate from the adder array that implements another data element size configuration, power can be saved by disabling the unused adder arrays. Only when the cooperative mode is selected to perform a full 64-bit multiplication are all of the adder arrays enabled. Thus, a multiplication circuit that supports multiple data element size configurations is possible with greater energy efficiency than using a single 64-bit single-block adder array. Figure 2 Figure 4 Figure 5 Thus, while the regions of the multiplication corresponding to "0" (or non-shaded) are shown for ease of understanding in FIG. 6, it should be understood that in the multiplication circuit 50 described above, none of these 0 regions need to have any addition occur at all; power can be saved by using smaller native-sized adder arrays and disabling other adder arrays that are not needed for the current data element size configuration selected for a given multiplication operation.

[0169] Figures 6A to 6C

[0170] Figure 7 ​​​​​The shared Booth encoding circuit 52 and the mapping of the respective bits [63:0] of the Booth number to the first operand src_a encoded by the Booth encoding circuit 52 are illustrated. The shared Booth encoding circuit 52 comprises 32 Booth encoders benc0 to benc31 each generating a corresponding partial product selection indicator indicative of a corresponding Booth number of the source operand src_a. The position of the bits encoded by each Booth encoder is shown in Figure 7 For example, the Booth encoder benc10 encodes bits [21:19] of src_a and the Booth encoder benc11 encodes bits [23:21] of src_a. It can be seen that there is an overlap of 1 bit between the respective Booth numbers encoded by consecutive Booth encoders.

[0171] At the lowest Booth number of each data element, the bottom bit of the Booth number is 0. When considering different element size configurations, the Booth numbers encoded by the Booth encoders benc4, benc8, benc12, benc20, benc24, benc28 cross the data element boundary for at least one data element size configuration, but are located at an intermediate position within the element for at least one other data element size configuration. Furthermore, when the cooperative 64-bit configuration is selected, the Booth number encoded by the Booth encoder benc16 crosses the data element boundary for each of the 8-bit, 16-bit and 32-bit configurations, but is at an intermediate position within the multiplied 64-bit value. Therefore, the Booth encoders marked with a dashed line have a selection circuit to select whether the bottom bit of this Booth number is a fixed value of 0b0 (for data element size configurations where the Booth number crosses the data element boundary) or is the top bit of the next least significant 8-bit portion of src_a (for data element size configurations where the Booth number is an intermediate position within the data element which does not cross the data element boundary). The text below each 8-bit portion in Figure 7 illustrates in more detail which configurations result in the selection of 0b0 and which result in the selection of the top bit of the next least significant 8-bit portion of scr_a. The other Booth encoders shown in solid line do not have such a selection circuit and can Booth encode a fixed subset of bits as shown by the mapping in Figure 7

[0172] For each Booth encoder benc0 to benc31, the Booth encoder generates a partial product selection indicator according to the Booth encoding rules for base 4 to indicate which multiple of the corresponding portion of the second operand should be selected as a partial product:

[0173] Selected bus number M i+1 , M i , M i-1 ]]> Multiples of R selected as partial products 000 0 001 +1 010 +1 011 +2 100 -2 101 -1 110 -1 111 0

[0174] ​If signed values represented in two's complement format are to be multiplied, then the top bit of each element will be the sign bit. Thus, for the example using radix-4, assuming the data element size has a number of bits that is a power of two, although each data element will nominally require encoding of one additional Booth numeral corresponding to the sign extension of the top bit of that data element, that additional Booth numeral will always have the value of 000 and 111 (since the sign extension causes the top bit of that data element to be copied to the other two bits of that Booth numeral), and since the Booth encoding rules shown above give a multiple of 0 for both 000 and 111, this will not require any additional partial products to be added. Thus, if only signed multiplication is to be supported, for the radix-4 example, it is sufficient to use the Booth encoders bencO to benc31 shown. Figure 7

[0175] However, if unsigned multiplication is to be supported for radix-4, then the top bit of each element can be either 0 or 1, and padded with 0s (non-sign extension bits) to form the top Booth numeral in each data element. This means that the possible values of the top Booth numeral in each element are 000 or 001, and since the 001 alternative requires a +1*R multiple of additions, additional partial products can be required. So, for the 8-bit example, to support unsigned multiplication with the minimum data element size configuration, each 8-bit portion of the first operand src_a is also provided with an additional unsigned Booth encoder "umul", so there are 8 such unsigned Booth encoders, which are labelled umul4, umul8, umul12, umul16, umul20, umul24, umul28, umul32 respectively (the numbers used to label the unsigned Booth encoders umul4 to umul28 are chosen to indicate that the Booth numerals encoded by these unsigned Booth encoders have the same effectiveness as the Booth numerals encoded by the Booth encoders benc4, benc8, benc12, benc16, benc20, benc24, benc28 respectively, and that umul32 encodes a more effective Booth numeral than the Booth numeral encoded by benc31). Each unsigned Booth encoder can have simpler hardware circuitry logic than the other Booth encoders bencO to benc31, since it only needs to select from two options (0 or +1*R) respectively based on whether the top bit of the corresponding 8-bit portion of src_a is 0 or 1.

[0176] ​On the other hand, if a higher base is used (e.g., base-8), or if base-4 is applied to data elements of a size other than exactly a power of 2, depending on the base used and the data element size, some Booth numbers can align with respect to the element boundary such that the sign extension of the top bit representing the data element can include two or more bits of src_a, rather than just one bit as in the base-4 example shown above. In this case, even for signed multiplication, the most significant Booth number in a given data element can have a Booth encoding corresponding to a multiple other than zero, and in this case, the unsigned Booth encoder umul for the most significant Booth number in each data element can be replaced with a full Booth encoder similar to the benc example shown for the less significant Booth numbers in the data element. Figure 7

[0177] Figure 8 is a table summarizing the partial product selection performed by the partial product selection circuit 54. For each adder array 56, the table indicates which bits of src_b to use in the non-cooperative mode (labeled "non-64b," referring to 8-bit, 16-bit, or 32-bit configurations) and the cooperative mode (labeled "64b") to form the partial products to be added by that adder array 56, respectively. It can be seen that each adder array has partial products selected based on different bit portions of src_b depending on whether the non-cooperative mode or the cooperative mode is used, and it can be seen that the selected portions of src_b correspond to the mappings shown in Figures 6A to 6D for the respective data element size configurations.

[0178] Figure 8 It is also indicated for each adder array which Booth encoders to use to select each partial product, both for signed multiplication (where the unsigned Booth encoder umul can be ignored by setting the corresponding partial product to 0) and unsigned multiplication (where each adder array also takes into account a partial product selected by one of the unsigned Booth encoders).

[0179] For example, for the adder array "8-bit mul 0," the Booth encoders used are benc 3, benc 2, benc 1, benc 0 (and for unsigned multiplication, then umul 4 is used). From Figure 7 it can be seen that this means that the partial product is selected based on the Booth encoding derived from src_a[7:0]. This corresponds to the "Byte0" region of the multiplication shown in Figure 6C (in the 8-bit size configuration where the partial product is based on src_b[7:0]), or Figure 6D (for the 64-bit cooperative size configuration where the partial product is based on src_b[63:56]). ​

[0180] It is noted that the extra-bit adder array 56 uses all of the Booth encoders benc31 to benc0 (as it is mapped to the "extra-bit mul" region that requires all 64 bits of src_a in Figure 6D For each other adder array 56, only an appropriate subset of the Booth encoders is used.

[0181] However, it can be seen that each of the other adder arrays can be considered to be part of a subset of adder arrays corresponding to a particular data element size configuration (e.g. an 8-bit adder array subset comprising 8-bit mul0 to 8-bit mul3, a 16-bit adder array subset comprising 16-bit mul0 to 16-bit mul3, and a 32-bit adder array subset comprising 32-bit mul0 and 32-bit mul1), and within a given one of these subsets, the adder arrays of that subset collectively use all of the Booth encoders benc31 to benc0, but any one of the adder arrays of that subset uses only some of these Booth encoders. Each of the 8-bit, 16-bit and 32-bit subsets of adder arrays also uses one unsigned Booth encoder to support unsigned multiplication.

[0182] It is also noted that the Booth encoders are shared between subsets of adder arrays. Any given one of the Booth encoders benc0 to benc31 is used by one adder array in each subset. For example, Booth encoder benc0 is used by 8-bit mul0, 16-bit mul0 and 32-bit mul0 (as well as by the extra-bit mul adder array). Booth encoder benc27 is used by 8-bit mul6, 16-bit mul3 and 32-bit mul1 (as well as by the extra-bit mul adder array). Thus, there are several adder arrays each receiving a respective partial product selected based on a completely identical partial product select indicator generated by the same Booth encoder (i.e. the same Booth encoder logic segment) of the shared Booth encoding circuitry. This greatly reduces the amount of hardware circuitry area required to implement the Booth encoders, compared to the non-shared approach shown in Figure 3

[0183] It is noted that, Figure 6D The particular mapping of adder arrays shown has the advantage that which bits of src_a are selected by the Booth encoders for a given adder array does not need to change depending on whether the cooperative or non-cooperative mode is selected. To avoid the selection circuitry at each Booth encoder needing to change the selection of src_a bits based on the cooperative / non-cooperative mode, it is also possible to select to retain Figure 6D the src_a allocation shown but to switch which parts of src_b are assigned to each adder array 56 (e.g. to switch Figure 6D ​other mappings (e.g., of the order of some of the rows shown).

[0184] Figure 9 The partial product addition performed by each of the 8-bit adder arrays is illustrated in more detail. Each of the 8-bit adder arrays has a carry-save-adder tree with a sufficient number of stages and a sufficient bit width to add 5 8-bit partial products (given the radix-4 Booth multiplication implementation, using a 2-bit offset in alignment between successive partial products, then this would be a 3-bit offset for radix-8) to produce a 16-bit result as the product representation value for that adder array (represented as a 16-bit carry term and a 16-bit sum term in carry-save representation). For each 8-bit adder array 56, Figure 9 The 5 partial products added at each row of the adder array are labeled in the figure, with notation ppi (i = 0...31) referring to partial products selected based on partial product selection indicators output by Booth encoder benci, and pp-umj referring to partial products selected based on partial product selection indicators output by unsigned Booth encoder umulj (j being one of 4, 8, 12, 16, 20, 24, 28, 32). The same convention is used in Figure 10A , Figure 10B , Figure 11A , Figure 11B and Figure 12 for partial product labels in each of the figures.

[0185] Figure 10A and Figure 10B The partial product addition performed by each of the 16-bit adder arrays is illustrated in more detail. Each of the 16-bit adder arrays has a carry-save-adder tree with a sufficient number of stages and a sufficient bit width to add 9 16-bit partial products (again using a 2-bit offset between adjacent partial products) to produce a 32-bit product representation value in carry-save representation.

[0186] Figure 11A and Figure 11B shows similar partial product addition for a 32-bit adder array, this time adding 17 32-bit partial products (with a 2-bit offset between each row) to produce a 64-bit carry-save result as the corresponding product representation value.

[0187] Figure 12 illustrates partial products for an extra-bit adder array that adds 33 partial products each including 8 bits, each column again having a 2-bit offset, to produce a 72-bit carry-save result as a product representation value.

[0188] It should be understoodFigures 9 to 12 The adder array diagram in Figure 8 illustrates the logical sequence of additions to be performed, where the carry-save result is represented by the sum of all bits in a given position within the "adder array diamond" representing the corresponding column within the diamond. However, it is not necessary to arrange the physical hardware circuitry into such a diamond, or to add the bits of the corresponding partial products in the same row order as shown in Figure 8. Any hardware circuit logic that generates equivalent results can be used (e.g., the sequence of bringing the partial products together can be reordered in some way). Furthermore, for a given column of the given adder array shown in Figure 8, the addition of the bits in different portions of that column can be performed in parallel before adding the results of those sub-column additions together, so it is not necessary to perform the partial product addition as a series of sequential additions each performed in row order. For example, for the 8-bit adder array 8-bit mulO, for the column that requires adding bits from ppO, pp1, pp2, pp3, this can be done as two parallel additions of ppO+pp1 and pp2+pp3 before adding the results together, or two parallel additions of ppO+pp2 and pp1+pp3 before adding the results together, or any other grouping of additions that gives the final result ppO+pp1+pp2+pp3. Thus, it should be understood that Figures 9 to 12 The logical sequence of additions used is shown, but the physical hardware can generate equivalent results in various ways. Figures 9 to 12 Figures 9 to 12 The alignment of the additions performed by the cooperative mode adder 58 in the 64-bit cooperative mode is illustrated. The cooperative mode adder 58 adds the respective product representation values 66 generated by each adder array 56 to produce a 128-bit carry-save result representing the numerical product result of the two 64-bit values represented by src_a[63:0] and src_b[63:0]. The bracketed terms [x:y] in Figure 8 illustrate the alignment of each product representation value with the bit lanes of the 128-bit addition, i.e., the top bits of the product representation value are aligned with bit lane x, and the bottom bits of the product representation value are aligned with bit lane y. The relative offset between the product representation values in the cooperative mode addition is selected based on the relative significance of those product representation values, as shown in Figure 9.

[0189] Figure 13 Figure 13 Figure 6D

[0190] Figure 14 ​​​A method of performing a multiplication operation using the previously described multiplication circuit 50 on a first operand src a and a second operand src b is illustrated. At step 300, the first operand src a is Booth encoded using the shared Booth encoding circuit 52 to generate partial product selection indicators 62 each corresponding to the Booth encoding of a respective Booth digit of the first operand. The shared Booth encoding circuit 52 is shared between the adder arrays 56, and the adder arrays 56 include separate instances of hardware circuitry, with at least two subsets of the adder arrays 56 having separate enable control signals 64 for independently controlling whether each subset of the adder arrays 56 is enabled or disabled. As previously mentioned, in some implementations, each adder array 56 can have its own independent enable control signal 64 separate from any other adder array 56 to be independently enabled / disabled.

[0191] At step 302, the partial product selection circuit 54 selects respective sets of partial products to be added by each adder array 56 based on the second operand src b and the partial product selection indicators 62 generated by the shared Booth encoding circuit 52. At least two of the adder arrays 56 receive partial products selected based on the same partial product selection indicator based on the Booth encoding of the same Booth digit of the first operand (e.g., as shown by the adder arrays 8-bit mul 1, 16-bit mul 0, 32-bit mul 0, and extra-bit mul all receiving partial products selected based on the partial product selection indicator indicating the Booth encoding of the same Booth digit src a[13:11] by the Booth encoder benc 6). Figure 7 and Figure 8 The adder arrays 8-bit mul 1, 16-bit mul 0, 32-bit mul 0, and extra-bit mul all receive partial products selected based on the partial product selection indicator indicating the Booth encoding of the same Booth digit src a[13:11] by the Booth encoder benc 6).

[0192] At step 304, any enabled adder arrays 56 (those for which the enable control circuit 60 has asserted the enable control signal 64 to enable the adder array) add their respective partial products to generate respective product representation values each representing the multiplication result of a respective pair of bit portions selected from the first operand src a and the second operand src b. Any adder arrays that are currently disabled (e.g., based on clocking the clock signal to a fixed value) do not add their partial products.

[0193] At step 306, if the current data element size configuration is based on a cooperative mode, then the cooperative mode adder 58 adds the respective product representation values obtained by each adder array 56. In a non-cooperative mode (e.g., the 8-bit, 16-bit, or 32-bit configurations shown above), step 306 is not performed, and instead the respective product representation values can be assembled into a result without further addition. In some implementations, the product representation values can be truncated to fit into a result of the same size as the input operands src a, src b. Alternatively, some variations can produce a result that is double the size of the input operands to preserve all bits of the product representation values.

[0194] At step 308, if a multiply-add operation is performed, then a further addition of the multiply result (either the result of a 64-bit multiplication produced at step 306 in a cooperative mode, or the assembled multiply result in a non-cooperative mode configuration) with a third operand is performed, e.g., using the vector adder 72 (the third operand can be a SIMD operand in a non-cooperative mode).

[0195] The concepts described herein can be embodied in computer readable code for manufacturing apparatus embodying the described concepts. For example, the computer readable code can be used in one or more stages of a semiconductor design and fabrication process, including electronic design automation stages, to manufacture integrated circuits including apparatus embodying the concepts. The computer readable code described above can additionally or alternatively implement definition, modeling, simulation, verification, and / or testing of apparatus embodying the concepts described herein.

[0196] For example, computer readable code for manufacturing apparatus embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representative of the concepts. For example, the code can define a register transfer level (RTL) abstraction defining one or more logic circuits for defining apparatus embodying the concepts. The code can define an HDL representative of one or more logic circuits embodying the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (very high speed integrated circuit hardware description language), as well as intermediate representations such as FIRRTL. The computer readable code can provide a definition embodying the concepts using a system level modeling language, such as SystemC and SystemVerilog, or other behavioral representations of the concepts that can be interpreted by a computer to implement simulation, functional and / or formal verification, and testing of the concepts.

[0197] Additionally or alternatively, the computer readable code can define a detailed description of an integrated circuit component implementing the concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. One or more netlists or other computer readable representations of the integrated circuit component can be generated by applying one or more logic synthesis processes to the RTL representation to generate definitions for manufacturing devices implementing the invention. Alternatively or additionally, one or more logic synthesis processes can generate a bitstream from the computer readable code that is loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA can be deployed for the purposes of verifying and testing concepts in integrated circuits prior to manufacturing, or the FPGA can be deployed directly in a product.

[0198] The computer readable code can include a mixture of code representations for manufacturing devices, such as a mixture including one or more of RTL representations, netlist representations, or another computer readable definition for semiconductor design and manufacturing processes to manufacture devices embodying the invention. Alternatively or additionally, the concepts can be defined in a combination of computer readable definitions for manufacturing devices in semiconductor design and manufacturing processes and computer readable code defining instructions for execution by the defined devices after manufacturing.

[0199] Such computer readable code can be provided in any known transient computer readable medium, such as wired or wireless transmission code over a network, or a non-transient computer readable medium, such as a semiconductor, magnetic or optical disk. Integrated circuits manufactured using the computer readable code can include components such as one or more of a central processing unit, a graphics processing unit, a neural processing unit, a digital signal processor, or other components individually or collectively embodying the concepts.

[0200] The concepts described herein can be embodied in a system including at least one packaged chip. The multiplication circuit described previously is implemented in at least one packaged chip (implemented in one particular chip of the system, or distributed over more than one packaged chip). The at least one packaged chip is assembled on a board with at least one system component. The chip-inclusive product can include a system assembled on a further board with at least one other product component. The system or chip-inclusive product can be assembled into a housing or onto a structural support such as a frame or a blade.

[0201] As Figure 15As shown, one or more packaged chips 400 are manufactured by a semiconductor chip manufacturer, with the multiplication circuit described above implemented on one chip or distributed across two or more chips. In some examples, the chip product 400 made by the semiconductor chip manufacturer can be provided as a semiconductor package including a protective enclosure (e.g., made of metal, plastic, glass, or ceramic) containing the semiconductor device implementing the multiplication circuit described above and connectors (such as flat, balls, or pins for connecting the semiconductor device to an external environment). In cases where more than one chip 400 is provided, these chips can be provided as separate integrated circuits (provided as separate packages), or can be packaged by the semiconductor provider into a multi-chip semiconductor package (e.g., using an interposer, or by using three-dimensional integration to provide a multi-layer chip product including two or more vertically stacked integrated circuit layers).

[0202] In some examples, a collection of small chips (i.e., small modular chips with specific functionality) can itself be referred to as a chip. The small chips can be individually packaged in semiconductor packages and / or packaged together with other small chips into a multi-small-chip semiconductor package (e.g., using an interposer, or by using three-dimensional integration to provide a multi-layer small-chip product including two or more vertically stacked integrated circuit layers).

[0203] The one or more packaged chips 400 are assembled on a board 402 along with at least one system component 404. For example, the board can comprise a printed circuit board. The board substrate can be made of any of a variety of materials, e.g., plastic, glass, ceramic, or a flexible substrate material such as paper, plastic, or fabric material. The at least one system component 404 comprises one or more external components that are not part of the one or more packaged chips 400. For example, the at least one system component 404 can comprise, e.g., any one or more of: another packaged chip (e.g., provided by a different manufacturer or produced at a different process node), an interface module, a resistor, a capacitor, an inductor, a transformer, a diode, a transistor, and / or a sensor.

[0204] A chip-inclusive product 410 is manufactured that includes the system 406 (including the board 402, one or more chips 400, and at least one system component 404) and one or more product components 412. The product components 412 include one or more additional components that are not part of the system 406. As an example, non-exhaustive list, the one or more product components 412 can include user input / output devices such as keypads, touchscreens, microphones, speakers, display screens, haptic devices, etc.; wireless communication transmitters / receivers; sensors; actuators for actuating mechanical motion; thermal control devices; additional packaged chips; interface modules; resistors; capacitors; inductors; transformers; diodes; and / or transistors. The system 406 and the one or more product components 412 can be assembled onto a further board 414.

[0205] The board 402 or the further board 414 can be provided on or in a device housing or other structural support such as a frame or a blade to provide a product that can be handled by a user and / or intended for operational use by an individual or a company.

[0206] The system 406 or the chip-inclusive product 416 can be at least one of an end-user product, a machine, a medical device, a computing or telecommunications infrastructure product, or an automation control system. For example, as an example, non-exhaustive list, the chip-inclusive product can be any of a telecommunications device, a mobile phone, a tablet computer, a laptop computer, a computer, a server (e.g., a rack server or a blade server), an infrastructure device, a network appliance, a vehicle or other automotive product, an industrial machine, a consumer device, a smart card, a credit card, smart glasses, avionics, a robotic device, a camera, a television, a smart television, a DVD player, a set-top box, a wearable device, a household appliance, a smart meter, a medical device, a heating / lighting control device, a sensor, and / or a control system for controlling a public infrastructure appliance such as a smart highway or a traffic light.

[0207] In this application, the term "configured to" is used to mean that an element of a device has a configuration able to perform the defined operation. In this context, a "configuration" means an arrangement or manner of interconnection of hardware or software. For example, the device can have special purpose hardware for providing the defined operation, or a processor or other processing device can be programmed to perform the function. "Configured to" does not imply that the device element needs to be changed in any way in order to provide the defined operation.

[0208] In this application, a feature list headed by the phrase "at least one of' means that one or more of the features in that list can be provided individually or in any combination. For example, "at least one of A, B, and C" covers the following options: A alone (not B or C), B alone (not A or C), C alone (not A or B), a combination of A and B (not C), a combination of A and C (not B), a combination of B and C (not A), or a combination of A, B, and C.

[0209] While the exemplary embodiments of the application have been described above in detail, it should be understood that the application is not limited to the precise embodiments, and as such, various changes and modifications can be made by those skilled in the art without departing from the scope of the application defined by the appended claims.

Claims

1. A multiplication circuit, the multiplication circuit comprising: Multiple adder arrays, each of which is used to add a corresponding set of partial products to generate a corresponding product representation value, the corresponding product representation value representing the multiplication result of a corresponding pair of bit portions selected from a first operand and a second operand, the multiple adder arrays including separate instances of hardware circuitry, the multiple adder arrays having at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled; A shared Booth coding circuit, shared among the plurality of adder arrays, is used to Booth code the first operand to generate a plurality of partial product selection indicators, each of which corresponds to the Booth code of a corresponding Booth number of the first operand. and A partial product selection circuit is configured to select, based on the second operand and the plurality of partial product selection indicators, the partial product to be added by the plurality of adder arrays; wherein: At least two of the adder arrays are configured to operate on corresponding partial products selected by the partial product selection circuit based on a common partial product selection indicator, which is generated by the shared Booth encoding circuit based on the Booth encoding of the same Booth number of the first operand.

2. The multiplication circuit of claim 1, wherein the multiplication circuit is configured to support at least two data element size configurations for multiplying corresponding pairs or more pairs of data elements selected from the first operand and the second operand, each data element size configuration corresponding to a different combination of data element sizes of the data elements selected from the first operand and the second operand.

3. The multiplication circuit according to claim 2, wherein the multiplication circuit includes an enable control circuit configured to select which adder arrays in the adder array will be enabled and which adder arrays in the adder array will be disabled based on the current data element size configuration to be used for the multiplication operation.

4. The multiplication circuit according to any one of claims 2 and 3, wherein the Booth encoding for a given Booth number of the first operand is: For at least one data element size configuration in the data element size configuration where the given Booth number crosses a data element boundary, the shared Booth coding circuit is configured to generate the Booth code for the given Booth number based on the least significant bit of the given Booth number being set to 0; and For at least one other data element size configuration in the data element size configuration where the given Booth number does not cross a data element boundary, the shared Booth encoding circuit is configured to generate the Booth code of the given Booth number based on the value of the least significant bit of the given Booth number being set to the corresponding bit of the first operand.

5. The multiplication circuit according to any one of claims 2 to 4, wherein the plurality of adder arrays comprises at least: A first subset of the adder array, each corresponding product representation value in the first subset representing the multiplication result of a corresponding pair of data elements selected from the first operand and the second operand according to the first data element size configuration; and A second subset of the adder array, each corresponding product representation value of the second subset representing the multiplication result of a corresponding pair of data elements selected from the first operand and the second operand according to the second data element size configuration.

6. The multiplication circuit of claim 5, wherein the at least two adder arrays in the adder arrays, wherein the corresponding partial products are selected based on the same partial product selection indicator, include the adder arrays in the first subset and the adder arrays in the second subset.

7. The multiplication circuit according to claim 6, wherein: When the multiplication is performed according to the configuration of the first data element size, the partial product of the adder array for the first subset of the first subset of the same partial product selection indicator is selected based on the first subset of the bits of the second operand; and When the multiplication is performed according to the second data element size configuration, the second subset of the adder array whose partial product is selected based on the same partial product selection indicator depends on the bits of the second operand.

8. The multiplication circuit according to any one of claims 5 to 7, the multiplication circuit comprising an enable control circuit, the enable control circuit being configured to: In response to the current data element size configuration information specifying that a given multiplication operation performed on the first operand and the second operand does not require the first data element size configuration, one or more enable control signals of the first subset of the adder array are set to disable the first subset of the adder array; and In response to the current data element size configuration information specifying that the given multiplication operation does not require the second data element size configuration, one or more enable control signals of the second subset of the adder array are set to disable the second subset of the adder array.

9. The multiplication circuit according to any one of claims 5 to 8, wherein the enable control circuit is configured to enable both the first subset and the second subset of the adder array by using the enable control signal to support the parallel execution of multiplication operations on the first operand and the second operand using both the first data element size configuration and the second data element size configuration.

10. The multiplication circuit according to any one of claims 5 to 9, wherein the plurality of adder arrays further comprises a third subset of the adder arrays, each corresponding product representation value of the third subset representing the multiplication result of a corresponding pair of data elements selected from the first operand and the second operand according to a third data element size configuration.

11. The multiplication circuit according to any one of claims 1 to 10, wherein the multiplication circuit comprises: A product-adding circuit, configured to add the corresponding product representations generated by the two or more adder arrays in the adder array for a multiplication operation to be performed in cooperative mode, to generate a product result value, the product result value representing the multiplication of the bit portion wider than the bit portion of the corresponding product representation value of any one of the two or more adder arrays in the adder array, for a multiplication operation in which the first operand and the second operand are used to generate the multiplication.

12. The multiplication circuit of claim 11, wherein for at least one adder array, the portion of the second operand used to form the partial product of the adder array is capable of changing depending on whether the multiplication operation will be performed in the cooperative mode.

13. The multiplication circuit of any one of claims 11 and 12, wherein the wider bit portion of the first operand and the second operand includes all magnitude indicator bits of the first operand and the second operand.

14. The multiplication circuit according to any one of claims 11 to 13, wherein the plurality of adder arrays comprises: Multiple subsets of adder arrays, the multiple subsets corresponding to different data element size configurations, each subset of the adder arrays comprising two or more adder arrays for generating a corresponding product representation value in a non-cooperative mode, the corresponding product representation value representing a multiplication of a corresponding data element pair selected from a first operand and a second operand according to the data element size configuration corresponding to the subset of the adder array, wherein in the cooperative mode, each subset of the multiple subsets of the adder arrays is assigned to a portion of the multiplication of the wider portion; and An additional adder array, used to generate additional product representations, representing the remainder of the multiplication of the wider portion excluding the portions assigned to the plurality of subsets of the adder array; and The multiplication circuit includes an enable control circuit for disabling the additional adder array in the non-cooperative mode.

15. The multiplication circuit according to any one of claims 1 to 14, wherein the plurality of adder arrays have individual enable control signals for independently controlling whether each adder array is enabled or disabled.

16. The multiplication circuit according to any one of claims 1 to 15, wherein the separate enable control signal comprises a separate clock signal.

17. An apparatus comprising: Processing circuitry, the processing circuitry being configured to perform data processing in response to instructions; The processing circuit includes a multiplication circuit according to any one of claims 1 to 16.

18. A system comprising: The multiplication circuit according to any one of claims 1 to 16 or the apparatus according to claim 17, wherein the multiplication circuit or the apparatus is implemented in at least one packaged chip; At least one system component; and plate, The at least one packaged chip and the at least one system component are assembled on the board.

19. A chip-containing product, the chip-containing product comprising the system of claim 18, the system being assembled on an additional board together with at least one other product component.

20. A method, the method comprising: The first operand is Booth-coded using a shared Booth coding circuit shared among multiple adder arrays to generate multiple sets of partial product selection indicators, each corresponding to the Booth coding of a corresponding Booth number of the first operand, wherein the multiple adder arrays include separate instances of hardware circuitry having at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled. The corresponding sets of partial products to be added by the multiple adder arrays are selected based on the second operand and the multiple partial product selection indicators. as well as The partial products are added using the enabled adder arrays of the plurality of adder arrays to generate corresponding product representations, each representing the result of multiplying a corresponding pair of bit portions selected from the first operand and the second operand. At least two of the adder arrays are configured to operate on corresponding partial products selected based on a common partial product selection indicator, which is generated by the shared Booth coding circuit based on the Booth code of the common Booth number of the first operand.

21. A computer-readable medium for storing computer-readable code for manufacturing a multiplication circuit, the multiplication circuit comprising: Multiple adder arrays, each of which is used to add a corresponding set of partial products to generate a corresponding product representation value, the corresponding product representation value representing the multiplication result of a corresponding pair of bit portions selected from a first operand and a second operand, the multiple adder arrays including separate instances of hardware circuitry, the multiple adder arrays having at least two separate enable control signals for independently controlling whether at least two subsets of the adder arrays are enabled or disabled; A shared Booth coding circuit, shared among the plurality of adder arrays, is used to Booth code the first operand to generate a plurality of partial product selection indicators, each of which corresponds to the Booth code of a corresponding Booth number of the first operand. and A partial product selection circuit is configured to select, based on the second operand and the plurality of partial product selection indicators, the partial product to be added by the plurality of adder arrays; wherein: At least two of the adder arrays are configured to operate on corresponding partial products selected by the partial product selection circuit based on a common partial product selection indicator, which is generated by the shared Booth encoding circuit based on the Booth encoding of the same Booth number of the first operand.