Method and apparatus for processing hypercomplex numbers
Patent Information
- Application Number
- PCT/RU2023/000195
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-05-22
AI Technical Summary
Current computers are inefficient in processing complex and dual numbers, as they emulate these numbers using real numbers, leading to slowed processing speed and increased power consumption, especially during multiplication and convolution operations.
A method and apparatus that treat hypercomplex numbers as compounds with two components, allowing for direct operations using a single instruction, which improves performance, reduces power consumption, and optimizes hardware resource utilization by reusing existing hardware resources for real-number operations.
This approach enhances processing speed and reduces power consumption for complex and dual number operations, achieving performance comparable to real-number operations while minimizing hardware requirements.
Smart Images

Figure RU2023000195_22052025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR PROCESSING HYPERCOMPLEXNUMBERSTECHNICAL FIELD
[0001] Embodiments of the present invention relate to the field of computer technologies, and more specifically, to a method and apparatus for processing hypercomplex numbers.BACKGROUND
[0002] There are three classes numbers that are widely used in technical fields such as engineering, science, and physics, and they are real numbers, complex numbers and dual numbers, respectively. The complex numbers are the second common class of numbers behind the real numbers, and are gaining popularity in neural networks in recent years. A complex number can be represented with a real part and an imaginary part, and complex-valued neural networks (CVNNs) can have better generalization characteristics, faster learning and higher robustness than real-valued neural networks (RVNNs). A dual number also consists of two parts, that is, a real part and a dual part, and has different properties compared to the complex number, which makes it particularly suited for use in neural networks. The dual numbers make it possible to automatically compute derivatives of functions, which is a huge advantage in deep learning, and in the training of neural networks in general. It has been demonstrated that dual-valued neural networks (DVNNs) can achieve better average precision (AP) than the CVNNs and the RVNNs.
[0003] However, up to this day, most computers work only with the real numbers, and when they need to work with the complex numbers or the dual numbers, arithmetic operations on these classes of numbers are emulated by using two real numbers. This emulation slows down the processing speed and increases power consumption. Multiplication or convolution of the complex numbers or the dual numbers is particularly expensive in these aspects.SUMMARY
[0004] Embodiments of the present application provide a method for processing hypercomplex numbers, which can improve processing speed and utilization of hardware resources, and reduce power consumption.
[0005] According to a first aspect, provided is a method for processing hypercomplex numbers, including: receiving a first instruction, where the first instruction indicates an operation, and the operation includes a reading operation, a writing operation or an arithmetic operation; performing, in response to the first instruction, the operation on an operand that includes at least two components, where a combined size of the at least two components is less than or equal to a size of a native central processing unit (CPU) data word, and the operand is taken as one compound with the at least the two components when the operation is performed.
[0006] According to the proposed solution of the present application, a number consisting of at least two components is perceived as one compound with the at least two components when an operation is performed on it. The compound representation of these numbers allows to access these numbers with one single instruction, and the instruction is operated directly on the at least two components. Compared with an operation on these numbers that are emulated with real numbers, the proposed solution obtains higher performance than performance of the emulation.
[0007] Besides, the proposed solution uses smaller executable code than emulation by real number due to fewer instructions. In addition, lower power consumption and faster processing speed than emulation by real numbers are realized due to few instructions executed for same algorithm.
[0008] In addition, operations for more than one class of numbers based on existing or few additional hardware resources are supported by the present application. Therefore, utilization of the hardware resources is improved.
[0009] In an embodiment of the first aspect, the operand includes a dual number, and the dual number consists of two components, where one of the two components is a real part and the other is a dual part, the two components are of the same size, and each of the two components is half the size of the native CPU data word.
[0010] In this embodiment, operations such as reading, writing and arithmetic for dual numbers are supported, especially hardware implementation of multiplication / convolution of dual numbers, and better performance can be obtained by the present application as described in the first aspect. Multiplication / convolution of dual numbers proposed by the present application gets a faster speed and requires less energy than multiplication / convolution by emulation with real numbers.
[0011] In an embodiment of the first aspect, the method is applied in a single instruction single data (SISD) type arithmetic logic unit (ALU); and the two components of the dual number are stored in one item that takes up a space of the native CPU data word.
[0012] In this embodiment, the real part and the dual part of the dual number are stored in one item, which achieves a more compact representation of the dual numbers, and the storage space for dual numbers is reduced. Besides, the real part and the imaginary part of the complex number are stored in the same way as the dual number, and the same technique effect will be achieved.
[0013] In an embodiment of the first aspect, the method is applied in a single-instruction-multiple-data (SIMD) type or a multiple-instruction-multiple-data (MIMD) type ALU; and the two components of the dual number are stored in a way such that the two components are next to each other in a storage device.
[0014] In this embodiment, the components of the dual number are grouped together, so that the two components of each dual number are next to each other, and this is very important when it comes to resource reuse of the hardware resources in SIMD type or MIMD type ALU. Besides, the components of the complex numbers are grouped in the same way as the dual number, and the same technique effect will be achieved. Alternatively, the SIMD is also called as a “vector processor”.
[0015] In an embodiment of the first aspect, the arithmetic operation includes vector multiplication; where the performing, in response to the first instruction, the operation on an operand that comprises at least two components, includes: performing, in response to the first instruction, the vector multiplication on a first vector and a second vectorto obtain a third vectorwhere the operand includes dual numbers in the first vector and the second vector,i < n, i and n are positive integers.
[0016] In this embodiment, the operation mentioned in the present application refers to a dual-valued multiplication operation, and the operand is dual numbers of the vectors on which the dual-valued multiplication needs to be performed. The existing hardware resources for real-valued multiplication are reused by the dual-valued multiplication, and the utilization of the hardware resources is improved.
[0017] In an embodiment of the first aspect, the operand includes a complex number, and the complex number consists of two components, one of the two components is a real part and the other is an imaginary part, the two components are of the same size, and each of the two components is half the size of the native CPU word.
[0018] In this embodiment, the multiplication / convolution of the complex numbers is supported by the present application, and better performance such as a faster processing speed can be got than through emulation by real numbers.
[0019] In an embodiment of the first aspect, the arithmetic operation includes a multiplication operation, and the method further includes: receiving a second instruction, where the second instruction indicates a multiplication mode, and the multiplication mode is a complex-valued multiplication mode or a dual-valued multiplication mode; where the performing, in response to the first instruction, the operation on an operand that includes at least two components, further includes:performing the multiplication operation on operands to obtain a first result that includes at least one compound, where a mode of the multiplication operation is the multiplication mode indicated by the second instruction.
[0020] In this embodiment, a combined implementation of multiplication / convolution for dual numbers and complex numbers is supported by the present application, and the combined implementation reuses the hardware resources for real numbers in the complex-valued multiplication, and adds no extra arithmetic functions for implementation of support for dualvalued multiplication. The proposed solution increases the performance of dual -valued multiplication and complex-valued multiplication to the same performance as real-valued multiplication. In this way, DVNNs and CVNNs will have an inference time comparable to RVNNs, but higher average precision (AP). The overall power consumption will also be reduced because fewer instructions need to be executed compared to emulation.
[0021] In an embodiment of the first aspect, the operand includes an algebraic or non-algebraic number of k dimensions, where the number of k dimensions includes k components, and k sizes of the k components are in one-to-one correspondence to that of k parts of the native CPU data word, respectively, where the native CPU data word is split into the k parts or w parts, where w is larger than k, and k and w are positive integers.
[0022] In an embodiment of the first aspect, the k parts or the w parts are obtained by an even or uneven splitting in size of the native CPU data word.
[0023] In an embodiment of the first aspect, the native CPU data word is split into the w parts, and remaining (w- k) parts from the w parts are unused parts.
[0024] According to these embodiments, the compound representation proposed by the present application also is applicable to some other classes of numbers, for example, the higher order algebras and numbers of non-algebraic system.
[0025] According to a second aspect, provided is an apparatus for processing hypercomplex numbers, including: multipliers and at least one adder, configured to: obtain operands, where a size of each operand is less than or equal to a size of a native CPU data word, and each operand is a dual number; perform a multiplication operation on the operands.
[0026] In an embodiment of the second aspect, where the dual number consists of two components, where one of the two components is a real part and the other is a dual part, the two components are of the same size, and each of the two components is half the size of the native CPU data word.
[0027] In this embodiment, one operand is one dual number, and the size of the operand is less than or equal to the size of the native CPU data word. In other words, n operands refer to n dual numbers in this embodiment, and the size of each operand is not more than the size of the native CPU data word. In addition, each operand is stored in one item that takes up aspace of one native CPU data word.
[0028] In an embodiment of the second aspect, the apparatus further includes a control logic instruction decoder, a multiplexer and at least one subtractor, where the control logic instruction decoder is configured to: output a signal to the multiplexer, where the signal indicates a multiplication mode, where the multiplication mode is a complex-valued multiplication mode or a dual-valued multiplication mode; where the multiplexer is configured to: choose a mode of the multiplication operation, where the mode of the multiplication operation is the multiplication mode indicated by the signal; where the multipliers, the at least one adder and the at least one subtractor are further configured to: perform the multiplication operation on operands based on the chosen mode of the multiplication operation, where each operand is taken as a compound with two components that are a real part and an imaginary part if the chosen mode is the complex-valued multiplication mode; or each operand is taken as a compound with the two components that are the real part and the dual part is the chosen mode is the dual-valued multiplication mode.
[0029] In an embodiment of the second aspect, the apparatus includes a SISD type ALU.
[0030] According to a third aspect, provided is an apparatus for processing hypercomplex numbers, including: multipliers and at least one adder, which are configured to: perform vector multiplication on a first vectorand a second vector to obtain a third vector,where the first vector, the second vector and the third vector are vectors of dual numbers, each of the dual numbers includes two components that are a real part and a dual part, and the two components of one dual number from any one of the vectors are stored in a way such that the two components are next to each other; and and n is are positive integers.
[0031] In an embodiment of the third aspect, the apparatus further includes a control logic instruction decoder, a multiplexer and at least one subtractor; where the control logic instruction decoder is configured to: output a signal to the multiplexer, where the signal indicates a multiplication mode, and the multiplication mode is a complex-valued multiplication mode or a dual-valued multiplication mode; where the multiplexer is configured to: choose a mode of the vector multiplication, where the mode of the vector multiplication is the multiplication modeindicated by the signa; where the multipliers, the at least one adder and the at least one subtractor are further configured to: perform the vector multiplication on the first vector and the second vector based on the chosen mode of the vector multiplication, where the first vector, the second vector and the third vector are vectors of dual numbers if the chosen mode is the dual-valued multiplication mode; or the first vector, the second vector and the third vector are vectors of complex numbers if the chosen mode is the complex-valued multiplication mode, and the two components of one complex number from any one of the vectors are stored in a way such that the two components are next to each other.
[0032] In an embodiment of the third aspect, the apparatus includes a S1MD type or a M1MD type ALU.
[0033] SIMD or M1MD implementation that combines the dual-valued multiplication and complex-valued multiplication of the present application gets higher performance and lower power consumption compared to emulation by real numbers. Further, the combined solution uses less hardware resources and less chip area, compared to a solution where the real-valued multiplication, the dual-valued multiplication and the complex-valued multiplication are implemented separately. This ultimately leads to lower cost of chip and lower cost-of-ownership of the overall solution.
[0034] The description of the technique effect of the apparatus in the second aspect or the third aspect can refer to that of the corresponding embodiment in the first aspect, and it will not be repeated herein.
[0035] According to a fourth aspect, a chip (or a chip system) is provided. The chip has a function of implementing the method in the first aspect and any embodiment of the first aspect. The function may be implemented by using hardware structures.
[0036] According to a fifth aspect, a chip (or a chip system) is provided. The chip includes multipliers and at least one adder. Further, the chip may include a control logic instruction decoder, a multiplexer and at least one subtractor. The chip has a function of implementing the method in the first aspect and any embodiment of the first aspect. The function may be implemented by using the foregoing hardware structures.DESCRIPTION OF DRAWINGS
[0037] One or more embodiments are exemplarily described by corresponding accompanying drawings, and these exemplary illustrations and accompanying drawings constitute no limitation on the embodiments. Elements with the same reference numerals in the accompanying drawings are illustrated as similar elements, and the drawings are not limiting to scale, in which:
[0038] FIG. 1 is an example of vector multiplication by a SIMD processor.
[0039] FIG. 2 is a flow chart of a method for processing hypercomplex numbers proposed by the present application.
[0040] FIG. 3 shows an organization of dual numbers in a UB.
[0041] FIG. 4 shows an example of hardware implementation of dual-valued multiplication proposed by the present application.
[0042] FIG. 5 shows a comparison of hardware implementation of dual-valued multiplication and complex-valued multiplication.
[0043] FIG. 6 shows an example of a circuit that combines dual-valued multiplication and complex-valued multiplication in an SISD type ALU.
[0044] FIG. 7 shows an example of a circuit for dual-valued multiplication in a SIMD type ALU.
[0045] FIG. 8 shows an example of a circuit that combines dual-valued multiplication and complex-valued multiplication in a SIMD type ALU.
[0046] FIG. 9 is a schematic block diagram of a chip (or a chip system) according to an embodiment of the present application.DESCRIPTION OF EMBODIMENTS
[0047] In order to understand features and technical contents of embodiments of the present application in detail, implementations of the embodiments of the present application will be described in detail below with reference to the accompanying drawings, and the attached drawings are only for reference and illustration purposes, and are not intended to limit the embodiments of the present applications. In the following technical descriptions, for ease of explanation, numerous details are set forth to provide a thorough understanding of the disclosed embodiments. One or more embodiments, however, may be practiced without these details. In other cases, well-known structures and apparatuses may be shown simplified in order to simplify the drawings.
[0048] Related technologies and concepts are introduced here firstly in order to better understand the technique solution proposed by the present application.
[0049] These notations will be used in the present application:A for a real number and VA for a vector of a real number;A for a dual number and VA for a vector of a dual number;A for a complex number and VA for a vector of a complex number.
[0050] It is known to all that a complex number consists of two components, and can be written as z = x + iy, wherex,y 6 R and i is an imaginary unit, which satisfies the equation i2= —1. A dual number also consists of two components, and can be written in the form d = x + ey, where s is a nilpotent element, which satisfies the relations s2= 0, s + 0.
[0051] Addition, subtraction and conjugation (a non-real part of a complex or dual number changes a sign) of the complex numbers or the dual numbers can easily be emulated by the real numbers on a single instruction single data (SISD) computer. Addition and subtraction require two instructions, which are used for the “real” part and the “dual” or the “imaginary” part, respectively. Conjugation requires an instruction that is on the “dual” or the “imaginary” part. In the case of a single instruction multiple data (SIMD) computer, only one instruction is required to emulate these three types of operations on vectors of complex numbers or dual numbers, because the SIMD instructions can be operated on the real part and the dual part (or the imaginary part) at the same time.
[0052] FIG. 1 is an example of vector multiplication by a SIMD processor. As shown in FIG. 1, the SIMD processor performs an operation on the ithelement from vector VA and the ithelement from vector VB, and writes the result to the ithelement of vector VC. For elements of vectors VA, VB and VB, a data type of float point (FP) and a size of 16-bit are just examples. A length of the vectors VA, VB and VC of 128 is also an example, and 128 pieces operations are correspondingly performed as shown in FIG. 1.
[0053] In SISD computers, multiplication of two complex numbers requires real-valued multiplication four times, addition once and subtraction once. Dual-valued multiplication requires multiplication three times and addition once of real values. In SIMD computers, such as a vector unit of which, the overhead becomes worse due to the fact that the real part and the imaginary part of the complex number (or the real part and the dual part of the dual number) are interleaved in a vector. The biggest problem is multiplication and convolution, and multiplication or convolution represents more than 99% of the computations in neural networks, fast Fourier transform (FFT), finite impulse response (FIR)-filters, infinite impulse response (IIR)filters, etc.
[0054] Multiplication and convolution of dual numbers or complex numbers by emulation with real numbers, are slower and require more energy than multiplication and convolution of real numbers. Emulation of SIMD operations on complexvalued and dual-valued numbers by real-valued numbers is comparatively slower than emulation on SISD architectures, due to the “channelized” or “vectorized” design of a SIMD unit, and it can refer to FIG. 1. In SISD, the complex-valued multiplication requires 6 instructions, and in some architectures of SIMD, the complex-valued multiplication requires 12 instructions. In SISD, the dual-valued multiplication requires 4 instructions, and in some architectures of SIMD, the dual-valued multiplication requires 12 instructions.
[0055] In addition, ideal performance implementation of hardware support for complex numbers requires roughly doubling the hardware resources, and support for dual numbers requires increase of hardware resources of 50%, even if theexisting hardware resources for real numbers are reused. Supposing that hardware area for real numbers only is 100%, and it will be up to 250% of that area for real numbers, dual numbers and complex numbers. There will be an increase of 150% (from 100% to 250%) of the hardware resources in the SIMD processor in order to support both classes of numbers. If the hardware resources for real numbers are not reused, it requires an increase of 350% (from 100% to 450%) in order to support all three classes of multiplication.
[0056] Due to the above technique status, the present application proposes a solution that supports hardware implementation for multiplication of dual numbers. In addition, the present application provides a combined hardware implementation for dual-valued multiplication and complex-valued multiplication, which reuses the hardware resources for real numbers, and adds no extra arithmetic function for implementation for support for dual-valued multiplication.
[0057] The performance of the dual-valued or complex-valued multiplication can be improved to be basically the same as that of the real- valued multiplication. In this way, the CVNNs and DVNNs will have an inference time comparable to RVNNs but higher AP. The overall power consumption of multiplication of dual numbers or complex numbers will also be reduced, compared to emulation with real numbers. Other types of algorithms, such as fast Fourier transform (FFT), will have also experienced significant speed up and reduced overall power consumption due to the direct execution of the dual-valued and complex-valued multiplication instead of emulation.
[0058] Any processor chip and any board level product, which uses these chips, can apply the solution of the present application. For example, the proposed solution of the present application can be applied in any SISD processor and SIMD processor including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a micro-processor unit (MPU) and neural networks (NN) accelerators, and the MPU is a CPU integrated with memory and peripheral interface functions. Besides, the proposed solution of the present application can also be applied in reduced instruction set computervariation (RISC-V), advanced RISC machines (ARM) and some other similar architectures.
[0059] The present application will be described in detail below.
[0060] FIG. 2 is a flow chart of a method for processing hypercomplex numbers proposed by the present application. The method (200) is performed by devices such as the above CPU, GPU, MPU or NN accelerators. The method (200) specifically includes the following steps.
[0061] Step 210: a first instruction is received, where the first instruction indicates an operation, and the operation includes a reading operation, a writing operation or an arithmetic operation.
[0062] Step 220: in response to the first instruction, the operation is performed on an operand that includes at least two components, where a combined size of the at least two components is less than or equal to a size of a native CPU data word, and the operand is taken as one compound with the at least two components when the operation is performed.
[0063] Specifically, the at least two components include two or more components.
[0064] For example, the operand is a dual number, and the dual number consists of two components that are a real part and a dual part, respectively. For another example, the operand is a complex number, and the complex number consists of two components that are a real part and an imaginary part, respectively.
[0065] From a performance point of view, the main problem in dealing with the dual number or the complex number mentioned in the background section, is the fact that it consists of two parts. The complex number consists of a real part and an imaginary part, and the dual number consists of a real part and a dual part. Even though most programming languages allow for user-defined data structures, which combine the two parts together in a source code, but an executable code still handles them as two different entities. Separate instruction steps are needed to be operated on the two parts of the complex number or the dual number.
[0066] For example, assuming that the native CPU data word is of 32-bit size, and if the real part and the dual part (or the imaginary part) are both 16-bit, each 16-bit entity is stored as 32-bit in known scenarios. In view of this, the present application proposes a new data type that composes the two parts of the dual number or the complex number in one native CPU data word item. In other words, the two components of the dual number or the complex number are stored in one item that takes up a space of the native CPU data word. For example, the two part of the dual number or the complex number are 16 -bit respectively, and the data type proposed by the present application composes the two 16-bit parts in one 32-bit item. The at least two components of the operand are stored in one data item that takes up a space of the native CPU data word, and the dual number and complex number are just examples of the operand.
[0067] Table 1 shows an example of a side-by-side representation in the known scenarios and a compound representation proposed by the present application.
[0068] As shown in Table 1 , in the side-by-side representation, if the CPU is 32-bit, every data item of the native CPU data word takes up 4 bytes, hence neighbor data items are apart with 4 bytes. A notation “addr” is short for the address of a data item. The notation “addr+OxOOOO” is so called as “base-address” with a “zero” offset, i.e., the base address with no offset. The notation “add+0x0004” is an address with an offset of 4 bytes, and this is the actual address of the “next” data item. A notation “Ox” is a programming notation for a number in a hexadecimal number-system. Compilers, like Keil C / C++ and IAR C / C++compilers, can actually map a number (for example, a dual number and a complex number) that consists of two or more components in a compound representation way proposed by the present application to achieve a more compact representation of the number. The compound representation of the number allows to access the number with one single instruction for data movement operations, such as reading and writing. As an embodiment, the single instruction for reading or writing could be the same as an instruction used for other types of data, which already exists in the CPU, and an instruction for arithmetic operations is a new instruction proposed by the present application. For example, in the present application, the instruction for arithmetic operation is operated directly on the two-part compound, such as a complex number and a dual number, and a result is written back as a two-part compound with a single instruction. The dual number or the complex number is taken as a compound with two components whose combined size is the same with that of the native CPU data word, and the size of each component is half the size of the native CPU data word. In other word, the two components of the dual number or the complex number are packed into one compound whose combined size is of a full size of the CPU data word.
[0069] Table 2 shows examples for compound representations in different native CPU data word sizes and data types.
[0070] Note that the compound representations in Table 1 and Table 2 are just examples of the compound representation proposed by the present application, the native CPU data word should not be restricted to 32-bit or 64-bit, and these values are shown only for the purpose of illustration of the concept. In addition, there is no restriction about the data type, and the data type could be any one of the following data types: signed integer, unsigned integer, fixed-point, floating-point (examples of IEEE 754), and any other proprietary or standardized number representation. For example, brain float (BF)16 is a propriety floating point format, which is also supported by several players in the CPU, GPU and general-purpose graphics processing unit (GPGPU)-marked. These descriptions are applicable in any one of examples below of the present application, and will not be repeated.
[0071] In addition, an order of the at least two components of the compound in the storage device is not restricted. In other words, the at least two components of the compound can be stored in any order. For example, the order of the real part and the imaginary part of the complex number, or the order of the real part and the dual part of the dual number can be stored in a way such that the real part is before or after the imaginary part / dual part.
[0072] The above compound representations of numbers in Table 1 and Table 2 are examples of data representation ofdual numbers or complex numbers in SISD type arithmetic logic unit (ALU).
[0073] Data representation of dual numbers and complex numbers proposed by the present application in a SIMD type or a MIMD type ALU will be described below.
[0074] In the SIMD type ALU or vector processor, the real-valued data will normally be organized as shown in Table 3.
[0075] Table 3 shows a real-valued vector whose size is n, that is to say, the real-valued vector includes n components, which are respectively, and n is a positive integer.
[0076] All the real numbers will be next to each other in order of appearance. They may not be related to each other from an algebraic point of view, but they may be related in an application domain, for example, a time-series sequence.
[0077] The represent application proposes an organization of dual numbers in the SIMD type or the MIMD type processor as shown in Table 4. Specifically, the two components of the dual number are stored in a way such that they are next to each other in a storage device. The “storage device” herein includes the SIMD processor, vector processor and so on.
[0078] Table 4 shows a dual-valued vector whose size is m, and m is a positive integer.
[0079] In table 4, m = n / 2. There will be half as many dual numbers in the dual-valued vector compared to the number of real numbers in the real- valued vector in table 3, because each dual number (or complex number) consists of two components. The elements in table 4 are grouped together such that two components of one dual number are next to each other. For example, the two components, which are aOrealand aoduai, of dual number a0, are next to each other, then the two components, which are alrealand alduahof dual number at, are next to each other, etc. The way to organize the dual numbers is very effective when it comes to resource sharing of the hardware resources in the SIMD type or MIMD type processors.
[0080] FIG. 3 shows an organization of dual numbers in a unified buffer (UB). As shown in FIG. 3, suppose that hardware resources of the UB can store 128 pieces (often abbreviated as “pcs”) of real numbers (for example, FP 16), and if the hardware resources are used to store the dual numbers in the way that is proposed by the present application as shown in table 4, 64 pcs of dual numbers can be stored in the UB. Note that, UB is an example of buffers memory, which holds data to be used by the “vector unit” in a SIMD computer or “scalar unit” in a SISD computer.
[0081] In addition, in method 200, the operand could be a number of an algebraic system or non-algebraic system, and this is not limited in the present application. Some examples will be given below.
[0082] As an embodiment of the present application, the operand includes an algebraic or non-algebraic number of k dimensions, where the number of k dimensions includes k components, and k sizes of the k components are in one-to-one correspondence to that of k parts of the native CPU data word, respectively, where the native CPU data word is split into the k parts or w parts, where w is larger than k, and k and w are integers.
[0083] Higher order unital algebras over fields of real numbers exist. They are the associative division algebras R (short for “real numbers”), C (short for “complex numbers”), H (short for “hypercomplex number”), O (short for “octonions”), which are of 1 , 2, 4, 8 dimensions, respectively, and dual numbers, which are also of two dimensions. Dual numbers are associative division algebras, but octonions are not the associative division algebras. These algebras can be written as follows:
[0084] The operand mentioned in the embodiments of the present application also covers these algebras. Further, the operand in the present application can be a hypercomplex number, for example, a hypercomplex number z of 4thorder, where
[0085] As an embodiment of the present application, the operand can be an algebraic number of k dimensions, and the algebraic number of k dimensions can be implemented in one m-bit native CPU data word by splitting the CPU data word into k fractions. The k fractions are in one-to-one correspondence to k components of the algebraic number of k dimensions.
[0086] Table 5 shows some examples of compound representations in different sizes and data types of the native CPU data word.
[0087] The 8 components d0~d7in Table 5 are 16-bit, respectively, and the 8 components are 128 bits in total.
[0088] In addition, there is no restriction on an order of the at least two components of the number, for example, the order of the 16-bit real part and the 16-bit dual part of a0, the order of the 16-bit real part and the 16-bit dual part, the order of the 4 components c0~c3of a 64-bit number, and the order of 8 components d0~d7of a 128-bit number are some examples of the order of the at least two components of the number of the present application. There is also no restriction on the data types of the components, and different data types are shown in Table 5 for purposes of illustration.
[0089] As mentioned above, one m-bit native CPU data word can be split into k fractions of even size, and each fraction corresponds to one component of the algebraic number of k dimensions. As another implementation, the splitting of the m-bitnative CPU data word could be an un-even splitting in size. In other words, the components of the algebraic number of k dimensions could be in different sizes. Besides, the algebraic number of k dimensions could be of any data type.
[0090] The only number systems, which are algebras, have dimensions which are power of two, for example, k 6 {1, 2, 4, 8}, where k is the dimension of the algebras. CPUs usually have ALU-width that is also power of two, for example, m G {8, 16, 32, 64}. Hence the division m / k will always be without residuals when m ^k. Alternatively, as another embodiment, the m-bit native CPU data word can be split into k fractions of uneven size, and there can be one or more residuals, which are not used.
[0091] Table 6 gives two examples for the algebraic number of k dimensions, one example is that the 32-bit data word is split into two fractions a0, atwith a residual, where a0, a±are of the same size. In this example, the operand can be an algebraic number of two dimensions, and hence the algebraic number includes two components, and the two components are in one-to-one correspondence to the two fractions a0, a±. For another example, the 32-bit CPU data size is split into two fractions b0, brwithout a residual, where b0, btare different in size. In this example, the operand can be an algebraic number of two dimensions, and the algebraic number includes two components, and the two components are in one-to-one correspondence to the two fractions b0, bt.
[0092] The ideas of the present application mentioned in above embodiments also can be applied in a non-algebraic number system. The m-bit native CPU data word is split into A: parts with the same or different sizes, with or without a residual.
[0093] Table 7 gives two examples, and one example is that a size of the 33-bit CPU data word is split into two fractions c0, Cj with a residual, where c0, cxare of the same size. In this example, the operand could be a non-algebraic number of two dimensions, and the non-algebraic number includes two components, and the two components are in one-to-one correspondence to the two fractions c0, Q. The two components are of the same size. For another example, a size of the 33- bit CPU data word is split into three fractions, d0, dr, d2without a residual, where d0, d1, d2are different in size. In this example, the operand could be a non-algebraic number of two dimensions, and the non-algebraic number includes two components, and the two components are in one-to-one correspondence to the fractions d0, d2, and dj is the fraction that willnot be used. Besides, the fraction that will not be used can be in any position of the CPU data word, and what is shown in Table 7 are just examples.
[0094] According to the examples shown in Table 6 and Table 7, the operand can be a number that includes at least two components, and the two components are in one-to-one correspondence to k parts of the native CPU data word. The k parts are all or some parts of the native CPU data word, and in a latter case, the native CPU data word is split into w parts, and k parts of the w parts are in one-to-one correspondence to k components from an algebraic or non-algebraic number of k dimensions, and (w- k) parts of the w parts of the native CPU data word are unused, where w is larger than k.
[0095] The following focuses on hardware acceleration application in dual numbers and complex numbers of the present application.
[0096] Hardware acceleration of dual number multiplication in a SISD type CPU or in a SIMD type or a MIMD type CPU will be descried in detail separately.
[0097] Hardware acceleration of dual number multiplication in the SISD type CPU is introduced firstly.
[0098] Consider the dual-valued equation (1) below:when this equation is emulated with real values, the result looks like equations in (2):when we combine the equations in (2) with the compound representation proposed by the present application, we get the following implementation of the hardware accelerated multiplication engine for dual numbers in the compound representation as shown in FIG. 4.
[0099] In FIG. 4, a and b are operand registers that are marked as 201 and 202, and c is a result register or destination register that is marked as 207. The registers can be physically separate registers, part of a so-called “register file” or memory locations. Circles with “X” are multipliers that are marked as 203, 204 and 205, and a circle with “+” is an adder that is marked as 206.
[0100] FIG. 4 shows how two operands a and b in the compound format are interpreted as dual numbers. Specifically, each of the two operands is perceived as a compound with two 16-bit components, not as a 32-bit value. Similarly, a result of the multiplication is returned, and the result also consists of two 16-bit components.
[0101] Note that, the data type and the width of the operands in FIG. 4 are not limited. The width of 16-bit of the operands shown in FIG. 4 is just for the purpose to give an example to describe the hardware implementation of multiplication for dual numbers.
[0102] Next, hardware implementation that combines the dual-valued and complex-valued multiplication in the SISD type or the MIMD type ALU is introduced.
[0103] Equations for emulating dual-valued multiplication with real values are shown in equation (2), and equations for emulating complex-valued multiplication with real values are shown in equations (3) as follow:
[0104] In the present application, the dual-valued and complex-valued multipliers are considered together.
[0105] FIG. 5 shows a comparison of hardware implementation of dual- valued and complex- valued multiplication. In the left figure of the FIG. 5, a and b are dual-valued operand registers, and c is a dual-valued result register or destination register. In the right figure of the FIG. 5, a and b are complex-valued operand registers, and c is a complex-valued result register or destination register. Reference can be made to the description of the FIG. 4 for the description of the left figure of FIG. 5. In the right figure of FIG. 5, the complex-valued operand registers are marked as 301 and 302, the complex-valued result register is marked as 309. Circles with “X” are multipliers that are marked as 303-306, a circle with is a subtractor that is marked as 307, and a circle with “+” is an adder that is marked with 308.
[0106] According to FIG. 5, it can be seen that the hardware implementation of the dual-valued multiplication and that of the complex-valued multiplication are very similar. Circuits for dual-valued multiplication are a subset of that for the complex-valued multiplication. The circuits in solid line are the circuits that are already part of the SIMD ALU, and the circuits in dash line are added circuits that are needed for dual-valued or complex- valued multiplication, respectively.
[0107] Based on the analysis of the hardware implementation shown in FIG. 5, the present application proposes hardware implementation that combines dual-valued and complex-valued multiplication together in a SISD ALU as shown in FIG. 6. In other words, the hardware implementation shown in FIG. 6 can perform the dual-valued multiplication and the complex-valued multiplication.
[0108] FIG. 6 shows an example of a circuit that combines dual-valued and complex-valued multiplication in a SISD ALU. Compared to FIG. 5, an added multiplexer shown in FIG. 6 makes it possible to select a dual-value mode or a complexvalue mode for the circuit. The dual-valued multiplication is performed by the circuit in the dual-value mode, and the complexvalued multiplication is performed in the complex-valued mode. Specifically, in the complex-valued mode, a product, that is, “im®xbtmg, is subtracted from cimg. In the dual-valued mode, “zero” is subtracted from cimg, that is to say, nothing is subtracted in the dual-valued mode. Note that, almg, bimg, and Cjmgare the imaginary parts of a, b and c, respectively.
[0109] The multiplexer is controlled by a functional block, and the functional block is called a “control logic instruction decoder” and will anyway be present in the ALU. Therefore, it just needs to add an additional signal as an output to thefunctional block, and the multiplexer receives the signal (or called as a second instruction) and selects the dual-valued mode or the complex-valued mode based on the signal from the functional block. Note that, FIG. 6 focuses on the multiplexer and the control logic instruction decoder, and others of FIG. 6 refer to that of FIG. 5.
[0110] In data representations of institute of electrical and electronics engineers (IEEE) 754 standard, in many data types, such as floating point, binary signed integer and binary two’ compliment, a value “zero” is represented by a bit-pattern, where all the bits are zero. As an example, the multiplexer shown in FIG. 6 can be implemented by a row of 2-input AND- gates. The multiplexer also can be implemented by other means of hardware, which is not limited herein.
[0111] Thirdly, hardware acceleration of dual number multiplication in the SIMD type or the MIMD type ALU is introduced.
[0112] By combining the compound representation and the hardware implementation for dual-valued multiplication provided by the present application, a circuit for dual-valued multiplication is get as shown in FIG. 7.
[0113] As shown in FIG. 7, a vectorconsists of dual values, that is, . A vector VBconsists of dual values, that is,. The dual numbers in the vector are multiplied by the dualnumbers in the vector, and a result consisting of dual numbers is written to a vector VC. Specifically, the dual number 40 is multiplied by the dual numberto yield a result etc. In general,are integers.
[0114] Note that, the individual components of the SIMD vector are no longer perceived as real values, but pairs of components are perceived as dual numbers. The grey lines in FIG. 7 show how the circuit is repeated for every pair of dual number in the SIMD type or the MIMD type ALU. The two multipliers in solid lines were already present in the real-value SIMD type or the MIMD type ALU, and the multiplier in dash line and an adder in dash line are new compared to the real- value SIMD type or the MIMD type ALU. In this way, the existing hardware resources for real-valued multiplication are reused by the dual-valued multiplication circuit.
[0115] Further, the present application provides an example of an instruction for vector multiplication of dual numbers: VDMUL. are integers, and ncorresponds to the number of dual numbers in the SIMD vector. The letter “D”, in the opcode VDMUL, signifies that the instruction VDMUL interprets the data as dual numbers.
[0116] Besides, hardware implementation that combines the dual-valued and complex-valued multiplication in the SIMD type or the MIMD type ALU is introduced.
[0117] By combining the compound representation and the hardware implementation for dual-valued and complexvalued multiplication provided by the present application, a circuit that combines the dual-valued multiplication and thecomplex-valued multiplication is get as shown in FIG. 8.
[0118] As shown in FIG. 8, the multiplexer chooses the dual-valued mode or the complex-valued mode according to a signal from the control logic instruction decoder. The circuit is replicated for each dual number or complex number (i.e., for every pair of two real values) in the SIMD vector. FIG. 8 shows how the dual-valued or complex-valued multiplication is performed for 710 and BO, and for 71(n-l) and B(n-l ), and the multiplications for other numbers in vectors VA and VB are similar to the shown ones, and they are not described one by one. The control logic instruction decoder decodes the instruction and controls the multiplexer to choose the dual-valued mode or the complex-valued mode.
[0119] Further, the present application provides an example of an instruction for vector multiplication of complex numbers:corresponds to the number of complex numbers in the SIMD vector. The letter “C”, in the opcode VCMUL, signifies that the instruction VCMUL interprets the data as complex numbers.
[0120] In addition, the instruction for dual-valued multiplication uses the same hardware resources, the instruction for dual-valued multiplication can refer to the instruction provided in above embodiment, and it is not repeated herein.
[0121] FIG. 9 is a schematic block diagram of a chip (or a chip system) 10 according to an embodiment of the present application. As shown in FIG. 9, the chip (or the chip system) may include multipliers 11 and at least one adder 12. Specifically, the multipliers 11 and the at least one adder 12 are configured to perform multiplication or convolution operations on operands proposed by the present application. The chip 10 may further include a control logic instruction decoder 14, a multiplexer 15 and at least one subtractor 13. The control logic instruction decoder 14 is configured to output a signal to the multiplexer, where the signal indicates a complex-valued multiplication mode or a dual-valued multiplication mode. The multiplexer 15 chooses the complex-valued multiplication mode or the dual-valued multiplication mode according to the signal. Further, the multipliers 11, the at least one adder 12 and the at least one subtractor 13 are configured to perform the complex-valued multiplication in a case that the multiplexer 15 chooses the complex-valued multiplication mode, where each of the complex numbers is taken as a compound with two components that are a real part and an imaginary part; or the multipliers 11, the at least one adder 12 and the at least one subtractor 13 are configured to perform the dual- valued multiplication in a case that the multiplexer 14 chooses the dual-valued multiplication mode, where each of the dual numbers is taken as a compound with two components that are a real part and a dual part. It can be seen that the at least one subtractor 13, the control logic instruction decoder 14 and the multiplexer 15 exist in a case that a combined implementation of multiplication for the complex numbers and the dual numbers is supported by the chip 10. The number of the devices, such as the multipliers 11, the at least one adder 12, and the at least one subtractor 13, depends on specific hardware implementation, for example, the number of adder 12 and the numberof subtractor 13 depend on the size of the vector of dual numbers or complex numbers. The reference can be made to the description of the method embodiments above for the specific process of the complex-valued or dual-valued multiplication, and it is not repeated herein.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed system and method may be implemented in other manners. For example, the described apparatus embodiment is merely an example. For example, the unit division is merely logical function division and may be other division in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.
[0123] The units described as separate parts may be or may not be physically separate, and parts displayed as units may be or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of the embodiments. In addition, functional units in the embodiments of this application may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.
[0124] The foregoing descriptions are merely specific implementations of this application, but are not intended to limit the protection scope of this application. Any variation or replacement readily figured out by a person skilled in the art within the technical scope disclosed in this application shall fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
CLAIMSWhat is claimed is:
1. A method for processing hypercomplex numbers, comprising: receiving a first instruction, wherein the first instruction indicates an operation, and the operation comprises a reading operation, a writing operation or an arithmetic operation; and performing, in response to the first instruction, the operation on an operand that comprises at least two components, wherein a combined size of the at least two components is less than or equal to a size of a native central processing unit (CPU) data word, and the operand is taken as one compound with the at least two components when the operation is performed.
2. The method according to claim 1, wherein the operand comprises a dual number, and the dual number consists of two components, wherein one of the two components is a real part and the other is a dual part, the two components are of the same size, and each of the two components is half the size of the native CPU data word.
3. The method according to claim 2, wherein the method is applied in a single-instruction-single-data (SISD) type arithmetic logic unit (ALU); wherein the two components of the dual number are stored in one item that takes up a space of the native CPU data word.
4. The method according to claim 2, wherein the method is applied in a single-instruction-multiple-data (SIMD) type or a multiple-instruction-multiple-data (MIMD) type ALU; wherein the two components of the dual number are stored in a way such that the two components are next to each other in a storage device.
5. The method according to claim 4, wherein the arithmetic operation comprises vector multiplication; wherein the performing, in response to the first instruction, the operation on an operand that comprises at least two components, comprises: performing, in response to the first instruction, the vector multiplication on a first vector VA and a second vector VB to obtain a third vector VC, wherein the operand comprises dual numbers in the first vector and the second vector, VA =for 0 < i < n, i and n are positive integers.
6. The method according to claim 1, wherein the operand comprises a complex number, and the complex number consists of two components, wherein one of the two components is a real part and the other is an imaginary part, the two componentsare of the same size, and each of the two components is half the size of the native CPU word.
7. The method according to any one of claims 3 to 6, wherein the arithmetic operation comprises a multiplication operation, and the method farther comprises: receiving a second instruction, wherein the second instruction indicates a multiplication mode, and the multiplication mode is a complex-valued multiplication mode or a dual-valued multiplication mode; wherein the performing, in response to the first instruction, the operation on an operand that comprises at least two components, farther comprises: performing the multiplication operation on operands to obtain a first result that comprises at least one compound, wherein a mode of the multiplication operation is the multiplication mode indicated by the second instruction.
8. The method according to claim 1 , wherein the operand comprises an algebraic or non-algebraic number of k dimensions, wherein the number of k dimensions comprises k components, and k sizes of the k components are in one-to-one correspondence to k sizes of k parts of the native CPU data word, respectively, wherein the native CPU data word is split into the k parts or w parts, wherein w is larger than k, and k and w are positive integers.
9. The method according to claim 8, wherein the k parts or the w parts are obtained by an even or uneven splitting in size of the native CPU data word.
10. The method according to claim 8 or 9, wherein the native CPU data word is split into the w parts, and remaining (w- k) parts from the w parts are unused parts.
11. An apparatus for processing hypercomplex numbers, comprising: multipliers and at least one adder, configured to: obtain operands, wherein a size of each operand is less than or equal to a size of a native central processing unit (CPU) data word, and each operand is a dual number; perform a multiplication operation on the operands.
12. The apparatus according to claim 11, wherein the dual number consists of two components, wherein one of the two components is a real part and the other is a dual part, the two components are of the same size, and each of the two components is half the size of the native CPU data word.
13. The apparatus according to claim 11 or 12, wherein the apparatus farther comprises a control logic instruction decoder, a multiplexer and at least one subtractor, wherein the control logic instruction decoder is configured to: output a signal to the multiplexer, wherein the signal indicates a multiplication mode, wherein the multiplication mode isa complex-valued multiplication mode or a dual-valued multiplication mode; wherein the multiplexer is configured to: choose a mode of the multiplication operation, wherein the mode of the multiplication operation is the multiplication mode indicated by the signal; wherein the multipliers, the at least one adder and the at least one subtractor are further configured to: perform the multiplication operation on the operands based on the chosen mode of the multiplication operation, wherein each operand is taken as a compound with two components that are a real part and an imaginary part if the chosen mode is the complex-valued multiplication mode; or each operand is taken as a compound with the two components that are the real part and the dual part if the chosen mode is the dual-valued multiplication mode.
14. The apparatus according to any one of claims 11 to 13, wherein the apparatus comprises a SISD type ALU.
15. An apparatus for processing hypercomplex numbers, comprising: multipliers and at least one adder, which are configured to: perform vector multiplication on a first vector KA and a second vector VB to obtain a third vector VC, wherein the first vector, the second vector and the third vector are vectors of dual numbers, each of the dual numbers comprises two components that are a real part and a dual part, and the two components of one dual number from any one of the vectors are stored in a way such that the two components are next to each other; and wherein VA = [AO,A1,A2, ....ATTT- 1)], VB = [ 80, Bl, B2, B'tn -' 1)], VC = [CO, Cl, C2, ... , C(iT- 1)], C[t] = A[i] * 8[i], for 0 < i < n, and n is an integer.
16. The apparatus according to claim 15, wherein the apparatus further comprises a control logic instruction decoder, a multiplexer and at least one subtractor; wherein the control logic instruction decoder is configured to: output a signal to the multiplexer, wherein the signal indicates a multiplication mode, and the multiplication mode is a complex-valued multiplication mode or a dual-valued multiplication mode; wherein the multiplexer is configured to: choose a mode of the vector multiplication, wherein the mode of the vector multiplication is the multiplication mode indicated by the signal; wherein the multipliers, the at least one adder and the at least one subtractor are further configured to: perform the vector multiplication on the first vector and the second vector based on the chosen mode of the vector multiplication, wherein the first vector, the second vector and the third vector are vectors of dual numbers if the chosen modeis the dual-valued multiplication mode; or the first vector, the second vector and the third vector are vectors of complex numbers if the chosen mode is the complex-valued multiplication mode, and two components of one complex number from any one of the vectors are stored in a way such that the two components are next to each other.
17. the apparatus according to claim 15 or 16, wherein the apparatus comprises a SIMD type or a MIMD type ALU.
18. A chip or a chip system, comprising an apparatus according to any one of claims 11 to 17.