System and method for rotating vector input
By incorporating a rotation vector register file and a MAC in computing devices, the processing bottleneck associated with conventional data transfer methods is mitigated, leading to improved efficiency and reduced circuit complexity for vector processing tasks.
Patent Information
- Application Number
- JP2024562331
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-02
- Filing Date
- 2023-04-25
- Publication Date
- 2025-05-27
AI Technical Summary
Existing computing devices face a processing bottleneck due to the conventional method of loading sub-vector values from memory to scalar registers, which involves multiple data transfers through various cache levels.
A processor is designed with a rotation vector register file and a multiply-accumulate circuit (MAC), where the rotation vector register file rotates data within its registers before inputting it to the MAC, reducing the need for multiple data transfers.
This approach enhances processing efficiency by minimizing data transfer overhead and enabling vector processing with reduced circuit complexity, thereby improving the performance of computationally intensive tasks like artificial neural network processing.
Smart Images

Figure 2025516160000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to commonly owned U.S. Patent Application No. 17 / 661,707, filed May 2, 2022, the entire contents of which are expressly incorporated herein by reference.
[0002] The present disclosure relates generally to data in vector registers for vector processing. [Background technology]
[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices that are small, lightweight, and easily carried by users, including wireless telephones such as mobile phones and smartphones, tablet computers, and laptop computers. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications, such as web browser applications that can be used to access the Internet. Thus, these devices can contain significant computing power.
[0004] The computing device may include one or more digital signal processors (DSPs), network processing units (NPUs), network signal processors (NSPs), image processors, or other processing devices that perform vector processing, which involves executing multiple instances of a common operation (e.g., a multiplication operation) to process multiple elements of vector data in parallel. For example, a vector may include multiple subvector values, such as 16 subvector values, each containing four halfword values. In an exemplary multiplication operation, for each subvector value of the vector, the first halfword is multiplied by a first one-byte value, the second halfword is multiplied by a second one-byte value, the third halfword is multiplied by a third one-byte value, and the fourth halfword is multiplied by a fourth one-byte value. The four multiplication products are added together, and the resulting sum is added to an accumulation vector register.
[0005] In some examples, each subvector value (e.g., four halfword values) may be read from a scalar register and provided as an input to a multiplier circuit. However, data transfer of halfword values from memory (e.g., dynamic random-access memory (DRAM), static random-access memory (SRAM), or another type of memory) to the scalar register can cause a processing bottleneck because the scalar register is loaded via a conventional processor operation that involves multiple transfers of data (e.g., loading the subvector values from memory to a second-level (L2) cache, from the L2 cache to a first-level (L1) cache, and from the L1 cache to a scalar register in a register file). Summary of the Invention
[0006] According to one implementation of the present disclosure, a device includes a processor including a rotating vector register file, a second vector register file, and multiply-accumulate circuitry (MAC). The rotating vector register file includes rotating vector registers. The rotating vector register file is configured to rotate data in the rotating vector registers. The second vector register file includes source vector registers. The MAC is configured to receive first input data from the rotating vector register file and second input data from the source vector registers.
[0007] According to another implementation of the present disclosure, a processor-implemented method includes rotating data in a rotating vector register of the rotating vector register file using a rotating vector register file. The processor-implemented method also includes receiving first input data from the rotating vector register file at a multiply-accumulate circuit (MAC). The processor-implemented method further includes receiving second input data at the MAC from a source vector register of a second vector register file.
[0008] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by a processor, cause the processor to rotate data in a rotating vector register of the rotating vector register file using a rotating vector register file. The instructions, when executed by the processor, also cause the processor to receive first input data from the rotating vector register file at a multiply-accumulate circuit (MAC). The instructions, when executed by the processor, further cause the processor to receive second input data from a source vector register of a second vector register file at the MAC.
[0009] According to another implementation of the present disclosure, an apparatus includes means for rotating data in a rotating vector register of a rotating vector register file. The apparatus also includes means, at a multiply-accumulate circuit (MAC), for receiving first input data from the rotating vector register file. The apparatus further includes means, at the MAC, for receiving second input data from a source vector register of a second vector register file.
[0010] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims. [Brief explanation of the drawings]
[0011] [Figure 1A] 1 is a block diagram of a particular illustrative aspect of a system operable to rotate a vector input, in accordance with certain examples of the present disclosure. [Figure 1B] 1B is a diagram of an example embodiment of data rotation of a rotating vector register of the system of FIG. 1A in accordance with some examples of the present disclosure. [Figure 2A] 1B is a diagram of an example embodiment of components of the system of FIG. 1A operable to rotate a vector input to execute an instruction, according to some examples of the present disclosure. [Figure 2B] 1B is a diagram of an exemplary aspect of a change in performance (e.g., count of operations) versus a change in operation intensity (e.g., operations per byte) of the system of FIG. 1A, in accordance with some examples of the present disclosure. [Figure 3] 2B is a diagram of an example embodiment of a first state of vector registers in the system of FIG. 1A before a first processing stage of execution of the instruction of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 4] 2B is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1A during a first sub-phase of a first processing phase of execution of the instructions of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 5]2B is a diagram of an example embodiment of a second state of a vector register in the system of FIG. 1A after a first sub-stage of a first processing stage of execution of the instruction of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 6] 2B is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1A during a second sub-phase of a first processing phase of execution of the instructions of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 7] 2B is a diagram of an example embodiment of a third state of a vector register in the system of FIG. 1A after a first processing stage of execution of the instruction of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 8] 2B is a diagram of an example embodiment of data rotation of a rotating vector register of the system of FIG. 1A during execution of the instruction of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 9] 2B is a diagram of an example embodiment of a fourth state of a vector register in the system of FIG. 1A before a second processing stage of execution of the instruction of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 10] 2B is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1A during a first sub-phase of a second processing phase of execution of the instructions of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 11] 2B is a diagram of an example embodiment of a fifth state of a vector register in the system of FIG. 1A after a first sub-stage of a second processing stage of execution of the instruction of FIG. 2A, in accordance with some examples of the present disclosure. [Figure 12] 1B is a diagram of an exemplary aspect of the operation of a selection circuit of the system of FIG. 1A, in accordance with some examples of the present disclosure. [Figure 13] 1B is a diagram of an exemplary aspect of the operation of the rotation circuit of the system of FIG. 1A, according to some examples of the present disclosure. [Figure 14] FIG. 1 illustrates an example of an integrated circuit including a rotating vector register file operable to rotate vector inputs, in accordance with some examples of the present disclosure. [Figure 15] FIG. 1 is a diagram of a mobile device including a rotating vector register file operable to rotate vector inputs, in accordance with some examples of the present disclosure. [Figure 16] FIG. 1 is a diagram of a headset including a rotating vector register file operable to rotate vector inputs, according to some examples of the present disclosure. [Figure 17] FIG. 1 is a diagram of a wearable electronic device including a rotating vector register file operable to rotate a vector input, according to some examples of the present disclosure. [Figure 18] FIG. 1 is a diagram of a voice-controlled speaker system including a rotating vector register file operable to rotate a vector input, according to some examples of the present disclosure. [Figure 19] 1 is a diagram of a camera including a rotation vector register file operable to rotate vector inputs, according to some examples of the present disclosure. [Figure 20] 1 is a diagram of a headset, such as a virtual reality, mixed reality, or augmented reality headset, including a rotation vector register file operable to rotate a vector input, according to some examples of the present disclosure. [Figure 21] 1 is a diagram of a first example vehicle including a rotating vector register file operable to rotate vector inputs, in accordance with some examples of the present disclosure. [Figure 22] FIG. 10 is a diagram of a second example vehicle including a rotating vector register file operable to rotate vector inputs, according to some examples of the present disclosure. [Figure 23] 1B is a diagram of a specific implementation of a method for rotating a vector input that may be performed by the device of FIG. 1A, in accordance with some examples of the present disclosure. [Figure 24] FIG. 1 is a block diagram of a particular illustrative example of a device including a rotating vector register file operable to rotate vector inputs, in accordance with certain examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] A vector register file is disclosed that includes a rotating vector register and a rotation circuit. The rotation circuit is operable to rotate vector data in the rotating vector register before providing subvector values of the rotated vector data as input to another circuit, such as a multiply-accumulate circuit (MAC). In conventional systems, multiple transfers of data to load subvector values from memory into scalar registers cause processing bottlenecks. Such processing bottlenecks are reduced (e.g., eliminated) by using a vector register file instead of a scalar register because the subvector values can be loaded into the rotating vector register with fewer data transfers (e.g., a single transfer).
[0013] According to some aspects, a vector load operation may be performed to load vector data into a rotating vector register. In some examples, the vector data is loaded from memory to a cache memory (e.g., an L2 cache) and from the cache memory to the rotating vector register. Any intermediate cache memory (e.g., an L1 cache) between the cache memory (e.g., the L2 cache) and the rotating vector register is bypassed to reduce overhead associated with loading data from the cache memory (e.g., the L2 cache) to the intermediate cache memory (e.g., the L1 cache) and from the intermediate cache memory to the rotating vector register. The vector data may include multiple subvector values (e.g., 16 subvector values). Each subvector value may include multiple elements (e.g., four halfwords).
[0014] Artificial neural network processing for machine learning tasks is often computationally intensive, requiring a large number of multiplication and addition operations. In some examples, a processor may execute a vector instruction (e.g., a multiply-accumulate instruction) to multiply each element of each subvector value stored in a rotating vector register by an element of vector data stored in a source vector register. For example, during a first processing stage of executing the vector instruction, the processor broadcasts elements (e.g., four halfwords) of the subvector values (v0-v3) of the rotating vector register to respective multipliers. Illustratively, the first halfword (v0), second halfword (v1), third halfword (v2), and fourth halfword (v3) of the subvector values are broadcast in parallel to the first multiplier, second multiplier, third multiplier, and fourth multiplier, respectively.
[0015] The source vector register stores sub-vector values (s0-s3), including a first 1-byte value (s0), a second 1-byte value (s1), a third 1-byte value (s2), and a fourth 1-byte value (s3). As used herein, "v" is used as a prefix to indicate a sub-vector value read from a rotating vector register, and "s" is used as a prefix to indicate a sub-vector value read from a source vector register.
[0016] During the first sub-stage of the first processing stage, each element of the sub-vector value in the rotating vector register is multiplied by the corresponding value of the sub-vector value in the source vector register. For example, the first multiplier multiplies a first halfword (v0) by a first one-byte value (s0) to generate a first multiplication product (v0s0), a second halfword (v1) by a second one-byte value (s1) to generate a second multiplication product (v1s1), a third halfword (v2) by a third one-byte value (s2) to generate a third multiplication product (v2s2), and a fourth halfword (v3) by a fourth one-byte value (s3) to generate a fourth multiplication product (v3s3). The four multiplication products are added together, and the resulting sum is added to an accumulation vector register.
[0017] During a second sub-stage of the first processing stage, each element of the sub-vector values (v0-v3) of the rotating vector register is multiplied by the corresponding value of the next sub-vector value (s4-s7) of the source vector register. The four multiplication products are added together, and the resulting sum is added to an accumulation vector register. In some aspects, additional sub-stages of the first processing stage are performed until the sub-vector values (v0-v3) of the rotating vector register have been multiplied with all of the sub-vector values of the source vector register.
[0018] After the first processing stage, the next subvector value (v4 through v7) in the rotating vector register is used for multiplication. Traditionally, enabling broadcasting of elements of subvector values from different portions of the vector register can increase circuit complexity. For example, a selection circuit must be able to read each element (e.g., each of the four halfword values) of each of the subvector values in the rotating vector register and provide the selected subvector value as an input to a broadcast circuit for broadcast to the appropriate multiplier.
[0019] To reduce such complexity, a rotation circuit is used to rotate data within a rotating vector register. For example, a specific portion of the rotating vector register is dedicated for broadcasting. During a first processing stage, subvector values (v0-v3) are read from the dedicated portion of the rotating vector register and broadcast to a multiplier. During a rotation stage after the first processing stage, the subvector values are rotated within the rotating vector register such that the next subvector value (v4-v7) is moved to the dedicated portion of the rotating vector register. During a second processing stage after the rotation stage, the next subvector value (v4-v7) is broadcast to a multiplier, which generates a multiplication product based on the next subvector value and updates an accumulation vector register.
[0020] In some aspects, additional rotation and processing steps may be performed until all sub-vector values in the rotating vector register have been rotated at least once into the dedicated portion, and then the next batch of vector data may be loaded into the rotating vector register. Enabling the selection circuitry to read data from elements in the dedicated portion of the rotating vector register reduces circuit complexity compared to enabling the selection circuitry to read data from all portions of the rotating vector register.
[0021] A sub-vector value of a rotating vector register containing four elements (e.g., four halfword values) broadcast to four multipliers is provided as an illustrative example. In other examples, the sub-vector value of a rotating vector register may contain fewer or more than four elements, which may be broadcast to the respective multipliers.
[0022] Certain aspects of the present disclosure are described below with reference to the drawings. In this description, common features are indicated by common reference numerals. Various terms used herein are used only for the purpose of describing particular implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly dictates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1A shows a device 102 including one or more processors ("processor(s)" 190 in FIG. 1A), indicating that in some implementations, the device 102 includes a single processor 190 and in other implementations, the device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and subsequently referred to in the singular unless an aspect relating to multiple features is described.
[0023] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. For example, referring to FIG. 1A, one or more rotators are shown and associated with reference numbers 172A and 172M. When referring to a particular one of these rotators, such as rotator 172A, the identifying letter "A" is used. However, when referring to any one of these rotators, reference number 172 is used without the identifying letter.
[0024] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” denotes an example, implementation, and / or aspect and should not be construed as limiting or as indicating a preferred or preferred implementation. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify an element, such as a structure, component, operation, etc., do not in themselves indicate a priority or order of the element with respect to other elements, but merely distinguish the element from other elements having the same name (apart from the use of ordinal terms). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.
[0025] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and also (or alternatively) may include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronic components, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) with no intervening components.
[0026] In this disclosure, terms such as "determining," "calculating," "estimating," "shifting," "adjusting," and the like may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generating," "calculating," "estimating," "using," "selecting," "accessing," and "determining" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component or device.
[0027] 1A, a particular exemplary embodiment of a system 100 configured to rotate a vector input is shown. System 100 includes a device 102 including a memory 132 coupled to one or more processors 190.
[0028] The one or more processors 190 include a multiply-accumulate circuit (MAC) 160 coupled to a vector register (VR) file 150 and a rotated VR file 140. The VR file 150 includes a source VR 154, an accumulation VR 156, one or more additional VRs, or a combination thereof. The rotated VR file 140 includes one or more rotated VRs 144 coupled to a rotation circuit 142. The one or more rotated VRs 144 include a rotated VR 144A, one or more additional rotated VRs, a rotated VR 144N, or a combination thereof, which are referred to herein as rotated VRs 144A-144N. In some examples, the rotated VRs 144A-144N include a single rotated VR 144A. In other examples, the rotated VRs 144A-144N include a rotated VR 144A and one or more additional rotated VRs.
[0029] Rotation circuit 142 includes one or more rotators, such as rotator 172A, one or more additional rotators, rotator 172M, or combinations thereof, referred to herein as rotators 172A-172M. In some examples, rotators 172A-172M include a single rotator 172A. In other examples, rotators 172A-172M include rotator 172A and one or more additional rotators. The count of rotators 172A-172M may be less than, equal to, or greater than the count of rotation VRs 144A-144N.
[0030] Each of rotators 172A-172M is configured to rotate data by a rotation amount. For example, rotator 172A is configured to receive data 183 from one or more of rotator VRs 144A-144N, perform a data rotation corresponding to rotation amount 174A to generate rotated data 177, and store rotated data 177 in one or more of rotator VRs 144A-144N, as further described with reference to FIG. 13 . Illustratively, performing the data rotation includes rotating at least a portion of data 183 by rotation amount 174A.
[0031] In some implementations, rotator 144A and rotator 172A together function as a shift register. For example, rotator 172A may include a multiplexer, dedicated wiring, or other bit-shifting hardware, causing bit values read from rotator VR 144A to be written back to rotator VR 144A at shifted positions corresponding to rotation amount 174A. Illustratively, rotator 172A performs a circular shift of the bit values read from rotator VR 144A based on rotation amount 174A.
[0032] In some alternative implementations, rotator VRs 144A-144N include shift registers, and rotators 172A-172M provide inputs (e.g., data advance signals) to rotator VRs 144A-144N based on the rotation amount to perform the data rotation. As an illustrative example, rotator VR 144A includes a shift register with a cascade of flip-flops, with the output of one flip-flop connected to the input of the next flip-flop and the last flip-flop connected to the first flip-flop. Upon receiving the data advance signal, the bit value stored in each flip-flop is shifted to the next flip-flop, and the bit value from the last flip-flop is shifted to the first flip-flop. In some aspects, rotator 172A provides the data advance signal to rotator VR 144A, and the count of the data advance signals provided to rotator VR 144A is based on rotation amount 174A.
[0033] In some implementations, variable target rotation may be achieved with multiple cascaded fixed rotations, as further described with reference to FIG. 13 . In an illustrative example, rotator 172A reads a bit value from rotator VR 144A, generates a first rotated bit value by rotating the bit value based on rotation amount 174A, and stores the first rotated bit value in a pipeline register. A second rotator among rotators 172A-172M reads the first rotated bit value from the pipeline register, generates a second rotated bit value based on the first rotated bit value, and writes the second rotated bit value to rotator VR 144A. In one example, the second rotator is configured to selectively rotate the bit value based on the second rotation amount. Illustratively, if the target rotation is the same as rotation amount 174A, the second rotator outputs the first rotated bit value (without rotation) as the second rotated bit value. If the target rotation is the same as the sum of rotation amount 174A and the second rotation amount, the second rotator rotates the first rotated bit value based on the second rotation amount. Thus, a variable target rotation (e.g., rotation amount 174A or the sum of rotation amount 174A and the second rotation amount) can be achieved using two cascaded rotations. The fixed rotation amount to be cascaded can be selected based at least in part on an instruction set architecture (ISA) specification associated with MAC 160, which can describe various rotation amounts to be supported. Using a pair of rotators, each configured to perform a predetermined rotation, can reduce hardware costs (e.g., power consumption, area, or both), improve performance (e.g., run faster), or a combination thereof, compared to a single rotator configured to perform a variable rotation.
[0034] Rotation VRs 144A-144N are also coupled to selection circuit 146. In some implementations, selection circuit 146 is configured to read data from a dedicated portion of each of rotation VRs 144A-144N and is not configured to read data from the remaining portions of each of rotation VRs 144A-144N. As an illustrative example, rotation VR 144A includes dedicated portion 134A that is readable by selection circuit 146. In some implementations, selection circuit 146 is configured to read data from dedicated portion 134A and is not configured to read data from the remaining portions of rotation VR 144A. For example, in these implementations, selection circuit 146 is coupled to dedicated portion 134A of rotation VR 144A and is not coupled to the remaining portions of rotation VR 144A.
[0035] In some implementations, the selection circuit 146 is configurable. For example, the selection circuit 146 can read data 145 from one or more sub-portions of the dedicated portion of the rotating VRs 144A-144N based on a configuration 148, as further described with reference to FIG.
[0036] MAC 160 is configured to receive data 145 from rotated VR file 140 and data 155 from source VR 154. MAC 160 is configured to generate output data 165 based on data 145 and data 155. For example, MAC 160 is configured to generate a multiplication product based on elements of data 145 and elements of data 155, generate output data 165 based on the multiplication product, and update accumulated data 167 in accumulation VR 156 based on the output data 165. By way of example, MAC 160 may be used in conjunction with artificial neural network processing, including multiplying weight values by corresponding activation values. In an example artificial neural network, data 145 may correspond to activation values, and data 155 may correspond to weight values. In some examples, MAC 160 is configured to store output data 165 in accumulation VR 156 as accumulated data 167. In some examples, MAC 160 is configured to update accumulated data 167 by adding the value indicated by output data 165 to the value indicated by accumulated data 167 and storing the sum as accumulated data 167 in accumulation VR 156.
[0037] In some implementations, device 102 corresponds to or is included in one of various types of devices. In an illustrative example, one or more processors 190 are integrated into a headset device, such as further described with reference to FIG. 16. In other examples, one or more processors 190 are integrated into at least one of a mobile phone or tablet computing device, such as described with reference to FIG. 15, a wearable electronic device, such as described with reference to FIG. 17, a voice-controlled speaker system, such as described with reference to FIG. 18, a camera device, such as described with reference to FIG. 19, or a virtual reality, mixed reality, or augmented reality headset, such as described with reference to FIG. 20. In another illustrative example, one or more processors 190 are integrated into a vehicle, such as further described with reference to FIGS. 21 and 22.
[0038] As will be further described with reference to FIG. 2A , in operation, vector 151 is loaded from memory 132 into source VR 154, and vector 171 is loaded from memory 132 into rotator VRs 144A-144N. In certain aspects, vector 171 includes multiple sub-vector values, such as sub-vector values (SV) 173A, SV173B, one or more additional SVs, SV173T, or combinations thereof, referred to herein as SV173A-173T. SV173A-173T are shown in FIG. 1B as data that can be rotated stored in rotator VR 144A, as will be further described below. In certain aspects, vector 151 includes multiple sub-vector values, such as SV175A, SV175B, one or more additional SVs, SV175S, or combinations thereof, referred to herein as SV175A-173S. The counts of SV173A-173T may be less than, equal to, or greater than the counts of SV175A-175S. In some examples, as will be further described with reference to FIG. 3, vector 151 is loaded into a single source vector register (e.g., source VR 154) and vector 171 is loaded into a single rotation VR (e.g., rotation VR 144A). In some examples, vector 171 is loaded into multiple rotation vector registers. For example, as will be further described with reference to FIGS. 12-13, a portion of vector 171 is loaded into rotation VR 144A and another portion of vector 171 is loaded into rotation VR 144B. In some examples, vector 151 may be loaded into multiple source vector registers, including source VR 154 and one or more additional source VRs.
[0039] The elements of vector 171 from rotator VRs 144A-144N are multiplied with the elements of vector 151 from source VR 154 to generate an output to be stored in accumulation VR 156. In an illustrative example, vector 171 is stored in rotator VR 144A, and the elements of vector 171 stored in rotator VR 144A are multiplied by MAC 160 with the elements of vector 151 stored in source VR 154. For example, during each of multiple processing stages, a different SV stored in rotator VR 144A is multiplied with the SV stored in source VR 154.
[0040] In the illustrative example, during the first processing stage, the selection circuit 146 reads the SV 173A from the dedicated portion 134A of the rotation VR 144A based on the configuration 148, as further described with reference to Figure 12. The selection circuit 146 provides the SV 173A to the MAC 160 as data 145.
[0041] In some implementations, each processing stage includes multiple sub-stages, where a different SV is provided from source VR 154 to MAC 160 as data 155 in each sub-stage. For example, during the first sub-stage of the first processing stage, MAC 160 receives SV 175A from source VR 154 as data 155. MAC 160 generates output data 165 based on data 145 (received from rotated VR file 140) and data 155 (received from source VR 154), as further described with reference to FIG. 4 . In some implementations, elements of data 145 are multiplied with elements of data 155. For example, elements of SV 173A received from rotated VR 144A are multiplied in parallel with elements of SV 175A received from source VR 154. Illustratively, the first element of SV173A is multiplied by the first element of SV175A, the second element of SV173A is multiplied by the second element of SV175A, and so on. In some examples, the elements of SV173A have the same size (e.g., include the same bit count) as the elements of SV175A. In other examples, the elements of SV173A have a different size (e.g., include a different bit count) than the elements of SV175A. MAC 160 stores output data 165 in accumulation VR 156 as accumulation data 167.
[0042] If vector 151 stored in source VR 154 includes multiple sub-vector values, each processing stage may include multiple sub-stages. For example, during the second sub-stage of the first processing stage, MAC 160 receives SV 175B (e.g., the next sub-vector value) from source VR 154 as data 155. Illustratively, elements of SV 173A received from rotator VR 144A are multiplied in parallel with elements of SV 175B received from source VR 154. MAC 160 generates output data 165 based on data 145 (e.g., SV 175A) and data 155 (e.g., SV 175B). In some aspects, additional sub-stages of the first processing stage are performed by MAC 160 processing SV 173A and each of the SVs of source VR 154 until SV 173A and SV 175S have been processed by MAC 160.
[0043] During a rotation stage after the first processing stage, rotators 172A-172M rotate the data in one or more of rotator VRs 144A-144N by a rotation amount, as further described with reference to FIG. 13 . For example, rotator 172A receives data 183 from rotator VR 144A, rotates data 183 by rotation amount 174A to generate rotated data 177, and stores rotated data 177 in rotator VR 144A. Referring briefly to FIG. 1B , in illustrative example 180, SV 173A is stored in dedicated portion 134A before rotation. Illustratively, data 183 corresponds to SV 173A and is followed by SV 173B, one or more additional SVs, and SV 173T. After rotation, SV 173B is stored in dedicated portion 134A of rotator VR 144A. Illustratively, rotation data 177 corresponds to SV173B, followed by one or more additional SVs, SV173T and SV173A.
[0044] 1A , in some embodiments, during the first rotation after vector 171 is loaded from memory 132 into rotation VR 144A, data 183 is the same as vector 171 retrieved from memory 132. In some embodiments, in subsequent rotations, data 183 corresponds to a rotated version of vector 171. For example, in these embodiments, data 183 corresponds to rotated data 177 generated in a previous rotation.
[0045] In some implementations, multiple rotators are used to generate rotated data 177. For example, rotator 172A rotates data 183 by rotation amount 174A to generate first rotated data, a second rotator of rotators 172A-172M rotates the first rotated data by a second rotation amount to generate second rotated data, and so on until a rotator of rotators 172A-172M rotates the rotated data received from the previous rotator to generate rotated data 177. Rotation circuit 142 stores rotated data 177 in rotation VR 144A.
[0046] During the second processing stage, selection circuit 146 retrieves data 145 (e.g., SV173B) from dedicated portion 134A and provides data 145 to MAC 160. During a first sub-stage of the second processing stage, MAC 160 receives SV175A from source VR 154 as data 155. MAC 160 generates output data 165 based on SV173B (e.g., data 145) and SV175A (e.g., data 155) and updates accumulated data 167 based on output data 165.
[0047] During a second sub-stage of the second processing stage, MAC 160 receives SV 175B from source VR 154 as data 155. MAC 160 generates output data 165 based on SV 173B (e.g., data 145) and SV 175B (e.g., data 155) and outputs accumulated data 167 based on output data 165. In some aspects, additional sub-stages of the second processing stage are performed by MAC 160 processing SV 173B and each of the SVs of source VR 154 until SV 173B and SV 175S have been processed by MAC 160. In some aspects, additional rotation and processing stages are performed until all SVs 173A-173T of vector 171 have been processed by MAC 160 to update accumulated VR 156.
[0048] MAC 160 multiplying elements of a single SV retrieved from rotation vector register 144A with elements of a single SV retrieved from source VR 154 during each processing stage is provided as an illustrative example. In some implementations, one or more processors 190 may include multiple copies of components of MAC 160 described herein that may be used to multiply elements of an SV retrieved from one of rotation VRs 144A-144N with elements of multiple SVs retrieved from source VR 154. For example, during a first processing stage, a first copy of the component may multiply elements of SV173A with elements of SV175A, and in parallel, a second copy of the component may multiply elements of SV173A with elements of SV175B. In some implementations, MAC 160 may multiply elements of SV173A with two or more elements of SV175A-175S in parallel. During a rotation stage after the first processing stage, SV173A is rotated to the end of rotation VR144A, and SV173B is rotated to dedicated portion 134A. During a second processing stage, a first copy of the component can multiply elements of SV173B with elements of SV175A, and in parallel, a second copy of the component can multiply elements of SV173B with elements of SV175B. In some implementations, MAC 160 can multiply elements of SV173B with two or more elements of SV175A-175S in parallel.
[0049] In some implementations, one or more processors 190 may include multiple copies of a component of MAC 160 described herein that may be used to multiply elements of multiple SVs derived from rotation VRs 144A-144N with elements of an SV derived from source VR 154. For example, during a first processing stage, a first copy of the component may multiply elements of SV173A with elements of SV175A, and in parallel, a second copy of the component may multiply elements of SV173B with elements of SV175A. In a first embodiment, SV173A and SV173B are derived from the same rotation VR (e.g., rotation VR 144A). In a second embodiment, SV173A is derived from rotation VR 144A, and SV173B is derived from rotation VR 144B. During a rotation stage after the first processing stage, the SVs provided to the components of MAC 160 during the first processing stage (e.g., SV173A, SV173B, one or more additional SVs, or a combination thereof) are rotated. In a first embodiment, SV173A, SV173B, one or more additional SVs, or a combination thereof are rotated to the end of rotation VR144A. In a second embodiment, SV173A is rotated to the end of rotation VR144A, SV173B is rotated to the end of rotation VR144B, and each of the one or more additional SVs is rotated to the end of a corresponding rotation VR, or a combination thereof.
[0050] Thus, system 100 enables vector processing with reduced complexity in selection circuitry 146. For example, selection circuitry 146 can read data from dedicated portion 134A of rotation VR 144A without needing to include circuitry to support reading from the remainder of rotation VR 144A.
[0051] Referring to Figure 2A, a diagram 200 of an exemplary embodiment of one or more processors 190 and memory 132 of system 100 of Figure 1A is shown. One or more processors 190 are operable to rotate vector inputs to execute instructions.
[0052] The one or more processors 190 include an instruction buffer 202 coupled to an instruction selector 204. The instruction selector 204 is coupled to a MAC 160 and a load / store unit 208. The load / store unit 208 is coupled to memory 132. In particular aspects, the load / store unit 208 is coupled to memory 132 via one or more caches. For example, the load / store unit 208 is coupled to an L2 cache 234 via an L1 cache 232, which is coupled to memory 132. In some aspects, the load / store unit 208 is also coupled to the L2 cache 234 (bypassing the L1 cache 232). The MAC 160 and the load / store unit 208 are each coupled to a source VR 154, a rotating VR file 140, and an accumulate VR 156.
[0053] In certain aspects, instructions are added to an instruction set architecture (ISA). The ISA can support instructions corresponding to various configurations of select circuit 146, rotate circuit 142, or both, various SV sizes for rotate VRs 144A-144N, various SV sizes for source VR 154, one or more of SVs with signed values, SVs with unsigned values, one or more source VRs for storing vector 151, rotate VRs 144A-144N for storing vector 171, or combinations thereof. In some implementations, one or more processors 190 correspond to vector processors that implement the ISA. For example, one or more processors 190 are configured to operate efficiently on vectors. In particular aspects, one or more processors 190 are configured to efficiently copy vectors (e.g., large one-dimensional data arrays) from memory 132 to vector registers and vice versa, and to perform parallel processing of multiple data values from the vector registers, such as using multiple parallel computation lanes in a single instruction multiple data (SIMD) configuration.
[0054] The instruction selector 204 uses various instruction selection techniques to select an instruction 280 (e.g., a multiply-accumulate instruction) from the instruction buffer 202 for processing. The instruction 280 includes an opcode 282 and a parameter 284. The opcode 282 indicates the type of the instruction 280 (e.g., multiply-accumulate). In a first implementation, the parameter 284 indicates a first memory address of the vector 171 and a second memory address of the vector 151. In the first implementation, during execution of the instruction 280, the vector 171 is loaded into at least one of the rotation VRs 144A-144N, and the vector 151 is loaded into the source VR 154. In a second implementation, before execution of the instruction 280, the vector 171 is loaded into at least one of the rotation VRs 144A-144N, and the vector 151 is loaded into the source VR 154. In a second implementation, the parameters 284 indicate at least one of the rotation VRs 144A-144N and the source VR 154.
[0055] The instruction selector 204 provides an instruction 280 (or data representing aspects of the instruction 280) to the MAC 160, the load / store unit 208, or both. In some aspects, the instruction selector 204 initiates execution of the instruction 280 by providing an opcode 282 to the MAC 160.
[0056] In a first implementation, parameters 284 indicate a first memory address of vector 171 and a second memory address of vector 151, and instruction selector 204 provides the first memory address and the second memory address to load / store unit 208. In response to receiving the first memory address, load / store unit 208 performs load operation 273 of vector 171 to rotated VR file 140. In a particular implementation, in response to receiving the first memory address of vector 171, load / store unit 208 loads vector 171 from L2 cache 234 to rotated VR file 140. If vector 171 is not available in L2 cache 234, load / store unit 208 copies vector 171 from the received first memory address of memory 132 to L2 cache 234 and copies vector 171 from L2 cache 234 to rotated VRs 144A-144N (e.g., bypassing L1 cache 232). Similarly, in response to receiving the second memory address, the load / store unit 208 performs a load operation 275 of the vector 151 to the source VR 154. In a particular implementation, in response to receiving the second memory address, the load / store unit 208 loads the vector 151 from the L2 cache 234 to the source VR 154. If the vector 151 is not available in the L2 cache 234, the load / store unit 208 copies the vector 151 from the received second memory address of the memory 132 to the L2 cache 234 and from the L2 cache 234 to the source VR 154 (e.g., bypassing the L1 cache 232).
[0057] One or more processors 190 execute instructions 280 to retrieve data 145 from rotated VR file 140, retrieve data 155 from source VR 154, process data 145 and data 155 in MAC 160 to generate output data 165, store output data 165 in accumulation VR 156, and rotate data in rotation VR 144A after data 145 is retrieved from rotated VR file 140. For example, executing instructions 280 may include performing multiple processing steps, one or more rotation steps, or a combination thereof, to update accumulated data 167, as described with reference to FIG.
[0058] In certain aspects, load / store unit 208 stores accumulated data 167 in memory 132 in response to determining that SVs 173A-173T of vector 171 have been processed by MAC 160. In some examples, vector 171 corresponds to a portion of vector data, and load / store unit 208 loads the next portion of vector data into rotator VR 144A as vector 171 after MAC 160 has processed vector 171. In these examples, additional processing steps, rotation steps, and loads are performed until the entire vector data has been processed.
[0059] Referring to FIG. 2B, a diagram 290 shows the change in performance (eg, count of operations) for varying operation intensity (eg, operations per byte) of MAC 160.
[0060] During the first processing stage, MAC 160 generates output data 165 based on vector 171 and vector 151 loaded into rotated VR file 140 and source VR 154, respectively. In some implementations, MAC 160 generates output data 165 once data 145 (e.g., SV 173A) and data 155 (e.g., SV 175A) are loaded. In these implementations, output data 165 is generated simultaneously with the loading of the remaining SVs of vector 171 and vector 151. In other implementations, MAC 160 generates output data 165 once all SVs of vector 171 and all SVs of vector 151 are loaded into rotated VR file 140 and source VR 154, respectively. The performance of MAC 160 corresponds to the bandwidth for loading vector 171 and vector 151 from memory 132 (or L2 cache).
[0061] Once vector 171 and vector 151 are loaded, the "optimal" computational intensity of MAC 160 is reached. For example, because the inputs to MAC 160 are available in rotated VR file 140 and source VR 154, the performance of MAC 160 is not limited by bandwidth. Thus, rotating the data in rotated VR file 140 to make the next subvector value available to MAC 160 allows MAC 160 to run at a higher (e.g., optimal) computational intensity compared to retrieving the next subvector value as a scalar value from memory 132 as input to MAC 160.
[0062] 3-11 show an example of executing instruction 280. Instruction 280 corresponds to a "vmpyaddhbz_x4" instruction, where "vmpyadd" indicates that a vector multiply-accumulate (e.g., add) algorithm is being implemented. Instruction 280 may have the following format: Vxx+=vmpyaddhbz_x4(Vuu,Zy4):sat:rot, where Vxx is the destination VR (e.g., accumulation VR 156), Vuu is the source VR (e.g., source VR 154), Zy4 is the rotate VR (e.g., rotate VR 144A), "sat" indicates that the value is saturated, and "rot" indicates that a rotate step is performed between processing steps.
[0063] Although examples herein describe an instruction 280 that performs a vector multiply-accumulate operation that reads data from a single source VR (e.g., source VR 154), other examples illustrate that an instruction 280 of the form Vxx+=vmpyaddzt_x8(Vuu,Vvv,Zy4):sat:rot can read data from multiple source VRs, such as a first source VR represented by Vuu and a second source VR represented by Vvv.
[0064] 3, there is shown a diagram 300 of an exemplary embodiment of a first state of Rotate VR 144A, Source VR 154, and Accumulate VR 156. In a particular embodiment, Rotate VR 144A, Source VR 154, and Accumulate VR 156 are in the first state prior to a first processing stage of execution of instruction 280 of FIG.
[0065] In a particular implementation, rotation VR 144A includes 64 elements (e.g., v0-v63), where elements v0-v63 each include a halfword value, and SV includes 8 elements. For example, SV 173A includes elements v0-v7, SV 173B includes elements v8-v15, and so on. In a first processing state, SV 173A is stored in dedicated portion 134A of rotation VR 144A. In particular aspects, SV 173A may include multiple SVs. For example, SV 173A includes SV 312A (e.g., v0-v3) and SV 312B (e.g., v4-v7).
[0066] In a particular implementation, source VR 154 includes 256 elements (e.g., s0 to s255), and each of elements s0 to s255 includes a byte value. In a particular aspect, the SV of source VR 154 includes four elements. For example, SV320A includes elements s0 to s3, and SV320B includes elements s128 to s131. In a particular aspect, SV175A includes SV320A and SV320B.
[0067] In a particular implementation, the accumulation VR 156 includes 64 elements (e.g., a0 to a63). For example, the accumulation VR 156 includes an SV 340A (e.g., a0), an SV 340B (e.g., a32), one or more additional SVs, or a combination thereof.
[0068] In a first state shown by diagram 300, before the first processing stage, the selection circuit 146 is configured to select SV173A stored in the dedicated portion 134A as data 145 for the first processing stage, which will be further described with reference to FIG. 4.
[0069] Referring to FIG. 4, a diagram 400 of an exemplary aspect of the operation of components of system 100 of FIG. 1A during a first sub-phase of a first processing phase of execution of instructions 280 is shown.
[0070] In a particular example, instruction 280 has the following format: Vxx+=vmpyaddhbz_x4(Vuu,Zy4):sat:rot, and the processing stages of execution of instruction 280 correspond to the following pseudocode for a loop: fHIDE(int i;)for(i=0;i<32;i++){ fHIDE(size8s_t acc[2];)acc[0]=Vxx.V32s[i]; acc[0]+=fMPY8SS(Vuu.V8s[4 * i+0],ZyV.V16s[0]); acc[0]+=fMPY8SS(Vuu.V8s[4 *i+1],ZyV.V16s[1]); acc[0]+=fMPY8SS(Vuu.V8s[4 * i+2],ZyV.V16s[2]); acc[0]+=fMPY8SS(Vuu.V8s[4 * i+3],ZyV.V16s[3]); acc[1]=Vxx.V32s[32+i]; acc[1]+=fMPY8SS(Vuu.V8s[128+4 * i+0],ZyV.V16s[4]); acc[1]+=fMPY8SS(Vuu.V8s[128+4 * i+1],ZyV.V16s[5]); acc[1]+=fMPY8SS(Vuu.V8s[128+4 * i+2],ZyV.V16s[6]); acc[1]+=fMPY8SS(Vuu.V8s[128+4 * i+3],ZyV.V16s[7]); Vxx.V32s[i]=fVSATN(32,acc[0]); Vxx.V32s[32+i]=fVSATN(32,acc[1]); } where Vxx corresponds to the accumulate VR 156, acc[0] corresponds to SV340A, acc[1] corresponds to SV340B, Vuu corresponds to the source VR 154, ZyV corresponds to the rotate VR 144A, V8s indicates a byte value, V16s indicates a halfword value, and fMPY8SS indicates a signed floating-point multiplication. Each iteration of the loop corresponds to a sub-stage of the processing stage.
[0071] During a first sub-stage of the first processing stage, selection circuit 146 selects SV 173A as data 145, as described with reference to FIGURE 1A. One or more processors 190 include a broadcast circuit 402 coupled to MAC 160. MAC 160 includes multiple multipliers, e.g., multiplier 410A, multiplier 410B, multiplier 410D, multiplier 410E, one or more additional multipliers, multiplier 410H, or combinations thereof, referred to herein as multipliers 410A-410H.
[0072] MAC 160 includes multiple inputs, where a pair of inputs from the multiple inputs is associated with a corresponding multiplier. In some aspects, MAC 160 includes inputs 407A-H configured to couple to broadcast circuit 402 and inputs 409A-H configured to couple to source VR 154. For example, input 407A of MAC 160 corresponds to (e.g., includes or is coupled to) the first input of multiplier 410A, input 407B of MAC 160 corresponds to (e.g., includes or is coupled to) the first input of multiplier 410B, and so on. As another example, input 409A of MAC 160 corresponds to (e.g., includes or is coupled to) the second input of multiplier 410A, input 409B of MAC 160 corresponds to (e.g., includes or is coupled to) the second input of multiplier 410B, and so on. In some aspects, inputs 407A-H and inputs 409A-H enable parallel multiplication of elements of data 145 with elements of data 155. For example, multiplier 410A multiplies the data element received from input 407A with the data element received from input 409A, multiplier 410B multiplies the data element received from input 407B with the data element of data 155 received from input 409B, and so on.
[0073] In particular implementations, the broadcast circuit 402 includes multiple outputs, each coupled to a corresponding input of the MAC 160 associated with a single multiplier. For example, a first output of the broadcast circuit 402 is coupled to an input 407A associated with the multiplier 410A and is not coupled to an input associated with any of the multipliers 410B-410H. As another example, a second output of the broadcast circuit 402 is coupled to an input 407B associated with the multiplier 410B and is not coupled to an input associated with any of the other multipliers 410A-410H. In particular aspects, one or more of the outputs of the broadcast circuit 402, one or more of the inputs of the MAC 160 (e.g., inputs 407A-H and inputs 409A-H), or a combination thereof, may include one or more bus interfaces, one or more latches, one or more flip-flops, one or more buffers, other data buffering circuitry, or a combination thereof.
[0074] In a particular aspect, selection circuit 146 provides data 145 (e.g., SV173A) to broadcast circuit 402. Data 145 (e.g., SV173A) includes a plurality of elements (e.g., v0-v7). Broadcast circuit 402 provides data 145 to multipliers 410A-410H. For example, broadcast circuit 402 provides, for each of the plurality of elements (e.g., v0-v7), the element to a respective separate input of inputs 407A-H of MAC 160. Illustratively, a first element (e.g., v0) is provided to input 407A associated with multiplier 410A via a first output of broadcast circuit 402 and is not provided by broadcast circuit 402 to any inputs associated with other multipliers of MAC 160. Similarly, the second element (e.g., v1) is provided to input 407B associated with multiplier 410B via a second output of broadcast circuit 402 and is not provided by broadcast circuit 402 to any input associated with any other multiplier of MAC160.
[0075] Data 155 (e.g., SV175A) includes a plurality of elements (e.g., s0-s3 and s128-s131). MAC 160 receives data 155 (e.g., SV175A) from source VR 154, as described with reference to FIG. 1A. For example, for each element of the plurality of elements (e.g., s0-s3 and s128-s131), the element is provided to a respective separate input of inputs 409A-H of MAC 160. Illustratively, a first element (e.g., s0) is provided to input 409A associated with multiplier 410A, a second element (e.g., s1) is provided to input 409B associated with multiplier 410B, and so on.
[0076] The multipliers 410A to 410H generate output data 165 based on data 145 (e.g., SV173A) and data 155 (e.g., SV175A). For example, the multipliers 410A to 410H generate multiplied data 455A to 455H as the output data 165. For example, the multiplier 410A receives a first element (e.g., v0) of the data 145 (e.g., SV173A) via the input 407A and receives a first element (e.g., s0) of the data 155 (e.g., SV175A) via the input 409A. The multiplier 410A generates multiplied data 455A (e.g., m0, output SV) by multiplying the first element (e.g., v0) of the data 145 (e.g., SV173A) by the first element (e.g., s0) of the data 155 (e.g., SV175A). Similarly, multiplier 410B receives the second element (e.g., v1) of data 145 (e.g., SV173A) via input 407B and the second element (e.g., s1) of data 155 (e.g., SV175A) via input 409B. Multiplier 410B generates multiplied data 455B (e.g., m1, output SV) by multiplying the second element (e.g., v1) of data 145 (e.g., SV173A) by the second element (e.g., s1) of data 155 (e.g., SV175A).
[0077] MAC 160 provides output data 165 to accumulate VR 156. In certain aspects, output data 165 is added to accumulate data 167, and the result is stored in accumulate VR 156. For example, MAC 160 includes multiple adders, such as adder 450A, adder 450B, adder 450D, adder 450E, adder 450H, one or more additional adders, or combinations thereof, which are referred to herein as adders 450A-450H. Adders 450A-450H receive at least a portion of accumulate data 167 from accumulate VR 156, generate a sum by adding output data 165 to accumulate data 167, and overwrite a portion of accumulate data 167 with the sum in accumulate VR 156.
[0078] For example, adder 450A receives SV340A (e.g., a value represented by a0) from accumulation VR 156 and receives multiplication data 455A (e.g., a value represented by m0) from multiplier 410A. Adder 450A generates addition data 457A (e.g., a value represented by p0) by adding SV340A (e.g., a0) and multiplication data 455A (e.g., m0), and stores addition data 457A (e.g., p0) in accumulation VR 156 as an updated value of SV340A. Similarly, adder 450B receives SV340A (e.g., p0) from accumulation VR 156 and receives multiplication data 455B (e.g., a value represented by m1) from multiplier 410B. The adder 450B generates addition data 457B (e.g., a value represented by p1) by adding SV340A (e.g., p0) and multiplication data 455B (e.g., m1), and stores the addition data 457B (e.g., p1) in the accumulation VR 156 as an updated value of SV340A. At the end of the first sub-stage, SV340A has been updated from a first value (e.g., a0) to a second value (e.g., a value represented by b0), and SV340B has been updated from a first value (e.g., a value represented by a32) to a second value (e.g., a value represented by b32).
[0079] 5, there is shown a diagram 500 of an exemplary embodiment of a second state of Rotate VR 144A, Source VR 154, and Accumulate VR 156. In a particular embodiment, Rotate VR 144A, Source VR 154, and Accumulate VR 156 are in the second state after a first sub-phase and before a second sub-phase of a first processing phase of execution of instruction 280 of FIG.
[0080] In the second state, SV173A remains in dedicated portion 134A. SV175B of source VR 154 includes SV520A and SV520B. Accumulation VR 156 includes SV540A (e.g., a1) and SV540B (e.g., a33). SV540A (e.g., a1) is the next of SV340A (e.g., b0) updated in the first sub-stage of the first processing stage. SV540B (e.g., a33) is the next of SV340B (e.g., b32) updated in the first sub-stage of the first processing stage.
[0081] Referring to FIG. 6, a diagram 600 of an exemplary aspect of the operation of components of system 100 of FIG. 1A during a second sub-phase of a first processing phase of execution of instructions 280 is shown.
[0082] During the second sub-stage of the first processing stage, MAC 160 receives SV 175B as data 155 from source VR 154, as described with reference to FIG. 1A. SV 175B includes a plurality of elements (e.g., s4 through s7 and s132 through s135). For each element of the plurality of elements (e.g., s4 through s7 and s132 through s135), the element is provided to a respective separate input of inputs 409A-H of MAC 160. Illustratively, a first element (e.g., s4) is provided to input 409A associated with multiplier 410A, a second element (e.g., s5) is provided to input 409B associated with multiplier 410B, and so on.
[0083] The multipliers 410A to 410H generate output data 165 based on data 145 (e.g., SV173A) and data 155 (e.g., SV175B). For example, the multipliers 410A to 410H generate multiplied data 455A to 455H as the output data 165. For example, the multiplier 410A generates multiplied data 455A (e.g., m0, output SV) by multiplying a first element (e.g., v0) of the data 145 (e.g., SV173A) by a first element (e.g., s4) of the data 155 (e.g., SV175B). Similarly, multiplier 410B generates multiplied data 455B (e.g., m1, output SV) by multiplying the second element (e.g., v1) of data 145 (e.g., SV173A) by the second element (e.g., s5) of data 155 (e.g., SV175B).
[0084] The MAC 160 provides output data 165 to the accumulation VR 156. For example, the adder 450A receives SV540A (e.g., a value represented by a1) from the accumulation VR 156 and receives multiplication data 455A (e.g., a value represented by m0) from the multiplier 410A. The adder 450A generates addition data 457A (e.g., a value represented by p0) by adding SV540A (e.g., a1) and multiplication data 455A (e.g., m0), and stores the addition data 457A (e.g., p0) in the accumulation VR 156 as an updated value of SV540A. Similarly, the adder 450B receives SV540A (e.g., p0) from the accumulation VR 156 and receives multiplication data 455B (e.g., a value represented by m1) from the multiplier 410B. The adder 450B generates addition data 457B (e.g., a value represented by p1) by adding SV540A (e.g., p0) and multiplication data 455B (e.g., m1), and stores the addition data 457B (e.g., p1) in the accumulation VR 156 as an updated value of SV540A. At the end of the second sub-stage, SV540A has been updated from a first value (e.g., a1) to a second value (e.g., a value represented by b1), and SV540B has been updated from a first value (e.g., a33) to a second value (e.g., a value represented by b33).
[0085] 7, there is shown a diagram 700 of an exemplary embodiment of a third state of Rotate VR 144A, Source VR 154, and Accumulate VR 156. In a particular embodiment, Rotate VR 144A, Source VR 154, and Accumulate VR 156 are in the third state after a first processing stage and before a second processing stage of execution of instruction 280 of FIG.
[0086] In the third state, SV173A remains in dedicated portion 134A. SV175S of source VR 154 includes SV720A and SV720B. All elements of accumulation VR 156 have been updated based on SV173A and SV175A-175S. For example, accumulation VR 156 includes SV740A, which has been updated from a first value (e.g., a31) to a second value (e.g., b31), and SV740B, which has been updated from a first value (e.g., a63) to a second value (e.g., b63).
[0087] 8, a diagram 800 of an exemplary embodiment of data rotation in rotate VR 144A of FIG. 1 during execution of instruction 280 is shown. For example, rotate circuit 142 of FIG. 1A rotates the data in rotate VR 144A as described with reference to FIGS. 1A and 1B. In a particular embodiment, the rotation is performed during a rotation stage after the first processing stage. During the rotation, SV 173A is moved to the end of rotate VR 144A and the remainder of the SV is shifted such that SV 173B is stored in dedicated portion 134A.
[0088] In a particular example, instruction 280 has the following form: Vxx+=vmpyaddhbz_x4(Vuu,Zy4):sat:rot, and the rotation phase of execution of instruction 280 corresponds to the following pseudocode for a loop: fHIDE(size1u_t tmp;)fHIDE(int k;)fHIDE(int a;)fHIDE(int b;)for(k=0;k<128-2 * 8;k++){ a=k% 128; b=(k+128-2 *8)% 128; tmp=ZyV.V8u[a]; ZyV.V8u[a]=ZyV.V8u[b]; ZyV.V8u[b]=tmp; } Here, ZyV corresponds to rotation VR144A.
[0089] 9, there is shown a diagram 900 of an exemplary embodiment of a fourth state of Rotate VR 144A, Source VR 154, and Accumulate VR 156. In a particular embodiment, Rotate VR 144A, Source VR 154, and Accumulate VR 156 are in the fourth state after the Rotate stage of FIG. 8 and before the second processing stage of execution of instruction 280 of FIG. 2A.
[0090] In the second processing state, SV173B is stored in dedicated portion 134A of rotation VR 144A. In certain aspects, SV173B may include multiple SVs. For example, SV173B includes SV912A (e.g., v8-v11) and SV912B (e.g., v12-v15).
[0091] Source VR 154 includes SV175A, which includes SV320A and SV320B. Accumulation VR 156 includes SV340A (e.g., b0) and SV340B (e.g., b32). In a particular aspect, selection circuit 146 is configured to select SV173B stored in dedicated portion 134A as data 145.
[0092] Referring to FIG. 10, a diagram 1000 of an exemplary aspect of the operation of components of system 100 of FIG. 1A during a first sub-phase of a second processing phase of execution of instructions 280 of FIG. 2A is shown.
[0093] During the first sub-stage of the second processing stage, operations similar to those described with reference to the first sub-stage of the first processing stage of FIG. 4 are performed. For example, selection circuit 146 provides data 145 (e.g., SV173B) to broadcast circuit 402. Data 145 (e.g., SV173A) includes multiple elements (e.g., v8 through v15). Broadcast circuit 402 provides data 145 (e.g., SV173B) to multipliers 410A through 410H. For example, a first element (e.g., v8) of data 145 (e.g., SV173B) is provided to input 407A associated with multiplier 410A via a first output of broadcast circuit 402, a second element (e.g., v9) of data 145 (e.g., SV173B) is provided to input 407B associated with multiplier 410B via a second output of broadcast circuit 402, and so on.
[0094] Data 155 (e.g., SV175A) includes multiple elements (e.g., s0 through s3 and s128 through s131). MAC 160 receives data 155 (e.g., SV175A) from source VR 154, as described with reference to FIG. 1A. For example, a first element (e.g., s0) is provided to input 409A associated with multiplier 410A, a second element (e.g., s1) is provided to input 409B associated with multiplier 410B, and so on.
[0095] The multipliers 410A to 410H generate output data 165 based on data 145 (e.g., SV173B) and data 155 (e.g., SV175A). For example, the multiplier 410A generates multiplied data 455A by multiplying a first element (e.g., v8) of the data 145 (e.g., SV173B) by a first element (e.g., s0) of the data 155 (e.g., SV175A). Similarly, the multiplier 410B generates multiplied data 455B by multiplying a second element (e.g., v9) of the data 145 (e.g., SV173B) by a second element (e.g., s1) of the data 155 (e.g., SV175A).
[0096] The MAC 160 provides output data 165 to the accumulation VR 156. The adder 450A receives SV340A (e.g., a value represented by b0) from the accumulation VR 156 and receives multiplication data 455A (e.g., a value represented by m0) from the multiplier 410A. The adder 450A generates addition data 457A (e.g., a value represented by p0) by adding SV340A (e.g., b0) and multiplication data 455A (e.g., m0) and stores the addition data 457A (e.g., p0) in the accumulation VR 156 as an updated value of SV340A. Similarly, the adder 450B receives SV340A (e.g., p0) from the accumulation VR 156 and receives multiplication data 455B (e.g., a value represented by m1) from the multiplier 410B. At the end of the first sub-phase, SV340A has been updated from a first value (e.g., b0) to a second value (e.g., the value represented by c0), and SV340B has been updated from a first value (e.g., the value represented by b32) to a second value (e.g., the value represented by c32).
[0097] 11 , a diagram 1100 of an exemplary embodiment of a fifth state of Rotate VR 144A, Source VR 154, and Accumulate VR 156 is shown. In a particular embodiment, Rotate VR 144A, Source VR 154, and Accumulate VR 156 are in the fifth state after a first sub-phase and before a second sub-phase of a second processing phase of execution of instruction 280.
[0098] In the fifth state, SV173B remains in the dedicated portion 134A. SV175B of the source VR 154 includes SV520A and SV520B. The accumulation VR 156 includes SV540A (e.g., b1) and SV540B (e.g., b33). SV540A (e.g., b1) is the next of SV340A (e.g., c0) updated in the first sub-stage of the second processing stage. SV540B (e.g., b33) is the next of SV340B (e.g., c32) updated in the first sub-stage of the second processing stage. In the second sub-stage of the second processing stage, SV540A and SV540B are updated based on SV173B and SV175B. In certain embodiments, additional sub-stages of the second processing stage are performed until accumulated data 167 is updated based on SV 173B and all of SVs 175A-175S. In certain embodiments, additional rotation and processing stages are performed until all of SVs 173A-173T have been processed by MAC 160 to update accumulated data 167.
[0099] 12, a diagram 1200 of an exemplary aspect of the operation of selection circuit 146 is shown. Selection circuit 146 is coupled to rotation VRs 144A-144N. For example, selection circuit 146 is coupled to rotation VR 144A and rotation VR 144B. Selection circuit 146 coupled to two rotation VRs is provided as an illustrative example. In other examples, selection circuit 146 may be coupled to fewer or more than two rotation VRs.
[0100] In certain aspects, selection circuit 146 is configured to access data stored in dedicated portions of rotation VRs 144A-144N and not access data from the remaining portions of rotation VRs 144A-144N. For example, selection circuit 146 is configured to access data from dedicated portion 134A of rotation VR 144A. As another example, selection circuit 146 is configured to access data from dedicated portion 134B of rotation VR 144B.
[0101] As an illustrative example, prior to or during a first processing stage of executing instruction 280 of FIG. 2A , write data 1271 (e.g., performed by load / store unit 208 of FIG. 2A ) stores a first portion of vector 171 as data portion 1211 in rotated VR 144A and a second portion of vector 171 as data portion 1213 in rotated VR 144B. In a particular aspect, the first portion of vector 171 includes alternating SVs of vector 171, and the second portion of vector 171 includes the remaining SVs of vector 171. For example, data portion 1211 includes SVs 173A of vector 171 (e.g., v0 through v7), and data portion 1213 includes SVs 173B of vector 171 (e.g., v8 through v15). Similarly, data portion 1211 contains SV173C (e.g., v16-v23), data portion 1213 contains SV173D (e.g., v24-v31), and so on. Thus, data portion 1211 contains half of vector 171 (e.g., 32 elements, i.e., v0-v7, v16-v23, v32-v39, and v48-v55), and data portion 1213 contains the other half of vector 171 (e.g., the other 32 elements, i.e., v8-v15, v24-v31, v40-v47, and v56-v63). Thus, rotator VR 144A is shown as storing 32 elements of vector 171, and rotator VR 144B is shown as storing the other 32 elements of vector 171. 3, each of Rotate VR 144A and Rotate VR 144B is sized to store 64 elements, and therefore some storage capacity of each of Rotate VR 144A and Rotate VR 144B is unused. In some implementations, rather than using some of Rotate VR 144A and Rotate VR 144B, data for the next vector to be processed can also be loaded into Rotate VR 144A and Rotate VR 144B (e.g., using the Next Write Data 1271 operation). In other implementations, each of Rotate VR 144A and Rotate VR 144B has a storage capacity of 32 elements, and therefore no storage capacity is unused.
[0102] It should be understood that each of Rotate VR 144A and Rotate VR 144B is shown as storing vector data (e.g., 1 column by 32 rows of elements) for ease of explanation. In other examples, each of Rotate Vector 144A or Rotate VR 144B can store data in various logical and / or physical arrangements, such as an array representation (e.g., 8 columns by 4 rows of elements or 4 columns by 8 rows of elements).
[0103] In a particular embodiment, SV1273A corresponding to the first eight elements of data portion 1211 is stored in dedicated portion 134A of rotator VR144A, and SV1273B corresponding to the first eight elements of data portion 1213 is stored in dedicated portion 134B of rotator VR144B.
[0104] In an illustrative example, the first eight elements of data portion 1211 correspond to SV173A (e.g., v0 through v7), and the first eight elements of data portion 1213 correspond to SV173B (e.g., v8 through v15). In this example, SV1273A corresponds to SV173A, and SV1273B corresponds to SV173B. In an example where the rotated data portion is written to rotation VR 144A, SV1273A includes the first eight elements of the rotated data portion (e.g., v16 through v23), as will be further described with reference to FIG.
[0105] Selection circuit 146 includes components that allow various portions of data from dedicated portion 134A of rotator VR 144A, dedicated portion 134B of rotator VR 144B, or both, to be output by selection circuit 146 based on one or more control signals. In a particular aspect, selection circuit 146 includes a multiplexer 1230 coupled to rotator VR 144B. Multiplexer 1232 is coupled to rotator VR 144A and multiplexer 1230. Delay element 1234 is coupled to multiplexer 1232. Combiner 1236 is coupled to rotator VR 144A and multiplexer 1230. Multiplexer 1238 is coupled to rotator VR 144A, delay element 1234, combiner 1236, and multiplexer 1230.
[0106] In some implementations, write data 1271 provides data portion 1213 to rotation VR 144B while simultaneously providing SV 1273B to be stored in dedicated portion 134B of rotation VR 144B to multiplexer 1230. In particular aspects, multiplexer 1230 is configured to select SV 1273B retrieved from dedicated portion 134B of rotation VR 144B or SV 1273B received via write data 1271. For example, multiplexer 1230 selects SV 1273B from dedicated portion 134B or from write data 1271 based on one or more control signals, such as register output indicator 1224, rotation indicator 1226, one or more additional control signals, or a combination thereof.
[0107] In certain aspects, a first value (e.g., 1) of rotation indicator 1226 indicates that write data 1271 (e.g., performed by rotation circuit 142 of FIG. 1A) is writing rotated data to at least one of rotation VR 144A or rotation VR 144B. A second value (e.g., 0) of rotation indicator 1226 indicates that rotated data is not being written to either rotation VR 144A or rotation VR 144B. In certain aspects, a first value (e.g., 1) of register output indicator 1224 indicates that an SV should be selected from write data 1271. A second value (e.g., 0) of register output indicator 1224 indicates that an SV should be selected from dedicated portion 134B.
[0108] In particular implementations, multiplexer 1230 selects SV1273B from write data 1271 in response to rotation indicator 1226 having a first value, register output indicator 1224 having a first value, or both. Alternatively, multiplexer 1230 selects SV1273B from dedicated portion 134B in response to rotation indicator 1226 having a second value, register output indicator 1224 having a second value, or both. In some aspects, selecting SV1273B from write data 1271 allows multiplexer 1230 to generate an output without waiting for write data 1271 to complete writing data portion 1213 to rotation VR 144B.
[0109] Multiplexer 1230 provides SV1231A (e.g., the first four elements) of SV1273B (e.g., selected SV1273B) to multiplexer 1238 and provides SV1231B (e.g., the last four elements) of SV1273B to multiplexer 1232. In the illustrative example, SV1273B corresponds to SV173B, SV1231A corresponds to SV912A (e.g., v8 to v11), and SV1231B corresponds to SV912B (e.g., v12 to v15).
[0110] In certain aspects, multiplexer 1230 provides SV1216 (e.g., the first two elements) of SV1273B (e.g., selected SV1273B) to combiner 1236. In an illustrative example, SV1273B corresponds to SV173B, and SV1216 corresponds to the first two elements (e.g., v8-v9) of SV173B.
[0111] Multiplexer 1232 receives SV1212B (e.g., the last four elements) of SV1273A from rotator VR 144A and receives SV1231B (e.g., the last four elements) of SV1273B (e.g., the selected SV1273B) from multiplexer 1230. Multiplexer 1232 outputs SV1212B or SV1231B based on a control signal (e.g., register indicator 1228). For example, multiplexer 1232 selects SV1212B (e.g., the last four elements) of SV1273A as multiplexer output 1233 in response to register indicator 1228 having a first value (e.g., 0). Alternatively, multiplexer 1232 selects SV1231B (eg, the last four elements) of SV1273B as multiplexer output 1233 in response to register indicator 1228 having a second value (eg, 1).
[0112] Combiner 1236 is configured to receive SV1214 (e.g., the first two elements) of SV1273A and SV1216 (e.g., the first two elements) of SV1273B. Combiner 1236 generates SV1218 by combining SV1214 and SV1216. In the illustrative example, SV1273A corresponds to SV173A, SV1273B corresponds to SV173B, and SV1218 corresponds to the combination of the first two elements (e.g., v0-v1) of SV173A and the first two elements (e.g., v8-v9) of SV173B.
[0113] Multiplexer 1238 receives SV1212A (e.g., the first four elements of SV1273A), SV1218 (e.g., the first two elements of SV1273A and the first two elements of SV1273B), and SV1231A (e.g., the first four elements of SV1273B) from rotator VR 144A. Based on a control signal (e.g., patterned control 1242), multiplexer 1238 selects one of SV1212A, SV1218, and SV1231A to output as the first portion of multiplexer output 1239.
[0114] Delay element 1234 receives multiplexer output 1233 (e.g., the last four elements of 1273A or the last four elements of 1273B) and provides multiplexer output 1233 from delay element 1234 to multiplexer 1238 after outputting the first portion of multiplexer output 1239. Multiplexer 1238 outputs multiplexer output 1233 (e.g., the last four elements of 1273A or the last four elements of 1273B) as the second portion of multiplexer output 1239. Data 145 includes the first portion of multiplexer output 1239 (e.g., the first four elements of SV1273A, the first two elements of SV1273A and the first two elements of SV1273B, or the first four elements of SV1273B) and a second portion of multiplexer output 1239 (e.g., the last four elements of 1273A or the last four elements of 1273B).
[0115] In the illustrative example, SV1273A corresponds to SV173A, and multiplexer 1238 outputs the first four elements of SV173A (e.g., v0 through v3) as the first portion of multiplexer output 1239 and the last four elements of SV173A (e.g., v4 through v7) as the second portion of multiplexer output 1239. During the first processing stage, data 145 includes the first four elements of SV173A (e.g., v0 through v3) and the last four elements of SV173A (e.g., v4 through v7). After the first processing stage, rotate circuit 142 rotates the data in rotator VR 144A during a first rotation stage, as will be further described with reference to FIG. 13 .
[0116] In an example having two rotation VRs, during the second processing stage, selection circuit 146 outputs the SV from dedicated portion 134B of rotation VR 144B. For example, multiplexer 1238 outputs the first four elements of SV 173B (e.g., v8 through v11) as the first portion of multiplexer output 1239 and the last four elements of SV 173B (e.g., v12 through v15) as the second portion of multiplexer output 1239. During the second processing stage, data 145 includes the first four elements of SV 173B (e.g., v8 through v11) and the first four elements of SV 173B (e.g., v12 through v15). After the second processing stage, rotation circuit 142 rotates the data in rotation VR 144B during the second rotation stage, as further described with reference to FIG. 13.
[0117] During the third processing stage, selection circuit 146 outputs the SV from dedicated portion 134A of rotator VR 144A that contains the rotated data. For example, multiplexer 1238 outputs the first four elements of SV1273A (e.g., v16-v19) as the first portion of multiplexer output 1239 and the last four elements of SV1273A (e.g., v20-v23) as the second portion of multiplexer output 1239. During the third processing stage, data 145 includes the first four elements of SV1273A (e.g., v16-v19) and the last four elements of SV1273A (e.g., v20-v23). After the third processing stage, rotation circuit 142 rotates the data in rotator VR 144A during a third rotation stage, as further described with reference to FIG. 13. In some aspects, additional processing and rotation steps are performed until all elements of data portion 1211 and data portion 1213 are output by selection circuit 146 and processed by MAC 160 .
[0118] The particular pattern of outputting the first eight elements of rotated VR 144A followed by the first eight elements of rotated VR 144B during each processing stage is provided as an illustrative example. In other examples, the elements stored in dedicated portion 134A, the elements stored in dedicated portion 134B, or a combination thereof, may be output by selection circuit 146 in various combinations based on configuration 148. For example, configuration 148 specifies the values of control signals (e.g., register output indicator 1224, rotation indicator 1226, register indicator 1228, patterned control 1242, or a combination thereof).
[0119] In some aspects, selection circuitry 146 can output data 145 that includes one or more elements from dedicated portion 134A and one or more elements from dedicated portion 134B in the same processing stage based on configuration 148. For example, if patterned control 1242 of configuration 148 indicates that data 145 should include SV1218, then data 145 includes the leading element stored in dedicated portion 134A and the leading element stored in dedicated portion 134B. As another example, based on the register indicator 1228 and the patterned control section 1242 of the configuration 148, the selection circuit 146 can output the leading element of one of the dedicated portions 134A (e.g., SV1212A) or 134B (e.g., SV1231A) in a first time step of a processing stage, and output the trailing element of the other of the dedicated portions 134A (e.g., SV1212B) or 134B (e.g., SV1231B) in a second time step of the same processing stage.
[0120] In certain aspects, configuration 148 is based on default data, configuration settings, user input, etc. In certain aspects, configuration 148 is based on data indicating particular values of control signals that map instructions 280 (e.g., opcodes 282, parameters 284, or both) to configuration 148.
[0121] 13, a diagram 1300 of an exemplary aspect of operation of rotation circuit 142 is shown. Rotation circuit 142 is coupled to rotation VRs 144A-144N. For example, rotation circuit 142 is coupled to rotation VR 144A and rotation VR 144B. Rotation circuit 142 coupled to two rotation VRs is provided as an illustrative example. In other examples, rotation circuit 142 may be coupled to fewer or more than two rotation VRs.
[0122] Rotate circuit 142 is configured to rotate data in rotator VR 144A or rotator VR 144B based on one or more control signals (e.g., register indicator 1324 and register indicator 1326) and one or more rotate amounts (e.g., rotate amount 174A and rotate amount 174B). Rotate circuit 142 includes logic gate 1302 (e.g., a two-input AND gate) and logic gate 1304 (e.g., a two-input AND gate). An input of logic gate 1302 is coupled to the output of rotator VR 144A, and another input of logic gate 1302 is configured to receive register indicator 1324. An input of logic gate 1304 (e.g., a two-input AND gate) is coupled to the output of rotator VR 144B, and another input of logic gate 1304 is configured to receive register indicator 1326.
[0123] Logic gate 1302 and logic gate 1304 are coupled to logic gate 1306 (e.g., a two-input OR gate). For example, one input of logic gate 1306 is coupled to the output of logic gate 1302, and another input of logic gate 1306 is coupled to the output of logic gate 1304. Logic gate 1306 is coupled to rotators 172A-172M of FIG. 1A.
[0124] In some implementations, each of rotators 172A-172M is configured to rotate data by a predetermined amount. In these implementations, rotation circuit 142 may include one or more of rotators 172A-172M. For example, the output of logic gate 1306 is coupled to the input of rotator 172A, and the output of rotator 172A is coupled to the input of rotator 172B. In some examples, the output of rotator 172A is coupled to the input of rotator 172B via a pipeline register. In particular aspects, rotator 172A is configured to rotate data by rotation amount 174A and store the rotated data in the pipeline register, and rotator 172B is configured to selectively rotate data retrieved from the pipeline register by rotation amount 174B based on a control signal. For example, rotator 172B rotates data by rotation amount 174B when the control signal has a first value (e.g., 1). Rotator 172B refrains from rotating the data when the control signal has a second value (e.g., 0). The output of Rotator 172B is coupled to the input of Rotate VR 144A and the input of Rotate VR 144B. In some alternative implementations, Rotator circuit 142 includes a single programmable rotator (e.g., Rotator 172A). In these implementations, the output of Logic gate 1306 is coupled to the input of Rotator 172A, and the output of Rotator 172A is coupled to the input of Rotate VR 144A and the input of Rotate VR 144B. In some examples, Rotator 172A and Rotate VR 144A together function as a variable shift register.
[0125] Logic gate 1302 provides data portion 1211 received from rotate VR 144A to logic gate 1306 in response to register indicator 1324 having a first value (e.g., 1). Logic gate 1304 provides data portion 1213 received from rotate VR 144B to logic gate 1306 in response to register indicator 1326 having a first value (e.g., 1). In some aspects (e.g., after a processing stage), one of register indicator 1324 or register indicator 1326 has a first value (e.g., 1) and the other of register indicator 1324 or register indicator 1326 has a second value (e.g., 0). In these aspects, either data portion 1211 or data portion 1213 is passed to logic gate 1306 for rotation. In some aspects (e.g., during a processing stage), register indicator 1324 and register indicator 1326 both have a second value (e.g., 0), such that neither data portion 1211 nor data portion 1213 is passed to logic gate 1306 and no rotation is performed.
[0126] Logic gate 1306 passes data portion 1211 or data portion 1213 received from logic gate 1302 or logic gate 1304, respectively, to rotator 172A as output 1307. For example, after a first processing stage, register indicator 1324 has a first value (e.g., 1), and logic gate 1302 passes data portion 1211 to logic gate 1306, which provides data portion 1211 to rotator 172A as output 1307. Rotator 172A rotates output 1307 by rotation amount 174A (e.g., two halfwords) to generate rotated data 1377 and provides rotated data 1377 to rotator 172B. In some examples, rotator 172A provides rotated data 1377 to a pipeline register, and rotator 172B retrieves rotated data 1377 from the pipeline register. Rotator 172B generates rotated data 177 by rotating rotated data 1377 based on rotate amount 174B (e.g., six halfwords). Rotate circuit 142 stores rotated data 177 in rotate VR 144A. For example, rotate circuit 142 stores rotated data 177 in rotate VR 144A in response to register indicator 1324 having a first value (e.g., 1). Alternatively, rotate circuit 142 stores rotated data 177 in rotate VR 144B in response to register indicator 1326 having a first value (e.g., 1).
[0127] A rotation mechanism including a pipeline register is provided as an illustrative example. In other implementations, Rotator 172A and Rotate VR 144A (or Rotate VR 144B) together function as a first shift register, and Rotator 172B and Rotate VR 144A (or Rotate VR 144B) together function as a second shift register. In one example, Rotator 172A rotates the bit values stored in Rotate VR 144A independently of (e.g., without) any additional registers, whereby rotated data 1377 is stored in Rotate VR 144A after the rotation. In another example, Rotator 172B rotates the bit values corresponding to rotated data 1377 stored in Rotate VR 144A independently of (e.g., without) any additional registers, whereby rotated data 177 is stored in Rotate VR 144A after the rotation.
[0128] In some examples, the rotate circuit 142 rotates data within a single rotation VR among the rotation VRs 144A-144N during a rotation phase. In some examples, the rotate circuit 142 rotates data within multiple rotation VRs among the rotation VRs 144A-144N during a rotation phase. For example, one of the register indicator 1324 or the register indicator 1326 has a first value (e.g., 1) during a first pass, and the other of the register indicator 1324 or the register indicator 1326 has a first value (e.g., 1) during a second pass.
[0129] In some aspects, the rotate circuit 142 is configurable. For example, one or more control signals, one or more rotate amounts, or a combination thereof, are based on a configuration 1348 of the rotate circuit 142. The configuration 1348 may be based on default data, user input, configuration settings, etc. In particular aspects, the configuration 1348 is based on data indicating particular values for one or more control signals, one or more rotate amounts, or both, that maps the instruction 280 (e.g., the opcode 282, the parameter 284, or both) to the configuration 1348. In some aspects, at least one of the rotate amounts is selectable. In one example, the rotate amount 174A is fixed (e.g., two halfwords), and the rotate amount 174B is selectable. Illustratively, the rotate amount 174B has a value (e.g., a selection of 0 halfwords or 6 halfwords) based on the configuration 1348.
[0130] In some aspects, the rotation circuit 142 enables rotating data within a single rotation VR, as described with reference to Figures 3-11. For example, the register indicator 1324 has a first value (e.g., 1) after each processing stage of the instruction. In some aspects, the rotation circuit 142 enables rotating data within alternating rotation VRs after processing stages, as described with reference to Figure 12. For example, the register indicator 1324 has a first value (e.g., 1) after the first processing stage and all odd processing stages after the first processing stage, and the register indicator 1326 has a first value (e.g., 1) after the second processing stage and all even processing stages after the second processing stage.
[0131] In a particular example, if configuration 148 of selection circuit 146 indicates that dedicated portion 134A of rotate VR 144A should be selected for a particular processing stage, then configuration 1348 indicates that the data in rotate VR 144A (e.g., not the data in any other rotate VR) should be rotated during a rotation stage following the particular processing stage. For example, register indicator 1324 has a first value (e.g., 1), and rotate amount 174B is selected to indicate six halfwords, resulting in a total rotation of eight halfwords, including the two halfwords indicated by rotate amount 174A.
[0132] In some examples, rotate circuit 142 allows for rotating data in multiple rotate VRs during a single rotate stage. For example, if configuration 148 of select circuit 146 indicates that the two leading elements of dedicated portion 134A of rotate VR 144A and the two leading elements of dedicated portion 134B of rotate VR 144B should be selected during a particular processing stage (e.g., as described with reference to SV1218 of FIG. 12), configuration 1348 indicates that the data in rotate VR 144A and the data in rotate VR 144B should be rotated during a rotate stage following the particular processing stage. For example, rotate amount 174B is selected to indicate zero halfwords, resulting in a total rotation of two halfwords, including the two halfwords indicated by rotate amount 174A. During the first pass of the rotate stage, register indicator 1324 has a first value (e.g., 1), and the data in rotate VR 144A is rotated by two halfwords. During the second pass of the rotate stage, register indicator 1326 has a first value (eg, 1) and the data in rotate VR 144B is rotated by two halfwords.
[0133] 14 shows one implementation 1400 of device 102 as an integrated circuit 1402 that includes one or more processors 190. One or more processors 190 include rotated VR file 140. In some aspects, one or more processors 190 also include MAC 160, VR file 150, or both. Integrated circuit 1402 also includes signal input 1404, such as one or more bus interfaces, one or more latches, one or more flip-flops, one or more buffers, other data buffering circuitry, or a combination thereof, that enables it to receive data 1428 for processing. In one example, data 1428 includes vector 171, vector 151, or both of FIG. 1A . Integrated circuit 1402 also includes signal output 1406, such as a bus interface, that enables it to transmit output signals, such as accumulated data 167. The integrated circuit 1402 enables implementations that rotate vector inputs as a component in a system such as the mobile phone or tablet shown in FIG. 15, the headset shown in FIG. 16, the wearable electronic device shown in FIG. 17, the voice-controlled speaker system shown in FIG. 18, the camera shown in FIG. 19, the virtual reality headset, mixed reality headset, or augmented reality headset shown in FIG. 20, or the vehicle shown in FIG. 21 or FIG. 22.
[0134] 15 illustrates an implementation 1500 in which the device 102 includes a mobile device 1502, such as a phone or tablet, as an illustrative, non-limiting example. The mobile device 1502 includes a display screen 1504. Components of one or more processors 190, including the rotation VR file 140, are integrated within the mobile device 1502 and are shown using dashed lines to indicate internal components that are not normally visible to a user of the mobile device 1502. In some aspects, the one or more processors 190 also include the MAC 160, the VR file 150, or both. In a particular example, the rotation VR file 140 operates to rotate data within the rotation VR during vector operations (e.g., matrix multiplication), and the results of the vector operations are then processed to perform one or more operations on the mobile device 1502, such as launching a graphical user interface or otherwise displaying other information associated with the vector operations on the display screen 1504 (e.g., via an integrated “smart assistant” application).
[0135] In some implementations, device 102 includes one or more other sensors or components that generate data that can be operated on by vector operations using instructions 280, such as, by way of illustrative and non-limiting examples, wireless network signal data, global positioning data or other location data, video or image data from one or more cameras, inertial measurements or other motion data from an inertial measurement unit (e.g., one or more gyroscopes, compasses, accelerometers, etc.), or health data such as heart rate data, oxygen level data, respiration data, etc. from one or more corresponding sensors. The vector operations generate output data that can be output or processed to generate processed data, either or both of which can be displayed via display screen 1504, output via a loudspeaker, transmitted over a wireless network to another device such as a wearable electronic device (e.g., a smartwatch or headset), or output via a haptic output signal, by way of illustrative and non-limiting examples.
[0136] 16 illustrates an implementation 1600 in which the device 102 includes a headset device 1602. Components of one or more processors 190, including the rotated VR file 140, are integrated into the headset device 1602. In some aspects, the one or more processors 190 also include the MAC 160, the VR file 150, or both. In particular examples, the rotated VR file 140 operates to rotate data in the rotated VR during vector operations (e.g., matrix multiplication), which can cause the headset device 1602 to perform one or more operations at the headset device 1602, transmit the data to a second device (not shown) for further processing, or a combination thereof.
[0137] 17 illustrates an implementation 1700 in which the device 102 includes a wearable electronic device 1702 depicted as a “smart watch.” A rotational VR file 140 is integrated into the wearable electronic device 1702. In some aspects, the MAC 160, the VR file 150, or both are also integrated into the wearable electronic device 1702. In a particular example, the rotational VR file 140 operates to rotate data in the rotational VR during the performance of a vector operation (e.g., matrix multiplication), thereby causing the wearable electronic device 1702 to perform one or more operations at the wearable electronic device 1702, such as launching a graphical user interface or otherwise displaying other information associated with the vector operation on a display screen 1704 of the wearable electronic device 1702. By way of example, the wearable electronic device 1702 may include a display screen configured to display notifications based on the vector operation performed by the wearable electronic device 1702. In particular examples, the wearable electronic device 1702 includes a haptic device that provides a tactile notification (e.g., vibrates) in response to the performance of an operation that uses vector operations, such as a neural network-based speech interface. For example, the tactile notification may direct the user's attention to the wearable electronic device 1702 to view a notification display indicating the performance of the vector operation (e.g., an action performed in response to the user's identification of a verbal query). Thus, the wearable electronic device 1702 may alert a user who is hearing impaired or who is wearing a headset about an operation that uses vector operations.
[0138] FIG. 18 illustrates an implementation 1800 in which device 102 includes a wireless speaker and a voice-activated device 1802. The wireless speaker and voice-activated device 1802 can have a wireless network connection and is configured to perform Assistant operations. One or more processors 190 including rotation VR file 140 are included in the wireless speaker and voice-activated device 1802. In some aspects, the one or more processors 190 also include MAC 160, VR file 150, or both. The wireless speaker and voice-activated device 1802 also includes a speaker 1804. In operation, the wireless speaker and voice-activated device 1802 can perform Assistant operations, for example, through execution of a voice-activated system (e.g., an integrated Assistant application). Assistant operations can include adjusting the temperature, playing music, turning on lights, etc. For example, an Assistant operation is performed in response to receiving a command after a keyword or key phrase (e.g., "Hello, Assistant"). In particular examples, the rotation VR file 140 operates to rotate data within the rotation VR during vector operations (e.g., matrix multiplication), which may cause one or more operations to be performed or may be part of a calculation associated with one or more operations in the wireless speaker and voice-activated device 1802.
[0139] 19 illustrates an implementation 1900 in which device 102 includes a portable electronic device corresponding to a camera device 1902. Rotation VR file 140 is included in camera device 1902. In some aspects, MAC 160, VR file 150, or both are also included in camera device 1902. In operation, camera device 1902 can perform actions in response to spoken user commands, such as adjusting image or video capture settings, image or video playback settings, or image or video capture instructions, as illustrative examples. In a particular example, rotation VR file 140 operates to rotate data in the rotation VR during vector operations (e.g., matrix multiplication associated with image filtering), which may cause one or more operations to be performed or may be part of computations associated with one or more operations in camera device 1902.
[0140] FIG. 20 illustrates an implementation 2000 in which the device 102 includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset 2002. The rotational VR file 140 is integrated into the headset 2002. In some aspects, the MAC 160, the VR file 150, or both are also integrated into the headset 2002. In a particular example, the rotational VR file 140 operates to rotate data in the rotational VR file during the execution of vector operations (e.g., matrix multiplications associated with image processing), which may cause one or more operations in the headset 2002 to be performed or may be part of a calculation associated with one or more operations. The visual interface device is positioned in front of the user's eyes to enable augmented reality, mixed reality, or virtual reality images or scenes to be displayed to the user while the headset 2002 is worn. In a particular example, the visual interface device is configured to display a notification indicating the execution of an operation based on the vector operations (e.g., an image processing operation). Illustratively, the visual interface device is configured to display one or more images generated by the image processing operation.
[0141] 21 illustrates an implementation 2100 in which the device 102 corresponds to or is integrated within a vehicle 2102, depicted as a manned or unmanned aerial device (e.g., a delivery drone). A rotational VR file 140 is integrated into the vehicle 2102. In some aspects, the MAC 160, the VR file 150, or both are also integrated into the vehicle 2102. In particular examples, the rotational VR file 140 operates to rotate data within the rotational VR during vector operations (e.g., matrix multiplication associated with image processing operations), which may cause one or more operations to be performed or may be part of a calculation associated with one or more operations in the vehicle 2102. For example, an image generated by the image processing operation may be displayed by the vehicle 2102 to provide assembly instructions to a package recipient.
[0142] 22 illustrates another implementation 2200 in which the device 102 corresponds to or is integrated within a vehicle 2202, shown as an automobile. The vehicle 2202 includes one or more processors 190 that include a rotational VR file 140. In some aspects, the one or more processors 190 also include a MAC 160, a VR file 150, or both. In a particular example, the rotational VR file 140 operates to rotate data in the rotational VR during vector operations (e.g., matrix multiplication), which can cause the vehicle 2202 to perform one or more operations on the vehicle 2202.
[0143] In certain aspects, the voice-activated system initiates one or more actions of the vehicle 2202 based on one or more keywords (e.g., "unlock," "start engine," "play music," "show weather," or another voice command) detected in the microphone output signal, such as by providing feedback or information via the display 2220 or one or more speakers. In some examples, the one or more keywords are detected in the output signal by performing operations, such as neural network calculations in an artificial intelligence (AI)-based automatic speech recognition system, based on vector operations using the rotated VR file 140.
[0144] 23, a particular implementation of a method 2300 for rotating a vector input is shown. In particular aspects, one or more operations of the method 2300 (e.g., a processor-implemented method) are performed by at least one of the rotators 172A-172M, the rotation circuit 142, the rotation VRs 144A-144N, the selection circuit 146, the rotation VR file 140, the MAC 160, the VR file 150, the source VR 154, the accumulation VR 156, one or more processors 190, the device 102, the system 100 of FIG. 1A, or a combination thereof.
[0145] The method 2300 includes, at 2302, rotating data in a rotating vector register of the rotating vector register file using a rotating vector register file. For example, the rotating VR file 140 of FIG. 1A rotates data in the rotating VR 144A of the rotating VR file 140, as described with reference to FIGS. 1A, 1B, and 13.
[0146] The method 2300 also includes receiving first input data from the rotating vector register file at a multiply-accumulate circuit (MAC) at 2304. For example, the MAC 160 of FIG. 1A receives data 145 from the rotating VR file 140, as described with reference to FIGS. 1A and 12.
[0147] The method 2300 further includes receiving second input data at the MAC from a source vector register of a second vector register file, at 2306. For example, the MAC 160 of FIG. 1A receives data 155 from the source VR 154 of the VR file 150, as described with reference to FIG. 1A.
[0148] In some examples, method 2300 includes generating output based on the first input data and the second input data using a MAC. For example, MAC 160 of FIG. 1A generates output data 165 based on data 145 and data 155, as described with reference to FIG. 1A. Method 2300 also includes storing the output in an accumulate vector register of a second vector register file. For example, MAC 160 of FIG. 1A stores output data 165 in accumulate VR 156 of VR file 150, as described with reference to FIG. 1A.
[0149] Method 2300 allows for vector processing with reduced complexity in selection circuit 146. For example, selection circuit 146 can read rotated data into dedicated portion 134A of rotation VR 144A without needing to include circuitry to support reading from the remainder of rotation VR 144A.
[0150] The method 2300 of Figure 23 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a graphics processing unit (GPU), a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 2300 of Figure 23 may be performed by a processor executing instructions, such as those described with reference to Figure 24.
[0151] 24, a block diagram of a particular example implementation of a device is shown, generally designated 2400. In various implementations, device 2400 may have more or fewer components than those shown in FIG. 24. In an example implementation, device 2400 may correspond to device 102. In an example implementation, device 2400 may perform one or more operations described with reference to FIGS. 1A-23.
[0152] In particular implementations, device 2400 includes a processor 2406 (e.g., a CPU). Device 2400 may include one or more additional processors 2410 (e.g., one or more DSPs, one or more GPUs, or a combination thereof). In particular aspects, one or more processors 190 of FIG. 1A correspond to processor 2406, processor 2410, or a combination thereof. Processor 2410 may include a speech and music coder-decoder (CODEC) 2408, including a voice coder (“vocoder”) encoder 2436, a vocoder decoder 2438, a rotated VR file 140, or a combination thereof.
[0153] Device 2400 may include memory 2486 and a CODEC 2434. Memory 2486 may include instructions 2456 executable by one or more additional processors 2410 (or processor 2406) to implement functionality described with reference to rotating VR file 140. Device 2400 may include a modem 2448 coupled to an antenna 2452 via a transceiver 2450.
[0154] The device 2400 may include a display 2428 coupled to a display controller 2426. One or more speakers 2492 and one or more microphones 2490 may be coupled to a CODEC 2434. The CODEC 2434 may include a digital-to-analog converter (DAC) 2402, an analog-to-digital converter (ADC) 2404, or both. In particular implementations, the CODEC 2434 may receive analog signals from the one or more microphones 2490, convert the analog signals to digital signals using the analog-to-digital converter 2404, and provide the digital signals to a speech and music codec 2408. The speech and music codec 2408 may process the digital signals. In particular implementations, the speech and music codec 2408 may provide the digital signals to the CODEC 2434. The CODEC 2434 may convert the digital signal to an analog signal using a digital-to-analog converter 2402 and may provide the analog signal to one or more speakers 2492 .
[0155] In certain implementations, the device 2400 may be included in a system-in-package or system-on-chip device 2422. In certain implementations, the memory 2486, the processor 2406, the processors 2410, the display controller 2426, the CODEC 2434, and the modem 2448 are included in the system-in-package or system-on-chip device 2422. In certain implementations, the input device 2430 and the power supply 2444 are coupled to the system-in-package or system-on-chip device 2422. Additionally, in certain implementations, the display 2428, the input device 2430, the one or more speakers 2492, the one or more microphones 2490, the antenna 2452, and the power supply 2444 are external to the system-in-package or system-on-chip device 2422, as shown in FIG. In particular implementations, each of the display 2428, input device 2430, one or more speakers 2492, one or more microphones 2490, antenna 2452, and power source 2444 may be coupled to a component of the system-in-package or system-on-chip device 2422, such as an interface or controller.
[0156] The device 2400 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, an extended reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an internet-of-things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.
[0157] In accordance with the described implementation, the apparatus includes means for rotating data in a rotated vector register of a rotated vector register file. For example, the means for rotating can correspond to one or more of rotators 172A-172M, rotation circuit 142, rotated VR file 140, one or more processors 190, device 102, system 100 of FIG. 1A , processor 2406, one or more processors 2410, device 2400, one or more other circuits or components configured to rotate data in a rotated vector register of a rotated vector register file, or any combination thereof.
[0158] The apparatus also includes means for receiving first input data from the rotating vector register file at the multiply-accumulate circuit (MAC). For example, the means for receiving first input data may correspond to one or more of inputs 407A-H of MAC 160, one or more processors 190, device 102, system 100 of FIG. 1A, processor 2406, one or more processors 2410, device 2400, one or more other circuits or components configured to receive input data from the rotating vector register file at the MAC, or any combination thereof.
[0159] The apparatus further includes means for receiving second input data at the MAC from a source vector register of a second vector register file. For example, the means for receiving second input data may correspond to one or more of inputs 409A-H of MAC 160, one or more processors 190, device 102, system 100 of FIG. 1A, processor 2406, one or more processors 2410, device 2400, one or more other circuits or components configured to receive input data from a source vector register at the MAC, or any combination thereof.
[0160] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 2486) includes instructions (e.g., instructions 2456) that, when executed by one or more processors (e.g., one or more processors 2410 or processor 2406), cause the one or more processors to rotate data in a rotated vector register (e.g., rotate VR 144A) of a rotated vector register file (e.g., rotate VR file 140) using a rotated vector register file. The instructions, when executed by one or more processors, also cause the one or more processors to receive first input data (e.g., data 145) from the rotated vector register file at a multiply-accumulate circuit (MAC) (e.g., MAC 160). The instructions, when executed by one or more processors, also cause the one or more processors to receive second input data (e.g., data 155) from a source vector register (e.g., source VR 154) of a second vector register file (e.g., VR file 150) at the MAC.
[0161] Certain aspects of the present disclosure are described below in a set of interrelated examples.
[0162] According to Example 1, a device includes a processor, the processor including: a rotating vector register file including a rotating vector register, the rotating vector register file configured to rotate data in the rotating vector register; a second vector register file including a source vector register; and a multiply-accumulate circuit (MAC) configured to receive first input data from the rotating vector register file and second input data from the source vector register.
[0163] Example 2 includes the device of example 1, further including: a broadcast circuit configured to: configure the first input data to include sub-vector values of a rotating vector register, the sub-vector values having a plurality of elements; and configure the processor to, for each element of the plurality of elements, provide the element to a respective separate input of a plurality of inputs of the MAC.
[0164] Example 3 includes the device of example 1 or example 2, wherein the second vector register file includes an accumulation vector register, and the MAC is configured to generate an output based on the first input data and the second input data, and store the output in the accumulation vector register.
[0165] Example 4 includes a device of any of Examples 1 to 3, and is configured such that the first input data includes a first input subvector value, the second input data includes a second input subvector value, the second vector register file includes an accumulation vector register, and the MAC generates a first output subvector value based on the first input subvector value and the second input subvector value, and stores the first output subvector value in the accumulation vector register.
[0166] Example 5 includes the device of example 4, wherein the MAC is configured to receive a third input subvector value from the source vector register, generate a second output subvector value based on the first input subvector value and the third input subvector value, and store the second output subvector value in the accumulate vector register.
[0167] A sixth example includes the device according to any one of the first to fifth examples, wherein the rotating vector register file is configured to rotate data by a certain rotation amount.
[0168] Example 7 includes the device of example 6, where the amount of rotation is selectable.
[0169] Example 8 includes the device of example 6 or example 7, wherein the amount of rotation is based on an opcode of the multiply-accumulate instruction or a parameter of the multiply-accumulate instruction.
[0170] Example 9 includes the device of any of Examples 6 to 8, wherein the rotated vector register file further includes a rotation circuit, the rotation circuit including: a first rotator coupled to an output of the rotated vector register, the first rotator configured to perform a first data rotation corresponding to the first rotation amount to generate first rotated data; and a second rotator coupled to the output of the first rotator, the second rotator configured to perform a second data rotation corresponding to the second rotation amount to generate second rotated data, wherein an output of the second rotator is coupled to an input of the rotated vector register to update data in the rotated vector register with the second rotated data.
[0171] Example 10 includes the device of example 9, wherein the second amount of rotation is selectable.
[0172] Example 11 includes the device of example 9 or example 10, wherein the MAC is configured to receive the second rotated data from the rotated vector register as the first input data.
[0173] Example 12 includes the device of any of Examples 1 to 11, wherein the rotating vector register file further includes a second rotating vector register and a configurable selection circuit, and the configurable selection circuit is configured to output a first subvector value from the rotating vector register as first input data in a first configuration, and to output a second subvector value from the rotating vector register and a third subvector value from the second rotating vector register as first input data in a second configuration.
[0174] Example 13 includes the device of any of Examples 1 to 12, wherein the processor is configured to execute a multiply-accumulate instruction to retrieve first input data from the rotating vector register file, retrieve second input data from the source vector register, and, in the MAC, process the first input data and the second input data and rotate the data in the rotating vector register.
[0175] Example 14 includes the device of example 13, wherein the data is rotated in the rotating vector register after the first input data is retrieved from the rotating vector register file.
[0176] Example 15 includes the device of any of Examples 1 to 14, wherein the processor is integrated into at least one of a mobile device, a headset device, a wearable electronic device, a wireless speaker and voice-activated device, a camera device, an extended reality headset, or a vehicle.
[0177] According to Example 16, a processor-implemented method includes using a rotating vector register file to rotate data in a rotating vector register of the rotating vector register file, receiving first input data from the rotating vector register file at a multiply-accumulate circuit (MAC), and receiving second input data from a source vector register of a second vector register file at the MAC.
[0178] Example 17 includes the processor-implemented method of example 16, wherein the first input data includes sub-vector values of a rotating vector register, the sub-vector values having a plurality of elements, and the method further includes, for each element of the plurality of elements, providing the element to a respective separate input of a plurality of inputs of the MAC using a broadcast circuit.
[0179] Example 18 includes the processor implementation method of example 16 or example 17, further including: generating an output based on the first input data and the second input data using the MAC; and storing the output in an accumulation vector register of a second vector register file.
[0180] Example 19 includes the processor implementation method of any of Examples 16 to 18, and further includes generating a first output subvector value in the MAC based on a first input subvector value and a second input subvector value, where the first input data includes the first input subvector value and the second input data includes the second input subvector value, and storing the first output subvector value in an accumulation vector register of a second vector register file.
[0181] Example 20 includes the processor-implemented method of Example 19, further including receiving a third input subvector value from the source vector register at the MAC, generating a second output subvector value at the MAC based on the first input subvector value and the third input subvector value, and storing the second output subvector value in an accumulation vector register.
[0182] Example 21 includes the processor implementation method of any of Examples 16 to 20, wherein a rotating vector register file is used to rotate data by a certain rotation amount.
[0183] Example 22 includes the processor implementation method of example 21, wherein the amount of rotation is selectable.
[0184] Example 23 includes the processor implementation method of example 21 or example 22, wherein the amount of rotation is based on an opcode of the multiply-accumulate instruction or a parameter of the multiply-accumulate instruction.
[0185] Example 24 includes the processor implementation method of any of Examples 21 to 23, and further includes: performing a first data rotation corresponding to the first rotation amount using a first rotator of a rotation circuit of the rotated vector register file to generate first rotated data; performing a second data rotation corresponding to the second rotation amount using a second rotator of the rotation circuit to generate second rotated data; and updating the data in the rotated vector register with the second rotated data.
[0186] Example 25 includes the processor implementation method of example 24, wherein the second rotation amount is selectable.
[0187] Example 26 includes the processor-implemented method of example 24 or example 25, wherein the second rotated data is received from the rotated vector register as the first input data at the MAC.
[0188] Example 27 includes the processor implementation method of any of Examples 16 to 26, and further includes, in the first configuration, outputting a first subvector value from the rotating vector register as first input data, and in the second configuration, outputting a second subvector value from the rotating vector register and a third subvector value from the second rotating vector register as first input data, and the rotating vector register file includes the second rotating vector register.
[0189] Example 28 includes the processor implementation method of any of Examples 16 to 27, and further includes executing a multiply-accumulate instruction including: retrieving first input data from a rotating vector register file; retrieving second input data from a source vector register; and, in the MAC, processing the first input data and the second input data; and rotating the data in the rotating vector register.
[0190] Example 29 includes the processor implemented method of example 28, wherein the data is rotated in the rotating vector register after the first input data is retrieved from the rotating vector register file.
[0191] Example 30 includes the processor implementation method of any of Examples 16 to 29, wherein the rotating vector register file, the MAC, and the second vector register file are integrated into at least one of a mobile device, a headset device, a wearable electronic device, a wireless speaker and voice-activated device, a camera device, an extended reality headset, or a vehicle.
[0192] According to Example 31, a device includes a memory configured to store instructions and a processor configured to execute instructions that implement the processor-implemented method of any of Examples 16 to 30.
[0193] According to Example 32, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform any of the processor-implemented methods of Examples 16 to 30.
[0194] According to a thirty-third embodiment, an apparatus includes means for executing the processor-implemented method according to any one of the sixteenth to thirtieth embodiments.
[0195] According to Example 34, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to use a rotating vector register file to rotate data in a rotating vector register of the rotating vector register file, receive first input data from the rotating vector register file at a multiply-accumulate circuit (MAC), and receive second input data from a source vector register of a second vector register file at the MAC.
[0196] Example 35 includes the non-transitory computer-readable medium of example 34, wherein the first input data includes sub-vector values of a rotating vector register, the sub-vector values having a plurality of elements, and the instructions, when executed by the processor, cause the processor, for each element of the plurality of elements, to provide the element to a respective separate input of a plurality of inputs of the MAC using a broadcast circuit.
[0197] According to Example 36, the apparatus includes means for rotating data in a rotating vector register of a rotating vector register file, means for receiving first input data from the rotating vector register file in a multiply-accumulate circuit (MAC), and means for receiving second input data from a source vector register of a second vector register file in the MAC.
[0198] Example 37 includes the apparatus of Example 36, wherein the rotating means, the means for receiving first input data, and the means for receiving second input data are integrated into at least one of a smart speaker, a speaker bar, a computer, a tablet, a display device, a television, a game console, a music player, a radio, a digital video player, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an extended reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, or a mobile device.
[0199] Those skilled in the art will further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0200] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.
[0201] The previous description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosed aspects. Various modifications of these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims.
Claims
1. A device comprising a processor, wherein the processor a rotation vector register file including a rotation vector register, configured to rotate data within the rotation vector register, a rotation vector register file; a second vector register file including a source vector register; a multiply-accumulate circuit (MAC) configured to receive first input data from the rotation vector register file and second input data from the source vector register; A device including.
2. The device according to claim 1, wherein the first input data includes a sub-vector value of the rotation vector register, the sub-vector value has a plurality of elements, and the processor further includes a broadcast circuit configured to provide the elements to respective separate inputs of a plurality of inputs of the MAC for each of the plurality of elements.
3. The second vector register file includes an accumulation vector register, and the MAC generates an output based on the first input data and the second input data, stores the output in the accumulation vector register, The device according to claim 1, configured as such.
4. The first input data includes a first input sub-vector value, the second input data includes a second input sub-vector value, the second vector register file includes an accumulation vector register, and the MAC generates a first output sub-vector value based on the first input sub-vector value and the second input sub-vector value, stores the first output sub-vector value in the accumulation vector register, The device according to claim 1, configured as such.
5. The MAC receives a third input sub-vector value from the source vector register, generates a second output sub-vector value based on the first input sub-vector value and the third input sub-vector value, stores the second output sub-vector value in the accumulation vector register, The device according to claim 4, configured as such.
6. The device according to claim 1, wherein the rotation vector register file is configured to rotate the data by a certain amount of rotation.
7. The device according to claim 6, wherein the amount of rotation is selectable.
8. The device according to claim 6, wherein the rotation amount is based on an opcode of a multiply-accumulate instruction or a parameter of the multiply-accumulate instruction. **Claim 9** The rotation vector register file further includes a rotation circuit, and the rotation circuit a first rotator coupled to an output of the rotation vector register, configured to perform a first data rotation corresponding to a first rotation amount to generate first rotation data; a first rotator a second rotator coupled to an output of the first rotator, configured to perform a second data rotation corresponding to a second rotation amount to generate second rotation data; and a second rotator The device according to claim 6, wherein an output of the second rotator is coupled to an input of the rotation vector register to update the data in the rotation vector register with the second rotation data. **Claim 10** The device according to claim 9, wherein the second rotation amount is selectable. **Claim 11** The device according to claim 9, wherein the MAC is configured to receive the second rotation data from the rotation vector register as the first input data. **Claim 12** The rotation vector register file a second rotation vector register; a configurable selection circuit; The rotation vector register file further includes: In a first configuration, output a first sub-vector value from the rotation vector register as the first input data; In a second configuration, output a second sub-vector value from the rotation vector register and a third sub-vector value from the second rotation vector register as the first input data; The device according to claim 1, configured as described above. **Claim 13** The processor retrieves the first input data from the rotation vector register file; retrieves the second input data from the source vector register; processes the first input data and the second input data in the MAC; rotates the data in the rotation vector register; The device according to claim 1, configured to execute a multiply-accumulate instruction. **Claim 14** The device according to claim 13, wherein the data is rotated in the rotation vector register after the first input data is retrieved from the rotation vector register file. **Claim 15** The device according to claim 1, wherein the processor is integrated into at least one of a mobile device, a headset device, a wearable electronic device, a wireless speaker and voice-activated device, a camera device, an extended reality headset, or a vehicle.
16. A processor implementation method, comprising: using a rotation vector register file to rotate data in a rotation vector register of the rotation vector register file; receiving first input data from the rotation vector register file in a multiply-accumulate circuit (MAC); receiving second input data from a source vector register of a second vector register file in the MAC; A processor implementation method, including the above.
17. The processor implementation method according to claim 16, wherein the first input data includes a sub-vector value of the rotation vector register, the sub-vector value has a plurality of elements, and the method further includes, for each element of the plurality of elements, using a broadcast circuit to provide the element to a separate input of each of the plurality of inputs of the MAC.
18. using the MAC to generate an output based on the first input data and the second input data; storing the output in an accumulation vector register of the second vector register file; The processor implementation method according to claim 16, further including the above.
19. generating a first output sub-vector value in the MAC based on a first input sub-vector value and a second input sub-vector value, wherein the first input data includes the first input sub-vector value and the second input data includes the second input sub-vector value; storing the first output sub-vector value in an accumulation vector register of the second vector register file; The processor implementation method according to claim 16, further including the above.
20. receiving a third input sub-vector value from the source vector register in the MAC; generating a second output sub-vector value in the MAC based on the first input sub-vector value and the third input sub-vector value; storing the second output sub-vector value in the accumulation vector register; The processor implementation method according to claim 19, further comprising
21. The processor implementation method according to claim 16, wherein the rotation vector register file is used to rotate the data by a certain amount of rotation.
22. The processor implementation method according to claim 21, wherein the amount of rotation is selectable.
23. The processor implementation method according to claim 21, wherein the amount of rotation is based on an opcode of a multiply-accumulate instruction or a parameter of the multiply-accumulate instruction.
24. To generate first rotation data, execute a first data rotation corresponding to a first amount of rotation using a first rotator of a rotation circuit of the rotation vector register file; To generate second rotation data, execute a second data rotation corresponding to a second amount of rotation using a second rotator of the rotation circuit; Update the data in the rotation vector register with the second rotation data; The processor implementation method according to claim 21, further comprising
25. The processor implementation method according to claim 24, wherein the second amount of rotation is selectable.
26. Fetch the first input data from the rotation vector register file; Fetch the second input data from the source vector register of the second vector register file; In the MAC, process the first input data and the second input data; Rotate the data in the rotation vector register; Further comprising executing a multiply-accumulate instruction including The processor implementation method according to claim 16.
27. A non-transitory computer-readable medium storing instructions, wherein when the instructions are executed by a processor, the processor is caused to Use a rotation vector register file to rotate data in a rotation vector register of the rotation vector register file; Receive first input data from the rotation vector register file in a multiply-accumulate circuit (MAC); Receive second input data from a source vector register of a second vector register file in the MAC.
28. The first input data includes sub-vector values of the rotation vector register, the sub-vector values have a plurality of elements, and when the instruction is executed by the processor, the processor provides each of the plurality of elements to separate inputs of the plurality of inputs of the MAC using a broadcast circuit. The non-transitory computer-readable medium according to claim 27. [
29. ] An apparatus comprising: means for rotating data in a rotation vector register of a rotation vector register file; means for receiving first input data from the rotation vector register file in a multiply-accumulate circuit (MAC); means for receiving second input data from a source vector register of a second vector register file in the MAC; An apparatus comprising the above. [
30. ] The means for rotating, the means for receiving the first input data, and the means for receiving the second input data are integrated in at least one of a smart speaker, a speaker bar, a computer, a tablet, a display device, a television, a game console, a music player, a radio, a digital video player, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, a communication device, a mono Internet of Things (IoT) device, a virtual reality (VR) device, a base station, or a mobile device. The apparatus according to claim 29.