Histogram operation

The dual scalar/vector data path architecture in digital signal processors addresses increasing workload demands and memory latency issues by optimizing data handling and routing, enhancing performance and efficiency in complex algorithms.

CN113924550BActive Publication Date: 2025-07-15TEXAS INSTRUMENTS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080038840.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-05-27
Publication Date
2025-07-15
Estimated Expiration
2040-05-27

AI Technical Summary

Technical Problem

Modern digital signal processors (DSPs) face challenges such as increased workloads, memory system delay affecting algorithm performance, reliability problems caused by transistor miniaturization, and wire routing congestion, especially in the case of memory and register unreliability and wide buses that are difficult to route, resulting in performance bottlenecks.

Method used

A digital data processor is adopted, including instruction memory, instruction decoder and computational unit, optimizes data processing by increasing histogram values at a specified position, and uses lookup table technology to perform fast data access, reduce memory latency and increase bandwidth.

Benefits of technology

By optimizing data processing flow and lookup table technology, the processing efficiency of DSP is improved, memory delay is reduced, data processing parallelism and throughput are enhanced, and memory system delay and line routing congestion are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113924550B_ABST
    Figure CN113924550B_ABST
Patent Text Reader

Abstract

A digital data processor (100) comprising: an instruction memory (121) storing instructions, each of the instructions specifying a data processing operation and at least one data operand field; an instruction decoder (113) coupled to the instruction memory for sequentially retrieving instructions from the instruction memory and determining the data processing operation and the at least one data operand; and at least one arithmetic unit (110) coupled to a data register file (123) and to the instruction decoder for performing a data processing operation on at least one operand corresponding to an instruction decoded by the instruction decoder and storing the result of the data processing operation. The arithmetic unit is configured to increment a histogram value in response to a histogram instruction by incrementing a bin entry at a specified location in at least one histogram of a specified number.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Modern digital signal processors (DSPs) face multiple challenges. Workloads continue to increase, thus requiring increased bandwidth. The size and complexity of systems on a chip (SOCs) continue to grow. Memory system latency severely affects certain classes of algorithms. As transistors become smaller, memories and registers become less reliable. As the software stack becomes larger, the number of possible interactions and errors becomes larger. Even the wires are increasingly becoming a challenge. Wide buses are difficult to route. Wire speed continues to lag behind transistor speed. Routing congestion is an ongoing challenge.

[0002] One technique available for filtering functions is table lookup. A data table is loaded into memory, which stores a set of results at a memory location corresponding to an input parameter. To perform the function, the input parameter is used to call the pre-computed result. This technique can be particularly valuable for rarely used and computationally difficult mathematical functions. Summary of the Invention

[0003] In some instances, a digital data processor includes: an instruction memory that stores instructions, each of the instructions specifying a data processing operation and at least one data operand field; an instruction decoder coupled to the instruction memory for sequentially calling instructions from the instruction memory and determining the data processing operation and the at least one data operand; and at least one arithmetic unit coupled to a data register file and coupled to the instruction decoder to perform a data processing operation on at least one operand corresponding to an instruction decoded by the instruction decoder and store the result of the data processing operation. The arithmetic unit is configured to increment a histogram value in response to a histogram instruction by incrementing a bin entry at a specified position in at least one of a specified number of histograms.

[0004] In some instances, a method includes incrementing a histogram value by an arithmetic unit coupled to a data register file and coupled to an instruction decoder by incrementing a bin entry at a specified position in at least one of a specified number of histograms in response to a histogram instruction. Brief Description of the Drawings

[0005] To describe the various instances in detail, reference will now be made to the accompanying drawings, in which:

[0006] Figure 1 Illustrates a dual scalar / vector data path processor according to one embodiment;

[0007] Figure 2 Description Figure 1 The registers and functional units in the dual scalar / vector data path processor illustrated in;

[0008] Figure 3 Illustrates the global scalar register file;

[0009] Figure 4 Describe the local scalar register file shared by the arithmetic functional units;

[0010] Figure 5 Describe the local scalar register file shared by the multiplication functional units;

[0011] Figure 6 Describe the local scalar register file shared by the load / store units;

[0012] Figure 7 Describe the global vector register file;

[0013] Figure 8 Describe the assertion register file;

[0014] Fig. 9 Describe the local vector register file shared by the arithmetic functional units;

[0015] Fig.10 Describe the local vector register file shared by the multiplication and related functional units;

[0016] Fig.11 Describe the pipeline stages of the central processing unit according to an example embodiment;

[0017] Fig.12 Describe the sixteen instructions of a single fetch packet;

[0018] Fig.13 Describe the instruction decoding example in one embodiment;

[0019] Fig.14 Describe the bit decoding of the condition code extension slot 0;

[0020] Fig.15 Describe the bit decoding of the condition code extension slot 1;

[0021] Fig.16 Describe the bit decoding of the constant extension slot 0;

[0022] Fig.17 Partial block diagram for illustrating constant extension;

[0023] Fig.18 Describe the carry control for SIMD operations according to an example embodiment;

[0024] Fig.19 Describe the data fields of the example lookup table configuration register;

[0025] Fig. 20 Describe the data fields in the example lookup table enable register that specify the types of operations permitted for a particular table set;

[0026] Fig.21 Describe the lookup table organization for one table of the table set;

[0027] Fig. 22 Describe the lookup table organization for two tables of the table set;

[0028] Fig.23 Describe the lookup table organization for four tables of the table set;

[0029] Fig.24 Describe the lookup table organization for eight tables of the table set;

[0030] Fig.25 Describe the lookup table organization for sixteen tables of the table set;

[0031] Fig.26 Describe the operations of the lookup table read instruction for four parallel tables, the data element size in bytes, and an example without promotion in the example embodiment;

[0032] Fig. 27 Describe the operations of the lookup table read instruction for four parallel tables, the data element size in bytes, and an example with 2x promotion in the example embodiment;

[0033] Fig.28 Describe the operations of the lookup table read instruction for four parallel tables, the data element size in bytes, and an example with 4x promotion in the example embodiment;

[0034] Fig.29A and 29B Collectively describe the example embodiments of the promotion implementation;

[0035] Fig.30 Describe Fig.29A Examples of the extended elements described in;

[0036] Fig.31 Describe the control of the multiplexing control encoder for the multiplexer described in Fig.29A and 29B ;

[0037] Fig.32 Describe the operations of the lookup table read instruction for four parallel tables, the data element size in words, and an example of 2-element interpolation in the example embodiment;

[0038] Fig.33 Describe the operations of the lookup table read instruction for four parallel tables, the data element size in words, and an example of 4-element interpolation in the example embodiment;

[0039] Fig.34An example of the operation of a multi - stage butterfly unit that re - orders data from a look - up table before writing the re - ordered data to a destination register in response to the execution of a look - up table read instruction in an example embodiment;

[0040] Fig.35 An example of the operation of a look - up table write instruction for four parallel tables with a data element size of words in an example embodiment;

[0041] Fig.36A and 36B An example of the operation of a look - up table initialization instruction in an example embodiment;

[0042] Fig.37 An example of the operation of a histogram instruction for four parallel histograms in an example embodiment; and

[0043] Fig.38 An example of the operation of a weighted histogram instruction for four parallel histograms in an example embodiment. Detailed Description

[0044] Figure 1 Describe a dual - scalar / vector data - path processor 100 according to some examples. The processor 100 includes a separate level - 1 instruction cache (L1I) 121 and a level - 1 data cache (L1D) 123. The processor 100 includes a level - 2 combined instruction / data cache (L2) 130 that stores both instructions and data. Figure 1 Describe the connection (bus 142) between the level - 1 instruction cache 121 and the level - 2 combined instruction / data cache 130. Figure 1 Describe the connection (bus 145) between the level - 1 data cache 123 and the level - 2 combined instruction / data cache 130. The level - 2 combined instruction / data cache 130 stores instructions to back up the level - 1 instruction cache 121 and stores data to back up the level - 1 data cache 123. The level - 2 combined instruction / data cache 130 is further connected to higher - level caches and / or main memory in a manner known in the art and Figure 1 not described herein. In one embodiment, the central processing unit core 110, the level - 1 instruction cache 121, the level - 1 data cache 123, and the level - 2 combined instruction / data cache 130 are formed on a single integrated circuit. This single integrated circuit optionally includes other circuits.

[0045] The central processing unit core 110 fetches instructions from the level-1 instruction cache 121 as controlled by the instruction fetch unit 111. The instruction fetch unit 111 determines the next set of instructions to be executed and invokes a set of fetch packet sizes for such instructions. The nature and size of the fetch packet are further detailed below. As is known in the art, instructions are directly fetched from the level-1 instruction cache 121 immediately after a cache hit (if these instructions are stored in the level-1 instruction cache 121). After a cache miss (the specified instruction fetch packet is not stored in the level-1 instruction cache 121), these instructions are searched for in the level-2 unified cache 130. In one embodiment, the size of a cache line in the level-1 instruction cache 121 is equal to the size of the fetch packet. The memory location of these instructions is a hit or miss in the level-2 unified cache 130. A hit is serviced by the level-2 unified cache 130. A miss is serviced by a higher-level cache (not shown) or by the main memory (not shown). As is known in the art, the requested instructions can be supplied to both the level-1 instruction cache 121 and the central processing unit core 110 simultaneously to speed up usage.

[0046] The central processing unit core 110 includes a plurality of functional units (also referred to as "execution units") for performing the data processing tasks specified by the instructions. The instruction dispatch unit 112 determines the target functional unit for each fetched instruction. In one embodiment, the central processing unit 110 operates as a very long instruction word (VLIW) processor, which is capable of simultaneously executing multiple instructions in corresponding functional units. The compiler can organize the instructions to be executed together in an execution packet. The instruction dispatch unit 112 directs each instruction to its target functional unit. In one embodiment, the functional unit assigned to an instruction is completely specified by the instruction generated by the compiler, as the hardware of the central processing unit core 110 has no part in this functional unit assignment. The instruction dispatch unit 112 can execute multiple instructions in parallel. The number of such parallel instructions is set by the size of the execution packet. This will be further detailed below.

[0047] One part of the dispatch task of the instruction dispatch unit 112 is to determine whether the instruction is to be executed on a functional unit on the scalar data path side A 115 or the vector data path side B 116. Instruction bits within each instruction, called s bits, determine which data path the instruction controls. This will be further detailed below.

[0048] The instruction decoding unit 113 decodes each instruction in the current execution packet. The decoding includes identifying the functional unit for the instruction, identifying the register that supplies data for the corresponding data processing operation from among the possible register files, and identifying the register destination for the result of the corresponding data processing operation. As described below, an instruction may include a constant field instead of a register number operand field. The result of this decoding is a signal for controlling the target functional unit to perform the data processing operation specified by the corresponding instruction on the specified data.

[0049] The central processing unit core 110 includes a control register 114. The control register 114 stores information for controlling the functional units in the scalar data path side A 115 and the vector data path side B 116. This information may include mode information or the like.

[0050] The decoded instructions from the instruction decoding unit 113 and the information stored in the control register 114 are supplied to the scalar data path side A 115 and the vector data path side B 116. Accordingly, the functional units within the scalar data path side A 115 and the vector data path side B 116 perform the data processing operations specified by the instructions on the instruction-specified data and store the results in one or more instruction-specified data registers. Each of the scalar data path side A 115 and the vector data path side B 116 includes a plurality of functional units that can operate in parallel. These will be described in further detail Figure 2 below. There is a data path 117 between the scalar data path side A 115 and the vector data path side B 116 that permits data exchange.

[0051] The central processing unit core 110 includes additional non-instruction-based modules. The emulation unit 118 permits determining the machine state of the central processing unit core 110 in response to an instruction. This capability can be used for algorithm development. The interrupt / exception unit 119 enables the central processing unit core 110 to respond to external asynchronous events (interrupts) and to respond to attempts to perform improper operations (exceptions).

[0052] The central processing unit core 110 includes a streaming engine 125. The streaming engine 125 supplies two data streams from a predetermined address that can be cached in the level 2 combined cache 130 to the register file of the vector data path side B. This provides controlled data movement directly from memory (such as cached in the level 2 combined cache 130) to the functional unit operand inputs. This is described in further detail below.

[0053] Figure 1Describe the data width of the bus between various parts for the example embodiment. The primary instruction cache 121 supplies instructions to the instruction fetch unit 111 via the bus 141. In this example embodiment, the bus 141 is a 512-bit bus. The bus 141 is unidirectional from the primary instruction cache 121 to the central processing unit 110. The secondary combined cache 130 supplies instructions to the primary instruction cache 121 via the bus 142. In this example embodiment, the bus 142 is a 512-bit bus. The bus 142 is unidirectional from the secondary combined cache 130 to the primary instruction cache 121.

[0054] The primary data cache 123 exchanges data with the register file in the scalar data path side A 115 via the bus 143. In this example embodiment, the bus 143 is a 64-bit bus. The primary data cache 123 exchanges data with the register file in the vector data path side B 116 via the bus 144. In this example embodiment, the bus 144 is a 512-bit bus. The buses 143 and 144 are described as bidirectional to support both data reading and data writing for the central processing unit 110. The primary data cache 123 exchanges data with the secondary combined cache 130 via the bus 145. In this example embodiment, the bus 145 is a 512-bit bus. The bus 145 is described as bidirectional to support cache services for both data reading and data writing for the central processing unit 110.

[0055] As is known in the art, the CPU data request is directly fetched from the primary data cache 123 immediately after a cache hit (if the requested data is stored in the primary data cache 123). After a cache miss (the specified data is not stored in the primary data cache 123), this data is searched for in the secondary combined cache 130. The memory location of this requested data is a hit or a miss in the secondary combined cache 130. A hit is serviced by the secondary combined cache 130. A miss is serviced by another level of cache (not shown) or by the main memory (not shown). As is known in the art, the requested instructions can be supplied to both the primary data cache 123 and the central processing unit core 110 simultaneously to accelerate usage.

[0056] The secondary combined cache 130 supplies data of a first data stream to the streaming engine 125 via the bus 146. In this example embodiment, the bus 146 is a 512-bit bus. The streaming engine 125 supplies data of this first data stream to the functional units of the vector data path side B 116 via the bus 147. In this example embodiment, the bus 147 is a 512-bit bus. The secondary combined cache 130 supplies data of a second data stream to the streaming engine 125 via the bus 148. In this example embodiment, the bus 148 is a 512-bit bus. The streaming engine 125 supplies data of this second data stream to the functional units of the vector data path side B 116 via the bus 149. In this example embodiment, the bus 149 is a 512-bit bus. In this example embodiment, the buses 146, 147, 148, and 149 are described as unidirectional from the secondary combined cache 130 to the streaming engine 125 and to the vector data path side B 116.

[0057] The streaming engine data requests are fetched directly from the secondary combined cache 130 immediately after a cache hit (if the requested data is stored in the secondary combined cache 130). After a cache miss (the specified data is not stored in the secondary combined cache 130), this data is searched for in another level of cache (not shown) or in the main memory (not shown). In some embodiments, the primary data cache 123 may cache data not stored in the secondary combined cache 130. If this operation is supported, then after a streaming engine data request that is a miss in the secondary combined cache 130, the secondary combined cache 130 may immediately snoop the primary data cache 123 for the data requested by the streaming engine. If the primary data cache 123 stores this data, then its snoop response will contain the data, which is then supplied to service the streaming engine request. If the primary data cache 123 does not store this data, then its snoop response will indicate this, and the secondary combined cache 130 will then service this streaming engine request from another level of cache (not shown) or from the main memory (not shown).

[0058] In one embodiment, both the level-1 data cache 123 and the level-2 combined cache 130 can be configured as a selected amount of cache or directly addressable memory in U.S. Patent No. 6,606,686, titled "UNIFIED MEMORY SYSTEM ARCHITECTURE INCLUDING CACHE AND DIRECTLY ADDRESSABLE STATIC RANDOM ACCESS MEMORY".

[0059] Figure 2 Describe additional details of the functional units and register files within the scalar data path side A 115 and the vector data path side B 116 in an example embodiment. The scalar data path side A 115 includes a global scalar register file 211, an L1 / S1 local register file 212, an M1 / N1 local register file 213, and a D1 / D2 local register file 214. The scalar data path side A 115 includes an L1 unit 221, an S1 unit 222, an M1 unit 223, an N1 unit 224, a D1 unit 225, and a D2 unit 226. The vector data path side B 116 includes a global vector register file 231, an L2 / S2 local register file 232, an M2 / N2 / C local register file 233, and an assertion register file 234. The vector data path side B 116 includes an L2 unit 241, an S2 unit 242, an M2 unit 243, an N2 unit 244, a C unit 245, and a P unit 246. These functional units can be configured to read from or write to certain register files, as will be detailed below.

[0060] The L1 unit 221 can accept two 64-bit operands and produce a 64-bit result. Each of the two operands is called from an instruction-specified register in the global scalar register file 211 or the L1 / S1 local register file 212. The L1 unit 221 can perform the following instruction-selected operations: 64-bit addition / subtraction operations; 32-bit minimum / maximum operations; 8-bit single instruction multiple data (SIMD) instructions, such as absolute value, minimum, and maximum determination sums; loop minimum / maximum operations; and various move operations between register files. The result produced by the L1 unit 221 can be written to an instruction-specified register in the global scalar register file 211, the L1 / S1 local register file 212, the M1 / N1 local register file 213, or the D1 / D2 local register file 214.

[0061] The S1 unit 222 can accept two 64-bit operands and produce a 64-bit result. Each of the two operands is called from an instruction-specified register in the global scalar register file 211 or the L1 / S1 local register file 212. In one embodiment, the S1 unit 222 can perform the same type of operations as the L1 unit 221. In other embodiments, there may be minor variations between the data processing operations supported by the L1 unit 221 and the S1 unit 222. The result produced by the S1 unit 222 can be written into an instruction-specified register in the global scalar register file 211, the L1 / S1 local register file 212, the M1 / N1 local register file 213, or the D1 / D2 local register file 214.

[0062] The M1 unit 223 can accept two 64-bit operands and produce a 64-bit result. Each of the two operands is called from an instruction-specified register in the global scalar register file 211 or the M1 / N1 local register file 213. The M1 unit 223 can perform the following instruction-selected operations: 8-bit multiplication operation; complex inner product operation; 32-bit bit count operation; complex conjugate multiplication operation; and bitwise logical operations, shifts, additions, and subtractions. The result produced by the M1 unit 223 can be written into an instruction-specified register in the global scalar register file 211, the L1 / S1 local register file 212, the M1 / N1 local register file 213, or the D1 / D2 local register file 214.

[0063] The N1 unit 224 can accept two 64-bit operands and produce a 64-bit result. Each of the two operands is called from an instruction-specified register in the global scalar register file 211 or the M1 / N1 local register file 213. The N1 unit 224 can perform the same type of operations as the M1 unit 223. There may be certain dual operations (called dual-issue instructions) that employ both the M1 unit 223 and the N1 unit 224 together. The result produced by the N1 unit 224 can be written into an instruction-specified register in the global scalar register file 211, the L1 / S1 local register file 212, the M1 / N1 local register file 213, or the D1 / D2 local register file 214.

[0064] The D1 unit 225 and the D2 unit 226 can each accept two 64-bit operands and each produce a 64-bit result. The D1 unit 225 and the D2 unit 226 can perform address calculations and corresponding load and store operations. The D1 unit 225 is used for 64-bit scalar load and store. The D2 unit 226 is used for 512-bit vector load and store. The D1 unit 225 and the D2 unit 226 can also perform: swapping, packing, and unpacking of loaded and stored data; 64-bit SIMD arithmetic operations; and 64-bit bitwise logical operations. The D1 / D2 local register file 214 generally stores the base address and offset address used in the address calculation for corresponding load and store. Each of the two operands is called from the instruction-specified register in the global scalar register file 211 or the D1 / D2 local register file 214. The result calculated by the D1 unit 225 and / or the D2 unit 226 can be written into the instruction-specified register in the global scalar register file 211, the L1 / S1 local register file 212, the M1 / N1 local register file 213, or the D1 / D2 local register file 214.

[0065] The L2 unit 241 can accept two 512-bit operands and produce a 512-bit result. Each of at most two operands is called from the instruction-specified register in the global vector register file 231, the L2 / S2 local register file 232, or the assertion register file 234. In addition to the wider 512-bit data, the L2 unit 241 can also perform instructions similar to those of the L1 unit 221. The result produced by the L2 unit 241 can be written into the instruction-specified register in the global vector register file 231, the L2 / S2 local register file 232, the M2 / N2 / C local register file 233, or the assertion register file 234.

[0066] The S2 unit 242 can accept two 512-bit operands and produce a 512-bit result. Each of at most two operands is called from the instruction-specified register in the global vector register file 231, the L2 / S2 local register file 232, or the assertion register file 234. In addition to the wider 512-bit data, the S2 unit 242 can also perform instructions similar to those of the S1 unit 222. The result produced by the S2 unit 242 can be written into the instruction-specified register in the global vector register file 231, the L2 / S2 local register file 232, the M2 / N2 / C local register file 233, or the assertion register file 234.

[0067] The M2 unit 243 can accept two 512-bit operands and produce a 512-bit result. Each of the two operands is called from an instruction-specified register in the global vector register file 231 or the M2 / N2 / C local register file 233. In addition to the wider 512-bit data, the M2 unit 243 can also perform instructions similar to those of the M1 unit 223. The result produced by the M2 unit 243 can be written into an instruction-specified register in the global vector register file 231, the L2 / S2 local register file 232, or the M2 / N2 / C local register file 233.

[0068] The N2 unit 244 can accept two 512-bit operands and produce a 512-bit result. Each of the two operands is called from an instruction-specified register in the global vector register file 231 or the M2 / N2 / C local register file 233. The N2 unit 244 can perform the same type of operations as the M2 unit 243. There may be certain dual operations (called dual-issue instructions) that employ both the M2 unit 243 and the N2 unit 244 together. The result produced by the N2 unit 244 can be written into an instruction-specified register in the global vector register file 231, the L2 / S2 local register file 232, or the M2 / N2 / C local register file 233.

[0069] The C unit 245 can accept two 512-bit operands and produce a 512-bit result. Each of the two operands is called from an instruction-specified register in the global vector register file 231 or the M2 / N2 / C local register file 233. The C unit 245 can perform: "Rake" and "Search" instructions; up to 512 2-bit PN * 8-bit multiplications; I / Q complex multiplications per clock cycle; 8-bit and 16-bit sum of absolute differences (SAD) calculations, up to 512 SADs per clock cycle; horizontal addition and horizontal minimum / maximum instructions; and vector permutation instructions. In one embodiment, the C unit 245 includes four vector control registers (CUCR0 to CUCR3) for controlling certain operations of the C unit 245 instructions. The control registers CUCR0 to CUCR3 are used as operands in certain C unit 245 operations. For example, the control registers CUCR0 to CUCR3 can be used in the control of the general permutation instruction (VPERM), or as masks for SIMD multiple DOT product operations (DOTPM) and SIMD multiple sum of absolute differences (SAD) operations. The control register CUCR0 can be used to store a polynomial for Galois Field multiplication operations (GFMPY). The control register CUCR1 can be used to store a Galois Field polynomial generator function.

[0070] The P unit 246 can perform basic logical operations on the registers of the local assertion register file 234. The P unit 246 has direct access to read from and write to the assertion register file 234. The operations performed by the P unit 246 can include AND, ANDN, OR, XOR, NOR, BITR, NEG, SET, BITCNT, RMBD, BIT Decimate, and Expand. One use of the P unit 246 can include manipulating SIMD vector comparison results for controlling additional SIMD vector operations.

[0071] Figure 3 Describe an example embodiment of the global scalar register file 211. In the illustrated embodiment, there are 16 independent 64-bit wide scalar registers labeled A0 to A15. Each register of the global scalar register file 211 can be read or written as 64-bit scalar data. All scalar data path side A 115 functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) can read from or write to the global scalar register file 211. The global scalar register file 211 can be read as 32 bits or 64 bits and can only be written as 64 bits. The executed instruction determines the read data size. Under the limitations detailed below, the vector data path side B 116 functional units (L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) can read from the global scalar register file 211 via the cross path 117.

[0072] Figure 4 Describe an example embodiment of the D1 / D2 local register file 214. In the illustrated embodiment, there are 16 independent 64-bit wide scalar registers labeled D0 to D16. Each register of the D1 / D2 local register file 214 can be read or written as 64-bit scalar data. All scalar data path side A 115 functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) can write to the D1 / D2 local register file 214. Only the D1 unit 225 and the D2 unit 226 can read from the D1 / D2 local register file 214. The data stored in the D1 / D2 local register file 214 can include a base address and an offset address used in address calculations.

[0073] Figure 5 Describe an example embodiment of the Ll / S1 local register file 212. Figure 5 The illustrated embodiment has 8 independent 64-bit wide scalar registers labeled AL0 to AL7. In certain instruction decoding formats (see Fig.13), the L1 / S1 local register file 212 may contain up to 16 registers. Figure 5 The illustrated embodiment implements only 8 registers to reduce circuit size and complexity. Each register of the Ll / S1 local register file 212 can read or write 64-bit scalar data. All scalar data path side A 115 functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) can write to the L1 / S1 local scalar register file 212. Only the L1 unit 221 and the S1 unit 222 can read from the L1 / S1 local register file 212.

[0074] Figure 6 Describe an example embodiment of the M1 / N1 local register file 213. Figure 6 The illustrated embodiment in has 8 independent 64-bit wide scalar registers labeled AM0 to AM7. In certain instruction decoding formats (see Fig.13 ), the M1 / N1 local register file 213 may contain up to 16 registers. Figure 6 The illustrated embodiment implements only 8 registers to reduce circuit size and complexity. Each register of the M1 / N1 local register file 213 can read or write 64-bit scalar data. All scalar data path side A 115 functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) can write to the M1 / N1 local register file 213. Only the M1 unit 223 and the N1 unit 224 can read from the M1 / N1 local register file 213.

[0075] Figure 7 Describe an example embodiment of the global vector register file 231. In the illustrated embodiment, there are 16 independent 512-bit wide vector registers. Each register of the global vector register file 231 can read or write 64-bit scalar data labeled B0 to B15. Each register of the global vector register file 231 can read or write 512-bit vector data labeled VB0 to VB15. The instruction type determines the data size. All vector data path side B 116 functional units (L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) can read or write to the global vector register file 231. Under the limitations to be described in detail below, the scalar data path side A 115 functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) can read from the global vector register file 231 via the cross path 117.

[0076] Figure 8Describe an example embodiment of the P local register file 234. In the illustrated embodiment, there are eight independent 64-bit wide registers labeled P0 through P7. Each register of the P local register file 234 can be read or written as 64-bit scalar data. The vector data path side B 116 functional units L2 unit 241, S2 unit 242, C unit 244, and P unit 246 can write to the P local register file 234. Only the L2 unit 241, S2 unit 242, and P unit 246 can read from the P local register file 234. The P local register file 234 can be used to: write a one-bit SIMD vector comparison result from the L2 unit 241, S2 unit 242, or C unit 245; manipulate the SIMD vector comparison result by the P unit 246; and use the manipulated result to control additional SIMD vector operations.

[0077] Fig. 9 Describe an example embodiment of the L2 / S2 local register file 232. Fig. 9 The illustrated embodiment has eight independent 512-bit wide vector registers. In certain instruction decoding formats (see Fig.13 ), the L2 / S2 local register file 232 can contain up to 16 registers. Fig. 9 The embodiment of only implements eight registers to reduce circuit size and complexity. Each register of the L2 / S2 local vector register file 232 can be read or written as 64-bit scalar data labeled BL0 through BL7. Each register of the L2 / S2 local vector register file 232 can be read or written as 512-bit vector data labeled VBL0 through VBL7. The instruction type determines the data size. All vector data path side B 116 functional units (L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) can write to the L2 / S2 local register file 232. Only the L2 unit 241 and S2 unit 242 can read from the L2 / S2 local vector register file 232.

[0078] Fig.10 Describe an example embodiment of the M2 / N2 / C local register file 233. Fig.10 The illustrated embodiment has eight independent 512-bit wide vector registers. In certain instruction decoding formats (see Fig.13 ), the M2 / N2 / C local register file 233 can contain up to 16 registers. Fig.10The embodiments of [description] implement only 8 registers to reduce circuit size and complexity. Each register of the M2 / N2 / C local vector register file 233 can be read or written with 64-bit scalar data labeled BM0 to BM7. Each register of the M2 / N2 / C local vector register file 233 can be read or written with 512-bit vector data labeled VBM0 to VBM7. All vector data path side B 116 functional units (L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) can write to the M2 / N2 / C local vector register file 233. Only the M2 unit 233, N2 unit 244, and C unit 245 can read from the M2 / N2 / C local vector register file 233.

[0079] Thus, in some embodiments, the global register file can be accessible by all functional units (e.g., scalar and vector) on one side, and the local register file can be accessible by only some of the functional units on one side. Some additional embodiments can be practiced with only one type of register file corresponding to the described global register file.

[0080] The cross path 117 permits limited data exchange between the scalar data path side A 115 and the vector data path side B 116. During each operation cycle, a 64-bit data word can be called from the global scalar register file 211 to be used as an operand by one or more functional units on the vector data path side B 116, and a 64-bit data word can be called from the global vector register file 231 to be used as an operand by one or more functional units on the scalar data path side A 115. Any scalar data path side A 115 functional unit (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) can read a 64-bit operand from the global vector register file 231. This 64-bit operand is the least significant bits of the 512-bit data in the accessed register of the global vector register file 231. The scalar data path side A 115 functional units can use the same 64-bit cross path data as the operand during the same operation cycle. However, in any single operation cycle, only one 64-bit operand is transferred from the vector data path side B 116 to the scalar data path side A 115. Any vector data path side B 116 functional unit (L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) can read a 64-bit operand from the global scalar register file 211. If the corresponding instruction is a scalar instruction, the cross path operand data is treated as any other 64-bit operand. If the corresponding instruction is a vector instruction, the upper 448 bits of the operand are filled with zeros. The vector data path side B 116 functional units can use the same 64-bit cross path data as the operand during the same operation cycle. In any single operation cycle, only one 64-bit operand is transferred from the scalar data path side A 115 to the vector data path side B 116.

[0081] In some cases, the streaming engine 125 transfers data. In Figure 1In an embodiment, the streaming engine 125 controls two data streams. A stream contains a series of elements of a particular type. A program operating on the stream reads the data sequentially and then operates on each element. Stream data can have the following basic characteristics: well-defined start and end times; fixed element size and type throughout the stream; and a fixed sequence of elements. Thus, a program cannot randomly search within the stream. Additionally, stream data is read-only only when active. Thus, a program cannot write to the stream while reading from it. Once the stream is opened, the streaming engine 125: calculates addresses; extracts the defined data type from the secondary unified cache 130 (which may require cache services from higher-level memory, i.e., in the case of a cache miss in the secondary unified cache 130); performs data type manipulation (e.g., zero extension, sign extension, and / or sorting / transposing of data elements such as matrix transposition); and delivers the data directly to the programmed data register file within the CPU 110. The streaming engine 125 is thus suitable for real-time digital filtering operations on benign data. The streaming engine 125 frees these memory extraction tasks from the corresponding CPU 110, thereby enabling the CPU 110 to perform other processing functions.

[0082] The streaming engine 125 provides several benefits. For example, the streaming engine 125 allows for multi-dimensional memory access, increases the available bandwidth of the functional units of the CPU 110, reduces the number of cache miss stalls due to the stream buffer bypassing the level-1 data cache 123, reduces the number of scalar operations required to maintain a loop, and manages address pointers. The streaming engine 125 can also handle address generation, which frees address generation instruction slots and the D1 unit 225 and D2 unit 226 for other computations.

[0083] The CPU 110 operates on an instruction pipeline. As further described below, instructions are fetched in fixed-length instruction packets. All instructions have the same number of pipeline stages for fetching and decoding, but can have different numbers of execution stages.

[0084] Fig.11 An example embodiment of an instruction pipeline having the following pipeline stages is described: a program fetch stage 1110, a dispatch and decode stage 1120, and an execution stage 1130. The program fetch stage 1110 includes three stages for all instructions. The dispatch and decode stage 1120 includes three stages for all instructions. The execution stage 1130 includes one to four stages depending on the instruction.

[0085] The fetch stage 1110 includes a program address generation stage 1111 (PG), a program access stage 1112 (PA), and a program reception stage 1113 (PR). During the program address generation stage 1111 (PG), the program address is generated in the CPU, and a read request is sent to the memory controller of the level 1 instruction cache L1I. During the program access stage 1112 (PA), the level 1 instruction cache L1I processes the request, accesses the data in its memory, and sends a fetch packet to the CPU boundary. During the program reception stage 1113 (PR), the CPU registers fetch the packet.

[0086] In an example embodiment, instructions are fetched as sixteen 32-bit wide time slots at a time, thereby forming a fetch packet. Fig.12 Describe one such embodiment, where a single fetch packet includes sixteen instructions 1201 to 1216. The fetch packet is aligned on a 512-bit (16-word) boundary. In one embodiment, the fetch packet has a fixed 32-bit instruction length. Fixed-length instructions are advantageous for several reasons. Fixed-length instructions enable a simple decoder alignment. Properly aligned instruction fetching can load multiple instructions into a parallel instruction decoder. Such properly aligned instruction fetching can be achieved through a predetermined instruction alignment when stored in a memory coupled to the fixed instruction packet fetch (fetch packet aligned on a 512-bit boundary). Aligned instruction fetching also permits the parallel decoder to operate on the fetched bits of a set instruction size. Variable-length instructions may require an initial step of locating each instruction boundary before they can be decoded. Fixed-length instruction sets generally permit a more regular layout of instruction fields. This simplifies the construction of each decoder, which is an advantage for a wide-issue VLIW central processing unit.

[0087] The execution of individual instructions is partially controlled by the p-bit in each instruction. This p-bit can be configured as bit 0 of a 32-bit wide time slot. The p-bit of an instruction determines whether the instruction is executed in parallel with the next instruction. Instructions are scanned from lower address to higher address. If the p-bit of an instruction is 1, then the next subsequent instruction (higher memory address) is executed in parallel with the instruction (in the same cycle as the instruction). If the p-bit of an instruction is 0, then the next subsequent instruction is executed in the cycle after the instruction.

[0088] The CPU 110 and the level 1 instruction cache L1I 121 pipeline are decoupled from each other. Depending on the external environment, such as whether there is a hit in the level 1 instruction cache 121 or a hit in the level 2 unified cache 130, the fetch packet returned from the level 1 instruction cache L1I may take a different number of clock cycles. Therefore, the program access stage 1112 (PA) may take several clock cycles instead of 1 clock cycle as in other stages.

[0089] Instructions executed in parallel form an execution packet. In one embodiment, an execution packet may contain up to sixteen instructions (time slots). No two instructions in an execution packet may use the same functional unit. A time slot may be one of five types: 1) an independent instruction executed on one of the functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, D2 unit 226, L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) of CPU 110; 2) a no-unit instruction, such as a no-operation (NOP) instruction or multiple NOP instructions; 3) a branch instruction; 4) a constant field extension; and 5) a conditional code extension. Some of these time slot types will be described below.

[0090] The dispatch and decode stage 1120 includes an instruction dispatch to an appropriate execution unit stage 1121 (DS), an instruction pre-decode stage 1122 (DC1); and an instruction decode, operand fetch stage 1123 (DC2). During the instruction dispatch to an appropriate execution unit stage 1121 (DS), the fetch packet is split into execution packets and assigned to appropriate functional units. During the instruction pre-decode stage 1122 (DC1), the source register, destination register, and associated paths are decoded for the instruction in the functional unit. During the instruction decode, operand fetch stage 1123 (DC2), more detailed unit decoding is performed, and operands are fetched from the register file.

[0091] The execution stage 1130 includes execution stages 1131 to 1135 (E1 to E5). Different types of instructions may require different numbers of these stages to complete their execution. These stages of the pipeline play an important role in understanding the device state at the CPU cycle boundary.

[0092] During the execution 1 stage 1131 (E1), the condition of the instruction is evaluated and operations are performed on the operands. As Fig.11 illustrated, the execution 1 stage 1131 may receive an operand from one of the stream buffer 1141 and the register file schematically shown as 1142. For load and store instructions, address generation is performed and the address modification is written to the register file. For branch instructions, the branch fetch packet in the PG stage is affected.

[0093] As Fig.11 illustrated, load and store instructions access memory, here schematically shown as memory 1151. For single-cycle instructions, the result is written to the destination register file. This assumes that any condition of the instruction evaluates to true. If the condition evaluates to false, then the instruction does not write any result or have any pipeline operations after the execution 1 stage 1131.

[0094] During the execution of stage 2 1132 (E2), load instructions send an address to the memory. Store instructions send an address and data to the memory. A single-cycle instruction that saturates the result sets the bit (SAT) in the control status register (CSR) when saturation occurs. For a 2-cycle instruction, the result is written to the destination register file.

[0095] During the execution of stage 3 1133 (E3), data memory accesses are performed. Any multiply instruction that saturates the result sets the SAT bit in the control status register (CSR) when saturation occurs. For a 3-cycle instruction, the result is written to the destination register file.

[0096] During the execution of stage 4 1134 (E4), load instructions bring data to the CPU boundary. For a 4-cycle instruction, the result is written to the destination register file.

[0097] During the execution of stage 5 1135 (E5), load instructions write data into a register. This is illustrated schematically by the input from the memory 1151 to the execution stage 5 1135. Fig.11 This is illustrated schematically by the input from the memory 1151 to the execution stage 5 1135.

[0098] Fig.13 An instruction decoding format 1300 for functional unit instructions in an exemplary embodiment is described. Those skilled in the art will recognize that other instruction decodings are possible and within the scope of this specification. In the illustrated embodiment, each instruction comprises 32 bits and controls the operation of one of individually controllable functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, D2 unit 226, L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246). The bit fields of the instruction decoding 1300 are defined as follows.

[0099] The creg field 1301 (bits 29 to 31) and the z bit 1302 (bit 28) are fields used in conditional instructions. These bits are used in conditional instructions to identify the assertion (also referred to as “conditional”) register and the condition. The z bit 1302 (bit 28) indicates whether the condition is based on zero or non-zero in the assertion register. If z = 1, then the test is for equality with zero. If z = 0, then the test is for non-zero. For unconditional instructions, both the creg field 1301 and the z bit 1302 are set to 0 to allow the unconditional instruction to execute. The creg field 1301 and the z field 1302 are encoded in the instruction as shown in Table 1.

[0100] Table 1

[0101]

[0102]

[0103] The execution of a conditional instruction is conditional on the value stored in a specified conditional data register. In the illustrated example, the conditional register is a data register in the global scalar register file 211. The "z" in the z column refers to the zero / non-zero comparison selection mentioned above, and "x" is an irrelevant state. In this example, using three bits of the creg field 1301 in this decoding allows only a subset (A0 to A5) of the 16 global registers in the global scalar register file 211 to be specified as the predicate register. This selection is made to conserve bits in the instruction decoding and reduce the opcode space.

[0104] The dst field 1303 (bits 23 to 27) specifies a register in the corresponding register file as the destination of the instruction result (e.g., where the result will be written).

[0105] The src2 / cst field 1304 (bits 18 to 22) can be interpreted in different ways depending on the instruction opcode field (bits 4 to 12 for all instructions and bits 28 to 31 for unconditional instructions). The src2 / cst field 1304 indicates a register from the corresponding register file or a second source operand as a constant depending on the instruction opcode field. Depending on the instruction type, when the second source operand is a constant, this can be treated as an unsigned integer and zero extended to the specified data length, or can be treated as a signed integer and sign extended to the specified data length.

[0106] The src1 field 1305 (bits 13 to 17) specifies a register in the corresponding register file as the first source operand.

[0107] The opcode field 1306 (bits 4 to 12) for all instructions (and bits 28 to 31 for unconditional instructions) specifies the type of the instruction and flags appropriate instruction options. This includes the identification of the functional unit used and the operation performed. Additional details regarding such instruction options are detailed below.

[0108] The e bit 1307 (bit 2) is used for immediate constant instructions where the constant can be extended. If e = 1, then the immediate constant is extended in the manner detailed below. If e = 0, then the immediate constant is not extended. In the latter case, the immediate constant is specified by the src2 / cst field 1304 (bits 18 to 22). The e bit 1307 can be used only for some types of instructions. Thus, through proper decoding, the e bit 1307 can be omitted from instructions that do not require it, and this bit can instead be used as an additional opcode bit.

[0109] The s-bit 1308 (bit 1) indicates either the scalar data path side A 115 or the vector data path side B 116. If s = 0, then the scalar data path side A 115 is selected, and the available functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) and register files (global scalar register file 211, Ll / S1 local register file 212, M1 / N1 local register file 213, and D1 / D2 local register file 214) will be the functional units and register files corresponding to the scalar data path side A 115 as described in Figure 2 Similarly, s = 1 selects the vector data path side B 116, and the available functional units (L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, C unit 245, and P unit 246) and register files (global vector register file 231, L2 / S2 local register file 232, M2 / N2 / C local register file 233, and predicate local register file 234) will be the functional units and register files corresponding to the vector data path side B 116 as described in Figure 2 The p-bit 1308 (bit 0) is used to determine whether an instruction is executed in parallel with the following instruction. The p-bit is scanned from the lower address to the higher address. If p = 1 for the current instruction, then the next instruction is executed in parallel with the current instruction. If p = 0 for the current instruction, then the next instruction is executed in the cycle after the current instruction. All instructions executed in parallel form an execution packet. In one exemplary embodiment, an execution packet can contain up to twelve instructions for parallel execution, where each instruction in the execution packet is assigned to a different functional unit.

[0110] In one exemplary embodiment of the processor 100, there are two different condition code extension slots (slot 0 and slot 1). In this exemplary embodiment, the condition code extension slots can be 32 bits, as in

[0111] The decoding format 1300 described. Each execution packet can contain each of these 32-bit condition code extension slots, and each slot contains a 4-bit creg / z field for the instructions in the same execution packet (e.g., similar to bits 28 to 31 of the decoding 1300). Fig.13 Exemplary decoding of condition code extension slot 0 is illustrated, and Fig.14 Exemplary decoding of condition code extension slot 1 is illustrated. Fig.15

[0112] Fig.14 ​Describe the example decoding 1400 of conditional code extension time slot 0. Field 1401 (bits 28 to 31) specifies 4 creg / z bits assigned to the L1 unit 221 instructions in the same execution packet. Field 1402 (bits 27 to 24) specifies 4 creg / z bits assigned to the L2 unit 241 instructions in the same execution packet. Field 1403 (bits 19 to 23) specifies 4 creg / z bits assigned to the S1 unit 222 instructions in the same execution packet. Field 1404 (bits 16 to 19) specifies 4 creg / z bits assigned to the S2 unit 242 instructions in the same execution packet. Field 1405 (bits 12 to 15) specifies 4 creg / z bits assigned to the D1 unit 225 instructions in the same execution packet. Field 1406 (bits 8 to 11) specifies 4 creg / z bits assigned to the D2 unit 226 instructions in the same execution packet. Field 1407 (bits 6 and 7) is unused / reserved. Field 1408 (bits 0 to 5) is decoded with a unique bit set (CCEX0) to identify conditional code extension time slot 0. Once this unique ID of conditional code extension time slot 0 is detected, the corresponding creg / z bits are used to control the conditional execution of any L1 unit 221, L2 unit 241, S1 unit 222, S2 unit 242, D1 unit 225, and D2 unit 226 instructions in the same execution packet. These creg / z bits are interpreted as shown in Table 1. If the corresponding instruction is conditional (e.g., the creg / z bits are not all 0), then the corresponding bits in conditional code extension time slot 0 override the conditional code bits (bits 28 to 31 of creg field 1301 and z bit 1302) in the instruction (e.g., decoded with decoding format 1300). In the example illustrated, no execution packet can have more than one instruction leading to a particular execution unit, and no instruction execution packet can contain more than one conditional code extension time slot 0. Therefore, the mapping of creg / z bits to functional unit instructions is unambiguous. As described above, setting the creg / z bits equal to "0000" makes the instruction unconditional. Thus, properly decoded conditional code extension time slot 0 can make some corresponding instructions conditional and some corresponding instructions unconditional.

[0113] Fig.15Describe the example decoding 1500 of the conditional code extension time slot 1. Field 1501 (bits 28 to 31) specifies 4 creg / z bits assigned to the M1 unit 223 instructions in the same execution packet. Field 1502 (bits 27 to 24) specifies 4 creg / z bits assigned to the M2 unit 243 instructions in the same execution packet. Field 1503 (bits 19 to 23) specifies 4 creg / z bits assigned to the C unit 245 instructions in the same execution packet. Field 1504 (bits 16 to 19) specifies 4 creg / z bits assigned to the N1 unit 224 instructions in the same execution packet. Field 1505 (bits 12 to 15) specifies 4 creg / z bits assigned to the N2 unit 244 instructions in the same execution packet. Field 1506 (bits 6 to 11) is unused / reserved. Field 1507 (bits 0 to 5) is decoded with a unique bit set (CCEX1) to identify the conditional code extension time slot 1. Once this unique ID of the conditional code extension time slot 1 is detected, the corresponding creg / z bits are used to control the conditional execution of any M1 unit 223, M2 unit 243, C unit 245, N1 unit 224, and N2 unit 244 instructions in the same execution packet. These creg / z bits are interpreted as shown in Table 1. If the corresponding instruction is conditional (e.g., the creg / z bits are not all 0), then the corresponding bits in the conditional code extension time slot 1 override the conditional code bits (bits 28 to 31 of the creg field 1301 and the z bit 1302) in the instruction (e.g., decoded with the decoding format 1300). In the illustrated example, no execution packet can have more than one instruction leading to a specific execution unit, and no instruction execution packet can contain more than one conditional code extension time slot 1. Therefore, the mapping of the creg / z bits to the functional unit instructions is unambiguous. As described above, setting the creg / z bits equal to "0000" makes the instruction unconditional. Therefore, the properly decoded conditional code extension time slot 1 can make some corresponding instructions conditional and some corresponding instructions unconditional.

[0114] As described above in connection with Fig.13 what was described, both the conditional code extension time slot 0 1400 and the conditional code extension time slot 1 can contain p bits for defining the execution packet. In one example embodiment, as Fig.14 and 15 illustrated, bit 0 of the code extension time slot 0 1400 and the conditional code extension time slot 1 1500 can provide the p bit. Assuming that the p bits for the code extension time slots 1400, 1500 are always encoded as 1 (parallel execution), then neither the code extension time slot 1400 nor the conditional code extension time slot 1500 should be the last instruction time slot of the execution packet.

[0115] In one example embodiment of the processor 100, there are two different constant extension time slots. Each execution packet may contain each of these unique 32-bit constant extension time slots, each of which contains 27 bits to be concatenated with a 5-bit constant field in instruction decoding 1300 as high-order bits to form a 32-bit constant. As mentioned in the instruction decoding 1300 description above, only some instructions define the 5-bit src2 / cst field 1304 as a constant rather than a source register identifier. At least some of those instructions may use the constant extension time slot to extend this constant to 32 bits.

[0116] Fig.16 Illustrate an example decoding 1600 of constant extension time slot 0. Each execution packet may contain an instance of constant extension time slot 0 and an instance of constant extension time slot 1. Fig.16 Illustrate constant extension time slot 0 1600 that includes two fields. Field 1601 (bits 5 to 31) constitutes the most significant 27 bits of the extended 32-bit constant having the destination instruction scr2 / cst field 1304 to provide the five least significant bits. Field 1602 (bits 0 to 4) is decoded a unique set of bits (CSTX0) to identify constant extension time slot 0. In an example embodiment, constant extension time slot 0 1600 is used to extend a constant of one of the L1 unit 221 instructions, data in the D1 unit 225 instructions, S2 unit 242 instructions, offset in the D2 unit 226 instructions, M2 unit 243 instructions, N2 unit 244 instructions, branch instructions, or C unit 245 instructions in the same execution packet. Constant extension time slot 1 is similar to constant extension time slot 0 except that bits 0 to 4 are decoded a unique set of bits (CSTX1) to identify constant extension time slot 1. In an example embodiment, constant extension time slot 1 is used to extend a constant of one of the L2 unit 241 instructions, data in the D2 unit 226 instructions, S1 unit 222 instructions, offset in the D1 unit 225 instructions, M1 unit 223 instructions, or N1 unit 224 instructions in the same execution packet.

[0117] Constant extension time slot 0 and constant extension time slot 1 are used as follows. The target instruction must be of a type that permits constant specification. As is known in the art, this is implemented by replacing an input operand register specification field with the least significant bits of the constant as described above with respect to the scr2 / cst field 1304. The instruction decoder 113 determines this from the instruction opcode bits, which is referred to as the immediate field. The target instruction also includes a constant extension bit (e-bit 1307) that is dedicated to signaling that the specified constant is not extended (e.g., constant extension bit = 0) or that the constant is extended (e.g., constant extension bit = 1). If the instruction decoder 113 detects constant extension time slot 0 or constant extension time slot 1, then it further examines the other instructions within the execution packet for the instruction corresponding to the detected constant extension time slot. Constant extension is performed when the corresponding instruction has a constant extension bit (e-bit 1307) equal to 1.

[0118] Fig.17 FIG. 1700 is a block diagram of constant extension logic that can be implemented in the processor 100. Fig.17 Assume that the instruction decoder 113 detects a constant extension time slot and the corresponding instruction within the same execution packet. The instruction decoder 113 supplies 27 extension bits (bit field 1601) from the constant extension time slot and 5 constant bits (bit field 1304) from the corresponding instruction to the concatenator 1701. The concatenator 1701 forms a single 32-bit word from these two parts. In the illustrated embodiment, the 27 extension bits (bit field 1601) from the constant extension time slot are the most significant bits, and the 5 constant bits (bit field 1305) are the least significant bits. This combined 32-bit word is supplied to one input of the multiplexer 1702. The 5 constant bits from the corresponding instruction field 1305 supply a second input to the multiplexer 1702. The selection of the multiplexer 1702 is controlled by the state of the constant extension bit. If the constant extension bit (e-bit 1307) is 1 (extended), then the multiplexer 1702 selects the concatenated 32-bit input. If the constant extension bit is 0 (not extended), then the multiplexer 1702 selects the 5 constant bits from the corresponding instruction field 1305. The multiplexer 1702 supplies this output to the input of the sign extension unit 1703.

[0119] The sign extension unit 1703 forms the final operand value from the input from the multiplexer 1703. The sign extension unit 1703 receives control input scalar / vector and data size. The scalar / vector input indicates whether the corresponding instruction is a scalar instruction or a vector instruction. In this embodiment, the functional units (L1 unit 221, S1 unit 222, M1 unit 223, N1 unit 224, D1 unit 225, and D2 unit 226) on the data path side A 115 may be limited to performing scalar instructions. Any instruction involving one of these functional units is a scalar instruction. The data path side B functional units L2 unit 241, S2 unit 242, M2 unit 243, N2 unit 244, and C unit 245 may perform scalar instructions or vector instructions. The instruction decoder 113 determines whether the instruction is a scalar instruction or a vector instruction based on the opcode bits. In this embodiment, the P unit 246 may only perform scalar instructions. The data size may be 8 bits (byte B), 16 bits (half word H), 32 bits (word W), 64 bits (double word D), quad word (128 bits) data, or half vector (256 bits) data.

[0120] Table 2 lists the operations of the sign extension unit 1703 for various options.

[0121] Table 2

[0122]

[0123] As described above in connection with Fig.13 As described, both the constant extension slot 0 and the constant extension slot 1 may contain p bits for qualifying the execution packet. In one example embodiment, as in the case of the condition code extension slot, bit 0 of the constant extension slot 0 and the constant extension slot 1 may provide the p bit. Assuming that the p bits of the constant extension slot 0 and the constant extension slot 1 are always encoded as 1 (parallel execution), then neither the constant extension slot 0 nor the constant extension slot 1 should be in the last instruction slot of the execution packet.

[0124] In some embodiments, an execution packet may include a constant extended slot 0 or 1 and more than one corresponding instruction flag with constant extension (e-bit = 1). For constant extended slot 0, this would mean that more than one of the L1 unit 221 instructions, the data in the D1 unit 225 instructions, the S2 unit 242 instructions, the offset in the D2 unit 226 instructions, the M2 unit 243 instructions, or the N2 unit 244 instructions in the execution packet have an e-bit of 1. For constant extended slot 1, this would mean that more than one of the L2 unit 241 instructions, the data in the D2 unit 226 instructions, the S1 unit 222 instructions, the offset in the D1 unit 225 instructions, the M1 unit 223 instructions, or the N1 unit 224 instructions in the execution packet have an e-bit of 1. In this case, in one embodiment, the instruction decoder 113 may determine this situation as an invalid and unsupported operation. In another embodiment, this combination may be supported by the extended bits of the constant extended slot for the constant extension of each corresponding functional unit instruction flag.

[0125] Special vector assertion instructions use registers in the assertion register file 234 to control vector operations. In the current embodiment, all SIMD vector assertion instructions operate on a selected data size. The data size may include byte (8-bit) data, half-word (16-bit) data, word (32-bit) data, double-word (64-bit) data, quad-word (128-bit) data, and half-vector (256-bit) data. Each bit of the assertion register controls whether a SIMD operation is performed on the corresponding byte of the data. The operation of the P unit 246 permits various composite vector SIMD operations based on more than one vector comparison. For example, range determination may be performed using two comparisons. A candidate vector is compared to a first vector reference having the minimum value of the range packed in a first data register. The candidate vector is compared to a second reference vector having the maximum value of the range packed in a second data register. The logical combination of the two resulting assertion registers will permit vector conditional operations to determine whether each data portion of the candidate vector is within the range or outside the range.

[0126] The L1 unit 221, S1 unit 222, L2 unit 241, S2 unit 242, and C unit 245 typically operate in a single instruction multiple data (SIMD) mode. In this SIMD mode, the same instruction is applied to packed data from two operands. Each operand holds multiple data elements placed in predetermined slots. SIMD operations are enabled by carry control at the data boundaries. This carry control enables operations on different data widths.

[0127] Fig.18Describe carry control. AND gate 1801 receives the carry output of bit N within the operand-wide arithmetic logic unit (64 bits for scalar data path side A 115 functional unit and 512 bits for vector data path side B 116 functional unit). AND gate 1801 also receives a carry control signal that will be further described below. The output of AND gate 1801 is supplied to the carry input of bit N+1 of the operand-wide arithmetic logic unit. For example, AND gates of AND gate 1801 are placed between each pair of bits at possible data boundaries. For example, for 8-bit data, such AND gates will be between bits 7 and 8, bits 15 and 16, bits 23 and 24, etc. Each such AND gate receives a corresponding carry control signal. If the data size is at a minimum, then each carry control signal is 0, effectively blocking carry transmission between adjacent bits. If the selected data size requires two arithmetic logic unit segments, then the corresponding carry control signal is 1. Table 3 below shows example carry control signals in the case of a 512-bit wide operand such as used by vector data path side B 116 functional unit, which operand can be divided into segments of 8 bits, 16 bits, 32 bits, 64 bits, 128 bits, or 256 bits. In Table 3, the upper 32 bits control the carry for the upper part (bits 128 to 511), and the lower 32 bits control the carry for the lower part (bits 0 to 127). The carry output of the most significant bit does not need to be controlled, so only 63 carry control signals are required.

[0128] Table 3

[0129]

[0130] Typically operates on data sizes that are integer powers of 2 (2 N ). However, this carry control technique is not limited to integer powers of 2. Those skilled in the art will understand how to apply this technique to other data sizes and other operand widths.

[0131] Processor 100 includes dedicated instructions for performing table lookup operations implemented via one of D1 unit 225 or D2 unit 226. The tables for these table lookup operations are mapped to memory directly addressable by the level 1 data cache 123. An example of such a table / cache configuration is described in U.S. Patent No. 6,606,686, titled "Unified Memory System Architecture Incorporating a Cache and Directly Addressable Static Random Access Memory". These tables can be loaded into the memory space containing the tables via normal memory operations by a dedicated LUTINIT instruction (described below) or just ordinary store instructions, e.g., via a direct memory access (DMA) port. In one example embodiment, processor 100 supports up to 4 separate sets of parallel lookup tables, and within a set, up to 16 tables can be looked up in parallel with byte, half-word, or word element sizes. In this embodiment, at least a portion of the level 1 data cache 123 dedicated to directly addressable memory has 16 sets. This permits parallel access to 16 memory locations and supports up to 16 tables per table set.

[0132] These lookup tables are accessed with independently specified base and index addresses. The lookup table base register (LTBR) specifies the base address for each set of parallel tables. Each lookup table instruction contains a set number identifying which base register will be used for that instruction. Based on the use of the directly addressable portion of the level 1 data cache 123, each base address should be aligned with the cache line size of the level 1 data cache 123. In one embodiment, the cache line size can be 128 bytes.

[0133] The set of lookup table configuration registers for each set of parallel tables controls the information going to the corresponding table set. Fig.19 Describe the data fields of the example lookup table configuration register 1900. The promotion field 1901 (bits 24 and 25) sets the type of promotion immediately after an element is stored into a vector register. The promotion field 1901 is decoded as shown in Table 4.

[0134] Table 4

[0135] promote describe 00 No improvement 01 2x boost 10 4x Boost 11 8x improvement

[0136] Decoding of field 1901 with a value of 00 indicates no promotion. Decoding of field 1901 with a value of 01 indicates 2x promotion. Decoding of field 1901 with a value of 10 indicates 4x promotion. Decoding of field 1901 with a value of 11 indicates 8x promotion. In one example embodiment, the promoted data is limited to a data size of at most double-word size (64 bits). Thus, in such an embodiment: 2x promotion is effective for data element sizes of byte (promoted from byte (8 bits) to half-word (16 bits)), half-word (promoted from half-word (16 bits) to word (32 bits)), and word (promoted from word (32 bits) to double-word (64 bits)); 4x promotion is effective for data element sizes of byte (promoted from byte (8 bits) to word (32 bits)) and half-word (promoted from half-word (16 bits) to double-word (64 bits)); and 8x promotion is effective only for data element sizes of byte (promoted from byte (8 bits) to double-word (64 bits)). Promotion will be described below.

[0137] The table size field 1902 (bits 16 to 23) sets the table size. The table size field 1902 is decoded as shown in Table 5.

[0138] Table 5

[0139]

[0140] The table base address stored in the corresponding lookup table base address register must be aligned with the table size specified in the lookup table configuration register.

[0141] The weight size (WSIZE) field 1903 (bits 11 to 13) indicates the size of the weight value in the source register used for weighted histogram operations, which is described in further detail below. The weight size field 1903 is decoded as shown in Table 6.

[0142] Table 6

[0143]

[0144] The interpolation field 1904 (bits 8 to 10) indicates the number of consecutive elements written to the destination register in response to a lookup table read instruction, which is described in further detail below. The interpolation field 1904 is decoded as shown in Table 7.

[0145] Table 7

[0146] interpolate describe 000 No interpolation, just write the indexed elements of each table 001 Return 2 elements per table 010 Return 4 elements per table 011 Return 8 elements per table 100-111 reserve

[0147] The interpolation field 1904 decoded as 000 indicates no interpolation, where only the indexed elements of each table are written from the lookup table to the destination register. The interpolation field 1904 decoded as 001 indicates interpolating 2 elements, where the indexed elements of each table and the additional neighboring elements are written from the lookup table to the destination register. The interpolation field 1904 decoded as 010 indicates interpolating 4 elements, where the indexed elements of each table and the additional 3 neighboring elements are written from the lookup table to the destination register. The interpolation field 1904 decoded as 011 indicates interpolating 8 elements, where the indexed elements of each table and the additional 7 neighboring elements are written from the lookup table to the destination register. In one example embodiment, the interpolation field 1904 (in combination with the number in the table number field 1908, described below) cannot exceed the maximum number of elements that can be returned by the L1D 123, which is 16 elements in one example. Thus, in such an embodiment, when the table number field 1908 indicates 16 tables, no interpolation is possible; when the table number field 1908 indicates 8 tables, the maximum of 2-element interpolation is possible; when the table number field 1908 indicates 4 tables, the maximum of 4-element interpolation is possible; and when the table number field 1908 indicates 2 tables or 1 table, the maximum of 8-element interpolation is possible.

[0148] The saturation (SAT) field 1905 (bit 7) indicates whether histogram bin entries are saturated to the minimum / maximum value in response to a histogram operation, which is described in further detail below. If the saturation field 1905 is 1, then the histogram bin entries are saturated to the minimum / maximum value of the element data type. For example, an unsigned byte is saturated to [0, 0xFF]; a signed byte is saturated to [0x80, 0x7F]; an unsigned halfword is saturated to [0, 0xFFFF]; a signed halfword is saturated to [0x8000, 0x7FFF]; an unsigned word is saturated to [0, 0xFFFF FFFF]; and a signed word is saturated to [0x8000 0000, 0x7FFF FFFF]. If the saturation field 1905 is 0, then the histogram bin entries are not saturated to the minimum / maximum value of the element data type, and instead will wrap around when incrementing or decrementing beyond the minimum or maximum value, respectively.

[0149] The signed field 1906 (bit 6) indicates whether the processor 100 treats the called lookup table elements as signed integers or unsigned integers. If the signed field 1906 is 1, then the processor 100 treats the lookup table elements as signed integers. If the signed field 1906 is 0, then the processor 100 treats the lookup table elements as unsigned integers.

[0150] The element size (ESIZE) field 1907 (bits 3 to 5) indicates the lookup table element size. The element size field 1907 is decoded as shown in Table 8.

[0151] Table 8

[0152]

[0153] The table number (NTBL) field 1908 (bits 0 to 2) indicates the number of tables to be searched in parallel. The table number field 1908 is decoded as shown in Table 9.

[0154] Table 9

[0155]

[0156] The lookup table enable register 2000 specifies the type of operations permitted for a particular set of tables. This is described in Fig. 20 and shown in Fig. 20 This example uses fields in a single register to control one of four sets of tables. Field 2001 controls table set 3. Field 2002 controls table set 2. Field 2003 controls table set 1. Field 2004 controls table set 0. The table enable fields 2001, 2002, 2003, and 2004 are decoded as shown in Table 10.

[0157] Table 10

[0158] Table Enable describe 00 No lookup table operation 01 Allow read operations 10 reserve 11 Allows read and write operations

[0159] If the table set field is 01, read operations from the corresponding lookup table base register and the corresponding lookup table configuration register are permitted. If the table set field is 11, read operations from the corresponding lookup table base register and the corresponding lookup table configuration register are permitted and write operations to the corresponding lookup table base register and the corresponding lookup table configuration register are permitted. If the table set field is 00, lookup table operations are not permitted.

[0160] Each lookup table instruction specifies a vector register as an operand. This vector register is considered a set of 32-bit lookup table indices for the lookup table operation corresponding to the instruction. Table 11 shows the decoding of the vector indices in the vector operand register based on the number of tables in the table set controlled by the table number field 1908 of the corresponding lookup table configuration register.

[0161] Table 11

[0162] Vector register bits index 1 table 2 tables 4 tables 8 tables 16 tables Vx[31:0] Index 0 efficient efficient efficient efficient efficient Vx[63:32] Index 1 - efficient efficient efficient efficient Vx[95:64] Index 2 - - efficient efficient efficient Vx[127:96] Index 3 - - efficient efficient efficient Vx[159:128] Index 4 - - - efficient efficient Vx[191:160] Index 5 - - - efficient efficient Vx[223:192] Index 6 - - - efficient efficient Vx[255:224] Index 7 - - - efficient efficient Vx[287:256] Index 8 - - - - efficient Vx[319:288] Index 9 - - - - efficient Vx[351:320] Index 10 - - - - efficient Vx[383:352] Index 11 - - - - efficient Vx[415:384] Index 12 - - - - efficient Vx[447:416] Index 13 - - - - efficient Vx[479:448] Index 14 - - - - efficient Vx[511:480] Index 15 - - - - efficient

[0163] Depending on the number of tables specified in the table number field 1908 of the corresponding look-up table configuration register 1900, the vector register bits specify various indexes. The address of the table element within the first table for look-up table operations is the base address stored in the base address register plus the index specified by bits 0 to 31 of the vector register value. The address of the table element within the second table (assuming at least two tables are specified by the table number field 1908) for look-up table operations is the base address stored in the base address register plus the index specified by bits 32 to 63 of the vector register. Similarly, the vector register specifies the offset for each of the specified tables.

[0164] Fig.21 Illustrates the look-up table organization when the corresponding look-up table configuration register 1900 specifies one table of the set of tables. The level 1 data cache 123 includes a portion of the directly addressable memory, including the portion 2111 below the look-up table, the look-up table, and the portion 2112 above the look-up table. The corresponding look-up table set base address register 2101 specifies the start of the table set. In Fig.21 the example, the table number field 1908 specifies a single table. The end of the memory allocated to the table set is specified by the table size field 1902. The index source register 2102 specifies a single offset (index iA) for addressing the table. As Fig.21 shown, the single table spans all 16 groups of the memory.

[0165] Fig. 22 Illustrates the look-up table organization when the corresponding look-up table configuration register 1900 specifies two tables for the set of tables. The look-up table set base address register 2101 specifies the start of the table set. The table number field 1908 specifies two tables. The end of the memory allocated to the table set is specified by the table size field 1902. The index source register 2102 specifies two offsets for addressing the tables. The first index iA addresses table 1 and the second index iB addresses table 2. The table 1 data is stored in memory groups 1 to 8. The table 2 data is stored in memory groups 9 to 16.

[0166] Fig.23 Illustrates the look-up table organization when the corresponding look-up table configuration register 1900 specifies four tables for the set of tables. The look-up table set base address register 2101 specifies the start of the table set. The table number field 1908 specifies four tables. The end of the memory allocated to the table set is specified by the table size field 1902. The index source register 2102 specifies four offsets for addressing the tables. The first index iA addresses table 1, the second index iB addresses table 2, the third index iC addresses table 3, and the fourth index iD addresses table 4. The table 1 data is stored in memory groups 1 to 4. The table 2 data is stored in memory groups 5 to 8. The table 3 data is stored in memory groups 9 to 12. The table 4 data is stored in memory groups 13 to 16.

[0167] Fig.24 Describe the lookup table organization when the corresponding lookup table configuration register 1900 specifies eight tables for the set of tables. The lookup table set base address register 2101 specifies the start of the table set. The table number field 1908 specifies eight tables. The end of the memory allocated to the table set is specified by the table size field 1902. The index source register 2102 specifies eight offsets for addressing the tables. The first index iA addresses table 1, the second index iB addresses table 2, the third index iC addresses table 3, the fourth index iD addresses table 4, the fifth index iE addresses table 5, the sixth index iF addresses table 6, the seventh index iG addresses table 7, and the eighth index iH addresses table 8. Table 1 data is stored in memory banks 1 and 2. Table 2 data is stored in memory banks 3 and 4. Table 3 data is stored in memory banks 5 and 6. Table 4 data is stored in memory banks 7 and 8. Table 5 data is stored in memory banks 9 and 10. Table 6 data is stored in memory banks 11 and 12. Table 7 data is stored in memory banks 13 and 14. Table 8 data is stored in memory banks 15 and 16.

[0168] Fig.25Describe the lookup table organization when the corresponding lookup table configuration register 1900 designates sixteen tables for the set of tables. The lookup table set base register 2101 designates the start of the table set. The table number field 1908 designates sixteen tables. The end of the memory allocated to the table set is designated by the table size field 1902. The index source register 2102 designates sixteen offsets for addressing the tables. The first index iA addresses table 1, the second index iB addresses table 2, the third index iC addresses table 3, the fourth index iD addresses table 4, the fifth index iE addresses table 5, the sixth index iF addresses table 6, the seventh index iG addresses table 7, the eighth index iH addresses table 8, the ninth index iI addresses table 9, the tenth index iJ addresses table 10, the eleventh index iK addresses table 11, the twelfth index iL addresses table 12, the thirteenth index iM addresses table 13, the fourteenth index iN addresses table 14, the fifteenth index iO addresses table 15, and the sixteenth index iP addresses table 16. Table 1 data is stored in memory bank 1. Table 2 data is stored in memory bank 2. Table 3 data is stored in memory bank 3. Table 4 data is stored in memory bank 4. Table 5 data is stored in memory bank 5. Table 6 data is stored in memory bank 6. Table 7 data is stored in memory bank 7. Table 8 data is stored in memory bank 8. Table 9 data is stored in memory bank 9. Table 10 data is stored in memory bank 10. Table 11 data is stored in memory bank 11. Table 12 data is stored in memory bank 12. Table 13 data is stored in memory bank 13. Table 14 data is stored in memory bank 14. Table 15 data is stored in memory bank 15. Table 16 data is stored in memory bank 16.

[0169] The following is the form of a lookup table read (LUTRD) instruction in an example embodiment.

[0170] LUTRD tbl_index,tbl_set,dst

[0171] Tbl_index is the instruction operand that designates a vector register (e.g., within the general vector register file 231) by register number. This is interpreted as the index number shown in Table 11. Tbl_set is the number [0:3] that designates the table set to be used in the instruction. This named table set number designates: the corresponding lookup table base register that stores the table base address, which can be a scalar register or a vector register; the corresponding lookup table configuration register ( Fig.19 ), which can be a scalar register or a vector register; and the lookup table enable register ( Fig. 20) corresponding operation part, which can be a scalar register or a vector register. The lookup table base address register corresponding to the named table set determines the base address of the table set. The index of the vector register named by Tbl_index is the offset from this table set base address. The lookup table configuration register corresponding to the named table set determines: the promotion mode (Table 4); the amount of memory allocated to the table size (Table 5); the weight size of the histogram operation (Table 6); the n-element interpolation of the lookup table read operation (Table 7); whether to treat the value as signed or unsigned; whether the histogram bin entry saturates to the minimum / maximum value; the data element size (Table 8); and the number of tables in the table set (Table 9). Dst is an instruction operand that designates a vector register (e.g., within the general vector register file 231) as the destination of the table lookup operation by register number. The data packed from the table called as specified by these other parameters is stored in this destination register. The promotion process adds no performance penalty.

[0172] Fig.26 An example illustrating the operation of the lookup table read instruction of the present invention. In Fig.26 In the example illustrated in, the corresponding lookup table configuration register (1900) specifies four parallel tables, a data element size of a byte (8 bits), and no promotion. For a lookup table operation, the lookup table enable register field (in register 2000) corresponding to the selected table set must enable either the read operation (01) or both the read and write operations (11).

[0173] The lookup table base address register 2101 corresponding to the specified table set stores the base address of the lookup table set, as Fig.26 schematically illustrated in. The table data is stored in a portion of the level 1 data cache 123 configured for direct access to memory, as described, for example, in U.S. Patent No. 6,606,686 titled "Unified Memory System Architecture Incorporating a Cache and Directly Addressable Static Random Access Memory".

[0174] Fig.26The example described herein has four tables: Table 1 2611; Table 2 2612; Table 3 2613; and Table 4 2614. As shown in Table 11, this instruction with these selected options treats the data stored in vector source index register 2102 as a collection of 4 32-bit fields specifying table offsets. The first field (bits Vx[31:0]) stores index iA into the first table 2611. In this example, this indexes to element A5. The second field (bits Vx[63:32]) stores index iB into the second table 2612. In this example, this indexes to element B1. The third field (bits Vx[95:64]) stores index iC into the third table 2613. In this example, this indexes to element C8. The fourth field (bits Vx[127:96]) stores index iD into the fourth table 2614. In this example, this indexes to element D10. The various indexes are memory address offsets from the base address for the table set to the specified data element. In the operation of this lookup table read instruction, the indexed data element in Table 1 2611 (A5) is stored in the first data slot in destination register 2120. The indexed data element in Table 2 2612 (B1) is stored in the second data slot in destination register 2120. The indexed data element in Table 3 2613 (C8) is stored in the third data slot in destination register 2120. The indexed data element in Table 4 2614 (D10) is stored in the fourth data slot in destination register 2120. In this example implementation, the other data slots in destination register 2120 are filled with zeros.

[0175] The lookup table read instruction directly maps the data called from the table to the vector lanes of destination register 2120. The instruction maps earlier elements to lower lane numbers and later elements to higher lane numbers. The lookup table read instruction deposits elements in the vector in increasing lane order. The lookup table read instruction fills each vector lane of destination register 2120 with the elements called from the table. If the data called is not equal to the vector length, then the lookup table read instruction fills the excess lanes of destination register 2120 with zeros.

[0176] When boost mode is enabled (boost field 1901 of corresponding lookup table configuration register 1900 ≠ 00), the lookup table read instruction boosts each called data element to a larger size. Fig. 27 An example illustrating the operation of the lookup table read instruction in the example implementation. Except that 2x boost is enabled (boost field 1901 is 01), Fig. 27 illustrates an operation similar to Fig.26 As in Fig.26 each of the four indexes in vector source index register 2102 calls a data element from the corresponding table. Fig. 27It is described that these are placed in the destination register 2120 in a different way from that in Fig.26 Each data element of each call, together with an equally sized extension, is stored in a time slot in the destination register 2120. This extension is formed corresponding to the signed / unsigned indication of the signed field 1906 of the corresponding lookup table configuration register. If the signed field 1906 indicates unsigned (0), the extension is filled with zeros. If the signed field 1906 indicates signed (1), the extension time slot is filled with the same value as the most significant bit (sign bit) of the corresponding data element. This treats the data element as a signed integer.

[0177] Fig.28 An example of the operation of the lookup table read instruction in the example embodiment is described. Except for enabling 4x boosting (boost field 1901 is 10), Fig.28 It is described that it is similar to Fig.26 Each of the four indexes of the vector source index register 2102 calls a data element from each table. Fig.28 It is described that these are placed in the destination register 2120 in a different way from that in Fig.26 Each data element of each call, together with three equally sized extensions, is stored in a time slot in the destination register 2120. These extensions are formed corresponding to the signed / unsigned indication of the signature field 1906 of the corresponding lookup table configuration register.

[0178] Those skilled in the art will understand that other data element sizes (e.g., half - word, word) will be implemented similarly. Additionally, other boosting factors, such as a boosting factor of 8x, will be implemented similarly. Those skilled in the art will understand how to apply the principles described in this specification to other numbers of lookup tables within a selected set of tables.

[0179] Fig.29A and 29BExample embodiments that collectively illustrate the enhanced implementation. The temporary register 2950 receives table data called from the level-1 data cache 123. The temporary register 2950 contains 16 bytes arranged in 16 one-byte block paths from path 1 to path 16. It should be noted that the length of each of these paths is equal to the minimum of the data sizes that can be specified in the element size field 1907. In this example, the minimum is 1 byte / 8 bits. The extended elements 2901 to 2908 form extensions to the respective paths 1 to 8. The plurality of multiplexers 2932 to 2946 couple the input paths from the temporary register 2950 to the corresponding paths of the destination register 2120. Not all input paths of the temporary register 2950 are coupled to each of the multiplexers 2932 to 2946. The plurality of multiplexers 2932 to 2946 also receive extended inputs and one or more of the extended elements 2901 to 2908. It should be noted that there is no multiplexer that supplies path 1 of the output register 2120. In the illustrated embodiment, path 1 of the destination register 2120 is always supplied by path 1 of the temporary register 2950.

[0180] Fig.30 Description Fig.29A of the exemplary extended element N. The sign bit (S) of the data element N of the temporary register 2950 supplies one input of the corresponding extended multiplexer 3001. The sign bit of a signed integer is as Fig.30The most significant bit of the value shown in . The constant 0 supplies the second input of multiplexer 3001. Multiplexer 3001 and other similar multiplexers corresponding to other input data elements are controlled by a signed / unsigned signal. This signed / unsigned signal is based on the signed field 1906 of the corresponding lookup table configuration register 1900. If the signed field 1906 is 0, then multiplexer 3001 (and the corresponding multiplexers for other input paths) selects the constant 0 input. If the signed field 1906 is 1, then multiplexer 3001 (and the corresponding multiplexers for other input paths) selects the sign bit input. The selected extension is supplied to expansion element 3002. Expansion element 3002 expands the bit selected by multiplexer 3001 to the path size. In this example, the path size equal to the minimum table data element size of 1 byte is selected. For a specified table data size equal to the path size and where the signed field 1906 is 0, the next data time slot is filled with 0 to achieve zero extension. If the signed field 1906 is 1, then the next data time slot is filled with the sign bit of the data element to achieve sign extension. Multiplexers 2932 to 2946 select the appropriate extension corresponding to the specified table data element size. If the selected table data element size is a half-word, then multiplexers 2932 to 2946 are controlled to select the extension from the alternate expansion elements. If the selected table data element size is a word, then multiplexers 2932 to 2946 are controlled to select the extension from every fourth expansion element.

[0181] Multiplexers 2932 to 2946 are controlled by Fig.31 the multiplexer control encoder 3110 illustrated in . The multiplexer control encoder 3110 receives the element data size input (element size field 1907), the promotion indication (promotion field 1901) and generates the corresponding control signals for multiplexers 2932 to 2946. Not all input bytes can supply each output byte. Table 12 illustrates this control. Table 12 shows the source data for each of the 16 paths of destination register 2120 for various data sizes and promotion modes in the example embodiment.

[0182] Table 12

[0183] 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 --1x 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 b-2x e8 8 e7 7 e6 6 e5 5 e4 4 e3 3 e2 2 e1 1 b-4x e4 e4 e4 4 e3 e3 e3 3 e2 e2 e2 2 e1 e1 e1 1 b-8x e2 e2 e2 e2 e2 e2 e2 2 e1 e1 e1 e1 e1 e1 e1 1 hw-2x e8 e8 8 7 e6 e6 6 5 e4 e4 4 3 e2 e2 2 1 hw-4x e4 e4 e4 e4 e4 e4 4 3 e2 e2 e2 e2 e2 e2 2 1 w-2x e8 e8 e8 e8 8 7 6 5 e4 e4 e4 e4 4 3 2 1

[0184] It should be noted that, regardless of the selected data size or promotion factor, path 1 of destination register 2120 is always the same as path 1 of temporary register 2950. The column dedicated to path 1 contains all 1s and Fig.29ADescribe the direct connection between path 1 of the temporary register 2950 and the destination register 2120. The first row of Table 12 shows that for no boost (1x), regardless of the data size of the selected table data element, each path of the destination register 2120 is the same as that of the temporary register 2950. For a data size of bytes and a boost factor of 2 (b-2x), the multiplexer 2932 selects the extension (e1) of path 1, the multiplexer 2933 selects the data of input path 2, the multiplexer 2934 selects the extension (e2) of path 2, the multiplexer 2935 selects the data of input path 3, the multiplexer 2936 selects the extension (e3) of path 3, and so on. For a data size of bytes and a boost factor of 4 (b-4x), the multiplexers 2932 to 2934 select the extension (e1) of path 1, the multiplexer 2935 selects the data of input path 2, and the multiplexers 2936 to 2938 select the extension (e2) of path 2, the multiplexer 2939 selects the data of input path 3, the multiplexers 2940 to 2942 select the extension (e3) of path 3, and so on. For a data size of bytes and a boost factor of 8 (b-8x), the multiplexers 2932 to 2938 select the extension (e1) of path 1, the multiplexer 2939 selects the data of input path 2, and the multiplexers 2940 to 2446 select the extension (e2) of path 2. For a data size of half-words and a boost factor of 2 (hw-2x), the multiplexer 2932 selects the data of path 2, the multiplexers 2933 and 2934 select the extension (e2) of path 2, the multiplexer 2935 selects the data of path 3, the multiplexer 2936 selects the data of path 4, the multiplexers 2937 and 2938 select the extension (e4) of path 4, and so on. For a data size of half-words and a boost factor of 4 (hw-4x), the multiplexer 2932 selects the data of path 2, the multiplexers 2933 to 2938 select the extension (e2) of path 2, the multiplexer 2939 selects the data of path 3, the multiplexer 2940 selects the data of path 4, and the multiplexers 2941 and 2946 select the extension (e4) of path 4. For a data size of words and a boost factor of 2 (w-2x), the multiplexer 2932 selects the data of path 2, the multiplexer 2933 selects the data of path 3, the multiplexer 2934 selects the data of path 4, the multiplexers 2935 to 2938 select the extension (e4) of path 4, the multiplexer 2939 selects the data of path 5, the multiplexer 2940 selects the data of path 6, the multiplexer 2941 selects the data of path 7, the multiplexer 2942 selects the data of path 8, and the multiplexers 2943 to 2946 select the extension of path 8. As previously described, these are all combinations of data sizes and boost factors supported in this example.

[0185] Fig.32Another example of the operation of the look-up table read instruction of this specification. In Fig.32 In the example described, the corresponding look-up table configuration register (1900) specifies four parallel tables, a data element size of a word (32 bits), and 2-element interpolation. For a look-up table operation, the look-up table enable register field corresponding to the selected table set (in register 2000) must enable either a read operation (01) or both a read and write operation (11).

[0186] As described above, the look-up table base address register 2101 corresponding to the specified table set stores the base address of the look-up table set, as Fig.32 Schematically illustrated in. The table data is stored in a portion of the level 1 data cache 123, which is configured as a directly addressable memory.

[0187] Fig.32 The example described in has four tables: Table 1 2611; Table 2 2612; Table 3 2613; and Table 4 2614. As shown in Table 11, this instruction with these selected options treats the data stored in the vector source index register 2102 as a set of 4 32-bit fields specifying the table offset. The first field (bits Vx[31:0]) stores the index iA into the first table 2611. In this example, this indexes to element A5. The second field (bits Vx[63:32]) stores the index iB into the second table 2612. In this example, this indexes to element B1. The third field (bits Vx[95:64]) stores the index iC into the third table 2613. In this example, this indexes to element C8. The fourth field (bits Vx[127:96]) stores the index iD into the fourth table 2614. In this example, this indexes to element D10. The various indexes are memory address offsets from the base address for the table set to the specified data element.

[0188] When the interpolation mode is enabled (interpolation field 1904 of the corresponding look-up table configuration register 1900 ≠ 00), the look-up table read instruction returns one or more additional data elements compared to the indexed elements of each table. Fig.32 An example of the operation of the look-up table read instruction in an example embodiment. In addition to enabling 2-element interpolation (interpolation field 1904 is 001), Fig.32 The operation is described as being similar to that described above Fig.26 As in Fig.26 Each of the four indexes in the vector source index register 2102 calls a data element from the corresponding table. Fig.32 These are described in the same way as Fig.26are placed in the destination register 2120 in different ways. Specifically, for each of the four indices, 2 data elements are returned: the data element called by the specific index and the next adjacent data element (without considering row boundaries). For example, if the data element called by the specific index is A0, then the next adjacent data element is A1; if the data element called by the specific index is A7, then the next adjacent data element is A8, but A8 is on a different row.

[0189] Each called data element is stored in a time slot in the destination register 2120. In the operation of this lookup table read instruction that enables 2-element interpolation, the indexed data element in Table 1 2611 (A5) and its next adjacent data element (A6) are stored in the first and second data time slots in the destination register 2120. The indexed data element in Table 2 2612 (B1) and its next adjacent data element (B2) are stored in the third and fourth data time slots in the destination register 2120. The indexed data element in Table 3 2613 (C8) and its next adjacent data element (C9) are stored in the fifth and sixth data time slots in the destination register 2120. The indexed data element in Table 4 2614 (D10) and its next adjacent data element (D11) are stored in the seventh and eighth data time slots in the destination register 2120. In this example implementation, the other data time slots of the destination register 2120 are filled with zeros.

[0190] Fig.33 An example illustrating the operation of the lookup table read instruction in the example embodiment. Except for enabling 4-element interpolation (the interpolation field 1904 is 010), Fig.33 illustrates an operation similar to Fig.32 As in Fig.32 each of the four indices in the vector source index register 2102 calls a data element from each table. Fig.33 illustrates that these are placed in the destination register 2120 in a different way from Fig.32 Specifically, for each of the four indices, 4 data elements are returned: the data element called by the specific index and the next 3 adjacent data elements (without considering row boundaries). For example, if the data element called by the specific index is A0, then the next 3 adjacent data elements are A1 to A3; if the data element called by the specific index is A7, then the next 3 adjacent data elements are A8 to A10, but A8 to A10 are on different rows.

[0191] Each data element of each call is stored in a time slot in destination register 2120. In the operation of this lookup table read instruction with 4-element interpolation enabled, the indexed data element in Table 1 2611 (A5) and its 3 neighboring data elements below (A6 to A8) are stored in the first 4 data time slots in destination register 2120. The indexed data element in Table 2 2612 (B1) and its 3 neighboring data elements below (B2 to B4) are stored in the second 4 data time slots in destination register 2120. The indexed data element in Table 3 2613 (C8) and its 3 neighboring data elements below (C9 to C11) are stored in the third 4 data time slots in destination register 2120. The indexed data element in Table 4 2614 (D10) and its 3 neighboring data elements below (D11 to D13) are stored in the fourth 4 data time slots in destination register 2120. In this example, with the data element size being a word, 16 call data elements completely fill destination register 2120; however, in other similar examples where the data element size is actually a half-word or a byte, 16 call data elements will not completely fill destination register 2120, and the remaining part of the data time slots in destination register 2120 is filled with zeros.

[0192] Those skilled in the art will understand how to apply the principles described in this specification to other numbers of lookup tables within a selected set of tables. For example, when the number of tables is 4 and the tables are arranged in the first-level data cache 123 in the manner shown in Fig.33 8-element interpolation is not possible because only one element of each memory bank can be accessed at a time. However, when the number of tables is 2 or 1, 8-element interpolation is possible.

[0193] Fig.34 An example embodiment illustrating the interpolation implementation described above. In the Fig.34 example, the intermediate register 3402 receives the table data called from the first-level data cache 123. In this example, the first-level data cache 123 containing Tables 2611, 2612, 2613, 2614 is configured to provide a data word from each of its 16 banks (described above) regardless of the data requested by the lookup read instruction. Without considering interpolation or other variables specified by the lookup table configuration register 1900, the data provided to the intermediate register 3402 is sorted when it appears in the 16 banks of the first-level data cache 123.

[0194] Continue Fig.33An example where the lookup table read instruction includes 4-element interpolation, and the intermediate register 3402 contains each of four indexed data elements (e.g., A5, B1, C8, D10) and the next 3 neighboring data elements from each of the indexed data elements. However, as described above, the intermediate register 3402 contains these data elements sorted when it appears in the 16 sets of the level-1 data cache 123, and thus not necessarily in the numerical order of the data elements, as shown in the destination register 2120 in Fig.33 above.

[0195] For example, the four elements from Table 1 2611 (A5 to A8) appear in the order A8, A5, A6, A7 because that is the order in which those elements appear in the first four sets of the level-1 data cache 123. On the other hand, the four elements from Table 3 2613 (C8 to C11) appear in the intermediate register 3402 in sequence because the element C8 is already in the lowest set in the level-1 data cache 123 corresponding to Table 3 2613.

[0196] To facilitate the proper sorting of the data elements written to the destination register 2120 (described above with respect to Fig.33 ), the examples of this specification include a multi-stage butterfly unit 3404 that receives the intermediate register 3402 as an input and produces an output written to the destination register 2120. The butterfly unit 3404 is configured to reorder the bits of the intermediate register 3402 according to various control signals. These control signals include (but are not necessarily limited to) all or a part of the addresses of the indexed data elements in each of Tables 2611, 2612, 2613, 2614. For example, decoding the index 3401 (or a part of the address of the indexed data element) indicates the position of the indexed data element within the table and thus whether reordering may be required. In this example, the address of the indexed data element A5 indicates that A5 is not aligned with the lowest set in the level-1 data cache 123 corresponding to Table 1 2611, and thus the bits of the intermediate register 3402 corresponding to Table 1 2611 should be reordered. On the other hand, the address of the indexed data element C8 indicates that C8 is aligned with the lowest set in the level-1 data cache 123 corresponding to Table 3 2613, and thus the bits of the intermediate register 3402 corresponding to Table 3 2613 do not need to be reordered.

[0197] To further facilitate the proper sorting of the data elements written to the destination register 2120, the multi-stage butterfly unit 3404 also receives all or a part of the lookup table configuration register 1900 as control signals. Specifically, the interpolation field 1904, the element size field 1907, and the table number field 1908 can affect the function of the multi-stage butterfly unit 3404. Although Fig.34 It is intended to illustrate how, in the example shown in Fig.34 , the butterfly unit 3404 processes a set of data elements from the lookup tables 2611, 2612, 2613, 2614, but the butterfly unit 3404 can process the data elements in several different possible settings in the lookup table configuration register 1900. Additional examples of the functionality of the butterfly unit 3404 are provided below.

[0198] In a first example, the control signals provided to the multi-stage butterfly network 3404 from the lookup table configuration register 1900 include an element size field 1907 indicating the size of the word, an interpolation field 1904 indicating no interpolation, and a table number field 1908 indicating two tables. In this example, for instance, according to the two-table embodiment shown in Fig. 22 , the multi-stage butterfly network 3404 receives one indexed element from each of groups 1 to 8 and groups 9 to 16 of the first-level data cache 123. Thus, in response to these control signals from the lookup table configuration register 1900, the multi-stage butterfly network 3404 places the indexed elements from each of groups 1 to 8 and groups 9 to 16 into the first two word-size paths of the destination register 2120, respectively. In one example, the remaining paths of the destination register 2120 are filled with zeros, while in other examples, the remaining paths of the destination register can be filled in different ways.

[0199] In a second example, the control signals provided to the multi-stage butterfly network 3404 from the lookup table configuration register 1900 include an element size field 1907 indicating the size of the word, an interpolation field 1904 indicating 2-element interpolation, and a table number field 1908 indicating two tables. In this example, the multi-stage butterfly network 3404 receives one indexed element and the next adjacent element (e.g., according to the 2-element interpolation described above) from each of groups 1 to 8 and groups 9 to 16 of the first-level data cache 123. Thus, in response to these control signals from the lookup table configuration register 1900, the multi-stage butterfly network 3404 places the indexed elements and the next adjacent elements from the groups 1 to 8 in order into the first two word-size paths of the destination register 2120. Similarly, the multi-stage butterfly network 3404 places the indexed elements and the next adjacent elements from the groups 9 to 16 in order into the second two word-size paths of the destination register 2120. In one example, the remaining paths of the destination register 2120 are filled with zeros, while in other examples, the remaining paths of the destination register can be filled in different ways.

[0200] In a third example, the control signals provided to the multi - stage butterfly network 3404 from the lookup - table configuration register 1900 include an element - size field 1907 that indicates the size of the words, an interpolation field 1904 that indicates 8 - element interpolation, and a table - number field 1908 that indicates the number of tables. In this example, the multi - stage butterfly network 3404 receives one indexed element and the next 7 adjacent elements (e.g., according to 8 - element interpolation) from each group of groups 1 to 8 and groups 9 to 16 of the first - level data cache 123. As described above, the next 7 adjacent elements can wrap around to subsequent rows. Thus, in response to these control signals from the lookup - table configuration register 1900, the multi - stage butterfly network 3404 places the indexed element and the next 7 adjacent elements from the groups of groups 1 to 8 in order in the first eight word - size paths of the destination register 2120. Similarly, the multi - stage butterfly network 3404 places the indexed element and the next 7 adjacent elements from the groups of groups 9 to 16 in order in the second eight word - size paths of the destination register 2120. In this example, assuming a 512 - bit destination register 2120, all paths are filled by the read operation and thus no additional filling is required.

[0201] Extensions of the above functionality of the multi - stage butterfly network 3404 in response to different combinations of control signals from the lookup - table configuration register 1900 would be apparent to one of ordinary skill in the art. For example, a change in the table - number field 1908 affects the "boundaries" at which elements can be reordered before being written to the destination register 2120 (e.g., a table - number field 1908 indicating four tables creates "boundaries" between groups 1 to 4, 5 to 8, 9 to 12, and 13 to 16, while a table - number field 1908 indicating eight tables creates "boundaries" between groups 1 to 2, 3 to 4, 5 to 6, 7 to 8, 9 to 10, 11 to 12, 13 to 14, and 15 to 16). As another example, a change in the element - size field 1907 affects how bits from the intermediate register 3402 are packed into the destination register 2120.

[0202] The following is the form of a lookup - table write (LUTWR) instruction in an example embodiment.

[0203] LUTWR tbl_index,tbl_set,write_data

[0204] Tbl_index is an instruction operand that specifies a vector register (e.g., within the general vector register file 231) by register number. This is interpreted as the index number as shown in Table 11 above. Tbl_set is the number [0:3] that specifies the table set to be used in the instruction. This named table set number specifies: the corresponding lookup table base register that stores the table base address, which can be a scalar register or a vector register; the corresponding lookup table configuration register ( Fig.19 ), which can be a scalar register or a vector register; and the corresponding operation part of the lookup table enable register ( Fig. 20 ), which can be a scalar register or a vector register. The lookup table base register corresponding to the named table set determines the base address of the table set. The index of the vector register named by Tbl_index is the offset from this table set base address. The lookup table configuration register corresponding to the named table set determines: the promotion mode (Table 4); the amount of memory allocated to the table size (Table 5); the weight size of the histogram operation (Table 6); the n-element interpolation for the lookup table read operation (Table 7); whether the value is considered signed or unsigned; whether the histogram bin entries saturate to the minimum / maximum values; the data element size (Table 8); and the number of tables in the table set (Table 9). Write_data is an instruction operand that specifies a vector register (e.g., within the general vector register file 231) by register number to provide the source data to be written to the lookup table, specifically, the source operand. As above, the number of tables, the size of the elements, and other parameters are specified in the lookup table configuration register 1900.

[0205] Fig.35 An example illustrating the operation of the lookup table write instruction of this specification. In the example illustrated in Fig.35 , the corresponding lookup table configuration register (1900) specifies four parallel tables and a data element size of a word (32 bits). For the lookup table operation, the lookup table enable register field (in register 2000) corresponding to the selected table set must enable both read and write operations (11).

[0206] As above, the lookup table base register 2101 corresponding to the specified table set stores the base address of the lookup table set, as schematically illustrated in Fig.35 . The table data is stored in a part of the level 1 data cache 123, which is configured as a directly accessible memory.

[0207] Fig.35The example described herein has four tables: Table 1 2611; Table 2 2612; Table 3 2613; and Table 4 2614. As shown in Table 11, this instruction with these selected options treats the data stored in the vector source index register 2102 as a set of 4 32-bit fields that specify table offsets. The first field (bits Vx[31:0]) stores the index iA into the first table 2611. In this example, this indexes to element A5 (e.g., between A4 and A6). The second field (bits Vx[63:32]) stores the index iB into the second table 2612. In this example, this indexes to element B1 (e.g., between B0 and B2). The third field (bits Vx[95:64]) stores the index iC into the third table 2613. In this example, this indexes to element C8 (e.g., between C7 and C9). The fourth field (bits Vx[127:96]) stores the index iD into the fourth table 2614. In this example, this indexes to element D10 (e.g., between D9 and D11). The various indexes are memory address offsets from the base address for the table set to the specified data element to be written.

[0208] The vector source data register 3502 contains the indexed data elements of data to be written to the lookup tables 2611, 2612, 2613, 2614. In this example, because the lookup table configuration register specifies a table number of four and an element size of a word, the first four words of the vector source data register 3502 are utilized by the lookup table write instruction, and the remainder of the vector source data register 3502 is not utilized by the lookup table write instruction. Specifically, the first word (W1) is written to the element specified by index iA or A5; the second word (W2) is written to the element specified by index iB or B1; the third word (W3) is written to the element specified by index iC or C8; and the fourth word (W4) is written to the element specified by index iD or D10.

[0209] In response to different parameters specified by the lookup table configuration register 1900, the extensions of the above lookup table write instruction will be apparent to those of ordinary skill in the art. For example, a change in the table number field 1908 affects the number of indexes specified by the vector source index register 2102 (e.g., as shown in Figures 21 to 25 . Similarly, a change in the element size field 1907 affects the portion of the vector source data register 3502 that provides the source data of the indexed data elements to be written to the lookup table.

[0210] The following is the form of a lookup table initialization (LUTINIT) instruction in an example embodiment.

[0211] LUTINIT tbl_index,tbl_set,write_data

[0212] Tbl_index, tbl_set, and write_data are generally similar to the Tbl_index, tbl_set, and write_data described above with respect to the look-up table write instruction. As above, the number of tables, the size of the elements, and other parameters are specified in the look-up table configuration register 1900. Different from the look-up table write instruction that only writes the source data to each table ( Fig.35 the tables 2611, 2612, 2613, 2614 in the example), the look-up table initialization instruction copies the source data to more efficiently write to a larger portion of the look-up table. As indicated by the instruction name, this copying of the source data is particularly useful when initializing the elements of the look-up table to a specific value.

[0213] The look-up table initialization instruction specifies the source data for a single table (e.g., the vector source index register 2102 contains only a single index, e.g., in the first field (bits Vx[31:0])). The size of the source data specified by the vector source data register 3502 varies depending on the number of tables specified by the look-up table configuration register 1900. In response to the execution of the look-up table initialization instruction, the level 1 data cache 123 internally copies the source data based on the number of tables specified by the look-up table configuration register 1900, and writes the resulting copied data to the locations specified by the index values in the vector source index register 2102 and the look-up table base register 2101. The copied data is written to the location specified by the index value in the vector source index register 2102 and the corresponding locations in each of the other tables in the table set.

[0214] Fig.36A and 36B shows an example of the initialization of the level 1 data cache 123 resulting from the execution of an exemplary look-up table initialization instruction. In Fig.36A the example, the number of tables is 16 and the element size is a word. The following are exemplary look-up table initialization instructions whose execution produces an instance group 3604 of the level 1 data cache 123:

[0215] LUTINIT D0,0,B0; where D0 = 0x00 (e.g., the index of the element in row 0, level 1 data cache 123 set 0), 0 is the specified table set, B0 = D[63:0]

[0216] LUTINIT D1,0,B1; where D1 = 0x02 (e.g., the index of the element in row 1, level 1 data cache 123 set 0), 0 is the specified table set, B1 = D[127:64]

[0217] LUTINIT D2,0,B2; where D2 = 0x04 (e.g., the index of the element in row 2, set 0 of the level-1 data cache 123), 0 is the specified table set, and B2 = D[191:128]

[0218] LUTINIT D3,0,B3; where D3 = 0x06 (e.g., the index of the element in row 3, set 0 of the level-1 data cache 123), 0 is the specified table set, and B3 = D[255:192]

[0219] As described above, the size of the source data specified by the vector source data register 3502 varies depending on the number of tables specified by the lookup table configuration register 1900. In this example, the level-1 data cache has a width of 1024 bits and there are 16 tables, and thus each table is 64 bits wide. Registers B0 to B3 are scalar registers and also have a size of 64 bits, and thus the lookup table initialization instruction utilizes the entire content of registers B0 to B3. In another example, the source data register is the vector source data register 3502, and thus the lookup table initialization instruction only utilizes a part of the content of the register (e.g., the first 64 bits).

[0220] The first lookup table initialization instruction causes 64 bits of source data from register B0 to be written to each of the 16 tables in table set 0, starting at the position specified by index 0x00 (plus the base address from the lookup table configuration register 1900), where the index 0x00 is the first element in row 0. The second lookup table initialization instruction causes 64 bits of source data from register B1 to be written to each of the 16 tables in table set 0, starting at the position specified by index 0x02 (plus the base address from the lookup table configuration register 1900), where the index 0x02 is the first element in row 1. The third lookup table initialization instruction causes 64 bits of source data from register B2 to be written to each of the 16 tables in table set 0, starting at the position specified by index 0x04 (plus the base address from the lookup table configuration register 1900), where the index 0x04 is the first element in row 2. The fourth lookup table initialization instruction causes 64 bits of source data from register B3 to be written to each of the 16 tables in table set 0, starting at the position specified by index 0x06 (plus the base address from the lookup table configuration register 1900), where the index 0x06 is the first element in row 3.

[0221] As will be further described below, in Fig.36AIn an example, since there are 16 tables and the element size is a word, there are two words per row per table. The allowed index values (e.g., in registers D0 to D3) correspond to the first element of each row and thus also increase by two (e.g., 0x00, 0x02, 0x04, 0x06, etc.). If the element size is changed to a half-word, there will be four half-words per row per table, and the allowed index values will increase by four (e.g., 0x00, 0x04, 0x08, 0x0C, and so on). Similarly, if the element size is changed to a byte, there will be eight bytes per row per table, and the allowed index values will increase by eight (e.g., 0x00, 0x08, 0x10, 0x18, and so on). This rule is reflected in Table 13 below, which provides examples of row addressing for each supported number of tables in the table ("way" in Table 13) and element size configurations. This row addressing can be employed by successive lookup table initialization instructions to fill the entire lookup table, where subsequent initialization instructions specify the index of the first element in the row being filled.

[0222] Table 13

[0223]

[0224] In Fig.36B the example, the number of tables is 1 and the element size is a byte. The following are exemplary lookup table initialization instructions whose execution produces an instance group 3608 of the level 1 data cache 123:

[0225] LUTINIT D0,0,VB0; where D0 = 0x00 (e.g., the index of the element in row 0, set 0 of the level 1 data cache 123), 0 is the specified table set, VB3 = DATA0[511:0]

[0226] LUTINIT D1,0,VB1; where D1 = 0x40 (e.g., the index of the element in row 0, set 8 of the level 1 data cache 123), 0 is the specified table set, VB1 = DATA1[511:0]

[0227] LUTINIT D2,0,VB2; where D2 = 0x80 (e.g., the index of the element in row 1, set 0 of the level 1 data cache 123), 0 is the specified table set, VB2 = DATA2[511:0]

[0228] LUTINIT D3,0,VB3; where D3 = 0xC0 (e.g., the index of the element in row 1, set 8 of the level 1 data cache 123), 0 is the specified table set, VB3 = DATA3[511:0]

[0229] As described above, the size of the source data specified by the vector source data register 3502 varies depending on the number of tables specified by the lookup table configuration register 1900. In this example, the level-1 data cache 123 has a width of 1024 bits, and there is a table with a maximum width of one vector or 512 bits, and thus each row of the level-1 data cache 123 contains two rows of the table. Since there is only one table, no replication occurs. Registers VB0 through VB3 are vector registers, also having a size of 512 bits, and thus the lookup table initialization instruction utilizes the entire content of registers VB0 through VB3.

[0230] The first lookup table initialization instruction causes 512 bits of source data from register VB0 to be written to the table in table set 0, starting at the location specified by index 0x00 (plus the base address from the lookup table configuration register 1900), where the index 0x00 is the first element in row 0. The second lookup table initialization instruction causes 512 bits of source data from register VB1 to be written to the table in table set 0, starting at the location specified by index 0x40 (plus the base address from the lookup table configuration register 1900), where the index 0x40 is the 65th element in row 0 (e.g., the first element in group 8). The third lookup table initialization instruction causes 512 bits of source data from register VB2 to be written to the table in table set 0, starting at the location specified by index 0x80 (plus the base address from the lookup table configuration register 1900), where the index 0x80 is the first element in row 1. The fourth lookup table initialization instruction causes 512 bits of source data from register VB3 to be written to the table in table set 0, starting at the location specified by index 0xC0 (plus the base address from the lookup table configuration register 1900), where the index 0xC0 is the 65th element in row 1 (e.g., the first element in group 8).

[0231] As described, the level-1 data cache 123 has a bandwidth of 1024 bits, but in Figure 1 the example embodiment shown, the data bus (e.g., 144) between the central processing unit core 110 and the level-1 data cache 123 is only 512 bits. Thus, 512 bits is the maximum size of write_data that can be provided (e.g., from the vector source data register 3502) to the level-1 data cache 123. Replication is not supported in the case where the number of tables is 1 because there is no other table to which the replicated source data can be written. In the example, when the number of tables is 1, the lookup table initialization instruction provides the same bandwidth as the lookup table write instruction (or vector store) described above. However, for a number of tables equal to 2, 4, 8, or 16, the write_data from the vector source data register 3502 is replicated according to Table 14 below.

[0232] Table 14

[0233]

[0234]

[0235] As shown, depending on the number of tables specified by the lookup table configuration register 1900, the portion of the vector source data register 3502 sent to the level-1 data cache 123 for replication changes, for example, from 64 bits when the number of tables is 16 to 512 bits when the number of tables is 2 (replication allowed) or 1 (replication not allowed). Additionally, as described with respect to Fig.36A In the instance where there are 16 tables, producing a table width of 64 bits, the scalar register may also contain the source data to be written to the table. However, in other instances, the scalar register cannot be used as a source register because more than 64 bits are required (e.g., for the 8-table, 4-table, 2-table, and 1-table instances).

[0236] The following are the forms of the histogram (HIST) instruction and the weighted histogram (WHIST) instruction in an example embodiment.

[0237] HIST hist_index,hist_set

[0238] WHIST hist_index,hist_set,hist_weights

[0239] Hist_index is an instruction operand that specifies a vector register (e.g., within the general vector register file 231) by register number. This is decoded as an index number as shown in Table 11 above. Hist_index is similar to Tbl_index described above, but is specifically for the histogram context rather than the lookup table context. Similarly, Hist_set is the number [0:3] that specifies the histogram set to be used in the instruction. This named histogram set number specifies: the corresponding lookup table (or histogram) base address register that stores the base address of the histogram set, which can be a scalar register or a vector register; the corresponding lookup table (or histogram) configuration register ( Fig.19 ), which can be a scalar register or a vector register; and the lookup table (or histogram) enable register ( Fig. 20) corresponding operation part, which can be a scalar register or a vector register. The base register corresponding to the named histogram set determines the base address of the histogram set. The index of the vector register named by Hist_index is the offset from this table set base address. The look-up table configuration register corresponding to the named table set determines: the boosting mode (Table 4); the amount of memory allocated to the table size (Table 5); the weight size of the histogram operation (Table 6); the n-element interpolation of the look-up table read operation (Table 7); whether the value is considered signed or unsigned; whether the histogram bin entry saturates to the minimum / maximum value; the data element size (Table 8); and the number of histograms in the histogram set (Table 9). For the weighted histogram (WHIST) instruction, hist_weights is an instruction operand that specifies a vector register (e.g., within the general vector register file 231) by register number to provide weights to increment the addressed bin entry in the histogram. Similarly, the number of histograms, the size of the bin entry, the size of the weights, whether saturation is used, and other parameters are specified in the configuration register 1900.

[0240] Fig.37 An example illustrating the operation of the histogram instructions of this specification. In Fig.37 In the example illustrated in, the corresponding configuration register (1900) specifies four parallel histograms 3711, 3712, 3713, 3714 and a data element size of a word (32 bits). For histogram operations, the enable register field (in register 2000) corresponding to the selected histogram set should enable both read and write operations (11).

[0241] As above, the base register 2101 corresponding to the specified histogram set stores the base address of the histogram set, as Fig.37 schematically illustrated in. The histogram data is stored in a part of the level-1 data cache 123, and the level-1 data cache 123 is configured as a directly accessible memory.

[0242] Fig.37The example described herein has four histograms: hist 1 3711; hist 2 3712; hist 3 3713; and hist 4 3714. As shown in Table 11, this instruction with these selected options treats the data stored in vector source index register 2102 as a collection of 4 32-bit fields that specify table offsets. The first field (bits Vx[31:0]) stores index iA into the first histogram 3711. In this example, this indexes to element A5 (e.g., between A4 and A6). The second field (bits Vx[63:32]) stores index iB into the second histogram 3712. In this example, this indexes to element B1 (e.g., between B0 and B2). The third field (bits Vx[95:64]) stores index iC into the third histogram 3713. In this example, this indexes to element C8 (e.g., between C7 and C9). The fourth field (bits Vx[127:96]) stores index iD into the fourth histogram 3714. In this example, this indexes to element D10 (e.g., between D9 and D11). The various indexes are memory address offsets from the base address for the histogram collection to the specified bin element to be modified.

[0243] In response to the execution of a histogram instruction according to Fig.37 the example, the indexed bin entries specified by base address register 2101 and vector source index register 2102 are incremented by 1. The bin entry specified by iA (e.g., A5) is incremented by 1; the bin entry specified by iB (e.g., B1) is incremented by 1; the bin entry specified by iC (e.g., C8) is incremented by 1; and the bin entry specified by iD (e.g., D10) is incremented by 1.

[0244] If the saturation field 1905 of configuration register 1900 is set, then in response to a histogram operation, the bin entries are saturated (e.g., limited) to the minimum / maximum values of the element data type. For example, an unsigned byte saturates to [0, 0xFF]; a signed byte saturates to [0x80, 0x7F]; an unsigned halfword saturates to [0, 0xFFFF]; a signed halfword saturates to [0x8000, 0x7FFF]; an unsigned word saturates to [0, 0xFFFF FFFF]; and a signed word saturates to [0x8000 0000, 0x7FFFFFFF]. If the saturation field 1905 is 0, then the histogram bin entries are not saturated to the minimum / maximum values of the element data type, but rather will wrap around when incremented beyond the maximum value or decremented beyond the minimum value.

[0245] Fig.38 An example illustrating the operation of the weighted histogram instruction of this specification. Fig.38 Similar to Fig.37, but includes the vector source data register 3502 described above. The vector source data register 3502 is specified by hist_weights and contains the weights for incrementing the addressed bin entries in histograms 3711, 3712, 3713, 3714.

[0246] Unlike the lookup table write instruction described above, for the weighted histogram instruction, the vector source data register 3502 contains data to be added to the indexed bin entries in histograms 3711, 3712, 3713, 3714. In some instances, the histogram weights are signed values and the weight magnitudes are limited to bytes and halfwords. Additionally, in some instances, the weight magnitude cannot be greater than the specified bin entry size. Thus, for byte-sized bins, only byte-sized weights are permitted; for halfword or word-sized bins, byte or halfword-sized weights are permitted. In certain instances, for word-sized bins, word-sized weights are also permitted. In this instance, because the lookup table configuration register specifies a histogram count of four and a weight magnitude of halfword (which is permitted due to specifying the bin entry size as word-sized), the first four halfwords of the vector source data register 3502 are utilized by the weighted histogram instruction, while the remainder of the vector source data register 3502 is not utilized by the weighted histogram instruction. Specifically, the first halfword (value +5) is added to the bin entry specified by index iA (e.g., A5 is incremented by 5); the second halfword (value +2) is added to the bin entry specified by index iB (e.g., B1 is incremented by 2); the third halfword (value -3) is added to the bin entry specified by index iC (e.g., C8 is decremented by 3); and the fourth halfword (value +7) is added to the bin entry specified by index iD (e.g., D10 is incremented by 7).

[0247] Similar to the above histogram instruction, if the saturation field 1905 of the configuration register 1900 is set, then in response to a histogram operation, the bin entries are saturated (e.g., limited) to the minimum / maximum values of the element data type. For example, an unsigned byte is saturated to [0, 0xFF]; a signed byte is saturated to [0x80, 0x7F]; an unsigned halfword is saturated to [0, 0xFFFF]; a signed halfword is saturated to [0x8000, 0x7FFF]; an unsigned word is saturated to [0, 0xFFFF FFFF]; and a signed word is saturated to [0x8000 0000, 0x7FFF FFFF]. If the saturation field 1905 is 0, then the histogram bin entries are not saturated to the minimum / maximum values of the element data type, but rather will wrap around when incremented beyond the maximum value or decremented beyond the minimum value.

[0248] Extensions of the above histogram and weighted histogram instructions in response to different parameters specified by configuration register 1900 would be obvious to one of ordinary skill in the art. For example, a change in the table (or histogram) number field 1908 affects the number of indices specified by vector source index register 2102 (e.g., as shown in Figures 21 to 25 ). Similarly, a change in the weight size field 1903 affects the portion of the source data in vector source data register 3502 that provides the indexed bin entries to be written to the histogram.

[0249] In the foregoing description and claims, the terms "comprising" and "including" are used in an open-ended manner and thus mean "including but not limited to...". Additionally, the term "couple / couples" means an indirect or direct connection. Thus, if a first device is coupled to a second device, the connection may be by a direct connection or by an indirect connection via other devices and connections. Similarly, a device that couples between a first component or location and a second component or location may be by a direct connection or by an indirect connection via other devices and connections. An element or feature "configured to" perform a task or function may be configured (e.g., programmed or structurally designed) to perform the function when manufactured by a manufacturer, and / or may be configured (or reconfigured) by a user after manufacture to perform the function and / or other additional or alternative functions. The configuration may be by firmware and / or software programming of the device, by the construction and / or layout of the hardware components and interconnections of the device, or a combination thereof. Further, the use of the phrase "ground" or the like in the foregoing description encompasses chassis ground, earth ground, floating ground, virtual ground, digital ground, common ground, and / or any other form of ground connection suitable for or adapted to the teachings of this specification. Unless otherwise stated, "about", "approximately", or "substantially" in front of a value means + / - 10 percent of the stated value.

[0250] The foregoing description illustrates the principles of this specification and various embodiments. Once the foregoing description is fully understood, many variations and modifications will become obvious to those of skill in the art. The appended claims cover all such variations and modifications.

Claims

1. A digital data processor, comprising: an instruction memory that stores instructions, each of the instructions specifying a data processing operation and at least one data operand field; an instruction decoder coupled to the instruction memory for sequentially invoking instructions from the instruction memory and determining the data processing operation and the at least one data operand; at least one arithmetic unit coupled to a data register file and coupled to the instruction decoder to perform a data processing operation on at least one operand corresponding to an instruction decoded by the instruction decoder and store a result of the data processing operation, wherein the at least one arithmetic unit is configured to increment a histogram value in response to a histogram instruction by: incrementing a bin entry at a specified position in at least one specified numbered histogram; and a primary data memory, comprising: a first portion that includes a primary data cache coupled to the at least one arithmetic unit, wherein the first portion is configured to store data for manipulation by the at least one arithmetic unit, and wherein the primary data cache services memory reads and writes of the at least one arithmetic unit; and a second portion that includes memory directly accessible via the at least one arithmetic unit, wherein the at least one histogram is stored in the second portion of the primary data memory; wherein a position at which the at least one arithmetic unit increments the bin entry in response to the histogram instruction is in the second portion of the primary data memory.

2. The digital data processor according to claim 1, wherein incrementing includes adding a value of one to the bin entry.

3. The digital data processor according to claim 1, wherein the histogram instruction specifies a look-up table configuration register having a saturation field indicating that a bin entry will saturate.

4. The digital data processor according to claim 1, wherein the at least one arithmetic unit is configured to increment the histogram value in response to a weighted histogram instruction by incrementing the bin entry by a weight value specified in a source data register at a specified position in at least one specified numbered histogram.

5. The digital data processor according to claim 4, wherein a size of the weight value is less than or equal to a size of the bin entry.

6. The digital data processor according to claim 4, wherein: the data register file includes a plurality of data registers identified by register numbers, each data register storing data; the weighted histogram instruction includes a first source operand field specifying a register number of one of the data registers in the data register file; and the instruction decoder is configured to decode the weighted histogram instruction to identify the data register having the register number of the first source operand field as the source data register.

7. The digital data processor according to claim 1, wherein: the data register file includes a plurality of data registers identified by register numbers, each data register storing data; The histogram instruction includes a source operand field specifying a register number of one of the data registers in the data register bank; and the instruction decoder is configured to decode the histogram instruction to use a portion of the data register having the register number of the source operand field as a pointer to the specified location to be incremented.

8. The digital data processor according to claim 7, wherein: the histogram instruction specifies a lookup table base address register storing a table base address; and the instruction decoder is configured to decode the histogram instruction to use the pointer as an offset from the table base address stored in the lookup table base address register.

9. The digital data processor according to claim 7, wherein the histogram instruction specifies a lookup table configuration register having a histogram bin entry size field indicating a specified size of the bin entry.

10. A method, comprising: providing a primary data memory including: a first portion including a primary data cache, wherein the first portion is configured to store data for manipulation by an arithmetic unit coupled to a data register bank and to an instruction decoder, wherein the primary data cache services memory reads and writes of the arithmetic unit; and a second portion including a memory directly accessible by the arithmetic unit; storing at least one histogram in the second portion of the primary data memory; incrementing a histogram value in response to a histogram instruction by the arithmetic unit by incrementing a bin entry at a specified location in at least one histogram of a specified number, wherein a location at which the arithmetic unit increments the bin entry in response to the histogram instruction is in the second portion of the primary data memory.

11. The method according to claim 10, wherein incrementing includes adding a value of one to the bin entry.

12. The method according to claim 10, wherein the histogram instruction specifies a lookup table configuration register having a saturation field indicating that a bin entry will saturate.

13. The method according to claim 10, further comprising incrementing the histogram value in response to a weighted histogram instruction by the arithmetic unit by incrementing a bin entry at a specified location in at least one histogram of a specified number by a weight value specified in a source data register.

14. The method according to claim 13, wherein a size of the weight value is less than or equal to a size of the bin entry.

15. The method according to claim 13, wherein: the data register bank includes a plurality of data registers identified by register numbers, each data register storing data; the weighted histogram instruction includes a first source operand field specifying a register number of one of the data registers in the data register bank; and the method further comprises decoding the weighted histogram instruction by the instruction decoder to identify the data register having the register number of the first source operand field as the source data register.

16. The method according to claim 10, wherein: the data register file includes a plurality of data registers identified by register numbers, each data register storing data; the histogram instruction includes a source operand field specifying a register number of one of the data registers in the data register file; and the method further includes decoding, by the instruction decoder, the histogram instruction to use a portion of the data register having the register number of the source operand field as a pointer to the incremented specified location.

17. The method according to claim 16, wherein the histogram instruction specifies a lookup table base address register storing a table base address, and the method further includes: decoding, by the instruction decoder, the histogram instruction to use the pointer as an offset to the table base address stored in the lookup table base address register.

18. The method according to claim 16, wherein the histogram instruction specifies a lookup table configuration register having a histogram bin entry size field indicating a specified size of the bin entry.

Citation Information

Patent Citations

  • Unified memory system architecture including cache and directly addressable static random access memory

    US6606686B1

  • Processor with inter-processing path communication

    US20130185538A1