Transposition circuit, transposition method, and related device

WO2026200607A1PCT designated stage Publication Date: 2026-10-01SHENZHEN MICROBT ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/083858
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-17
Publication Date
2026-10-01

Smart Images

  • Figure CN2026083858_01102026_PF_FP_ABST
    Figure CN2026083858_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a transposition circuit, a transposition method, and a related device. The transposition circuit comprises: an input module, configured to receive first data of N rows by M columns; a storage module, connected to the input module and comprising a first plurality of storage units of N rows by M columns, wherein the storage module is configured to write each column of data of the first data into a corresponding column of storage units among the first plurality of storage units in a corresponding cycle of a first loop; a selection module, connected to the storage module and configured to output, in a corresponding cycle of a second loop after the first loop, data read from each row of storage units among the first plurality of storage units; and an output module, connected to the selection module and configured to output data received from the selection module as transposed first data, where N and M are positive integers and at least one of N and M is greater than 1.
Need to check novelty before this filing date? Find Prior Art

Description

Transpose circuit, transpose method and related devices

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to CN application number 202510346811.5, filed on March 24, 2025, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure generally relates to the field of electronic circuit technology, and more specifically, to a transpose circuit and a transpose method, and also to a two-dimensional (2D) Discrete Cosine Transform (DCT) device, processor, and computing device comprising such a transpose circuit. Background Technology

[0004] 2D Directional Transform (DCT) is widely used in image and video processing, primarily for signal compression and feature extraction. In 2D DCT, the intermediate results of the transformation need to be transposed. Specifically, transposition means converting row-based input data into column-based output, or vice versa. The transposition circuit is a crucial component in the hardware implementation of 2D DCT. For example, 2D DCT typically consists of two 1D DCT steps: row transformation followed by column transformation. However, the data access patterns for row and column transformations differ. The transposition circuit rearranges the data between the row and column transformations to ensure that the data is input in the correct order. Summary of the Invention

[0005] According to a first aspect of this disclosure, a transpose circuit is provided, comprising: an input module configured to receive N rows × M columns of first data; a storage module connected to the input module and including a first plurality of storage cells of N rows × M columns, the storage module being configured to write each column of the first data into a corresponding column of the first plurality of storage cells in a corresponding cycle of a first loop; a selection module connected to the storage module and configured to output data read from each row of the first plurality of storage cells in a corresponding cycle of a second loop after the first loop; and an output module connected to the selection module and configured to output the data received from the selection module as transposed first data, wherein N and M are positive integers and at least one of N and M is greater than 1.

[0006] In some embodiments, the selection module includes a combination of multiplexers having N inputs and one output, each of the N inputs being connected to a corresponding row of storage cells in the plurality of storage cells, and the one output being connected to the output module.

[0007] In some embodiments, the input module is further configured to receive second data in N rows × M columns, and the storage module further includes a second plurality of storage units in N rows × M columns and is configured to write each column of the second data into a corresponding column of the second plurality of storage units in a corresponding cycle of the second loop, wherein the selection module is further configured to output data read from each row of storage units in the second plurality of storage units in a corresponding cycle of the third loop after the second loop, wherein the output module is further configured to output the data received from the selection module as transposed second data.

[0008] In some embodiments, the input module is further configured to receive N rows × M columns of third data, the storage module is further configured to write each column of the third data into a corresponding column of the first plurality of storage units in a corresponding cycle of the third cycle, the selection module is further configured to output the data read from each row of the first plurality of storage units in a corresponding cycle of the fourth cycle after the third cycle, and the output module is further configured to output the data received from the selection module as transposed third data.

[0009] In some embodiments, the selection module includes a combination of multiplexers having 2N inputs and 1 output, each of the 2N inputs being connected to a corresponding row of storage cells in the N rows of the first plurality of storage cells and the N rows of storage cells in the second plurality of storage cells, and the 1 output being connected to the output module.

[0010] In some embodiments, the selection module includes: a first selection submodule, comprising a first multiplexer combination having N inputs and y outputs, each of the N inputs of the first multiplexer combination being connected to a corresponding row of storage cells in the x rows of the first plurality of storage units and the Nx rows of storage cells in the second plurality of storage units; and a second selection submodule, comprising a second multiplexer combination having N inputs and z outputs, each of the N inputs of the second multiplexer combination being connected to the remaining Nx rows of storage cells in the first plurality of storage units. The storage unit and the remaining x rows of storage units of the second plurality of storage units; the third selection submodule includes a third multiplexer combination having y+z input terminals and 1 output terminal, wherein the y output terminals of the first multiplexer combination and the z output terminals of the second multiplexer combination are respectively connected to the corresponding input terminals of the y+z input terminals of the third multiplexer combination, and the 1 output terminal of the third multiplexer combination is connected to the output module, wherein x is an integer and 0≤x≤N, y is an integer and 0<y<N, and z is an integer and 0<z<N.

[0011] In some embodiments, x = 0 or N, y = z = 1.

[0012] In some embodiments, x = y = z = N / 2, wherein the first multiplexer combination comprises N / 2 2-to-1 multiplexers, each 2-to-1 multiplexer having a first input connected to a corresponding row of storage cells in the x rows of storage cells of the first plurality of storage units and a second input connected to a corresponding row of storage cells in the Nx rows of storage cells of the second plurality of storage units; and the second multiplexer combination comprises N / 2 2-to-1 multiplexers, each 2-to-1 multiplexer having a first input connected to a corresponding row of storage cells in the remaining Nx rows of storage cells of the first plurality of storage units and a second input connected to a corresponding row of storage cells in the remaining x rows of storage cells of the second plurality of storage units.

[0013] In some embodiments, the input module includes a first trigger configured to output each column of received data to a corresponding column of storage cells based on a received clock signal. In some embodiments, the output module includes a second trigger configured to output data read from each row of storage cells based on a received clock signal.

[0014] In some embodiments, the storage unit includes: a latch unit; or a dynamic random access memory (DRAM) unit; or a static random access memory (SRAM) unit.

[0015] In some embodiments, N equals M.

[0016] In some embodiments, the data is image data.

[0017] According to a second aspect of this disclosure, a two-dimensional discrete cosine transform apparatus is provided, including the transpose circuit described in the first aspect of this disclosure, wherein N equals M.

[0018] According to a third aspect of this disclosure, a processor is provided, including a transpose circuit according to a first aspect of this disclosure or a two-dimensional discrete cosine transform device according to a second aspect of this disclosure.

[0019] According to a fourth aspect of this disclosure, a computing device is provided, including a transpose circuit according to a first aspect of this disclosure, a two-dimensional discrete cosine transform device according to a second aspect of this disclosure, or a processor according to a third aspect of this disclosure.

[0020] According to a fifth aspect of this disclosure, a transposition method is provided, comprising: receiving N rows × M columns of first data and writing each column of the first data into a corresponding column of a first plurality of storage units of N rows × M columns in a corresponding cycle of a first loop; outputting data read from each row of storage units in the first plurality of storage units in a corresponding cycle of a second loop after the first loop; receiving N rows × M columns of second data and writing each column of the second data into a corresponding column of a second plurality of storage units of N rows × M columns in a corresponding cycle of the second loop; and outputting data read from each row of storage units in the second plurality of storage units in a corresponding cycle of a third loop after the second loop, wherein N and M are positive integers and at least one of N and M is greater than 1.

[0021] In some embodiments, the transposition method further includes: receiving N rows × M columns of third data and writing each column of the third data into a corresponding column of the first plurality of storage units in a corresponding cycle of the third cycle; and outputting data read from each row of the first plurality of storage units in a corresponding cycle of the fourth cycle after the third cycle.

[0022] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0023] The accompanying drawings, which form part of this specification, illustrate embodiments of the present disclosure and, together with the specification, serve to explain the principles of the present disclosure. The present disclosure can be more clearly understood from the following detailed description, with reference to the accompanying drawings, wherein:

[0024] Figure 1 illustrates a transpose circuit according to a comparative example of the present disclosure;

[0025] Figure 2 illustrates a transpose circuit according to some embodiments of the present disclosure;

[0026] Figure 3 illustrates a transpose circuit according to some other embodiments of the present disclosure;

[0027] Figure 4 shows an example implementation of the selection module of the transpose circuit shown in Figure 3;

[0028] Figure 5 shows another example implementation of the selection module of the transpose circuit shown in Figure 3;

[0029] Figure 6 illustrates an example implementation of the data writing portion of a transpose circuit according to some embodiments of the present disclosure;

[0030] Figure 7 illustrates an example implementation of the data readout portion of a transpose circuit according to some embodiments of the present disclosure;

[0031] Figure 8 illustrates another example implementation of the data readout portion of the transpose circuit according to some embodiments of the present disclosure;

[0032] Figure 9 is a flowchart illustrating a transposition method according to some embodiments of the present disclosure;

[0033] Figure 10 illustrates an example implementation of a storage cell of a transpose circuit according to some embodiments of the present disclosure.

[0034] Note that in the embodiments described below, the same reference numerals are sometimes used across different figures to denote the same parts or parts having the same function, and repeated descriptions are omitted. In this specification, similar reference numerals and letters are used to denote similar items; therefore, once an item is defined in one figure, it need not be further discussed in other figures unless otherwise stated.

[0035] For ease of understanding, the positions, dimensions, and extents of the structures shown in the accompanying drawings and other materials may not represent actual positions, dimensions, and extents. Therefore, the disclosed invention is not limited to the positions, dimensions, and extents disclosed in the accompanying drawings and other materials. Furthermore, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components. Detailed Implementation

[0036] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0037] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this disclosure or its application or use. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods in this disclosure. However, those skilled in the art will understand that they merely illustrate exemplary ways that can be used to implement this disclosure, and not exhaustive ways. Furthermore, the drawings are not necessarily drawn to scale, and some features may be enlarged to show details of specific components.

[0038] In addition, techniques, methods and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods and equipment should be considered part of the specification.

[0039] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0040] Figure 1 illustrates a transpose circuit 10 according to a comparative example of the present disclosure, which can be used to perform a transpose operation on 4×4 two-dimensional data (e.g., an image). As shown in Figure 1, the transpose circuit 10 includes a column input module 11, a column output module 12, a row input module 13, a row output module 14, and a storage module 15. Specifically, the storage module 15 consists of 16 register groups arranged in 4 rows and 4 columns, each register group capable of storing one piece of data from the two-dimensional data (e.g., one pixel in an image). The number of registers included in each register group can be determined by the bit width of the data to be stored in that register group. The column input module 11 is connected to the input terminals of the first column register group p00 to p30. The output terminals of the first column register group p00 to p30 are respectively connected to the input terminals of the second column register group p01 to p31. The output terminals of the second column register group p01 to p31 are respectively connected to the input terminals of the third column register group p02 to p32. The output terminals of the third column register group p02 to p32 are respectively connected to the input terminals of the fourth column register group p03 to p33. The output terminals of the fourth column register group p03 to p33 are connected to the column output module 12. The row input module 13 is connected to the input terminals of the first row register group p00 to p03. The output terminals of the first row register group p00 to p03 are respectively connected to the input terminals of the second row register group p10 to p13. The output terminals of the second row register group p10 to p13 are respectively connected to the input terminals of the third row register group p20 to p23. The output terminals of the third row register group p20 to p23 are respectively connected to the input terminals of the fourth row register group p30 to p33. The output terminals of the fourth row register group p30 to p33 are connected to the row output module 14.

[0041] For example, in the first loop, two-dimensional data The data is input column by column to the storage module 15 via the column input module 11. Specifically, in the first cycle of the first loop, the first column of the two-dimensional data A is input. The data is stored in the first column register group p00 to p30; in the second cycle of the first loop, the first column data a_0 of the two-dimensional data A is shifted to the right to be stored in the second column register group p01 to p31, and the second column data of the two-dimensional data A... The data is stored in the first column register group p00 to p30; in the third cycle of the first loop, the first column data a_0 and the second column data a_1 of the two-dimensional data A are shifted to the right to be stored in the third column register group p02 to p32 and the second column register group p01 to p31, respectively, and the third column data of the two-dimensional data A... The data is stored in the first column register group p00 to p30; in the fourth cycle of the first loop, the first column data a_0, the second column data a_1, and the third column data a_2 of the two-dimensional data A are shifted to the right to be stored in the fourth column register group p03 to p33, the third column register group p02 to p32, and the second column register group p01 to p31, respectively. The fourth column data of the two-dimensional data A... The data is stored in the first column register group, p00 to p30. After four cycles, the storage module 15 is filled with two-dimensional data A.

[0042] Next, in the second loop, the two-dimensional data A is output row by row from the storage module 15 via the row output module 14, while the two-dimensional data... The data is input row by row to the storage module 15 via the row input module 13. Specifically, in the first cycle of the second loop, the fourth row data a3_ = (a30 a31 a32 a33) of two-dimensional data A is output from the fourth row register group p30 to p33 to the row output module 14. The third row data a2_ = (a20 a21 a22 a23), the second row data a1_ = (a10 a11 a12 a13), and the first row data a0_ = (a00 a01 a02 a03) of two-dimensional data A are shifted down to be stored in the fourth row register group p30 to p33, the third row register group p20 to p23, and the second row register group p10 to p13, respectively. The fourth row data b3_ = (b30 b31 b32) of two-dimensional data B is also shifted down to be stored in the fourth row register group p30 to p33, the third row register group p20 to p23, and the second row register group p10 to p13, respectively. b33) is stored in the first row register group p00 to p03; in the second cycle of the second loop, the third row data a2_ of two-dimensional data A is output from the fourth row register group p30 to p33 to the row output module 14, the second row data a1_ and the first row data a0_ of two-dimensional data A and the fourth row data b3_ of two-dimensional data B are shifted down to be stored in the fourth row register group p30 to p33, the third row register group p20 to p23 and the second row register group p10 to p13 respectively, and the third row data b2_ of two-dimensional data B = (b20 b21 b22) b23) is stored in the first row register group p00 to p03; in the third cycle of the second loop, the second row data a1_ of two-dimensional data A is output from the fourth row register group p30 to p33 to the row output module 14, the first row data a0_ of two-dimensional data A and the fourth row data b3_ and third row data b2_ of two-dimensional data B are shifted down to be stored in the fourth row register group p30 to p33, the third row register group p20 to p23 and the second row register group p10 to p13 respectively, and the second row data b1_ of two-dimensional data B = (b10 b11 b12) b13) is stored in the first row register group p00 to p03; in the fourth cycle of the second loop, the first row data a0_ of two-dimensional data A is output from the fourth row register group p30 to p33 to the row output module 14. The fourth row data b3_, the third row data b2_, and the second row data b1_ of two-dimensional data B are shifted down to be stored in the fourth row register group p30 to p33, the third row register group p20 to p23, and the second row register group p10 to p13, respectively. The first row data b0_ of two-dimensional data B = (b00 b01 b02 b03) is stored in the first row register group p00 to p03. Thus, after 4 cycles, the transposed two-dimensional data A is completely output from the storage module 15, and the storage module 15 is filled again with two-dimensional data B.

[0043] In the next third cycle, two-dimensional data B can be output column-wise from storage module 15 via column output module 12, while two-dimensional data C can be input column-wise to storage module 15 again via column input module 11, just like two-dimensional data A. The transpose circuit 10 cycles between column input, row output and row input, column output, so that a transpose result of two-dimensional data can be output in each cycle.

[0044] As can be seen from the above process, because the transpose circuit 10 continuously performs data shifting operations during the transpose operation, all registers need to undergo data toggling in each cycle (a toggle represents a signal moving from 0 to 1 or from 1 to 0). This high toggle rate results in very high dynamic power consumption for the transpose circuit 10. Furthermore, the registers consume a large amount of circuit area. However, since the transpose operation of the transpose circuit 10 depends on data shifting operations, and shifting operations require precise timing control, edge-triggered registers are needed to provide strict timing control.

[0045] To address this, this disclosure provides a transpose circuit that can perform a transpose operation without performing a data shift operation, thereby achieving a reduced flip-flop rate and consequently reduced dynamic power consumption. Various embodiments of the transpose circuit according to this disclosure will now be described in detail with reference to the accompanying drawings. It should be understood that actual circuits may include other components, but to avoid obscuring the essential points of this disclosure, these other components will not be discussed herein and are not shown in the drawings.

[0046] Figure 2 illustrates a transpose circuit 100 according to some embodiments of the present disclosure. As shown in Figure 2, the transpose circuit 100 includes an input module 110, a storage module 120, a selection module 130, and an output module 140. It should be understood that although the input module 110 is depicted as a column input module and the output module 140 as a row output module in the figures, and the description herein primarily uses the transpose operation of converting column-input data into row-output data as an example, this is merely exemplary and not restrictive. In this document, "row" and "column" are relative concepts and can be used interchangeably.

[0047] Input module 110 is configured to receive N rows × M columns of first data. N and M are positive integers, and at least one of N and M is greater than 1. In some embodiments, N is equal to M. For example, N and M can each be equal to 2. iWhere i is a positive integer. In the embodiment shown in the accompanying drawings, N and M are each equal to 4, thus providing a 4×4 transpose unit. It is understood that this is merely exemplary and not limiting, and the present disclosure may similarly provide transpose units of different specifications such as 2×2, 8×8, 16×16, and 32×32. N and M may also have other values ​​as needed. In some embodiments, N may not be equal to M. In this document, for ease of description, the example of N = M = 4 will be used primarily. In addition, the data processed by the present disclosure can be any form of two-dimensional data, such as, but not limited to, image data.

[0048] Input module 110 includes the input terminal of transpose circuit 100. In some embodiments, input module 110 may include a first flip-flop (illustrated later), which may be, for example, a D flip-flop (DFF). The first flip-flop is capable of cutting off the timing outside transpose circuit 100, thereby ensuring the timing inside transpose circuit 100.

[0049] Storage module 120 is connected to input module 110 and includes a first plurality of storage units 122 arranged in N rows × M columns. Since no data shifting operation is required, the storage units 122 do not need to be connected to each other. The input of each storage unit 122 can be connected only to input module 110 without connecting to other storage units 122. The input of each storage unit 122 can receive data to be written only from input module 110 without shifting data from other storage units 122 into it. Moreover, since no data shifting operation is required, the storage units 122 do not need to provide strict timing control and can be formed by any circuit element capable of providing storage functionality, without necessarily requiring registers. Storage units 122 can utilize circuit elements with the smallest possible area while satisfying the storage function. For example, storage cell 122 may include storage elements selected from the group consisting of: latch cells; dynamic random access memory cells (DRAM cells); or static random access memory cells (SRAM cells). The number of storage elements included in storage cell 122 can be determined by the bit width of the data to be stored in storage cell 122. These implementations of storage cell 122 offer greater area advantages than register implementations. For example, latch cells can be either dynamic or static latch cells. Of course, the storage elements of storage cell 122 can also be registers.

[0050] Storage module 120 is configured to write each column of the first data into a corresponding column of storage cells 122 in a first cycle of the first plurality of storage cells 122. Specifically, during the writing process, no data shifting operation occurs in storage module 120; instead, a new column of data is stored into a new column of storage cells 122 each time. For example, referring to FIG2, in the first cycle, two-dimensional data... The data is input column-by-column to the storage module 120 via the input module 110. Specifically, in the first cycle of the first loop, the first column of the two-dimensional data A is input. Store in the first column of storage cells 1221, 1225, 1229, and 122. 13 In the second cycle of the first loop, the second column of two-dimensional data A. Stored in the second column storage units 1222, 1226, and 122. 10 122 14 In the third cycle of the first loop, the third column of two-dimensional data A. Stored in the third column storage units 1223, 1227, and 122 11 122 15 In the fourth cycle of the first loop, the fourth column of two-dimensional data A... Stored in the fourth column of storage cells 1224, 1228, and 122. 12 122 16 Thus, after four cycles, the first plurality of storage cells 122 of storage module 120 are filled with two-dimensional data A. It should be understood that the above sequence is merely exemplary and not restrictive. Because it is not limited by shift operations, each column of the first data can actually be stored in any of the storage cells 122 that have not yet been filled with new data in this cycle.

[0051] For example, writing to each memory cell 122 can be controlled by a clock signal. For non-limiting illustrative purposes, Figure 10 shows an example implementation of the memory cell. As shown in (A) of Figure 10, the memory cell may have an input terminal D, an output terminal Q, and a clock control terminal G. The clock control terminal G is configured to receive a clock signal clk. When the clock signal clk is configured to cause the clock control terminal G of the memory cell to flip, the data wd received at the input terminal D of the memory cell will be written to the memory cell. When the clock signal clk is configured not to cause the clock control terminal G of the memory cell to flip, the data wd received at the input terminal D of the memory cell will not be written to the memory cell. Alternatively, writing to each memory cell 122 can be controlled by a selection signal. As shown in (B) of Figure 10, a multiplexer MUX can be arranged before the memory cell, with a first input terminal for receiving data wd and a second input terminal for receiving the output at the output terminal Q of the memory cell. The output terminal of the multiplexer MUX is coupled to the input terminal D of the memory cell. In this configuration, the clock signal clk can always be configured to toggle the clock control terminal G of the memory cell. Simultaneously, the selection signal sel selects whether to write data wd to the memory cell as the output of the multiplexer MUX, or to prevent data wd from being written to the memory cell by using the existing data of the memory cell as the output of the multiplexer MUX. In these implementations, data wd can always be supplied to the input terminal D of the memory cell.

[0052] In some implementations, the input module 110 may include N buffer units, each buffer unit being used to temporarily store a corresponding row of data (i.e., a single data item) from the currently received column of data. Each buffer unit may be connected to the respective storage units 122 in the corresponding row storage unit 122. For example, referring to FIG2, in the first cycle of the first loop, the four buffer units (not shown) of the input module 110 may respectively store each row of data a00, a10, a20, a30 in the first column of data a_0 of the two-dimensional data A. The first buffer unit storing data a00 may be connected to the first row storage units 1221, 1222, 1223, 1224, such that data a00 can be provided as wd to the input terminal D of each storage unit 122 in the first row storage units 1221, 1222, 1223, 1224. At this time, referring to (A) in Figure 10, the clock control terminal G of the storage cell 1221 in the first column can be flipped to write data a00 into storage cell 1221, while the clock control terminals G of the storage cells 1222, 1223, and 1224 in other columns are prevented from flipping to write data a00 into storage cells 1222, 1223, and 1224. Similarly, the input terminals D of the second, third, and fourth row storage cells are also connected to the second buffer cell storing data a10, the third buffer cell storing data a20, and the fourth buffer cell storing data a30, respectively, to receive the corresponding data a10, a20, and a30 from them. At this time, the clock control terminal G of the storage cell in the first column can also be flipped to write the corresponding data a10, a20, and a30 into it, while the clock control terminals G of the storage cells in other columns are prevented from flipping to write the corresponding data a10, a20, and a30 into it.

[0053] The above writing process does not involve data shifting operations. Only one column of storage cells 122 undergoes data flipping in each cycle. This low flip rate results in low dynamic power consumption.

[0054] Selection module 130 is connected to storage module 120. Output module 140 is connected to selection module 130 and configured to output data received from selection module 130 as transposed first data. Output module 140 includes the output of transpose circuit 100. In some embodiments, input module 110 may include M buffer units (not shown), each buffer unit for temporarily storing a corresponding column of data (i.e., one data) from the currently received row of data. In some embodiments, output module 140 may include a second flip-flop (shown later), which may be, for example, in the form of a DFF. The second flip-flop is capable of cutting off timing outside transpose circuit 100, thereby ensuring timing inside transpose circuit 100.

[0055] The selection module 130 is configured to output data read from a corresponding row of storage cells 122 in the first plurality of storage cells 122 in a corresponding cycle of the second cycle after the first cycle. Specifically, no data shifting operation occurs in the storage module 120 during the readout process. For example, referring to FIG2, in the second cycle, two-dimensional data A is output from the storage module 120 row by row via the selection module 130 and the output module 140. Specifically, in the first cycle of the second cycle, the fourth row of two-dimensional data A, a3_ = (a30 a31 a32 a33), is output from the fourth row storage cell 122. 13 122 14 122 15 122 16 Output; in the second cycle of the second loop, the third row data a2_=(a20 a21 a22 a23) of the two-dimensional data A is output from the third row storage units 1229, 122. 10 122 11 122 12 The data is output; in the third cycle of the second loop, the second row of data A, a1_ = (a10 a11 a12 a13), is output from the second row storage units 1225, 1226, 1227, and 1228; in the fourth cycle of the second loop, the first row of data A, a0_ = (a00 a01 a02 a03), is output from the first row storage units 1221, 1222, 1223, and 1224. Thus, after four cycles, the transposed two-dimensional data A is completely output from the first plurality of storage units 122 of the storage module 120. It is understood that the above sequence is merely exemplary and not restrictive. Because it is not limited by shift operations, data in any row of storage units 122 that has not yet been read in the current cycle can actually be read in each cycle.

[0056] The above readout process does not involve data shifting operations, and no memory cell 122 undergoes data flipping in each cycle. This low flip rate results in low dynamic power consumption.

[0057] In some embodiments, the selection module 130 may include a combination of multiplexers having N inputs and one output, each of the N inputs being connected to a corresponding row of memory cells 122 in a first plurality of memory cells 122, and the one output being connected to an output module 140. For example, referring to FIG2, the selection module 130 may be formed by a single 4-to-1 multiplexer (not shown), or by a 4-to-2 multiplexer and a 2-to-1 multiplexer (not shown) arranged in two stages, or by a pair of 2-to-1 multiplexers and a 2-to-1 multiplexer (not shown) arranged in two stages, or by a 3-to-1 multiplexer and a 2-to-1 multiplexer (not shown) arranged in two stages, and so on. When N is large, N-to-1 multiplexers may be difficult to obtain. In this case, a larger-scale multiplexing function can be achieved by cascading multiple smaller-scale multiplexers.

[0058] When the storage module 120 of the transpose circuit 100 includes only the first plurality of storage cells 122, each write process (such as the first loop, which is an odd-numbered loop) requires M cycles, and each read process (such as the second loop, which is an even-numbered loop) requires N cycles. By alternating between the write and read processes, a transpose of the two-dimensional data can be obtained every two loops (M+N cycles).

[0059] Furthermore, in each write cycle, only one column of memory cells 122 in the transpose circuit 100 undergoes data flipping, and in each read cycle, no memory cells 122 undergo data flipping. This low flip rate results in lower dynamic power consumption, especially lower than that of the transpose circuit 10, where all registers undergo data flipping in each cycle. Moreover, when used as transpose units of the same specifications, even if the number of memory cells in the transpose circuit 100 is the same as the number of registers in the transpose circuit 10, the transpose circuit 100 has a significant area advantage over the transpose circuit 10 because the memory cells in the transpose circuit 100 do not need to be implemented as registers, but can instead use circuit elements with the smallest possible area (such as latch units) that satisfy the storage function.

[0060] Figure 3 illustrates a transpose circuit 200 according to some other embodiments of the present disclosure. As shown in Figure 3, the transpose circuit 200 differs from the transpose circuit 100 in that its storage module 120 further includes a second plurality of storage cells 124 in N rows × M columns. Since no data shifting operation is required, the storage cells 124 do not need to be connected to each other, nor to storage cells 122. The input of each storage cell 124 can be connected only to the input module 110, without being connected to other storage cells 124 or 122. The input of each storage cell 124 can receive the data to be written only from the input module 110, without shifting and storing data from other storage cells 124 or 122. The configuration of the storage cells 124 can be similar to the configuration of the storage cells 122, and will not be described in detail here.

[0061] For example, the input module 110 can also be configured to receive N rows × M columns of second data, and the storage module 120 can be configured to write each column of the second data into a corresponding column of the second plurality of storage cells 124 in a corresponding cycle of the second loop. In other words, while reading the first data from the first plurality of storage cells 122, the second data can be written to the second plurality of storage cells 124.

[0062] Specifically, during the process of writing second data to the second plurality of storage cells 124, no data shifting operation occurs in the storage module 120; instead, a new column of data is stored into a new column of storage cells 124 each time. For example, referring to FIG3, the write and read processes occurring at the first plurality of storage cells 122 in the first and second cycles are the same as those described above with respect to FIG2. In addition, in the second cycle, the two-dimensional data... The data is input column-by-column to the storage module 120 via the input module 110. Specifically, in the first cycle of the second loop, the first column of the two-dimensional data B is input. Store in the first column of storage cells 1241, 1245, 1249, and 124. 13 In the second cycle of the second loop, the second column of two-dimensional data B. Stored in the second column of storage cells 1242, 1246, and 124. 10 124 14 In the third cycle of the second loop, the third column of the two-dimensional data B. Stored in the third column of storage cells 1243, 1247, and 124. 11 124 15 In the fourth cycle of the second loop, the fourth column of the two-dimensional data B. Stored in the fourth column of storage cells 1244, 1248, and 124. 12 124 16Thus, after four cycles, the second plurality of storage units 124 of storage module 120 are filled with two-dimensional data B. It should be understood that the above sequence is merely exemplary and not restrictive. Because it is not limited by shift operations, each column of the second data can actually be stored in any of the storage units 124 that have not yet been filled with new data in this cycle.

[0063] Similarly, the writing to each storage unit 124 can be controlled by various suitable methods such as clock signals or selection signals, which will not be elaborated here. In some embodiments, the input module 110 may include N buffer units, each buffer unit being used to temporarily store a corresponding row of data (i.e., one data) in the currently received column of data. Each buffer unit can be connected to each storage unit 122 in the corresponding row of storage units 122 and to each storage unit 124 in the corresponding row of storage units 124. For example, referring to FIG3, in the first cycle of the second loop, the four buffer units (not shown) of the input module 110 can respectively store each row of data b00, b10, b20, b30 in the first column of data b_0 of the two-dimensional data B. The first buffer unit storing data b00 can be connected to the first row of storage units 1221, 1222, 1223, 1224 and 1241, 1242, 1243, 1244, so that data b00 can be provided as wd to the input terminal D of each of the first row of storage units 1221, 1222, 1223, 1224 and 1241, 1242, 1243, 1244. At this time, referring to (A) in Figure 10, the clock control terminal G of the storage cell 1241 in the first column of the second N×M matrix can be flipped to write the data b00 into the storage cell 1241. At the same time, the clock control terminals G of the storage cells 1221, 1222, 1223, 1224 in the first N×M matrix and the storage cells 1242, 1243, 1244 in the other columns of the second N×M matrix are prevented from flipping to write the data b00 into the storage cells 1221, 1222, 1223, 1224 and 1242, 1243, 1244. Similarly, the input terminals D of the second, third, and fourth row storage units are also connected to the second buffer unit storing data b10, the third buffer unit storing data b20, and the fourth buffer unit storing data b30, respectively, to receive the corresponding data b10, b20, and b30 from them. At this time, the clock control terminal G of the storage unit located in the first column of the second N×M matrix can be flipped to write the corresponding data b10, b20, and b30 into it, and the clock control terminal G of the storage unit located in the first N×M matrix and the storage units in other columns of the second N×M matrix cannot be flipped to prevent the corresponding data b10, b20, and b30 from being written into it.

[0064] The above writing process does not involve data shifting operations. Only one column of storage cells 124 is flipped in each cycle. This low flip rate results in low dynamic power consumption.

[0065] The selection module 130 is also configured to output data read from each row of storage cells in the second plurality of storage cells 124 in a corresponding cycle of the third cycle after the second cycle. The output module 140 is also configured to output the data received from the selection module 130 as transposed second data. Specifically, no data shifting operation occurs in the storage module 120 during the reading of the second data from the second plurality of storage cells 124. For example, referring to FIG3, in the third cycle, the two-dimensional data B is output row by row from the storage module 120 via the selection module 130 and the output module 140. Specifically, in the first cycle of the third cycle, the fourth row data b3_ = (b30 b31 b32 b33) of the two-dimensional data B is output from the fourth row storage cell 124. 13 124 14 124 15 124 16 Output; in the second cycle of the third loop, the third row data b2_ = (b20 b21 b22 b23) of the two-dimensional data B is output from the third row storage units 1249, 124. 10 124 11 124 12 The data is output; in the third cycle of the third loop, the second row of data B, b1_ = (b10 b11 b12 b13), is output from the second row storage units 1245, 1246, 1247, and 1248; in the fourth cycle of the third loop, the first row of data B, b0_ = (b00 b01 b02 b03), is output from the first row storage units 1241, 1242, 1243, and 1244. Thus, after four cycles, the transposed two-dimensional data B is completely output from the second plurality of storage units 124 of the storage module 120. It is understood that the above sequence is merely exemplary and not restrictive. Because it is not limited by shift operations, data in any row of storage units 124 that has not yet been read in the current cycle can actually be read in each cycle.

[0066] The above readout process does not involve data shifting operations, and no memory cell 124 is flipped in each cycle. This low flip rate results in low dynamic power consumption.

[0067] In some embodiments, the input module 110 is further configured to receive N rows × M columns of third data, the storage module 120 is further configured to write each column of the third data into a corresponding column of storage cells 122 in a corresponding cycle of the third loop, the selection module 130 is further configured to output the data read from a corresponding row of storage cells 122 in a corresponding cycle of the fourth loop after the third loop, and the output module 140 is further configured to output the data received from the selection module 130 as transposed third data. In other words, while reading the second data from the second plurality of storage cells 124, the third data can be written to the first plurality of storage cells 122. Similarly, while reading the third data from the first plurality of storage cells 122, new data can be written to the second plurality of storage cells 124.

[0068] In the case where the storage module 120 of the transpose circuit 200 includes both a first plurality of storage units 122 and a second plurality of storage units 124, each write process (such as a loop with an odd ordinal number, like the first loop) requires M cycles, and each read process (such as a loop with an even ordinal number, like the second loop) requires N cycles. Each loop can include MAX(N,M) cycles (the MAX function is used to take the maximum value between N and M). By alternating between the write and read processes, each loop (MAX(N,M) cycles) yields a transposed result of two-dimensional data. Therefore, compared to the transpose circuit 100, the transpose circuit 200 can have an improved data throughput. In addition, compared to the transpose circuit 10, the transpose circuit 200 can have a comparable data throughput.

[0069] It is understandable that when N equals M, the number of cycles required for the write and read processes is the same. Therefore, each loop can consist of N cycles, and in each of these N cycles, a write to the corresponding column of storage cells 122 and a read from the corresponding row of storage cells 124 can be performed simultaneously, or a read from the corresponding row of storage cells 122 and a write to the corresponding column of storage cells 124 can be performed simultaneously. When N is not equal to M, for example, assuming N is greater than M, each loop can consist of N cycles, and in each of these N cycles, a read from the corresponding row of storage cells 122 can be performed, while a write to the corresponding column of storage cells 124 can be performed in each of the M selected cycles within these N cycles, or a read from the corresponding row of storage cells 124 can be performed in each of these N cycles, while a write to the corresponding column of storage cells 122 can be performed in each of the M selected cycles within these N cycles. The case where N is less than M is similar and will not be elaborated further here.

[0070] Furthermore, even under full load (i.e., receiving new data to be transposed in each cycle), at most one column of memory cells 122 or 124 in the transpose circuit 200 undergoes data flipping in each cycle. This low flip rate results in lower dynamic power consumption, especially lower than that of the transpose circuit 10, where all registers undergo data flipping in each cycle. Moreover, even though the number of memory cells in the transpose circuit 200 is twice the number of registers in the transpose circuit 10 when used as transpose units of the same specifications, the transpose circuit 200 still has a significant area advantage over the transpose circuit 10 because the memory cells in the transpose circuit 200 do not need to be implemented as registers, but can instead use circuit elements with the smallest possible area (such as latch units) that satisfy the storage function.

[0071] In the embodiment of FIG3, the first plurality of storage units 122 and the second plurality of storage units 124 share the same input module 110. This can advantageously save circuit area. In some embodiments, the input module 110 may also include a first input module separately configured for the first plurality of storage units 122 and a second input module separately configured for the second plurality of storage units 124. When the first input module is configured as a column input module as described above with respect to input module 110, the second input module may be configured as either a column input module or a row input module. That is, the first plurality of storage units 122 and the second plurality of storage units 124 may both be used to perform column-to-row output conversion, or one may be used to perform column-to-row output conversion while the other performs row-to-column output conversion. In the latter case, the corresponding modules of the circuit can be adaptively adjusted, which will not be elaborated here.

[0072] In some embodiments, the selection module 130 of the transpose circuit 200 may include a combination of multiplexers having 2N inputs and 1 output. Each of the 2N inputs is connected to an N-row of memory cells in a first plurality of memory cells and a corresponding row of memory cells in the N-row of memory cells in a second plurality of memory cells 124. The output is connected to the output module 140. Similarly, when a 2N-to-1 multiplexer is available, it can be used directly. Larger-scale multiplexing functionality can also be achieved by cascading multiple smaller-scale multiplexers, especially when 2N-to-1 multiplexers may be difficult to obtain.

[0073] Further, the selection module 130 of the transpose circuit 200 may include first to third selection submodules. The first selection submodule may include a first multiplexer combination having N inputs and y outputs, where y is an integer and 0 < y < N. Each of the N inputs of the first multiplexer combination can be connected to an x-row of storage cells in the first plurality of storage cells 122 and a corresponding row of storage cells in the Nx-row of storage cells in the second plurality of storage cells 124, where x is an integer and 0 ≤ x ≤ N. The second selection submodule may include a second multiplexer combination having N inputs and z outputs, where z is an integer and 0 < z < N. Each of the N inputs of the second multiplexer combination is connected to the remaining Nx-row of storage cells in the first plurality of storage cells 122 and a corresponding row of storage cells in the remaining x-row of storage cells in the second plurality of storage cells. The third selection submodule may include a third multiplexer combination having y+z inputs and 1 output. The y outputs of the first multiplexer combination and the z outputs of the second multiplexer combination can be connected to the corresponding inputs of the y+z inputs of the third multiplexer combination, and one output of the third multiplexer combination can be connected to the output module 140.

[0074] In some embodiments, x equals 0 or N. In such cases, one of the first selection submodule and the second selection submodule is dedicated to controlling the sequential output of data in the first plurality of storage units 122, while the other of the first selection submodule and the second selection submodule is dedicated to controlling the sequential output of data in the second plurality of storage units 124, thereby decoupling the operation of the first plurality of storage units 122 from that of the second plurality of storage units 124. For example, Figure 4 illustrates an example implementation 130A of the selection module 130, wherein the first selection submodule 132 is implemented as a 4-to-1 multiplexer (MUX 4-1) and is used to selectively output data read from one row of storage cells 122 in the first plurality of storage cells 122; the second selection submodule 134 is also implemented as a 4-to-1 multiplexer (MUX 4-1) and is used to selectively output data read from one row of storage cells 124 in the second plurality of storage cells 124; and the third selection submodule 136 is implemented as a 2-to-1 multiplexer (MUX 2-1) and is used to selectively output the output of one of the selection submodules, the first selection submodule 132 and the second selection submodule 134. Specifically, referring to Figure 4, the first input terminal of the first selection submodule 132 is connected to the first row of storage units 1221, 1222, 1223, and 1224 in the first plurality of storage units 122; the second input terminal is connected to the second row of storage units 1225, 1226, 1227, and 1228; and the third input terminal is connected to the third row of storage units 1229 and 1220. 10 12211 122 12 The fourth input terminal is connected to the fourth row of storage unit 122. 13 122 14 122 15 122 16 The output terminal is connected to one input terminal of the third selection submodule 136. Additionally, the first input terminal of the second selection submodule 134 is connected to the first row of storage units 1241, 1242, 1243, and 1244 in the second plurality of storage units 124; the second input terminal is connected to the second row of storage units 1245, 1246, 1247, and 1248; and the third input terminal is connected to the third row of storage units 1249 and 1240. 10 124 11 124 12 The fourth input terminal is connected to the fourth row of storage unit 124. 13 124 14 124 15 124 16 The output is connected to another input of the third selection submodule 136. In the example shown in Figure 4, y = z = 1, but this is not restrictive; y and z do not have to be equal to 1, as long as a unique output is ultimately achieved through the third selection submodule 136.

[0075] In other embodiments, x is not equal to 0 or N. In such cases, the first selection submodule and the second selection submodule each receive data read from one or more rows of storage cells 122 in the first plurality of storage cells 122, and also receive data read from one or more rows of storage cells 124 in the second plurality of storage cells 124. Since the reading processes of the first plurality of storage cells 122 and the reading processes of the second plurality of storage cells 124 do not occur simultaneously, this configuration can still operate correctly.

[0076] In some examples, x = y = z = N / 2. For example, a first multiplexer combination may include N / 2 2-to-1 multiplexers, each with a first input connected to a corresponding row of storage cells 122 in the x rows of storage cells 122 of a first plurality of storage cells 122 and a second input connected to a corresponding row of storage cells 124 in the Nx rows of storage cells 124 of a second plurality of storage cells 124. A second multiplexer combination may also include N / 2 2-to-1 multiplexers, each with a first input connected to a corresponding row of storage cells 122 in the remaining Nx rows of storage cells 122 of the first plurality of storage cells 122 and a second input connected to a corresponding row of storage cells 124 in the remaining x rows of storage cells 124 of the second plurality of storage cells 124. For example, Figure 5 illustrates another example implementation 130B of the selection module 130. The first selection submodule 132 is implemented as two 2-to-1 multiplexers (MUX 2-1), one of which is used to selectively output data read from the first row of storage cells 122 (e.g., storage cells 1221, 1222, 1223, 1224 as shown) in the first plurality of storage cells 122 or from the first row of storage cells 124 (e.g., storage cells 1241, 1242, 1243, 1244 as shown), and the other 2-to-1 multiplexer is used to selectively output data read from the second row of storage cells 122 (e.g., storage cells 1225, 1226, 1227, 1228 as shown) in the first plurality of storage cells 122 or from the second row of storage cells 124 (e.g., storage cells 1245, 1246, 1247, 1248 as shown). The second selection submodule 134 is also implemented as two 2-to-1 multiplexers (MUX 2-1), one of which is used to selectively output from the third row of memory cells 122 (e.g., memory cells 1229, 122 as shown in the figure) from the first plurality of memory cells 122. 10 122 11 122 12 Data read from the second or more storage units 124 or from the third row of storage units 124 (e.g., storage units 1249, 124 as shown in the figure). 10 124 11 124 12 The data read from the first plurality of storage cells 122, and another 2-to-1 multiplexer is used to selectively output the data from the fourth row of storage cells 122 (e.g., storage cell 122 as shown in the figure). 13 122 14 12215 122 16 Data read from the second or more storage cells 124 or from the fourth row of storage cells 124 (e.g., storage cell 124 as shown in the figure). 13 124 14 124 15 124 16 The data read from the first selection submodule 132 and the second selection submodule 134. The third selection submodule 136 is implemented as a 4-to-1 multiplexer (MUX 4-1) and is used to selectively output the output of one of the 2-to-1 multiplexers in the first selection submodule 132 and the second selection submodule 134. In the example shown in Figure 5, y = z = 2, but this is not restrictive, and y and z can be other than 2, as long as a single output is ultimately achieved through the third selection submodule 136.

[0077] It is understood that the above implementation of dividing the selection module 130 into first to third selection sub-modules is a logical functional division, which is exemplary and not restrictive. Other suitable division methods are also feasible.

[0078] To further illustrate the write and read logic of the transpose circuit 200, Figure 6 shows a non-limiting example implementation of the data write portion of the transpose circuit 200, and Figures 7 and 8 show various non-limiting example implementations of the data read portion of the transpose circuit 200.

[0079] As shown in Figure 6, each long horizontal bar represents a column of memory cells. A first flip-flop (DFF1) is arranged upstream of these memory cells, configured to output each column of received data to the corresponding column of memory cells based on the received clock signal. The first flip-flop (DFF1) can be always enabled; its function is to cut off external timing and ensure internal timing, rather than to control the selective writing of data to specific memory cells. A column of data received via the first flip-flop (DFF1) can be written to the corresponding column of memory cells in the eight columns of memory cells of the transpose circuit 200's memory module 120 according to corresponding control information (e.g., by controlling the signal toggling at the clock control terminal of the memory cell). One column of data can be written in each cycle, and at most one column of memory cells will experience data toggling in each cycle.

[0080] As shown in Figures 7 and 8, each long horizontal bar represents a row of memory cells. Multiple 2-to-1 multiplexers are cascaded together to provide selection from eight inputs to one output. Downstream of these multiplexers is a second flip-flop DFF2, configured to output the data read from each row of memory cells based on a received clock signal. The second flip-flop DFF2 can be always enabled; its function is to cut off external timing and ensure internal timing, rather than to control the selective reading of data from specific memory cells. A row of data is read from the corresponding row of memory cells in the eight rows of memory cells of the transpose circuit 200's memory module 120 using these multiplexers according to the corresponding control information (e.g., via a selection signal, etc.) and output via the second flip-flop DFF2. At most one row of data can be read in each cycle, and the read operation does not cause data flipping in the memory cells. The difference between Figures 7 and 8 is that in Figure 7, the reading of each row of memory cells 122 is decoupled from the reading of each row of memory cells 124.

[0081] Accordingly, this disclosure also provides a transposition method. Figure 9 illustrates a transposition method 300 according to some embodiments of this disclosure. As shown in Figure 9, the transposition method 300 includes: in step S302, receiving N rows × M columns of first data and writing each column of the first data into a corresponding column of a first plurality of storage units of N rows × M columns in a corresponding cycle of a first loop; in step S304, outputting data read from each row of storage units in the first plurality of storage units in a corresponding cycle of a second loop after the first loop, and receiving N rows × M columns of second data and writing each column of the second data into a corresponding column of a second plurality of storage units of N rows × M columns in a corresponding cycle of the second loop; in step S306, outputting data read from each row of storage units in the second plurality of storage units in a corresponding cycle of a third loop after the second loop, wherein N and M are positive integers and at least one of N and M is greater than 1.

[0082] In some embodiments, step S306 further includes receiving N rows × M columns of third data and writing each column of the third data into a corresponding column of storage cells in a corresponding cycle of the third loop. In some embodiments, the transpose method 300 further includes outputting data read from each row of storage cells in the first plurality of storage cells in a corresponding cycle of the fourth loop after the third loop. Various embodiments of the transpose method 300 can be referred to similarly to the various embodiments described above with respect to the transpose circuit, and will not be repeated here.

[0083] In another aspect, this disclosure also provides a two-dimensional discrete cosine transform apparatus, which includes a transpose circuit according to any embodiment of this disclosure, wherein N equals M.

[0084] In another aspect, this disclosure also provides a processor that includes a transpose circuit according to any embodiment of this disclosure or a two-dimensional discrete cosine transform device according to any embodiment of this disclosure. For example, such a processor can be a neural network processor, a central processing unit, a coprocessor, a digital signal processor, an image processor, a video processor, a dedicated instruction processor, or various other processors.

[0085] This disclosure also provides a computing device that includes a transposed circuit according to any embodiment of this disclosure, a two-dimensional discrete cosine transform device according to any embodiment of this disclosure, or a processor according to any embodiment of this disclosure. Examples of such computing devices may include, but are not limited to, consumer electronics products, components of consumer electronics products, electronic testing equipment, cellular communication infrastructure such as base stations, etc. Examples of computing devices may include, but are not limited to, mobile phones such as smartphones, wearable computing devices such as smartwatches or headphones, telephones, televisions, computer monitors, computers, modems, handheld computers, laptop computers, tablet computers, personal digital assistants (PDAs), microwave ovens, refrigerators, in-vehicle electronic systems such as automotive electronic systems, stereo systems, DVD players, CD players, digital music players such as MP3 players, radios, portable video cameras, cameras such as digital cameras, portable storage chips, washing machines, dryers, washer / dryer systems, peripheral devices, clocks, etc. Furthermore, computing devices may include incomplete products.

[0086] The terms “left,” “right,” “front,” “back,” “top,” “bottom,” “upper,” “lower,” “high,” “lower,” etc., used in the specification and claims, if present, are for descriptive purposes and not necessarily for describing unchanging relative positions. It should be understood that such terms are interchangeable where appropriate, so that embodiments of this disclosure described herein can operate, for example, in orientations different from those shown or otherwise described herein. For example, when the device in the drawings is reversed, a feature previously described as “above” other features may now be described as “below” other features. The device may also be oriented in other ways (rotated 90 degrees or in other orientations), in which case the relative spatial relationships will be interpreted accordingly.

[0087] In the specification and claims, when an element is described as being "on top of," "attached to," "connected to," "coupled to," or "in contact with" another element, the element may be directly located on top of, directly attached to, directly connected to, directly coupled to, or directly in contact with the other element, or one or more intermediate elements may be present. Conversely, when an element is described as being "directly" located on top of, directly attached to, directly connected to, directly coupled to, or directly in contact with another element, no intermediate elements are present. In the specification and claims, when a feature is arranged "adjacent" to another feature, it may mean that a feature has a portion overlapping with the adjacent feature or a portion located above or below the adjacent feature.

[0088] As used herein, the term "exemplary" means "serving as an example, instance, or illustration," and not as a "model" to be precisely copied. Any implementation described herein by example is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, this disclosure is not limited to any theory expressed or implied as given in the art, background, summary of the invention, or detailed description.

[0089] As used herein, the term "substantially" means any minor variation resulting from design or manufacturing defects, device or component tolerances, environmental influences, and / or other factors. The term "substantially" also allows for differences from the perfect or ideal situation due to parasitic effects, noise, and other practical considerations that may exist in the actual implementation.

[0090] Additionally, terms such as “first,” “second,” etc., may be used in this document for reference purposes only and are not intended to be limiting. For example, unless the context clearly indicates otherwise, the words “first,” “second,” and other such numerical terms relating to structures or elements do not imply order or sequence.

[0091] It should also be understood that the term "including / comprises" as used herein indicates the presence of the indicated feature, whole, step, operation, unit, and / or component, but does not preclude the presence or addition of one or more other features, wholes, steps, operations, units, and / or components, and / or combinations thereof. In this disclosure, the term "provide" is used broadly to cover all ways of obtaining an object; therefore, "providing an object" includes, but is not limited to, "purchasing," "preparing / manufacturing," "arranging / setting," "installing / assembling," and / or "ordering" an object.

[0092] As used herein, the term “and / or” includes any and all combinations of one or more of the listed items in association. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise.

[0093] The same or similar parts between the various embodiments of this disclosure can be referred to mutually, and each embodiment focuses on describing the differences from other embodiments. In the description of this disclosure, the reference to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," "exemplary," etc., means that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this disclosure, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this disclosure and the features of the different embodiments or examples.

[0094] Additionally, when used in this disclosure, the terms “here,” “above,” “below,” “this,” “the following,” “the text,” “the preceding,” and similar terms should refer to the entire disclosure and not any particular part of it. Furthermore, unless expressly stated otherwise or otherwise understood in the context in which they are used, conditional language used herein, such as “may,” “possibly,” “for example,” “like,” etc., is generally intended to express that certain embodiments include, while other embodiments do not, certain features, elements, and / or states. Therefore, such conditional language is not generally intended to imply that one or more embodiments require features, elements, and / or states in any way, or whether such features, elements, and / or states are included or performed in any particular embodiment.

[0095] Those skilled in the art will recognize that the boundaries between the above operations are merely illustrative. Multiple operations may be combined into a single operation, a single operation may be distributed among additional operations, and operations may be performed with at least partial overlap in time. Moreover, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be changed in various other embodiments. However, other modifications, variations, and substitutions are equally possible. Aspects and elements of all the embodiments disclosed above may be combined in any way and / or in combination with aspects or elements of other embodiments to provide multiple additional embodiments. Therefore, this specification and the accompanying drawings should be considered illustrative rather than restrictive.

[0096] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. The various embodiments disclosed herein can be combined in any way without departing from the spirit and scope of this disclosure. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A transpose circuit, comprising: The input module is configured to receive the first data in N rows × M columns; A storage module, connected to the input module and comprising a first plurality of storage units with N rows × M columns, the storage module being configured to write each column of the first data into a corresponding column of the first plurality of storage units in a corresponding cycle of the first loop. The selection module is connected to the storage module and configured to output data read from each row of storage cells in the first plurality of storage cells in a corresponding cycle of the second cycle after the first cycle; An output module, connected to the selection module and configured to output the data received from the selection module as transposed first data, Where N and M are positive integers, and at least one of N and M is greater than 1.

2. The transpose circuit according to claim 1, wherein, The selection module includes a combination of multiplexers with N inputs and 1 output. Each of the N inputs is connected to a corresponding row of storage cells in the first plurality of storage cells, and the 1 output is connected to the output module.

3. The transpose circuit according to claim 1, wherein, The input module is further configured to receive second data in N rows × M columns, and the storage module further includes a second plurality of storage units in N rows × M columns and is configured to write each column of the second data into a corresponding column of the second plurality of storage units in a corresponding cycle of the second loop. The selection module is further configured to output the data read from each row of storage cells in the second plurality of storage cells in a corresponding cycle of the third cycle after the second cycle. The output module is further configured to output the data received from the selection module as transposed second data.

4. The transpose circuit according to claim 3, wherein, The input module is also configured to receive third data in N rows × M columns. The storage module is further configured to write each column of the third data into a corresponding column of the first plurality of storage units in a corresponding cycle of the third cycle. The selection module is also configured to output the data read from each row of storage cells in the first plurality of storage cells in a corresponding cycle of the fourth cycle following the third cycle. The output module is also configured to output the data received from the selection module as transposed third data.

5. The transpose circuit according to claim 3, wherein, The selection module includes a multiplexer combination with 2N input terminals and 1 output terminal. Each of the 2N input terminals is connected to a corresponding row of storage cells in the N rows of the first plurality of storage cells and the N rows of storage cells in the second plurality of storage cells. The 1 output terminal is connected to the output module.

6. The transpose circuit according to claim 5, wherein, The selection module includes: The first selection submodule includes a first multiplexer combination having N input terminals and y output terminals, wherein each of the N input terminals of the first multiplexer combination is connected to a corresponding row of storage units in the x rows of the first plurality of storage units and the Nx rows of storage units in the second plurality of storage units; The second selection submodule includes a second multiplexer combination having N inputs and z outputs, wherein each of the N inputs of the second multiplexer combination is connected to a corresponding row of storage cells in the remaining Nx rows of the first plurality of storage cells and the remaining x rows of storage cells in the second plurality of storage cells. The third selection submodule includes a third multiplexer combination having y+z input terminals and 1 output terminal. The y output terminals of the first multiplexer combination and the z output terminals of the second multiplexer combination are respectively connected to the corresponding input terminals of the y+z input terminals of the third multiplexer combination. The 1 output terminal of the third multiplexer combination is connected to the output module. Where x is an integer and 0 ≤ x ≤ N, y is an integer and 0 < y < N, and z is an integer and 0 < z < N.

7. The transpose circuit according to claim 6, wherein, x = 0 or N, y = z = 1.

8. The transpose circuit according to claim 6, wherein, x=y=z=N / 2, and where, The first multiplexer combination includes N / 2 2-to-1 multiplexers, each 2-to-1 multiplexer having a first input connected to a corresponding row of memory cells in the x rows of the first plurality of memory cells and a second input connected to a corresponding row of memory cells in the Nx rows of memory cells in the second plurality of memory cells. The second multiplexer combination includes N / 2 2-to-1 multiplexers, each of which has a first input connected to a corresponding row of the remaining Nx rows of the first plurality of storage units and a second input connected to a corresponding row of the remaining x rows of the second plurality of storage units.

9. The transpose circuit according to any one of claims 1 to 8, wherein the transpose circuit comprises at least one of the following: Included in the input module is a first trigger, configured to output each column of received data to a corresponding column of storage cells based on a received clock signal; or The output module includes a second flip-flop configured to output the data read from each row of memory cells based on the received clock signal.

10. The transpose circuit according to any one of claims 1 to 8, wherein, The storage unit includes: latch unit; or Dynamic Random Access Memory (DRAM) cells; or Static Random Access Memory (SRAM) unit.

11. The transpose circuit according to any one of claims 1 to 8, wherein, N equals M.

12. The transpose circuit according to any one of claims 1 to 8, wherein, The data is image data.

13. A two-dimensional discrete cosine transform device, comprising a transpose circuit according to any one of claims 1 to 12, wherein N equals M.

14. A processor comprising a transpose circuit according to any one of claims 1 to 12 or a two-dimensional discrete cosine transform device according to claim 13.

15. A computing device comprising a transpose circuit according to any one of claims 1 to 12, a two-dimensional discrete cosine transform device according to claim 13, or a processor according to claim 14.

16. A transpose method, comprising: Receive the first data in N rows × M columns and write each column of the first data into the corresponding column of the first plurality of storage units in the first cycle of the first loop. The data read from each row of storage cells in the first plurality of storage cells is output in a corresponding cycle of the second cycle after the first cycle, and the second data of N rows × M columns is received and each column of the second data is written into a corresponding column of storage cells in the second plurality of storage cells of N rows × M columns in a corresponding cycle of the second cycle. The data read from each row of storage cells in the second plurality of storage cells is output in the corresponding cycle of the third loop after the second loop. Where N and M are positive integers, and at least one of N and M is greater than 1.

17. The transposition method according to claim 16, further comprising: Receive N rows × M columns of third data and write each column of the third data into the corresponding column of the first plurality of storage units in the corresponding cycle of the third cycle; The data read from each row of storage cells in the first plurality of storage cells is output in a corresponding cycle of the fourth cycle after the third cycle.