Data processing method and device, electronic equipment, chip and storage medium
By adopting specific data writing and reading rules in the high-throughput FFT scenario and using pseudo-dual-port random access memory RAM, the problem of too many RAM blocks in FFT operations is solved, and the RAM area is reduced and data processing efficiency is improved.
Patent Information
- Application Number
- CN202411545630.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-07-25
AI Technical Summary
In high throughput scenarios, fast Fourier transform (FFT) operations require a large number of random access memory (RAM) blocks, resulting in large RAM area consumption.
By dividing the data into multiple data groups in multiple write cycles and writing them into multiple memory units in each cycle, combining specific read and write rules to reduce the number of RAM blocks, and using pseudo-dual-port random access memory RAM to perform data splicing and processing.
In the high throughput FFT scenario, the number of RAM blocks is reduced, the area consumption of RAM is reduced, and data processing efficiency is improved.
Smart Images

Figure CN120371200A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and particularly to a data processing method, apparatus, electronic device, chip, and storage medium. Background Art
[0002] The Fourier transform theory establishes the relationship between the time domain and the frequency domain, and becomes a basic analysis theory in fields such as signal analysis, linear systems, and probability theory. With the development of large-scale integrated circuits and digital signal processing technologies, the Fast Fourier Transform (FFT) algorithm, as a key technology in the field of digital signal processing, plays an irreplaceable role and is widely used in fields such as communication, radar, and image processing.
[0003] For FFT operations in high-throughput scenarios, a large number of Random Access Memory (RAM) blocks are required, consuming a large amount of RAM area. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, electronic device, chip, and storage medium to solve the problems in the related art.
[0005] In a first aspect embodiment of the present disclosure, a data processing method is proposed. The method includes: obtaining a data sequence to be processed, where the data sequence to be processed includes N data; in each of multiple writing cycles, respectively writing M data out of the N data into P storage units, and the storage capacity corresponding to each address of each storage unit is Q; in each of multiple reading cycles, respectively reading Q / P data from the P storage units, splicing them, and processing the spliced Q data, where N, M, P, and Q are positive integers, M = P × Q, and P < M.
[0006] In some embodiments of the present disclosure, writing M data out of the N data into P storage units in each of multiple writing cycles includes: dividing the N data into K data groups, where K is a positive integer and K is equal to the number of cycles of the multiple writing cycles, and each data group includes M data; determining the writing rule for the M data in each data group; based on the writing rule, in each writing cycle, writing the M data in each data group into the P storage units.
[0007] In some embodiments of the present disclosure, determining the writing rule for the M data in each data group includes: determining the reading rule for reading the N data according to at least one of N, M, P, and Q; and determining the writing rule according to the reading rule.
[0008] In some embodiments of the present disclosure, N = 256, M = 64, P = 4, Q = 16. Among them, the writing rule is as follows: for the first writing cycle, write 64 data in the first data group among the K data groups to the first address of the first storage unit, the second address of the second storage unit, the third address of the third storage unit, and the fourth address of the fourth storage unit; for the second writing cycle, write 64 data in the second data group among the K data groups to the third address of the first storage unit, the fourth address of the second storage unit, the first address of the third storage unit, and the second address of the fourth storage unit; for the third writing cycle, write 64 data in the third data group among the K data groups to the fourth address of the first storage unit, the first address of the second storage unit, the second address of the third storage unit, and the third address of the fourth storage unit; for the fourth writing cycle, write 64 data in the fourth data group among the K data groups to the second address of the first storage unit, the third address of the second storage unit, the fourth address of the third storage unit, and the first address of the fourth storage unit.
[0009] In some embodiments of the present disclosure, N = 256, M = 64, P = 4, Q = 16. Among them, the reading rule is as follows:
[0010] In 4 read cycles of the first read, 64 data are read from the first addresses of P storage units respectively. Among them, in the first read cycle of the first read, the first group of 4 data of the first address of the first storage unit, the second group of 4 data of the first address of the second storage unit, the third group of 4 data of the first address of the third storage unit, and the fourth group of 4 data of the first address of the fourth storage unit are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the first read are completed, obtaining 4 data sequences of the first read, and each data sequence includes 16 data; in 4 read cycles of the second read, 64 data are read from the second addresses of P storage units respectively. Among them, in the first read cycle of the second read, the fifth group of 4 data of the second address of the second storage unit, the sixth group of 4 data of the second address of the third storage unit, the seventh group of 4 data of the second address of the fourth storage unit, and the eighth group of 4 data of the second address of the first storage unit are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the second read are completed, obtaining 4 data sequences of the second read, and each data sequence includes 16 data; in 4 read cycles of the third read, 64 data are read from the third addresses of P storage units respectively. Among them, in the first read cycle of the third read, the ninth group of 4 data of the third address of the third storage unit, the tenth group of 4 data of the third address of the fourth storage unit, the eleventh group of 4 data of the third address of the first storage unit, and the twelfth group of 4 data of the third address of the second storage unit are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the third read are completed, obtaining 4 data sequences of the third read, and each data sequence includes 16 data; in 4 read cycles of the fourth read, 64 data are read from the fourth addresses of P storage units respectively. Among them, in the first read cycle of the fourth read, the thirteenth group of 4 data of the fourth address of the fourth storage unit, the fourteenth group of 4 data of the fourth address of the first storage unit, the fifteenth group of 4 data of the fourth address of the second storage unit, and the sixteenth group of 4 data of the fourth address of the third storage unit are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the fourth read are completed, obtaining 4 data sequences of the fourth read, and each data sequence includes 16 data.
[0011] In some embodiments of the present disclosure, in each of a plurality of read cycles, Q / P data are respectively read from P memory cells, and processing the Q / P data includes: performing a first-level radix-16 fast Fourier transform (FFT) operation on a first data sequence read out in each read cycle to obtain a first-level output sequence, where both the first data sequence and the first-level output sequence include Q data; writing the first-level output sequence back to the write-back addresses of the P memory cells, where the write-back addresses are the same as the read addresses of the first data sequence; reading the first-level output sequence and performing a second-level radix-16 FFT operation to obtain a second-level output sequence.
[0012] In some embodiments of the present disclosure, writing the first-level output sequence back to the write-back addresses of the P memory cells includes: determining, when reading the first data sequence, Q / P data bits of each memory cell, where the Q / P data bits of each memory cell are the read addresses when reading the first data sequence; adding a mask to the Q - Q / P data bits of each memory cell; and writing the Q data in the first-level output sequence back to the Q / P data bits of each of the P memory cells respectively.
[0013] In some embodiments of the present disclosure, N = 256, M = 64, P = 4, Q = 16; and / or the memory cell is a pseudo dual-port random access memory (RAM); and / or the number of write cycles for processing N data at a time is K, the number of read cycles is L, and the total number of cycles for processing H times of N data < H×(K + L), and the total number of cycles for processing H times of N data ≥ K + L+(H - 1)×L / 2.
[0014] A second aspect embodiment of the present disclosure provides a data processing device, which includes: an acquisition module for acquiring a data sequence to be processed, where the data sequence to be processed includes N data; a writing module for writing M data out of the N data into P memory cells respectively in each of a plurality of write cycles, where the storage capacity corresponding to each address of each memory cell is Q; a reading and processing module for reading Q / P data from the P memory cells respectively in each of a plurality of read cycles and processing the Q / P data, where N, M, P, and Q are positive integers, M = P×Q, and P < M.
[0015] A third aspect embodiment of the present disclosure provides a storage device, which includes: at least one storage module, where each storage module includes P memory cells, and the storage capacity corresponding to each address of each memory cell is Q; where, when the data sequence to be processed includes N data, each storage module can write M data in each of a plurality of write cycles, and N, M, P, and Q are positive integers, M = P×Q, and P < M.
[0016] In some embodiments of the present disclosure, N = 256, M = 64, P = 4, Q = 16; the storage unit is a pseudo dual-port random access memory (RAM).
[0017] In some embodiments of the present disclosure, the number of at least one storage module is 2. Among them, the number of write cycles for processing N data at a time is K, and the read cycle is L. The total number of cycles for processing H times of N data < H×(K + L), and the total number of cycles for processing H times of N data ≥ K + L+(H - 1)×L / 2.
[0018] An embodiment of the fourth aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the embodiment of the first aspect of the present disclosure.
[0019] An embodiment of the fifth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in the embodiment of the first aspect of the present disclosure.
[0020] An embodiment of the sixth aspect of the present disclosure provides a chip, characterized in that it includes at least one processor and a communication interface; the communication interface is used to receive signals input to the chip or signals output from the chip, and the processor communicates with the communication interface and implements the method described in the embodiment of the first aspect of the present disclosure through logic circuits or by executing code instructions.
[0021] In some embodiments of the present disclosure, the above-mentioned chip further includes the storage device of the third aspect.
[0022] In summary, the data processing method proposed by the present disclosure can, by determining the writing and reading rules of data in a high-throughput FFT scenario, reduce the number of RAM blocks and the area consumption of RAM while completing the FFT operation in a high-throughput FFT scenario.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0025] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure;
[0026] Figure 2 Flow schematic diagram of a data processing method provided by an embodiment of the present disclosure;
[0027] Figure 3 Flow schematic diagram of a data processing method provided by an embodiment of the present disclosure;
[0028] Figure 4 Schematic diagram of a 2-level R16 structure provided by an embodiment of the present disclosure;
[0029] Figure 5 Flow schematic diagram of input data provided by an embodiment of the present disclosure;
[0030] Figure 6 Schematic diagram of a data storage method provided by an embodiment of the present disclosure;
[0031] Figure 7 Schematic diagram of a data splicing method provided by an embodiment of the present disclosure;
[0032] Figure 8 Schematic diagram of a data reading method provided by an embodiment of the present disclosure;
[0033] Figure 9 Another schematic diagram of a data splicing method provided by an embodiment of the present disclosure;
[0034] Figure 10 Another schematic diagram of a data reading method provided by an embodiment of the present disclosure;
[0035] Figure 11 Schematic diagram of a structure of a data processing device provided by an embodiment of the present disclosure;
[0036] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure; [[ID=4,2]]
[0037] Figure 13 Schematic diagram of a chip structure provided by an embodiment of the present disclosure. Detailed implementation manners
[0038] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0039] The related technology is the radix-16 FFT fast calculation method. The following is the scheme principle of the radix-16 FFT method.
[0040] Given a finite-length sequence of length N, the defined N-point Discrete Fourier Transform (DFT) is as follows:
[0041]
[0042] In Equation 1 represents the twiddle factor. By continuously decomposing the DFT operation of a long sequence into several short-sequence DFTs and utilizing the periodicity and symmetry, the number of addition and multiplication operations in the DFT can be reduced. The FFT can be divided into the Decimation In Time FFT (DIT FFT) and the Decimation In Frequency FFT (DIF FFT).
[0043] The radix-16 (R16) operation in this article is based on the original radix-2 DIT FFT operation. First, the decomposition of the radix-2 DIT is introduced below. Assume that the sequence length N in the formula satisfies N = 2 M , where M is a natural number. Decompose x(n) into two subsequences of N / 2 points according to the parity of n:
[0044]
[0045] Then the DFT of x(n) is given by Equation 4 below:
[0046]
[0047]
[0048] Since both x1(k) and x2(k) in Equation 4 are periodic with N / 2, and Therefore, X(k) can be expressed as:
[0049]
[0050] In this way, the N-point DFT is decomposed into 2 N / 2-point DFTs. If the process from Equation 4 to Equation 6 is continued, it is easy to obtain 4 N / 4-point DFTs, until finally N / 2 2-point DFTs are obtained, which completes all the radix-2 DIT FFT decompositions.
[0051] Based on the principle of the above radix-2 FFT, the principle of the radix-16 FFT can be determined. The input index of the radix-16 DIT represents the number of the input, and the scrambling rule of the output index is the result of the binary bit-reversal of the sequential number. For example: (1, 0001) after bit-reversal is (8, 1000), and (2, 0010) after bit-reversal is (4, 0100).
[0052] Optionally, the 16-point FFT can be decomposed into a 4-stage radix-2 DIT. Each 2-stage operation with 4 input numbers represents a Radix-4 calculation. Then, it can be decomposed into a 2-stage radix-4 DIT in total, and there are 4 Radix-4 modules in each stage. The inputs of the second-stage radix-4 DIT are the combinations of the four 0 outputs, the four 1 outputs, the four 2 outputs, and the four 3 outputs of the four Radix-4 modules in the first stage, respectively.
[0053] By adopting the radix-16 structure, the number of stages of the FFT operation is reduced, that is, the complete number of times of reading the ram required for the FFT operation is reduced. However, for scenarios with a high input data throughput rate, the related technology requires a large number of RAM blocks, consuming a large amount of RAM area. For example, when performing the FFT operation on 256-point data in the related technology, 16 RAM blocks are required for data storage.
[0054] Therefore, to solve the above problems, the present disclosure proposes a data processing method that can reduce the number of RAM blocks used for the FFT operation in high-throughput scenarios by determining the read and write rules of the data.
[0055] The specific content of this method is as follows.
[0056] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure. As Figure 1 shown, this method may include the following steps.
[0057] Step 101: Obtain a data sequence to be processed.
[0058] In some embodiments, the data sequence to be processed includes N data, where N is a positive integer.
[0059] In some embodiments, the data sequence to be processed may be data that needs to be stored in a storage module, for example, data that needs to be stored in a RAM.
[0060] Step 102: In each of multiple write cycles, write M data out of the N data into P storage units respectively.
[0061] In some embodiments, the storage capacity corresponding to each address of each storage unit is Q. In other words, each address of each storage unit can store Q data.
[0062] In some embodiments, optionally, the storage unit may be a pseudo dual-port random access memory RAM.
[0063] In some embodiments, N data can be written into P storage units in cycles, where each address of each storage unit can store Q data. In each of multiple write cycles, writing M of the N data into the P storage units respectively includes: dividing the N data into K data groups, where K is a positive integer and K is equal to the number of cycles of the multiple write cycles, and each data group includes M data; determining the writing rules for the M data in each data group; and based on the writing rules, in each write cycle, writing the M data in each data group into the P storage units, where N, M, P, and Q are positive integers, M = P × Q, and P < M.
[0064] In other words, N data can be input into the storage units in K cycles, with M data input in each cycle, that is, N = M × K. In each cycle, the M data are respectively written into P storage units, and Q data are written into each storage unit in each cycle. After K cycles, each storage unit has K × Q data.
[0065] Optionally, grouping the N data can be performed in the order of the data index. That is, each data can have a corresponding index value, and grouping can be performed in order according to the size of the index value. For example, when N = 256 and K = 4, the data with index values from 0 to 63 can be grouped into one group and written in the first write cycle; the data with index values from 64 to 127 can be grouped into one group and written in the second write cycle; the data with index values from 128 to 191 can be grouped into one group and written in the third write cycle; and the data with index values from 192 to 255 can be grouped into one group and written in the fourth write cycle.
[0066] Optionally, the grouping of the data can also be other grouping methods, such as non-sequential grouping according to the index value; the writing order can be other orders, which can be determined according to the actual application scenario, and the present disclosure does not limit this.
[0067] Step 103, in each of multiple read cycles, read Q / P data from the P storage units respectively and splice them, and process the spliced Q data.
[0068] In some embodiments, data can be read out in multiple read cycles and the read data can be processed. For example, processing the data can be performing FFT calculation on the data.
[0069] Optionally, the data read from the P storage units can be spliced. For example, the Q / P data read from each storage unit can be spliced, and after splicing, it is Q data. For example, when P = 4 and Q = 16, 16 data can be obtained after splicing, and a R16 FFT operation can be performed using these 16 data.
[0070] In summary, in the above embodiments of the present disclosure, in scenarios with a large amount of input data, the bit width of the storage unit can be increased, the number of storage blocks can be reduced, and by determining the writing and reading rules of the data, data can be read and processed conveniently and efficiently, the data processing efficiency can be improved, and at the same time, the area consumption of the storage unit can be reduced.
[0071] Figure 2 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure. As Figure 2 shown, based on Figure 1 the embodiments shown, the method includes the following steps.
[0072] Step 201, determine the reading rule for reading N data according to at least one of N, M, P, and Q.
[0073] In some embodiments, the data reading rule can be determined according to at least one of the quantity of data to be processed, the quantity of data written in each cycle, the number of storage units, and the bit width of the storage unit. Optionally, the read data can be used for performing a radix-16 FFT operation, that is, the read data is the input data of the radix-16 FFT operation. Therefore, it is possible to first determine the input data required in each calculation cycle and determine the rule for taking out the input data from the storage unit, which is the above-mentioned data reading rule. Therefore, the reading rule for reading N data can be determined according to the input data required for data operations and at least one of the above N, M, P, and Q.
[0074] Optionally, the reading of N data can be divided into four reading processes, that is, each reading process can read M data, each reading can include four reading cycles, and each reading cycle can read Q data. The Q data read out can be used for one radix-16 FFT operation. For example, in the case of N = 256, M = 64, P = 4, and Q = 16, the reading rule is as follows:
[0075] In the 4 reading cycles of the first reading, 64 data are read from the first addresses of P storage units. Among them, in the first reading cycle of the first reading, the first group of 4 data at the first address of the first storage unit, the second group of 4 data at the first address of the second storage unit, the third group of 4 data at the first address of the third storage unit, and the fourth group of 4 data at the first address of the fourth storage unit are read, and the four groups of data read out are concatenated in sequence until the 4 reading cycles in the first reading are completed, obtaining 4 data sequences in the first reading, and each data sequence includes 16 data;
[0076] In 4 read cycles of the second read, 64 data are read from the second addresses of each of the P memory cells. Among them, in the first read cycle of the second read, the fifth group of 4 data of the second address of the second memory cell, the sixth group of 4 data of the second address of the third memory cell, the seventh group of 4 data of the second address of the fourth memory cell, and the eighth group of 4 data of the second address of the first memory cell are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the second read are completed, obtaining 4 data sequences of the second read, and each data sequence includes 16 data;
[0077] In 4 read cycles of the third read, 64 data are read from the third addresses of each of the P memory cells. Among them, in the first read cycle of the third read, the ninth group of 4 data of the third address of the third memory cell, the tenth group of 4 data of the third address of the fourth memory cell, the eleventh group of 4 data of the third address of the first memory cell, and the twelfth group of 4 data of the third address of the second memory cell are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the third read are completed, obtaining 4 data sequences of the third read, and each data sequence includes 16 data;
[0078] In 4 read cycles of the fourth read, 64 data are read from the fourth addresses of each of the P memory cells. Among them, in the first read cycle of the fourth read, the thirteenth group of 4 data of the fourth address of the fourth memory cell, the fourteenth group of 4 data of the fourth address of the first memory cell, the fifteenth group of 4 data of the fourth address of the second memory cell, and the sixteenth group of 4 data of the fourth address of the third memory cell are read, and the four groups of read data are concatenated in sequence until the 4 read cycles in the fourth read are completed, obtaining 4 data sequences of the fourth read, and each data sequence includes 16 data.
[0079] In other words, for the FFT operation of 256 data, it can be read in four times during the read process. Each time of reading can obtain 4 data sequences, and each data sequence includes 16 data, that is, 64 data can be read each time of reading. Each read data sequence can be used for a radix-16 FFT operation. Optionally, the 256 data can be used to determine the storage address corresponding to the data through two-level R16 butterfly operations. The number of operations of each level of R16 can be 16 times, that is, for the 256 data, the 16 data sequences read four times can be used for the 16 R16 butterfly operations of the first level or the second level.
[0080] Optionally, each read can include 4 read cycles. In each read cycle, partial data at a certain address of P memory cells can be read, and the data can be spliced to obtain a data sequence. For example, the position of the data can be represented by (ram_idx, addr), which can indicate writing data to which address of which memory cell, or can indicate reading data from which address of which memory cell. Here, ram_idx is the index value of the memory cell, and addr is the address in the memory cell. For example, (0, 0) can represent reading data from the first address of the memory cell with an index value of 0. In this example, the first address indicates address 0.
[0081] Specifically, when N = 256, a two-stage R16 structure can be used for FFT calculation. Therefore, data needs to be read from the memory cells before both stages of calculation. When reading data in each stage, four reads are required, and each read includes four read cycles. The 256 data read in the first stage are used for the first-stage data operation, and the 256 data read in the second stage are used for the second-stage data operation. Optionally, the following is the specific method for reading data during the first-stage operation:
[0082] During the first read, the data of (0, 0), (1, 0), (2, 0), and (3, 0) can be read respectively, that is, 64 data can be read from the first address of the first memory cell, the first address of the second memory cell, the first address of the third memory cell, and the first address of the fourth memory cell. Optionally, in the first read cycle, the first 4 data from left to right of the first address of the first memory cell, the first 4 data from left to right of the first address of the second memory cell, the first 4 data from left to right of the first address of the third memory cell, and the first 4 data from left to right of the first address of the fourth memory cell can be spliced to obtain a data sequence. Optionally, after the first data read from left to right is spliced in the order of the memory cells, the second data read can be spliced in the order of the memory cells, and the spliced data can be spliced after the splicing result of the first data, that is, the data can be spliced column by column from left to right.
[0083] Similarly, in the second read cycle, a data sequence can be obtained by concatenating the 5th to 8th data from left to right of the first address of the first storage unit, the 5th to 8th data from left to right of the first address of the second storage unit, the 5th to 8th data from left to right of the first address of the third storage unit, and the 5th to 8th data from left to right of the first address of the fourth storage unit. In the third read cycle, similarly, a data sequence can be obtained by concatenating the 9th to 12th data from left to right of the first address of the first storage unit, the 9th to 12th data from left to right of the first address of the second storage unit, the 9th to 12th data from left to right of the first address of the third storage unit, and the 9th to 12th data from left to right of the first address of the fourth storage unit. In the fourth read cycle, similarly, a data sequence can be obtained by concatenating the 13th to 16th data from left to right of the first address of the first storage unit, the 13th to 16th data from left to right of the first address of the second storage unit, the 13th to 16th data from left to right of the first address of the third storage unit, and the 13th to 16th data from left to right of the first address of the fourth storage unit.
[0084] Specifically, in the second read, the data of (1, 1), (2, 1), (3, 1), and (0, 1) can be read respectively, that is, 64 data can be read from the second address of the second storage unit, the second address of the third storage unit, the second address of the fourth storage unit, and the second address of the first storage unit. The data concatenation method for the four cycles in the second read is the same as that of the four cycles in the above first read, which will not be elaborated here.
[0085] Specifically, in the third read, the data of (2, 2), (3, 2), (0, 2), and (1, 2) can be read respectively, that is, 64 data can be read from the third address of the third storage unit, the third address of the fourth storage unit, the third address of the first storage unit, and the third address of the second storage unit. The data concatenation method for the four cycles in the third read is the same as that of the four cycles in the above first read, which will not be elaborated here.
[0086] Specifically, in the fourth read, the data of (3, 3), (0, 3), (1, 3), and (2, 3) can be read respectively, that is, 64 data can be read from the fourth address of the fourth storage unit, the fourth address of the first storage unit, the fourth address of the second storage unit, and the fourth address of the third storage unit. The data concatenation method for the four cycles in the fourth read is the same as that of the four cycles in the above first read, which will not be elaborated here.
[0087] Optionally, the following is the specific method for reading data during the second-level operation:
[0088] During the first read, the data at (0, 0), (2, 2), (1, 1), and (3, 3) can be read separately. That is, 64 data can be read from the first address of the first storage unit, the second address of the second storage unit, the third address of the third storage unit, and the fourth address of the fourth storage unit. Optionally, in the first read cycle, the 1st, 9th, 5th, and 13th data from left to right of the first address of the first storage unit, the 1st, 9th, 5th, and 13th data from left to right of the second address of the second storage unit, the 1st, 9th, 5th, and 13th data from left to right of the third address of the third storage unit, and the 1st, 9th, 5th, and 13th data from left to right of the fourth address of the fourth storage unit can be concatenated to obtain a data sequence. The data can be concatenated column by column from left to right.
[0089] Similarly, in the second read cycle, the 3rd, 11th, 7th, and 15th data from left to right of the first address of the first storage unit, the 3rd, 11th, 7th, and 15th data from left to right of the second address of the second storage unit, the 3rd, 11th, 7th, and 15th data from left to right of the third address of the third storage unit, and the 3rd, 11th, 7th, and 15th data from left to right of the fourth address of the fourth storage unit can be concatenated to obtain a data sequence. In the third read cycle, similarly, the 2nd, 10th, 6th, and 14th data from left to right of the first address of the first storage unit, the 2nd, 10th, 6th, and 14th data from left to right of the second address of the second storage unit, the 2nd, 10th, 6th, and 14th data from left to right of the third address of the third storage unit, and the 2nd, 10th, 6th, and 14th data from left to right of the fourth address of the fourth storage unit can be concatenated to obtain a data sequence. In the fourth read cycle, similarly, the 4th, 12th, 8th, and 16th data from left to right of the first address of the first storage unit, the 4th, 12th, 8th, and 16th data from left to right of the second address of the second storage unit, the 4th, 12th, 8th, and 16th data from left to right of the third address of the third storage unit, and the 4th, 12th, 8th, and 16th data from left to right of the fourth address of the fourth storage unit can be concatenated to obtain a data sequence.
[0090] Specifically, during the second read, the data at (2, 0), (0, 2), (3, 1), and (1, 3) can be read separately. That is, 64 data can be read from the first address of the third storage unit, the third address of the first storage unit, the second address of the fourth storage unit, and the fourth address of the second storage unit. The method of concatenating the data in the four cycles of the second read is the same as that in the four cycles of the above first read, which will not be elaborated here.
[0091] Specifically, during the third read, the data of (1, 0), (3, 2), (2, 1), and (0, 3) can be read respectively, that is, 64 data can be read from the first address of the second storage unit, the third address of the fourth storage unit, the second address of the third storage unit, and the fourth address of the first storage unit. The data splicing method for the four cycles in the second read is the same as that for the four cycles in the first read above, which will not be elaborated here.
[0092] Specifically, during the fourth read, the data of (3, 0), (1, 2), (0, 1), and (2, 3) can be read respectively, that is, 64 data can be read from the first address of the fourth storage unit, the third address of the second storage unit, the second address of the first storage unit, and the fourth address of the third storage unit. The data splicing method for the four cycles in the fourth read is the same as that for the four cycles in the first read above, which will not be elaborated here.
[0093] Step 202: Determine the writing rule according to the reading rule.
[0094] In some embodiments, after determining the reading rule of the data, the writing rule of the data can be determined according to the reading rule. For example, the reading position of the input data required for calculation can be determined, and when writing, the data can be written to the above reading position, which is convenient for subsequent extraction for calculation.
[0095] In some embodiments, the determined writing rule can be the writing rule when the data is first written into the storage unit, that is, the writing rule before the first-level data operation. After the data completes the first-level data operation, the first-level output sequence corresponding to the data sequence to be processed can be obtained, and the first-level output sequence can be written into the storage unit according to the in-place storage principle, that is, the data writing before the second-level data operation can be the same as the writing rule before the first-level data operation.
[0096] In some embodiments, optionally, when N = 256, M = 64, P = 4, and Q = 16, the writing rule is as follows: for the first writing cycle, write 64 data in the first data group among the K data groups to the first address of the first storage unit, the second address of the second storage unit, the third address of the third storage unit, and the fourth address of the fourth storage unit; for the second writing cycle, write 64 data in the second data group among the K data groups to the third address of the first storage unit, the fourth address of the second storage unit, the first address of the third storage unit, and the second address of the fourth storage unit; for the third writing cycle, write 64 data in the third data group among the K data groups to the fourth address of the first storage unit, the first address of the second storage unit, the second address of the third storage unit, and the third address of the fourth storage unit; for the fourth writing cycle, write 64 data in the fourth data group among the K data groups to the second address of the first storage unit, the third address of the second storage unit, the fourth address of the third storage unit, and the first address of the fourth storage unit.
[0097] In other words, when the number of data in the data sequence is 256, the 256 data can be written in four cycles. For example, in the first writing cycle, write data 0 - 63 to (0, 0)(1, 1)(2, 2)(3, 3); in the second writing cycle, write data 64 - 127 to (0, 2)(1, 3)(2, 0)(3, 1); in the third writing cycle, write data 128 - 191 to (0, 3)(1, 0)(2, 1)(3, 2); in the fourth writing cycle, write data 192 - 255 to (0, 1)(1, 2)(2, 3)(3, 0).
[0098] In summary, in the above embodiments of the present application, the reading rule and writing rule of data can be determined, and data can be read more conveniently, which is convenient for subsequent data reading and data processing. For example, when N = 256, through the writing and reading rules of this data, the number of required storage units can be reduced. Only 4 blocks of RAM are needed to store the data and perform FFT calculations. When the storage unit is a RAM storage unit, compared with splitting more blocks of RAM, the present disclosure can reduce the area consumption of RAM.
[0099] Figure 3 It is a schematic flow chart of a data processing method provided by an embodiment of the present disclosure. As Figure 3 shown, based on Figure 1 the embodiments shown, the method includes the following steps.
[0100] Step 301, perform a first-level radix-16 fast Fourier transform (FFT) operation on the first data sequence read out for each reading cycle to obtain a first-level output sequence.
[0101] In some embodiments, both the first data sequence and the first-level output sequence include Q data.
[0102] In other words, the obtained data sequences read and spliced above can be respectively subjected to the first-level radix-16 fast Fourier transform (FFT) operation, and each data sequence can be subjected to the first-level radix-16 fast Fourier transform (FFT) operation once.
[0103] In some embodiments, when performing the first-level radix-16 fast Fourier transform (FFT) operation, numbers can be read from the storage unit, and the read data can be used as the input data for the first-level radix-16 fast Fourier transform (FFT) operation. After each read data is subjected to the first-level radix-16 fast Fourier transform (FFT) operation, a corresponding first-level output data can be obtained.
[0104] In some embodiments, when N = 256, before the first-level radix-16 fast Fourier transform (FFT) operation, four data reads can be performed. The data read each time can be spliced into four data sequences, that is, 16 data sequences can be obtained after four reads are completed. Each data sequence can be subjected to the first-level radix-16 fast Fourier transform (FFT) operation once, and an output sequence corresponding to each data sequence can be obtained. An output sequence can include 16 first-level output data. Optionally, after each read data is subjected to the first-level radix-16 fast Fourier transform (FFT) operation, it can have a corresponding first-level output data, and the index values of the data and the corresponding first-level output data are the same. Therefore, after the first-level radix-16 fast Fourier transform (FFT) operation on 256 data, 256 first-level output data can be obtained, that is, after 256-point data is subjected to 16 first-level radix-16 fast Fourier transform (FFT) operations, 16 first-level output sequences can be obtained.
[0105] Step 302: Write the first-level output sequence back to the write-back address of P storage units, and the write-back address is the same as the read address of the first data sequence.
[0106] In some embodiments, after performing the first-level radix-16 fast Fourier transform (FFT) operation to obtain the first-level output sequence, the first-level output sequence can be written back to the storage positions of N data, that is, the first-level output sequence can be written back to the storage unit according to the in-place storage principle.
[0107] In some embodiments, the write-back address for writing the first-level output sequence back to P storage units includes: determining Q / P data bits of each storage unit when reading the first data sequence, where the Q / P data bits of each storage unit are the read addresses when reading the first data sequence; adding a mask to the Q - Q / P data bits of each storage unit; and writing the Q data in the first-level output sequence back to the Q / P data bits of each of the P storage units respectively.
[0108] In other words, after obtaining a first-level output sequence by completing a first-level radix-16 fast Fourier transform (FFT) operation, the first-level output sequence can be written back in-place storage. Therefore, the storage location of the first data sequence corresponding to the first-level output sequence can be determined first, that is, the location when reading the first data sequence. Since Q data are stored in P storage units, each storage unit can correspondingly store Q / P first-level output sequences. To avoid affecting the previously written first-level output sequences, the mask feature of the RAM storage unit can be used to add a mask to the locations of the already written first-level output sequences, or add a mask to all the remaining locations that do not need to be written, that is, a mask can be added to the Q - Q / P data bits of each storage unit. At this time, only the data of Q / P bits are updated, and all other locations are masked.
[0109] Step 303: Read the first-level output sequence and perform a second-level radix-16 FFT operation to obtain a second-level output sequence.
[0110] In some embodiments, after writing back the first-level output sequence, the rule for reading the second-level data can be determined according to the requirements of the second-level radix-16 FFT operation, and the first output sequence can be read out. The first-level output sequence can be used as the input data for the second-level radix-16 FFT operation. Optionally, referring to the specific method of reading data during the second-level operation in step 201, the method of reading out the first-level output sequence can be determined, which will not be elaborated here.
[0111] In some embodiments, N = 256, M = 64, P = 4, Q = 16; and / or the storage unit is a pseudo dual-port random access memory (RAM); and / or the number of write cycles for processing N data at a time is K, the read cycle is L, and the total number of cycles for processing H times of N data < H×(K + L), and the total number of cycles for processing H times of N data ≥ K + L+(H - 1)×L / 2.
[0112] In the above embodiments, the read cycle L can be the total cycle of two reads. Optionally, the read cycle can be the total cycle of radix-16 reading and calculation at two levels. For example, when N = 256, the first-level radix-16 can process 16 data per cycle, and the second-level radix-16 can also process 16 data per cycle. Then the total number of cycles spent by the first-level radix-16 and the second-level radix-16 is Then the read cycle L is 32 at this time.
[0113] In this embodiment, by using two copies of storage resources for ping-pong switching, the data throughput rate is greatly improved.
[0114] In some embodiments, when N = 256, it takes 4 clock cycles to write to the storage unit, and 64 data can be written in each clock cycle. When performing the first-level radix-16 FFT, 16 data are processed in each cycle for the first-level radix-16 FFT, that is, the first-level radix-16 FFT requires 16 clock cycles. Similarly, the second-level radix-16 FFT requires 16 clock cycles. Therefore, the total number of clock cycles consumed for 256 data from writing to the storage unit to completing the second-level radix-16 FFT operation and obtaining the second-level output sequence is 36 clock cycles.
[0115] Optionally, in order to increase the pipelining efficiency of data processing, a ping-pong ram mechanism can be added. When the pong ram starts the second-level radix-16 FFT operation, the ping ram can receive the next set of 256-point data and calculate the first-level radix-16 FFT operation. Eventually, the pipelining efficiency can reach receiving a 256-point data input every 20 cycles.
[0116] In other words, when the first set of 256-point data is performing the first-level radix-16 FFT, the second set of 256-point data can start to be written to the storage unit. When the first set completes the first-level radix-16 FFT, the second set finishes writing at the same time. The second-level radix-16 FFT of the first set can be performed simultaneously with the first-level radix-16 FFT of the second set, which can improve the data processing efficiency when the data volume is large.
[0117] In summary, in the above embodiments of the present disclosure, by performing the first-level radix-16 fast Fourier transform (FFT) operation on the read data to obtain the first-level output sequence, and storing the first-level output sequence in place in the storage unit, the storage unit can be reused, reducing the area consumption of the storage unit. By using the mask feature to write back the first-level output sequence, the impact on other data can be avoided. By performing two-level radix-16 FFT operations, the FFT operation of 256 data can be completed using only 4 storage units; by adopting the ping-pong ram mechanism, the data processing efficiency can be improved.
[0118] The technical solution of the present disclosure will be further described in detail below in combination with specific application embodiments.
[0119] The following is an address mapping method for a high-throughput 256-point FFT provided by an embodiment of the present disclosure. This method can be used in a specific throughput scenario, that is, when the bus bandwidth meets the requirement of transmitting 64-point FFT data each time, a radix-16 basic operation structure is used to implement FFT calculation. This method can be implemented using a two-stage R16 structure as shown in Figure 4 . The main introduction of this method is the data storage rule at each stage of R16 when 64 data come each time, so as to minimize the number of RAMs used and reduce the number of times of reading and writing RAM. The complete content of this method is as follows.
[0120] Data access rule for the first-stage radix-16 (stage0)
[0121] Two points need to be considered for the writing of stage0:
[0122] As shown in Figure 5 , 64 data are sequentially input in each clock cycle, each data is 32 bits (bits), and 1 256-point input is completed every 4 clock cycles.
[0123] In this method, the storage of input data only uses 4 pieces of pseudo dual-port RAM (Simple Dual Port RAM). Each piece of pseudo dual-port RAM can be written and read at the same time. Its bit width is set to 16x32 bits, that is, each piece of pseudo dual-port RAM can write 16 data at a time. The specific writing rule should ensure that 16 data required by the R16 structure are included in the 64 data read out from the 4 pieces of pseudo dual-port RAM subsequently once.
[0124] Optionally, the following uses (ram_idx, addr) to represent which ram and which address the data is written to, where ram_idx is the index of the RAM and addr is the address in the RAM. It is possible to write 256-point data to 4 pieces of pseudo dual-port RAM in stage0. As shown in Figure 6 , wr_cycle0 (0 to 63) corresponds to (0, 0), (1, 1), (2, 2), (3, 3) respectively; wr_cycle1 (64 to 127) corresponds to (0, 2), (1, 3), (2, 0), (3, 1) respectively; wr_cycle2 (128 to 191) corresponds to (0, 3), (1, 0), (2, 1), (3, 2) respectively; wr_cycle3 (192 to 255) corresponds to (0, 1), (1, 2), (2, 3), (3, 0) respectively. That is, in the writing cycle 0 (wr_cycle0), data 0 to 63 are written to address 0 of RAM0, address 1 of RAM1, address 2 of RAM2, and address 3 of RAM3, and so on.
[0125] After 256 data are written in a total of 4 cycles, start reading data. The data reading rule is as follows:
[0126] The read cycle 0 (rd_cycle0) corresponds to the readouts of (0, 0), (1, 0), (2, 0), and (3, 0) respectively. A total of 64 data are read out, and 4 calculation cycles can be used. In cycle 0 (cycle0), take out Figure 6 0 to 3 data of each of the 4 rams in Figure 7 From left to right, perform column-wise splicing to obtain the 16 data used in the first calculation cycle, as Figure 6 shown. In cycle 1, take out Figure 6 4 to 7 data of each of the 4 rams in Figure 6 From left to right, perform column-wise splicing; in cycle 2, take out
[0127] 8 to 11 data of each of the 4 rams in Figure 8 From left to right, perform column-wise splicing; in cycle 3, take out Figure 8 12 to 15 data of each of the 4 rams in Figure 4 From left to right, perform column-wise splicing.
[0128] The second-level radix-16 data access rule (stage1)
[0129] The write in stage1 is the calculation result of stage0. After each R16 calculation is completed, write the calculation result back to the original position Figure 6 , Figure 6 The storage rule of
[0130] still meets the requirements of the data read out in stage1. Due to the characteristics of the pseudo-dual-port RAM, here the calculation result can be written while reading the input data of the subsequent R16, so as to reuse the same pseudo-dual-port RAM. Since only a partial position of a certain address of each RAM will be written each time, the byte mask (bytemask) feature of the RAM needs to be enabled.
[0131] Rd_cycle0 corresponds to reading out (0, 0), (2, 2), (1, 1), (3, 3) respectively, and a total of 64 data are read out. 4 calculation cycles can be used. In cycle0, take out Figure 6 the data at the (0, 8, 4, 12) positions of the 4 rams respectively. From left to right, splice by column to obtain the calculation data for the second-level R16 operation, as Figure 9 shown.
[0132] In cycle1, take out Figure 6 the data at the (2, 10, 6, 14) positions of the 4 rams respectively. From left to right, splice by column to obtain; in cycle2, take out Figure 6 the data at the (1, 9, 5, 13) positions of the 4 rams respectively. From left to right, splice by column to obtain; in cycle3, take out Figure 6 the data at the (3, 11, 7, 15) positions of the 4 rams respectively. From left to right, splice by column to obtain; Rd_cycle1 corresponds to reading out (2, 0), (0, 2), (3, 1), (1, 3) respectively, and a total of 64 data are read out. 4 calculation cycles can be used. Similar to Rd_cycle0, it will not be elaborated; Rd_cycle2 corresponds to reading out (1, 0), (3, 2), (2, 1), (0, 3) respectively, and a total of 64 data are read out. 4 calculation cycles can be used. Similar to Rd_cycle0, it will not be elaborated; Rd_cycle3 corresponds to reading out (3, 0), (1, 2), (0, 1), (2, 3) respectively, and a total of 64 data are read out. 4 calculation cycles can be used. Similar to Rd_cycle0, it will not be elaborated. The finally read data is as Figure 10 shown.
[0133] Figure 10 Each column is Figure 4 the input data for each R16 calculation in stage1 of
[0134]
[0135] Finally, a 256-point FFT calculation result is obtained. If the R16 calculation delay of the pipeline is not considered, the total number of clock cycles consumed by the overall scheme from the start of input data to the last stage is:
[0136] In summary, in the above examples of the present disclosure, in the scenario of inputting 64 data, the single RAM bit width is increased to 16 data, thereby reducing the number of RAM blocks to 4 and reducing the RAM area consumption; the specific address mapping is applicable to the 2nd level of radix-16 of 256 points, and the calculation result can be stored back in the RAM in-place, thereby reusing one set of RAM and reducing the RAM area consumption; a specific address mapping guarantee is provided to ensure that the 64 data read each time can be sequentially used for 4 cycles in the R16 calculation, thereby achieving the minimum number of RAM read and write operations and reducing the power consumption waste of reading useless data.
[0137] Figure 11 FIG. 1100 is a schematic structural diagram of a data processing device 1100 provided by an embodiment of the present disclosure. As Figure 11 shown, the device includes: an acquisition module 1110, configured to acquire a data sequence to be processed, where the data sequence to be processed includes N data; a writing module 1120, configured to write M data out of the N data into P storage units respectively in each of multiple writing cycles, and the storage amount corresponding to each address of each storage unit is Q; a reading and processing module 1130, configured to read Q / P data from the P storage units respectively in each of multiple reading cycles, and process the Q / P data, where N, M, P, and Q are positive integers, M = P×Q, and P < M.
[0138] In some embodiments, the writing module is further configured to divide the N data into K data groups, where K is a positive integer and K is equal to the number of cycles of the multiple writing cycles, each data group includes M data; determine a writing rule for the M data in each data group; and based on the writing rule, write the M data in each data group into the P storage units in each writing cycle.
[0139] In some embodiments, the data processing device further includes a determination module, configured to determine a reading rule for reading the N data according to at least one of N, M, P, and Q; and determine a writing rule according to the reading rule.
[0140] In some embodiments, N = 256, M = 64, P = 4, Q = 16, and the writing rule is as follows: for the first writing cycle, 64 data in the first data group among the K data groups are written at the first address of the first storage unit, the second address of the second storage unit, the third address of the third storage unit, and the fourth address of the fourth storage unit; for the second writing cycle, 64 data in the second data group among the K data groups are written at the third address of the first storage unit, the fourth address of the second storage unit, the first address of the third storage unit, and the second address of the fourth storage unit; for the third writing cycle, 64 data in the third data group among the K data groups are written at the fourth address of the first storage unit, the first address of the second storage unit, the second address of the third storage unit, and the third address of the fourth storage unit; for the fourth writing cycle, 64 data in the fourth data group among the K data groups are written at the second address of the first storage unit, the third address of the second storage unit, the fourth address of the third storage unit, and the first address of the fourth storage unit.
[0141] In some embodiments, N = 256, M = 64, P = 4, Q = 16. The reading rule is as follows: In the first four reading cycles of the first read, 64 data are read from the first addresses of each of the P memory cells. Specifically, in the first reading cycle of the first read, the first set of 4 data of the first address of the first memory cell, the second set of 4 data of the first address of the second memory cell, the third set of 4 data of the first address of the third memory cell, and the fourth set of 4 data of the first address of the fourth memory cell are read, and the four sets of read data are concatenated in sequence until the four reading cycles in the first read are completed, obtaining 4 data sequences for the first read, with each data sequence including 16 data. In the four reading cycles of the second read, 64 data are read from the second addresses of each of the P memory cells. Specifically, in the first reading cycle of the second read, the fifth set of 4 data of the second address of the second memory cell, the sixth set of 4 data of the second address of the third memory cell, the seventh set of 4 data of the second address of the fourth memory cell, and the eighth set of 4 data of the second address of the first memory cell are read, and the four sets of read data are concatenated in sequence until the four reading cycles in the second read are completed, obtaining 4 data sequences for the second read, with each data sequence including 16 data. In the four reading cycles of the third read, 64 data are read from the third addresses of each of the P memory cells. Specifically, in the first reading cycle of the third read, the ninth set of 4 data of the third address of the third memory cell, the tenth set of 4 data of the third address of the fourth memory cell, the eleventh set of 4 data of the third address of the first memory cell, and the twelfth set of 4 data of the third address of the second memory cell are read, and the four sets of read data are concatenated in sequence until the four reading cycles in the third read are completed, obtaining 4 data sequences for the third read, with each data sequence including 16 data. In the four reading cycles of the fourth read, 64 data are read from the fourth addresses of each of the P memory cells. Specifically, in the first reading cycle of the fourth read, the thirteenth set of 4 data of the fourth address of the fourth memory cell, the fourteenth set of 4 data of the fourth address of the first memory cell, the fifteenth set of 4 data of the fourth address of the second memory cell, and the sixteenth set of 4 data of the fourth address of the third memory cell are read, and the four sets of read data are concatenated in sequence until the four reading cycles in the fourth read are completed, obtaining 4 data sequences for the fourth read, with each data sequence including 16 data.
[0142] In some embodiments, the reading and processing module may also be used to perform a first - level radix - 16 fast Fourier transform (FFT) operation on the first data sequence read out for each reading cycle to obtain a first - level output sequence. Both the first data sequence and the first - level output sequence include Q data; write the first - level output sequence back to the write - back addresses of P storage units, where the write - back addresses are the same as the read - out addresses of the first data sequence; read the first - level output sequence and perform a second - level radix - 16 FFT operation to obtain a second - level output sequence.
[0143] In some embodiments, the writing module is further used to determine, when reading the first data sequence, Q / P data bits of each storage unit, where the Q / P data bits of each storage unit are the read - out addresses when reading the first data sequence; add a mask to the Q - Q / P data bits of each storage unit; write the Q data in the first - level output sequence back to the Q / P data bits of each of the P storage units respectively.
[0144] In some embodiments, N = 256, M = 64, P = 4, Q = 16; and / or the storage unit is a pseudo - dual - port random access memory (RAM); and / or the number of writing cycles for processing N data at a time is K, the number of reading cycles is L, and the total number of cycles for processing H times of N data < H×(K + L), and the total number of cycles for processing H times of N data ≥ K + L+(H - 1)×L / 2.
[0145] In summary, the data processing device 1100 can write and read data according to rules, can implement FFT operations in scenarios with a large amount of input data, while reducing the number of RAM blocks and the RAM area consumption.
[0146] In the above - mentioned embodiments provided by the present application, the methods and devices provided by the embodiments of the present application are introduced. To implement the various functions in the methods provided by the embodiments of the present application, the electronic device may include a hardware structure, software modules, and implement the above - mentioned various functions in the form of a hardware structure, software modules, or a combination of a hardware structure and software modules. A certain function among the above - mentioned various functions can be executed in the form of a hardware structure, software module, or a combination of a hardware structure and software module.
[0147] Figure 12 FIG. is a block diagram of an electronic device 1200 for implementing the above - mentioned method shown according to an exemplary embodiment. For example, the electronic device 1200 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0148] Refer to Figure 12, the electronic device 1200 may include one or more of the following components: a processing component 1202, a memory 1204, a power component 1206, a multimedia component 1208, an audio component 1210, an input / output (I / O) interface 1212, a sensor component 1214, and a communication component 1216.
[0149] The processing component 1202 generally controls the overall operation of the electronic device 1200, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1202 may include one or more modules to facilitate the interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate the interaction between the multimedia component 1208 and the processing component 1202.
[0150] The memory 1204 is configured to store various types of data to support the operation of the electronic device 1200. Examples of these data include instructions for any application or method operating on the electronic device 1200, contact data, phone book data, messages, pictures, videos, etc. The memory 1204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0151] The power component 1206 provides power to various components of the electronic device 1200. The power component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 1200.
[0152] The multimedia component 1208 includes a screen that provides an output interface between the electronic device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the electronic device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0153] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 further includes a speaker for outputting audio signals.
[0154] The I / O interface 1212 provides an interface between the processing component 1202 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0155] The sensor component 1214 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 1200. For example, the sensor component 1214 can detect the on / off state of the electronic device 1200, the relative positioning of components, such as the display and the keypad of the electronic device 1200. The sensor component 1214 can also detect a change in the position of the electronic device 1200 or a component of the electronic device 1200, the presence or absence of user contact with the electronic device 1200, the orientation or acceleration / deceleration of the electronic device 1200, and a change in the temperature of the electronic device 1200. The sensor component 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1214 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1214 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0156] The communication component 1216 is configured to facilitate communication between the electronic device 1200 and other devices in a wired or wireless manner. The electronic device 1200 can access a communication standard-based wireless network, such as WiFi, 2G or 3G, 4G LTE, 5G NR (New Radio), or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0157] In an exemplary embodiment, the electronic device 1200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0158] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1204 including instructions, and the above instructions can be executed by a processor 1220 of the electronic device 1200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0159] Embodiments of the present disclosure also propose a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in the above embodiments of the present disclosure.
[0160] Embodiments of the present disclosure also propose a storage device, the device includes at least one storage module, each storage module includes P storage units, and the storage capacity corresponding to each address of each storage unit is Q; wherein, when the data sequence to be processed includes N data, each storage module can write M data in each of multiple write cycles, N, M, P, Q are positive integers, M = P × Q, and P < M.
[0161] For the above storage device. Optionally, N = 256, M = 64, P = 4, Q = 16; the storage unit is a pseudo dual-port random access memory RAM.
[0162] Figure 13It is a schematic structural diagram of a chip 1300 for implementing the above method shown according to an exemplary embodiment. Referring to Figure 13 , the chip 1300 includes a communication interface 1301 and at least one processor 1302. The communication interface 1301 is used to receive signals input to the chip 1300 or signals output from the chip 1300. The processor 1302 communicates with the communication interface 1301 and implements the method described in the above embodiments of the present disclosure through logic circuits or by executing code instructions.
[0163] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0164] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in at least one embodiment or example.
[0165] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a manner other than shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the technical field to which the embodiments of the present invention belong.
[0166] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having at least one wiring (control method), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0167] It should be understood that various parts of the embodiments of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0168] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0169] In addition, each functional unit in various embodiments of the present invention may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, or the like.
[0170] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining a data sequence to be processed, where the data sequence to be processed includes N data; In each of multiple writing cycles, respectively writing M of the N data into P storage units, and the storage amount corresponding to each address of each storage unit is Q; In each of multiple reading cycles, respectively reading Q / P data from the P storage units, splicing them, and processing the spliced Q data, where N, M, P, and Q are positive integers, M = P × Q, and P < M.
2. The method according to claim 1, characterized in that, The step of, in each of multiple writing cycles, respectively writing M of the N data into P storage units includes: Dividing the N data into K data groups, where K is a positive integer and K is equal to the number of cycles of the multiple writing cycles, and each data group includes M data; Determining the writing rule for the M data in each data group; Based on the writing rule, in each writing cycle, writing the M data in each data group into the P storage units.
3. The method according to claim 2, wherein The step of determining the writing rule for the M data in each data group includes: Determining a reading rule for reading the N data according to at least one of N, M, P, and Q; Determining the writing rule according to the reading rule.
4. The method according to claim 3, wherein N = 256, M = 64, P = 4, Q = 16, where The writing rule is: For the first writing cycle, writing 64 data in the first data group among the K data groups at the first address of the first storage unit, the second address of the second storage unit, the third address of the third storage unit, and the fourth address of the fourth storage unit; For the second writing cycle, writing 64 data in the second data group among the K data groups at the third address of the first storage unit, the fourth address of the second storage unit, the first address of the third storage unit, and the second address of the fourth storage unit; For the third writing cycle, writing 64 data in the third data group among the K data groups at the fourth address of the first storage unit, the first address of the second storage unit, the second address of the third storage unit, and the third address of the fourth storage unit; For the fourth writing cycle, writing 64 data in the fourth data group among the K data groups at the second address of the first storage unit, the third address of the second storage unit, the fourth address of the third storage unit, and the first address of the fourth storage unit.
5. The method according to claim 3, characterized in that N = 256, M = 64, P = 4, Q = 16, where The reading rule is: In 4 read cycles of the first read, 64 data are read from the first addresses of the respective P memory cells. Among them, in the first read cycle of the first read, the first set of 4 data of the first address of the first memory cell, the second set of 4 data of the first address of the second memory cell, the third set of 4 data of the first address of the third memory cell, and the fourth set of 4 data of the first address of the fourth memory cell are read, and the four sets of read data are concatenated in sequence until the 4 read cycles in the first read are completed, obtaining 4 data sequences of the first read, and each data sequence includes 16 data; In 4 read cycles of the second read, 64 data are read from the second addresses of the respective P memory cells. Among them, in the first read cycle of the second read, the fifth set of 4 data of the second address of the second memory cell, the sixth set of 4 data of the second address of the third memory cell, the seventh set of 4 data of the second address of the fourth memory cell, and the eighth set of 4 data of the second address of the first memory cell are read, and the four sets of read data are concatenated in sequence until the 4 read cycles in the second read are completed, obtaining 4 data sequences of the second read, and each data sequence includes 16 data; In 4 read cycles of the third read, 64 data are read from the third addresses of the respective P memory cells. Among them, in the first read cycle of the third read, the ninth set of 4 data of the third address of the third memory cell, the tenth set of 4 data of the third address of the fourth memory cell, the eleventh set of 4 data of the third address of the first memory cell, and the twelfth set of 4 data of the third address of the second memory cell are read, and the four sets of read data are concatenated in sequence until the 4 read cycles in the third read are completed, obtaining 4 data sequences of the third read, and each data sequence includes 16 data; In 4 read cycles of the fourth read, 64 data are read from the fourth addresses of the respective P memory cells. Among them, in the first read cycle of the fourth read, the thirteenth set of 4 data of the fourth address of the fourth memory cell, the fourteenth set of 4 data of the fourth address of the first memory cell, the fifteenth set of 4 data of the fourth address of the second memory cell, and the sixteenth set of 4 data of the fourth address of the third memory cell are read, and the four sets of read data are concatenated in sequence until the 4 read cycles in the fourth read are completed, obtaining 4 data sequences of the fourth read, and each data sequence includes 16 data.
6. The method according to any one of claims 1 to 5, characterized in that, In each of the multiple read cycles, Q / P data are respectively read from the P memory cells and concatenated, and the processing of the concatenated Q data includes: For the first data sequence read in each read cycle, a first-level radix-16 fast Fourier transform (FFT) operation is performed to obtain a first-level output sequence, and both the first data sequence and the first-level output sequence include Q data; Write the first-level output sequence back to the write-back addresses of the P storage units, where the write-back addresses are the same as the read addresses of the first data sequence; Read the first-level output sequence and perform a second-level radix-16 FFT operation to obtain a second-level output sequence.
7. The method according to claim 6, wherein The write-back addresses for writing the first-level output sequence back to the P storage units include: Determine Q / P data bits of each storage unit when reading the first data sequence, where the Q / P data bits of each storage unit are the read addresses when reading the first data sequence; Add a mask to the Q - Q / P data bits for each storage unit; Write the Q data in the first-level output sequence to the Q / P data bits of each of the P storage units respectively.
8. The method according to claim 1, wherein N = 256, M = 64, P = 4, Q = 16; and / or The storage unit is a pseudo-dual-port random access memory RAM; and / or The number of write cycles for processing N data at one time is K, the number of read cycles is L, the total number of cycles for processing H times of N data < H×(K + L), and the total number of cycles for processing H times of N data ≥ K + L+(H - 1)×L / 2.
9. A data processing device, characterized in that, The apparatus includes: An acquisition module for acquiring a data sequence to be processed, where the data sequence to be processed includes N data; A writing module for writing M data out of the N data into P storage units respectively in each of multiple writing cycles, and the storage capacity corresponding to each address of each storage unit is Q; A reading and processing module for reading Q / P data from the P storage units respectively in each of multiple reading cycles and processing the Q / P data, where N, M, P, Q are positive integers, M = P×Q, and P < M.
10. A storage device, characterized in that, It includes: At least one storage module, each storage module includes P storage units, and the storage capacity corresponding to each address of each storage unit is Q; where, in the case where the data sequence to be processed includes N data, each storage module can write M data in each of multiple writing cycles, and N, M, P, Q are positive integers, M = P×Q, and P < M.
11. The storage device according to claim 10, characterized in that, N = 256, M = 64, P = 4, Q = 16; The storage unit is a pseudo-dual-port random access memory RAM.
12. The storage device according to claim 10 or 11, characterized in that, The number of the at least one storage module is 2, where the number of write cycles for processing N data at one time is K, the number of read cycles is L, the total number of cycles for processing H times of N data < H×(K + L), and the total number of cycles for processing H times of N data ≥ K + L+(H - 1)×L / 2.
13. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 - 8.
14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 - 8.
15. A chip, characterized in that, Comprising at least one processor and a communication interface; the communication interface is used for receiving signals input to the chip or signals output from the chip, and the processor communicates with the communication interface and implements the method according to any one of claims 1 to 8 through logic circuits or by executing code instructions.
16. The chip according to claim 15, wherein Further comprising a storage device according to any one of claims 10 to 12.
Citation Information
Cited By
Image data storage method and device, electronic equipment and storage medium
CN121785549A