Radar imaging large matrix signal DSP transposition method and device

By dividing the large matrix into sub-arrays and transpose operations on cached data blocks, the problem of inefficient transposition of big data in DSP is solved, efficient matrix transposition and storage space optimization are achieved, and the processing capability of the radar imaging system is improved.

CN120336228APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510404293.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The current method of big data transposition in DSP is less efficient and takes up a large storage space. Especially during matrix transposition operations, DMA is inefficient in data transfer under address jumps, requiring additional storage space.

Method used

The matrix to be transposed is divided into M×N sub-arrays, and the sub-array data is transported to the cached data block through the kernel for transposition operation, and the calculated sub-array is put back into the transposition position to reduce the number of address jumps, and the assembly function is used to optimize the transposition process.

Benefits of technology

It improves transposition efficiency, reduces storage space requirements, is suitable for large-scale matrix transposition, and improves the processing efficiency of radar imaging systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336228A_ABST
    Figure CN120336228A_ABST
Patent Text Reader

Abstract

The invention discloses a radar imaging large matrix signal DSP transposition method and device, and the method comprises the steps: dividing a to-be-transposed matrix into M * N to-be-transposed sub-arrays according to the size of a cache data block; carrying out transposition operation on each to-be-transposed sub-array through the kernel until all to-be-transposed sub-arrays complete transposition operation; wherein the process of transposing the i row and j column of to-be-transposed sub-arrays comprises the following steps: carrying data of the i row and j column of to-be-transposed sub-arrays to a cache data block, and performing transposing operation on the to-be-transposed sub-arrays on the cache data block to obtain i row and j column of operated sub-arrays; and carrying the subarrays after calculation in the ith row and the jth column to corresponding transposition positions of the subarrays to be transposed in the ith row and the jth column to obtain the subarrays after transposition in the ith row and the jth column, and completing the transposition operation on the subarrays to be transposed in the ith row and the jth column. The transposition device is high in transposition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar data processing, and particularly relates to a DSP transposition method and device for large matrix signals in radar imaging. Background Art

[0002] With the rapid development of engineering technologies such as large-capacity fast storage, high-speed digital signal processors, and multi-core parallel architectures, the real-time performance of synthetic aperture radar (SAR) signal processing has been greatly improved. As a result, the imaging resolution is getting larger and the imaging swath is getting wider. Therefore, a larger number of azimuth data accumulation points are required. With the increase in the number of range bins and azimuth bins, in the SAR imaging process, not only processing such as range migration correction and Doppler centroid compensation in the range direction is required, but also processing such as high-order phase filtering in the azimuth direction and azimuth time-domain dechirping is required.

[0003] After the large squint (the squint angle is generally greater than 65 degrees) spotlight echo data in the SAR signal passes through digital down-conversion and range pulse compression in a field-programmable gate array (FPGA), it is usually continuously sent to a double data rate (DDR) processor in a digital signal processor (DSP) through a serial rapid input / output (SRIO) high-speed interface in the range direction. For DDR reading, the reading speed of consecutive addresses is much faster than that of discrete addresses. However, there is no matrix concept inside the DSP, only a one-dimensional concept. Therefore, when performing data processing in different dimensions, problems will be encountered where the data storage method is inconsistent with the current data fetching dimension, especially when performing matrix transposition operations.

[0004] Currently, the large data transposition operation inside the DSP is usually implemented through a direct memory access (DMA) transfer function. However, only when the data storage method is consistent with the current data fetching dimension, can the DMA continuously transfer data. Otherwise, the DMA needs to adopt an address hopping method to transfer data. When the DMA transfers data in the case of address hopping, the amount of transferred data is limited, the efficiency will be greatly reduced, and additional storage space needs to be allocated.

[0005] Therefore, the current method for large data transposition inside the DSP has low efficiency and occupies a large amount of storage space. Summary of the Invention

[0006] An embodiment of the present invention provides a method and device for transposing large matrix signals in radar imaging by DSP, which can solve the problems of low efficiency and large storage space occupation of the method for transposing large data in the current DSP.

[0007] In a first aspect, a method for transposing large matrix signals in radar imaging by DSP provided by an embodiment of the present invention can be applied to a DSP. The method includes:

[0008] Dividing the matrix to be transposed stored in the DSP into M×N sub-matrices to be transposed according to the size of the cache data block, where both M and N are positive integers;

[0009] Performing a transposition operation on each sub-matrix to be transposed through a kernel until all sub-matrices to be transposed have completed the transposition operation to obtain a transposed matrix stored on the DSP. The transposed matrix is composed of M×N transposed sub-matrices;

[0010] Among them, the process of performing a transposition operation on the sub-matrix to be transposed in the i-th row and j-th column includes: moving the data of the sub-matrix to be transposed in the i-th row and j-th column to the cache data block, and performing a transposition operation on the sub-matrix to be transposed on the cache data block to obtain an operation result sub-matrix in the i-th row and j-th column; moving the operation result sub-matrix in the i-th row and j-th column to the corresponding transposed position of the sub-matrix to be transposed in the i-th row and j-th column to obtain a transposed sub-matrix in the i-th row and j-th column, and completing the transposition operation on the sub-matrix to be transposed in the i-th row and j-th column.

[0011] In a second aspect, an embodiment of the present invention provides a device for transposing large matrix signals in radar imaging by DSP, including: a preprocessing module, a cache data block, and multiple kernels;

[0012] The preprocessing module is configured to divide the matrix to be transposed stored in the DSP into M×N sub-matrices to be transposed according to the size of the cache data block, where both M and N are positive integers;

[0013] The kernel is configured to perform a transposition operation on each sub-matrix to be transposed until all sub-matrices to be transposed have completed the transposition operation to obtain a transposed matrix stored on the DSP. The transposed matrix is composed of M×N transposed sub-matrices;

[0014] Among them, the kernel is specifically configured to: move the data of the sub-matrix to be transposed in the i-th row and j-th column to the cache data block, and perform a transposition operation on the sub-matrix to be transposed on the cache data block to obtain an operation result sub-matrix in the i-th row and j-th column; move the operation result sub-matrix in the i-th row and j-th column to the corresponding transposed position of the sub-matrix to be transposed in the i-th row and j-th column to obtain a transposed sub-matrix in the i-th row and j-th column, and complete the transposition operation on the sub-matrix to be transposed in the i-th row and j-th column.

[0015] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: According to the method provided by the present invention, by dividing a large matrix into multiple sub-matrices, moving the sub-matrices to the cache data blocks according to the data storage order of the sub-matrices, performing transpose operations on the cache blocks, and then putting them back to the transposed positions; instead of directly reading data according to the transposed data order, it can reduce the number of address jumps during the transpose operation and improve the transpose efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of a matrix to be transposed provided by an embodiment of the present invention;

[0017] Figure 2 It is a flowchart of the implementation of a method for transposing a large matrix signal DSP in radar imaging provided by an embodiment of the present invention;

[0018] Figure 3 It is an engineering implementation code diagram for sub-matrix division provided by an embodiment of the present invention;

[0019] Figure 4 It is a schematic diagram of a scenario for performing transpose operations on a sub-matrix to be transposed provided by an embodiment of the present invention;

[0020] Figure 5 It is a schematic diagram of the storage address of an assembly function provided by an embodiment of the present invention;

[0021] Figure 6 It is a schematic diagram of the processing order of a sub-matrix to be transposed provided by an embodiment of the present invention;

[0022] Figure 7 It is a schematic diagram of the structure of a DSP provided by an embodiment of the present invention;

[0023] Figure 8 It is a schematic diagram for comparing the processing efficiencies of different transpose functions provided by an embodiment of the present invention;

[0024] Figure 9 It is a schematic diagram for comparing the processing efficiencies of different transpose methods provided by an embodiment of the present invention;

[0025] Figure 10 It is a schematic diagram of an SAR imaging processing flow provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to provide a thorough understanding of the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0027] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0028] It should also be understood that the term "and / or" as used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] As used in the specification and claims of the present invention, the term "if" can be interpreted according to the context as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted according to the context as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".

[0030] In addition, in the description of the specification and claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0031] The reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0032] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0033] As an example, refer to Figure 1 , taking a 4×4 matrix as an example, if the radar data is stored in the range direction and also read in the range direction, the traditional DMA-based transpose method can read the data in the storage order: data 1, data 2, data 3... data 16. At this time, the addresses of each data are continuous, and no address jumps are required, so the reading speed is relatively fast. If matrix transpose operation is to be performed, at this time, the traditional method will directly read the data in the azimuth direction, and its reading order is: data 1, data 5, data 9, data 13, data 2, data 6... data 16. At this time, the addresses between each data are not continuous, and 3 address jumps are required for each column of 4 data read, resulting in a significant reduction in the efficiency of data processing.

[0034] When the present invention performs the transpose operation, it first divides the large matrix into multiple sub-matrices, and then transposes the data of the sub-matrices after moving them to the cache data block. Inside the sub-matrix, the data can still be read in the storage direction of the data. For example, when dividing a 4×8 matrix into two 4×4 sub-matrices, the present invention reads a sub-matrix in the order: data 1, data 2, data 3, data 4, data 9, data 10, data 11... data 28. Only 1 address jump is required for each column of 4 data read, and the number of jumps is 1 / 3 of the traditional method.

[0035] In fact, the data scale of the large matrix signal of radar imaging far exceeds Figure 1 the example in. If the traditional method is used to transpose a matrix of size a×b, ab - 1 address jumps are required. If the present invention divides the a×b matrix into M×N sub-matrices, only aN - 1 address jumps are required, and the data processing efficiency is much higher than that of the traditional method.

[0036] The DSP transpose method for radar imaging large matrix signals provided by the embodiments of the present invention can be applied to radar imaging large matrix signal DSP transpose devices such as DSP, and the embodiments of the present invention do not impose any restrictions on the specific model of the DSP.

[0037] Figure 2 The flowchart of implementing a DSP transpose method for radar imaging large matrix signals provided by the embodiments of the present invention is shown. As an example but not a limitation, the method may include steps S201 - S206, and each step will be described below.

[0038] S201, divide the matrix to be transposed stored in the DSP into M×N sub-matrices to be transposed according to the size of the cache data block.

[0039] In a possible implementation manner, based on the transpose principle of the matrix (refer to the following formula (1.1)), the matrix to be transposed can be divided into M×N sub-matrices to be transposed.

[0040] Exemplarily, both M and N are positive integers. The magnitudes of M and N can be determined by the sizes of the matrix to be transposed and the cache data block.

[0041] Exemplarily, the transpose principle of a matrix can be expressed as:

[0042]

[0043] Among them, A, B, C, and D are four sub-matrices to be transposed, and T is the transpose symbol.

[0044] In one example, the cache data block can be set in the L2 module of the DSP.

[0045] Exemplarily, to meet most applicable situations and take into account the memory of L2, the size of the cache data block can be 128×128.

[0046] In one example, the present invention does not limit the data size of the sub-matrix to be transposed. Therefore, the sub-matrix to be transposed can be divided into a complete sub-matrix, a row remainder sub-matrix, a column remainder sub-matrix, and a row-column remainder sub-matrix according to the data type.

[0047] Exemplarily, referring to Figure 1 the sub-matrix boxed by the box 101 in, the number of rows and columns of the complete sub-matrix can be the same as the number of rows and columns of the cache data block respectively.

[0048] Exemplarily, referring to Figure 1 the sub-matrix boxed by the box 102 in, the number of columns of the column remainder sub-matrix is less than the number of columns of the cache data block.

[0049] Exemplarily, referring to Figure 1 the sub-matrix boxed by the box 103 in, the number of rows of the row remainder sub-matrix is less than the number of rows of the cache data block.

[0050] Exemplarily, referring to Figure 1 the sub-matrix boxed by the box 104 in, the number of rows and columns of the row-column remainder sub-matrix are less than the number of rows and columns of the cache data block respectively.

[0051] Specifically, the process of dividing the sub-matrix can be through, such as Figure 3The engineering implementation code shown. If nrn×nan is not an integer multiple of 128×128, the matrix to be transposed with nrn distance points and nan azimuth points is divided into blockRow_num (equal to M - 1 at this time) × blockCol_num (equal to N - 1 at this time) complete sub - matrices, M - 1 column - remainder sub - matrices with blockRow_rem columns, N - 1 row - remainder sub - matrices with blockCol_rem rows, and 1 row - and - column - remainder sub - matrix. Among them, blockRow_num is the integer part of nrn divided by 128, blockRow_rem is the remainder part, blockCol_num is the integer part of nrn divided by 128, and blockCol_rem is the remainder part. If nrn×nan is an integer multiple of 128×128, the matrix to be transposed can be divided into blockRow_num (equal to M at this time) × blockCol_num (equal to N at this time) complete sub - matrices.

[0052] S202, transfer the data of the sub - matrix to be transposed at the i - th row and j - th column to the cache data block through the kernel.

[0053] Exemplarily, i is a positive integer less than or equal to M, and j is a positive integer less than or equal to N.

[0054] In one example, see Figure 4 , the kernel of the DSP can transfer the data of the sub - matrix at the i - th row and j - th column to the cache data block in L2 according to the original data order.

[0055] For example, see Figure 1 the sub - matrix boxed by box 101 in, and its storage format in the cache data block is still: the first row, data 1, data 2, data 3; the second row, data 9, data 10, data 11.

[0056] Exemplarily, the kernel can transfer the data of the sub - matrix to be transposed from the DDR to L2 through DMA.

[0057] S203, perform a transpose operation on the sub - matrix to be transposed at the i - th row and j - th column on the cache data block to obtain the sub - matrix after the operation at the i - th row and j - th column.

[0058] In one example, a small - matrix transpose function optimized by an assembly function can be used to perform a transpose operation on the sub - matrix to be transposed at the i - th row and j - th column to obtain the sub - matrix after the operation at the i - th row and j - th column.

[0059] Specifically, see Figure 5 , an assembly function can be used to replace the inline function of the traditional method for transpose operation, and all the assembly programs are placed in a file with the suffix ".sa".

[0060] S204. Transfer the processed sub-array at the \(i\)-th row and \(j\)-th column to the corresponding transposed position of the sub-array to be transposed at the \(i\)-th row and \(j\)-th column through the kernel, to obtain the transposed sub-array at the \(i\)-th row and \(j\)-th column.

[0061] Exemplarily, the corresponding transposed position of the sub-array to be transposed at the \(i\)-th row and \(j\)-th column is the original position of the sub-array at the \(j\)-th row and \(i\)-th column in the matrix to be transposed on the matrix to be transposed.

[0062] For example, referring to Figure 4 , where the row number \(i\) and column number \(j\) of the orange-filled sub-array are equal, which is a diagonal sub-array, so the positions before and after transposition are the same. The row number \(i\) and column number \(j\) of the green, dark blue, and light blue-filled sub-arrays are not equal, which are non-diagonal sub-arrays. After transposition, their row numbers and column numbers are swapped, transferring from the \(i\)-th row and \(j\)-th column to the \(j\)-th row and \(i\)-th column.

[0063] In a possible implementation, the DSP may include multiple parallel processing kernels, such as 8. Each kernel can simultaneously complete the transposition operation on different sub-arrays to be transposed through the above steps S202 - S204.

[0064] In an example, a quarter of the kernels can be responsible for the transposition operation of the diagonal matrix, and the remaining kernels are responsible for the transposition operation of the non-diagonal matrix.

[0065] For example, the 8 kernels of the DSP can be respectively denoted as kernel 0 to kernel 7. Kernel 1 and kernel 3 are responsible for the transposition operation of the diagonal matrix, and they can transpose the diagonal matrix and put it back in place, or at the corresponding position of another DDR. The remaining kernels are responsible for the transposition operation of the non-diagonal matrix. For example, Figure 4 in, kernel 1 and kernel 2 are responsible for the transposition operation of the sub-array filled with green, kernel 4 and kernel 5 are responsible for the transposition operation of the sub-array filled with dark blue, and kernel 6 and kernel 7 are responsible for the transposition operation of the sub-array filled with light blue.

[0066] Specifically, if the transposed sub-array is to be stored in the DDR where the matrix to be transposed is located, then when processing the non-diagonal matrix, two kernels need to simultaneously process the sub-array to be transposed at the \(i\)-th row and \(j\)-th column and the sub-array to be transposed at the \(j\)-th row and \(i\)-th column, so as to avoid occupying the storage space where the sub-array to be transposed at the \(j\)-th row and \(i\)-th column is located after the sub-array to be transposed at the \(i\)-th row and \(j\)-th column is processed, resulting in data loss of the sub-array to be transposed at the \(j\)-th row and \(i\)-th column.

[0067] S205. Determine whether there is a sub-array to be transposed that has not been transposed.

[0068] In a possible implementation, if there are still uncompleted sub-arrays, the size of \(i\) and / or \(j\) can be increased, and starting from step S202, continue to execute the transposition operation.

[0069] Exemplarily, for the kernel responsible for processing diagonal matrices, i can be set to i + 1 and j can be set to j + 1 to process the next diagonal matrix; for the kernel responsible for processing non - diagonal matrices, i can be set to i + 1 or j can be set to j + 1 to process the next non - diagonal matrix.

[0070] In one example, the priority order of sub - arrays to be transposed with different data types from largest to smallest can be: complete sub - arrays, row remainder sub - arrays, column remainder sub - arrays, and row - column remainder sub - arrays.

[0071] Exemplarily, referring to Figure 6 , when performing the transposition operation, the complete sub - arrays can be processed preferentially. After the complete sub - arrays are processed, according to the number of row remainder sub - arrays and column remainder sub - arrays, the type with a larger number is processed first, and finally the row - column remainder sub - arrays are processed.

[0072] Specifically, “function name.cproc formal parameter” can be defined at the beginning of the assembly function, and at the end of the function, the “.endproc” statement can be used. The matrix transposition is implemented in assembly language by dividing it into two loops, namely the outer loop and the inner loop, which traverse the rows and columns of the matrix respectively, calculate the addresses of the input matrix and the output matrix, exchange matrix elements through LDB and STB instructions, and in addition, performance optimization is achieved by using parallel instructions of the DSP such as LDDW and STDW to load and store multiple elements at one time.

[0073] In another possible implementation, after processing all sub - arrays to be transposed, step S206 can be performed to obtain the transposed matrix stored on the DSP, which is composed of the subsequent M × N transposed sub - arrays.

[0074] According to the method provided by the present invention, by dividing a large matrix into multiple sub - arrays, transporting the sub - arrays to the cache data block according to the data storage order of the sub - arrays, performing the transposition operation on the cache block and then putting it back to the transposed position; instead of directly reading data according to the transposed data order, the number of address jumps during the transposition operation can be reduced, and the transposition efficiency can be improved. Moreover, by transporting the data to the cache data block for operation, only the data of the sub - array size needs to be stored each time, and there is no need to allocate new space to store the data of the entire matrix to be transposed, which can reduce the occupied storage space.

[0075] Furthermore, through the design of sub - arrays with different data types, such as row remainder sub - arrays, column remainder sub - arrays, and row - column remainder sub - arrays, the present invention removes the limitation on the transposition length and does not require the length of the matrix to be transposed to be an integer power of 2. In addition, through the parallel processing of multiple kernels, the transposition efficiency can be further improved.

[0076] Figure 7The following is a schematic structural diagram of a DSP transpose device for large matrix signals in radar imaging provided by an embodiment of the present invention. By way of example and not limitation, the device may include a cached data block 710, a preprocessing module 720, and a plurality of cores 730 (only one is shown here).

[0077] The preprocessing module 720 is configured to divide the matrix to be transposed stored in the DSP into M×N sub-matrices to be transposed according to the size of the cached data block 710, where M and N are both positive integers; the core 730 is configured to perform a transpose operation on each sub-matrix to be transposed until all sub-matrices to be transposed have completed the transpose operation to obtain the transposed matrix stored on the DSP, and the transposed matrix is composed of M×N transposed sub-matrices; specifically, the core 730 is configured to: move the data of the sub-matrix to be transposed in the i-th row and j-th column to the cached data block 710, and perform a transpose operation on the sub-matrix to be transposed on the cached data block to obtain the sub-matrix after operation in the i-th row and j-th column; move the sub-matrix after operation in the i-th row and j-th column to the corresponding transposed position of the sub-matrix to be transposed in the i-th row and j-th column to obtain the transposed sub-matrix in the i-th row and j-th column, and complete the transpose operation on the sub-matrix to be transposed in the i-th row and j-th column.

[0078] According to the method provided by the present invention, by dividing a large matrix into multiple sub-matrices, moving the sub-matrices to the cached data block according to the data storage order of the sub-matrices, performing a transpose operation on the cached block and then putting them back to the transposed positions; instead of directly reading data according to the data order after transpose, the number of address jumps during the transpose operation can be reduced, and the transpose efficiency can be improved.

[0079] Exemplarily, the device provided by the present invention can be mounted on SAR real-time imaging systems of various platforms such as airborne and missile-borne platforms to perform high-resolution and wide-swath real-time imaging and optimize the processing efficiency of these systems. Refer to Figure 8 It can be seen that in the case of single-DSP imaging processing, for an image of 1024*2048 points, the ground range projection of 512*512 points only takes about 0.172 s, which proves the practicability of the method of the present invention.

[0080] Furthermore, the present invention uses an assembly function instead of an inline function in the traditional method for transpose operation, which can further improve the transpose efficiency. Specifically, refer to Figure 9 In comparison with traditional DMA matrix transpose, optimized inline function transpose, and optimized assembly matrix transpose of the present invention, all transpose a 1024*2048 point matrix. When the DSP main frequency is 1 GHz, the assembly takes about 3.6 ms with the highest efficiency, the inline function takes about 5.47 ms, and the traditional DMA transpose takes about 25.7 ms with the lowest efficiency. During the SAR imaging processing, refer to Figure 10, it will involve three large matrix transpositions. Taking 1024*2048 points as an example, the implementation time of the large matrix assembly is about 3.6 ms. For the ground distance projection of 512*512 points, the SAR imaging time is only about 0.172 s. The present invention greatly reduces the transposition time and improves the real-time processing efficiency.

[0081] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0082] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

Claims

1. A method for transposing large matrix signals of radar imaging by DSP, characterized in that, The method comprises: According to the size of the cache data block, the matrix to be transposed stored in the DSP is divided into M×N sub-matrices to be transposed, where M and N are both positive integers; Performing a transposition operation on each of the sub-matrices to be transposed by the kernel until all the sub-matrices to be transposed have completed the transposition operation to obtain a transposed matrix stored on the DSP, wherein the transposed matrix is composed of M×N transposed sub-matrices; The process of performing a transposition operation on the i-th row and j-th column to-be-transposed sub-matrix includes: moving data of the i-th row and j-th column to-be-transposed sub-matrix to the cache data block, and performing a transposition operation on the i-th row and j-th column to-be-transposed sub-matrix on the cache data block to obtain the i-th row and j-th column operated sub-matrix; moving the i-th row and j-th column operated sub-matrix to the corresponding transposition position of the i-th row and j-th column to-be-transposed sub-matrix to obtain the i-th row and j-th column transposed sub-matrix, and completing the transposition operation on the i-th row and j-th column to-be-transposed sub-matrix.

2. The method according to claim 1, wherein The cache data block is arranged in the L2 module of the DSP.

3. The method according to claim 2, wherein The size of the cache data block is 128×128.

4. The method according to claim 1, wherein The data types of the sub-matrix to be transposed include: a complete sub-matrix, a row-rematrix, a column-rematrix, and a row-column-rematrix sub-matrix; the number of rows and the number of columns of the complete sub-matrix are respectively the same as the number of rows and the number of columns of the cache data block, the number of rows of the row-rematrix sub-matrix is smaller than the number of columns of the cache data block, the number of rows of the column-rematrix sub-matrix is smaller than the number of columns of the cache data block, and the number of rows and the number of columns of the row-column-rematrix sub-matrix are respectively smaller than the number of rows and the number of columns of the cache data block.

5. The method according to claim 4, wherein The priority order of the submatrices to be transposed of different data types from high to low is: complete submatrix, row and column remainder submatrix, and row and column remainder submatrix.

6. The method according to claim 1, wherein The corresponding transposed position of the i-th row and j-th column to-be-transposed submatrix is the original position of the j-th row and i-th column to-be-transposed matrix on the to-be-transposed matrix.

7. The method according to claim 1, characterized in that, The DSP includes a plurality of cores for parallel processing, wherein one quarter of the cores are used to perform the transposition operation on the submatrix to be transposed whose position type is a diagonal submatrix, and the remaining cores are used to perform the transposition operation on the submatrix to be transposed whose position type is a non-diagonal submatrix; The row number i and column number j of the submatrix to be transposed whose position type is a diagonal submatrix are equal, and the row number i and column number j of the submatrix to be transposed whose position type is a non-diagonal submatrix are not equal.

8. A DSP transposition device for radar imaging large matrix signals, characterized in that include: Preprocessing modules, cache blocks, and multiple cores; The preprocessing module is used to divide the matrix to be transposed stored in the DSP into M×N sub-matrices to be transposed according to the size of the cache data block, where M and N are both positive integers; The kernel is used to perform a transposition operation on each of the sub-matrices to be transposed, until all the sub-matrices to be transposed have completed the transposition operation to obtain a transposed matrix stored on the DSP, wherein the transposed matrix is composed of M×N transposed sub-matrices; Among them, the kernel is specifically used for: moving the data of the sub-array to be transposed at the i-th row and j-th column to the cache data block, and performing a transposition operation on the sub-array to be transposed on the cache data block to obtain the sub-arrays after operation at the i-th row and j-th column; moving the sub-arrays after operation at the i-th row and j-th column to the corresponding transposed positions of the sub-arrays to be transposed at the i-th row and j-th column to obtain the transposed sub-arrays at the i-th row and j-th column, thereby completing the transposition operation on the sub-arrays to be transposed at the i-th row and j-th column.