Cross-module parallel high-speed space-time anti-interference hardware algorithm
By enabling cross-module parallel computing through field-programmable gate arrays, the computational resource and speed issues of space-time anti-interference algorithms under conditions of large array element number and delay tap number are solved, thereby improving the real-time performance of anti-interference systems for satellite communication and satellite navigation.
Patent Information
- Application Number
- CN202510857153.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-31
AI Technical Summary
Existing space-time anti-interference algorithms, under conditions of large array elements and delayed taps, involve large matrix operations and numerous memory accesses, which limits hardware implementation resource utilization and computation speed, making it difficult to meet real-time requirements.
Field-programmable gate arrays (FPGAs) are used to achieve cross-module parallel computing. Complex matrices are stored and read out in column parallelism, and autocorrelation matrices and Koleski decompositions are calculated in block parallelism to optimize the calculation order and achieve parallelization within modules and cross-module parallel computing.
It significantly improves the speed of anti-interference weight update and reduces computational latency, making it suitable for high-speed anti-interference requirements in fields such as satellite communication and satellite navigation.
Smart Images

Figure CN120880595A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to space-time anti-interference algorithms, specifically to a high-speed space-time anti-interference hardware algorithm that operates across modules in parallel. Background Technology
[0002] In modern wireless and satellite communication applications, electromagnetic spectrum resources are becoming increasingly scarce, and channel environments are becoming more complex. Communication systems face dual threats from natural environments such as multipath fading and rain obstruction, as well as from human interference such as intentional interference, signal interception, and deception interference. Anti-jamming communication technology, as a key means to improve the reliability, security, and stability of communication systems, has become a research hotspot in recent years. Its main goal is to ensure the smooth operation of communication links and the effective transmission of information under hostile or interference conditions. Among anti-jamming technologies, array signal processing-based anti-jamming algorithms have been widely used. These algorithms, particularly those based on spatiotemporal anti-jamming, can suppress interference based on the direction of arrival and process the interference signal in the spectrum using delayed taps. Therefore, they can effectively suppress broadband interference. Furthermore, the anti-jamming capability of the system can be enhanced by increasing the number of tap delays, demonstrating high flexibility.
[0003] Space-time anti-interference algorithms can be categorized into various criteria based on different optimization objectives, such as the minimum mean square error criterion, the signal-to-interference-plus-noise ratio (SIR) maximization criterion, the power inversion criterion, and the minimum variance distortionless response criterion. Among these, the minimum variance distortionless response criterion and the power inversion criterion are the most commonly used. The former aims to minimize the total power while maintaining the target signal response unchanged, but it requires the direction of arrival of the target signal as a prerequisite. The latter aims to minimize the total power while ensuring the power of a single reference channel remains constant, requiring no prior information, but it is only effective for cases with a high SIR. Although these two criteria are widely used, the required matrix operations increase quadratically when the number of array elements and taps is large. The need for intermediate storage of matrix elements also increases the number of memory accesses, posing challenges to resource utilization and computational speed in the hardware implementation of the algorithm. Therefore, to ensure the real-time performance of the anti-interference system and the speed of weight updates, it is still necessary to research and develop efficient and feasible high-speed hardware implementation algorithms. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention proposes a cross-module parallel high-speed space-time anti-interference hardware algorithm based on field-programmable gate arrays (FPGAs). This algorithm significantly improves the anti-interference weight update speed under conditions of large array element number and delay tap number, enhances the real-time performance of the hardware system, and is applicable to fields such as satellite communication and satellite navigation.
[0005] The high-speed, space-time anti-interference hardware algorithm of the present invention, which is modularly parallel, includes the following steps:
[0006] 1) Column-parallel storage and retrieval of matrices:
[0007] Hardware computation is performed based on field-programmable gate arrays (FPGAs). The complex matrices involved in the computation are stored and read out according to the following rules: all complex matrices are stored in column-parallel format using random access memory (RAM), and each column is read out in parallel when reading out the complex matrix.
[0008] 2) Each parallel computing module:
[0009] a) Autocorrelation matrix solving module:
[0010] Satellite communication or satellite navigation receives communication signals and preprocesses the communication signals to obtain a complex sampling matrix;
[0011] Each data vector is read in parallel from the complex sampling matrix, and then the outer product of the vectors is performed to obtain the autocorrelation matrix.
[0012] R xx :
[0013] In the process of vector outer product operation, the autocorrelation matrix is first divided into sub-blocks, and then parallel computation is performed in each sub-block; the reading and storage principle adopts column parallelism.
[0014] b) Koleski decomposition module:
[0015] Performing Koleski decomposition on the autocorrelation matrix decomposes the complex conjugate symmetric autocorrelation matrix into a lower triangular matrix L.
[0016] Each element in the lower triangular matrix is solved iteratively;
[0017] In the iterative process of solving the lower triangular matrix, one column of elements of the lower triangular matrix is solved each time, and then the other elements of the lower triangular matrix are updated by using the outer product of the column of elements. The solution and update process is repeated until all elements are iterated.
[0018] c) Forward iterative solution module:
[0019] The forward iterative equation LT = b is solved by the forward iterative solution module to obtain the forward solution T, where b is the anti-interference criterion vector. During the forward iterative solution process, one element of the forward solution is solved each time, and this element is used to update the other elements of the forward solution. The solution and update process is repeated until all elements of the forward solution are iterated.
[0020] 3) Cross-module parallel computing:
[0021] In step 2)a), the autocorrelation matrix solving module is computed in parallel using a block-based approach. Within a single sub-block, the autocorrelation matrix elements of multiple columns are computed in parallel. In step 2)b), the Koleski decomposition module also computes the elements of the lower triangular matrix in a column-parallel manner, and the computation of each column only requires the lower triangular elements of the current column's autocorrelation matrix. In step 2)c), the forward iterative solving module, when solving for an element of a forward solution, also only requires the lower triangular elements of the current column's lower triangular matrix. Thus, the three modules achieve cross-module parallel computation by marking the completion of each column's computation, without waiting for the previous module to complete its computation before starting the computation of the next module.
[0022] 4) Backward Iterative Solution Module:
[0023] After the three modules have completed their parallel computations, the backward iterative equation L is solved by the backward iterative solution module. H w tmp =T obtains the intermediate variable w tmp L H It is an upper triangular matrix;
[0024] In the process of solving the backward iterative equation in the backward iterative solution module, each time an element of the intermediate variable is solved, this element is used to update the other elements of the intermediate variable. The solution and update process is repeated until all elements are iterated.
[0025] Then, according to the weighting formula, the intermediate variable w tmp Normalization yields the following spatiotemporal weights w:
[0026] w = w tmp / b H w tmp
[0027] Among them, b H This is the conjugate transpose of the anti-interference criterion vector.
[0028] In step 1), the number of array elements of the receiving communication signal array antenna is M, the number of delays per unit time for each received communication signal is defined as the number of taps, and the number of taps is P. The received communication signal is sampled in a snapshot, and the number of data vectors in the snapshot is defined as the number of snapshots, and the number of snapshots is N. Each received communication signal is sampled, and the resulting complex sampling matrix is M×P rows and N columns. The algorithm involves storing an M×P row and N column complex sampling matrix, and the calculation process also involves storing three M×P row and M×P column complex intermediate matrices, including an autocorrelation matrix, an upper triangular matrix, and a lower triangular matrix. When the values of M, P, and N are all large, the storage and reading of the complex matrix will bring a huge time overhead. This invention employs column-parallel storage of all complex matrices mentioned above using Random Access Memory (RAM). Specifically, each matrix is stored using M×P RAM memories, with each memory storing one row of data for the complex matrix. Each complex element uses the same RAM memory space, with the least significant bit representing the imaginary part and the most significant bit representing the real part. Furthermore, when reading the complex matrix, each column is read in parallel. M is a natural number ≥ 8, P is a natural number ≥ 4, and N is a power of 2 ≥ 128.
[0029] In step 2)a), the communication signal is preprocessed to obtain a complex sampling matrix, including the following steps: digital quadrature downconversion to obtain a complex signal, which is then sampled and stored in parallel to obtain a complex sampling matrix, and stored in the complex sampling matrix storage RAM.
[0030] Autocorrelation matrix R xx It has complex conjugate symmetry properties, and its autocorrelation matrix R xxWhen calculating the elements, only the lower triangular elements need to be calculated. When M and P are large, the limited FPGA computing resources cannot support the parallel computation of all elements. Therefore, the lower triangular elements of the autocorrelation matrix are decomposed into K sub-blocks, and the elements of each sub-block in the K sub-blocks are calculated in parallel. In the vector outer product operation, the autocorrelation matrix is first divided into sub-blocks. The specific division method of the autocorrelation matrix is as follows: the lower triangular elements of the autocorrelation matrix are sorted in column-first traversal and stretched into a row vector. Then, the row vector is divided into K sub-blocks on an equal basis, where K is a natural number from 4 to 16. The number of elements in each sub-block is C = (M×P×(M×P+1)) / (2×K). Then, parallel computation is performed in each sub-block. The elements in a sub-block are obtained by traversing all N sampling snapshots. Finally, the column-parallel fixed-to-floating module of the field-programmable gate array is called to perform column-parallel fixed-to-floating calculation on the elements in the calculated sub-blocks. The parallel computing scheme within each sub-block is as follows: First, two address-mapped read-only memories (ROMs) are used to store the mapping relationship between the two-dimensional row coordinates and column coordinates of the lower triangular elements of the autocorrelation matrix and the one-dimensional row vector index, respectively. The two-dimensional row coordinates and column coordinates are the stored values of the address-mapped ROMs, and the one-dimensional row vector index is the address of the address-mapped ROMs. Second, C complex multiplication accumulators are constructed, the number of which is equal to the number of elements C contained in each sub-block. These C complex multiplication accumulators calculate each element of the autocorrelation matrix according to the parallel computing method within the sub-block. In this process, C complex multiplication accumulators select a subset of sampled snapshot elements for calculation (referring to elements in a vector). The selected subset of sampled snapshot elements is determined by the number of sub-blocks being calculated, k = 1, ..., K. The specific mapping relationship is as follows: the one-dimensional row vector index corresponding to the autocorrelation matrix element calculated by the c-th complex multiplication accumulator is C × k + c, c = 1, ..., C. This index is used as the address input to the address mapping ROM, and the resulting two-dimensional row coordinate i and column coordinate j are the indexes of the sampled snapshots in the above formula.
[0031] The principle of column-parallel read and storage is as follows: When calculating each sub-block element, all sampling snapshots in the complex sampling matrix must be read in parallel. The read operation and complex multiplication and accumulation operation are performed in parallel according to the pipeline to improve the calculation speed of each sub-block element. When the sampling snapshot is traversed, the autocorrelation matrix of the k-th sub-block is also calculated at the same time. Then, the calculated C elements are written into the first autocorrelation matrix storage RAM in the form of a two-dimensional matrix. After writing, the sub-block number is incremented by 1 to start the parallel operation of the next sub-block until all sub-blocks are calculated and stored.
[0032] In step 2)b), if the current calculation column is the first column, there is no need to update the autocorrelation matrix elements; otherwise, after each autocorrelation matrix element is read in, the autocorrelation matrix elements are first updated by parallel column subtraction, and then the diagonal elements are obtained by square root operation in a separate calculation cycle, and the off-diagonal elements are obtained by parallel division operation. This calculation order makes it possible to calculate only the elements of the autocorrelation matrix in the same column when calculating the lower triangular matrix of each column. After the calculation of a column of elements is completed, the elements of the column are written back to the lower triangular matrix storage RAM. After the calculation of a column is completed, the column elements are multiplied by the outer product and the previous update result is accumulated to obtain the current update result.
[0033] The specific method for outer product is as follows: Parallel multiplication calculation is performed by calling the multiplication module built into the field-programmable gate array (FPGA). l ik Let be the element in the i-th row and k-th column of the lower triangular matrix. The element l in the j-th row and k-th column of the lower triangular matrix jk The conjugate of . The parallel method uses fixed indices k and i, calculating the results corresponding to all indices j. If the current calculation column is the first column, the multiplication result is directly written back to the lower triangular matrix storage RAM with (i,j) as the index; otherwise, the corresponding element of a column needs to be read in parallel from the intermediate storage RAM with (i,j) as the index, and then added in parallel with the column element calculated by multiplication. When all Once the calculation is complete, you can begin calculating the next column of lower triangular matrix elements, repeating the column operations and updates until all lower triangular matrix elements have been calculated.
[0034] In step 2)c), under the minimum variance distortion-free response criterion and the power inversion criterion, the spatiotemporal weight calculation formula is unified into the following form:
[0035]
[0036] The space-time anti-interference algorithm has two accurate criteria: the minimum variance distortionless response criterion and the power inversion criterion. For the minimum variance distortionless response criterion, the anti-interference criterion vector b is taken as the space-time steering vector determined by the incoming wave direction and tap delay. For the power inversion criterion, the anti-interference criterion vector b takes the value [1,0,0,……0]. An intermediate variable w is defined. tmp The spatiotemporal weights are solved using equations:
[0037]
[0038] intermediate variable w tmp Consider it as equation R xx w tmp =L(L H w tmpThe solution to ) = b; define LT = w tmp T is the forward solution, L H Since the matrix is an upper triangular matrix, it is the conjugate transpose of the lower triangular matrix, thus we get w = w tmp / b H w tmp .
[0039] The forward iterative solution process of the equation is described by the following formula:
[0040]
[0041] From the above equation, we obtain: Calculate the i-th element t of the forward solution i Only the elements of the lower triangular matrix L up to the i-th column need to be known; the process of solving each forward solution of the equation is as follows: First, the i-th element b of the anti-interference criterion vector is obtained by column-parallel subtraction. i The update process involves updating the first element b1 of the anti-interference criterion vector without requiring column-parallel subtraction; then, the element t of the current forward solution is obtained through division. i Then use the lower triangular elements l of all lower triangular matrices after the i-th column and i-th row. ki With the i-th element t of the forward solution i Multiply, where k > i; if i is 1, store the multiplication result directly; otherwise, add the multiplication result to the element stored in the corresponding position before updating; solve each forward solution element in turn according to the above process until all elements are found.
[0042] Step 3) specifically includes the following steps: The autocorrelation matrix solving module outputs the column index of the currently written autocorrelation matrix storage RAM to the Kolesky decomposition module. If the column index of the currently to be calculated lower triangular matrix is less than the column index output by the autocorrelation matrix solving module, the Kolesky decomposition module starts to calculate the lower triangular elements of the current column lower triangular matrix; otherwise, it waits. The Kolesky decomposition module outputs the column index of the currently written autocorrelation matrix storage RAM to the forward iteration solving module. If the index of the element of the currently to be calculated forward solution is less than the output column index of the Kolesky decomposition module, the forward equation iteration solving module starts to solve the elements of the current forward solution; otherwise, it waits.
[0043] In step 4), the backward iterative solution process of the equation is described by the following formula:
[0044]
[0045] The coefficient matrix of the backward iteration equation is composed of the upper triangular matrix L H The backward iterative equation requires the M×P-th element w of the intermediate variable. tmpM×P Start by solving in reverse order until the first element w of the intermediate variable is reached.tmp1 And calculate the i-th element w of the intermediate variable. tmpi Only the upper triangular matrix L from the i-th row onwards needs to be known. H The elements of the matrix; therefore, to facilitate backward iteration, after reading the elements of each column of the upper triangular matrix L in parallel during forward iteration, the conjugate of the elements of that column is taken and serially stored into the upper triangular matrix storage RAM; after the forward iteration is completed, backward iteration is performed in the same parallel computing manner; and, for each intermediate variable, the i-th element w is calculated... tmpi That is, the i-th element of the normalization factor Perform a multiplication and accumulation. The conjugate of the i-th element of the anti-interference criterion vector; after the forward and backward equations are solved iteratively, the intermediate variable w is obtained simultaneously. tmp and normalization factor b H w tmp Finally, the intermediate variable w tmp Regarding the normalization factor b H w tmp Normalization by division yields the spatiotemporal weight w.
[0046] Normalization is achieved through normalized inner product and normalized division.
[0047] Advantages of this invention:
[0048] This invention decomposes a space-time anti-interference algorithm based on the minimum variance distortion-free criterion and the power inversion criterion into an autocorrelation matrix solving module, a Kolesky decomposition module, and a weight calculation module. By optimizing the calculation order of matrix elements and the storage method of the matrix in the three modules, and by using a field-programmable gate array (FPGA), cross-module parallel computing of the three modules is realized, significantly improving the system computing speed under conditions of high matrix element number and high delay tap number. It is suitable for high-speed anti-interference scenarios such as satellite communication and satellite navigation. The specific advantages are as follows:
[0049] (1) By storing the intermediate matrix in parallel RAM, the extra time overhead caused by serial memory access of a large number of matrix elements is avoided.
[0050] (2) By rationally designing the computation order within each module, highly parallel computation within the module was achieved;
[0051] (3) The unified planning of each module is to perform calculations in a column-parallel manner. By passing column indexes between modules, cross-module parallel calculations of the three modules are realized, which significantly reduces the overall algorithm's computational latency. Attached Figure Description
[0052] Figure 1 This is a flowchart of the high-speed space-time anti-interference hardware algorithm for modular parallelism of the present invention.
[0053] Figure 2 This is a decomposition diagram of the autocorrelation matrix solving module of the high-speed space-time anti-interference hardware algorithm of the present invention, which is in modular parallelism.
[0054] Figure 3 This is a diagram showing the calculation order of the autocorrelation matrix elements of the modularly parallel high-speed space-time anti-interference hardware algorithm of the present invention.
[0055] Figure 4 This is a schematic diagram of the Koleski decomposition module of the modular parallel high-speed space-time anti-interference hardware algorithm of the present invention.
[0056] Figure 5 This is a decomposition diagram of the forward iterative solution module of the modular parallel high-speed space-time anti-interference hardware algorithm of the present invention;
[0057] Figure 6 This is a decomposed schematic diagram of the backward iterative solution module of the high-speed space-time anti-interference hardware algorithm of the present invention, which is in modular parallelism. Detailed Implementation
[0058] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0059] This embodiment features a module-parallel high-speed space-time anti-interference hardware algorithm and a space-time anti-interference algorithm under the power inversion criterion, such as... Figure 1 As shown, it includes the following steps:
[0060] 1) Column-parallel storage and retrieval of matrices:
[0061] Hardware computation is performed using a field-programmable gate array (FPGA), employing four sets of complex column parallel storage RAMs. The array antenna for receiving communication signals has 16 elements (M=16). The number of time units of delay applied to each received communication signal is defined as the number of taps (P=4). The received communication signals are sampled in snapshots, and the number of data vectors in each snapshot is defined as the number of snapshots (N=256). The real and imaginary parts of each complex sampled data are represented by 16-bit fixed-point numbers. The complex sampling matrix storage RAM group consists of 64 RAMs with a depth of 256. Each RAM cell is 32 bits, with the high 16 bits representing the imaginary part and the low 16 bits representing the real part. The autocorrelation matrix storage RAM group consists of 64 RAMs with a depth of 64. To avoid data overflow, each storage cell uses an 80-bit width, with the high 40 bits representing the fixed-point imaginary part and the low 40 bits representing the fixed-point real part. The remaining two RAM groups each consist of 64 RAMs with a depth of 64 bits. Each RAM has a storage width of 64 bits, with the high 32 bits storing the imaginary part of the floating-point number and the low 32 bits storing the real part of the floating-point number.
[0062] 2) Each parallel computing module:
[0063] a) Autocorrelation matrix solving module, such as Figure 2 As shown:
[0064] The communication signal is preprocessed to obtain a complex sampling matrix, including the following steps: digital quadrature downconversion to obtain a complex signal, which is then sampled and stored in parallel to obtain a complex sampling matrix, and stored in the complex sampling matrix storage RAM;
[0065] Each data vector, i.e., each sampling snapshot, is read in parallel from the complex sampling matrix, and then the vector outer product is performed according to the following formula:
[0066]
[0067] Among them, R xx X(n) represents the autocorrelation matrix, and X(n) represents the nth data vector read in parallel from the complex sampling matrix, i.e., the sampling snapshot, which is a column vector of length M××P.
[0068] Autocorrelation matrix R xx It has complex conjugate symmetry properties, and its autocorrelation matrix R xx When calculating the elements, only the lower triangular elements need to be computed. When M and P are large, the limited FPGA computing resources cannot support the parallel computation of all elements. Therefore, the lower triangular elements of the autocorrelation matrix are decomposed into K sub-blocks, with K set to 8. Parallel computation is then performed on each element in each of the K sub-blocks. During the vector outer product operation, the autocorrelation matrix is first divided into sub-blocks. The specific division method is as follows: the lower triangular elements of the autocorrelation matrix are sorted according to column-first traversal and stretched into a row vector. Then, the row vector is divided into 8 sub-blocks, each containing 260 multiply-accumulate modules. Parallel computation is then performed within each sub-block by traversing all... 256 sample snapshots are taken and averaged to obtain 260 elements in each sub-block. Finally, the column-parallel fixed-to-float module built into the field-programmable gate array (FPGA) is called to perform column-parallel fixed-to-float calculations on the calculated elements in the sub-block. The parallel calculation scheme within each sub-block is as follows: First, two ROMs are used to store the mapping relationship between the two-dimensional row coordinates and column coordinates of the triangular elements under the autocorrelation matrix and the one-dimensional row vector index, respectively. The two-dimensional row coordinates and column coordinates are the stored values in the ROM, and the one-dimensional row vector index is the address of the ROM. Second, C complex multiplication-accumulation groups are constructed. These C complex multiplication-accumulation groups calculate each element of the autocorrelation matrix according to the parallel calculation method within the sub-block.
[0069]
[0070] Among them, R xx(i,j) is the element in the i-th row and j-th column of the autocorrelation matrix, and X(n,i) is the i-th element of the n-th sampling snapshot. * (n,j) is the complex conjugate of the j-th element of the n-th sampling snapshot; the input sampling snapshot elements of the C complex multiplication-accumulation groups are determined by the number of sub-blocks currently being computed, k = 1, ..., K. Specifically, the one-dimensional row vector index corresponding to the autocorrelation matrix element computed by the C-th complex multiplication-accumulation group is C × k + c, c = 1, ..., C. This index is used as the address to input to the ROM, and the resulting two-dimensional row coordinate i and column coordinate j are the sampling snapshot indices in the above formula; the division method of the lower triangular elements is as follows... Figure 3 As shown, following the column-first traversal method, first traverse the first 260 elements corresponding to the solid arrows, which is the first block, then traverse the 260 elements corresponding to the dashed arrows, which is the second block, and so on until the segmentation is completed.
[0071] The principle of column-parallel read and store is as follows: When calculating each sub-block element, all sampling snapshots in the complex sampling matrix must be read in parallel. The read operation and complex multiplication and accumulation operation are performed in parallel according to the pipeline to improve the calculation speed of each sub-block element. When the sampling snapshot is traversed, the autocorrelation matrix of the k-th sub-block is also calculated at the same time. Then, the calculated C elements are written into the first autocorrelation matrix storage RAM in the form of a two-dimensional matrix. After writing, the sub-block number is incremented by 1, i.e., k = k + 1, and the parallel operation of the next sub-block begins until all sub-blocks have been calculated and stored.
[0072] b) Koleski decomposition module:
[0073] Performing Koleski decomposition on the autocorrelation matrix decomposes the complex conjugate symmetric autocorrelation matrix into a lower triangular matrix L:
[0074] R xx =LL H
[0075] Among them, L H It is an upper triangular matrix, which is the conjugate transpose of the lower triangular matrix;
[0076] Each element in the lower triangular matrix is solved iteratively:
[0077]
[0078] Among them, l ij Let i be the element in the i-th row and j-th column of the lower triangular matrix. Let be the complex conjugate of the element in the j-th row and k-th column of the lower triangular matrix;
[0079] During the iterative solution of the lower triangular matrix, the elements of the lower triangular matrix are calculated column by column, that is, all elements in the i-th row are calculated in parallel while keeping j constant; if the current calculation column is the first column, there is no need to update the elements of the autocorrelation matrix. and If all values are 0, then the autocorrelation matrix elements are updated first using column-parallel subtraction, requiring 64 floating-point complex subtractors. Then, for diagonal elements, a separate computation cycle is used to obtain them via square root operations, requiring one floating-point square root operator and 63 floating-point complex dividers. For off-diagonal elements, parallel division operations are performed in parallel. For diagonal elements... ii It requires a separate calculation cycle to obtain the result via square root operation, for off-diagonal elements l ij This can be achieved using division operations in parallel; this computational order means that only the elements of the autocorrelation matrix in the same column are needed when calculating each column of the lower triangular matrix; after the calculation of one column of elements is completed, the elements of that column are written back to the lower triangular matrix storage RAM until all outer product operations are completed, then the calculation of the next column of the next matrix begins; this step requires 63 floating-point complex multipliers and floating-point complex adders; the column index written to the L matrix storage RAM can be used to indicate whether the forward iterative solution module of the subsequent equations should start the solution of the current element, such as... Figure 4 As shown;
[0080] Parallel multiplication calculations are performed by calling the multiplication module built into the field-programmable gate array (FPGA). l ik Let be the element in the i-th row and k-th column of the lower triangular matrix. The element l in the j-th row and k-th column of the lower triangular matrix jk The conjugate of the multiplication operation; the parallel method is to fix the k and i indices and calculate the results corresponding to all j indices; if the current calculation column is the first column, the multiplication result is directly written back to the lower triangular matrix storage RAM with (i,j) as the index; otherwise, it is necessary to read the corresponding element of a column in parallel with (i,j) as the index from the intermediate storage RAM and then add it in parallel with the column element calculated by multiplication; when all Once the calculation is complete, you can begin calculating the next column of lower triangular matrix elements, repeating column operations and updates until all lower triangular matrix elements have been calculated.
[0081] c) Forward iterative solution module:
[0082] Under the minimum variance distortionless response criterion and the power inversion criterion, the formula for calculating the spatiotemporal weight w is unified as follows:
[0083]
[0084] The space-time anti-interference algorithm has two accurate criteria: the minimum variance distortionless response criterion and the power inversion criterion. For the minimum variance distortionless response criterion, the anti-interference criterion vector b is taken as the space-time steering vector determined by the incoming wave direction and tap delay. For the power inversion criterion, b takes the value [1,0,0,……0]. An intermediate variable w is defined. tmp Solving the equations is the solution method.
[0085] Weight:
[0086]
[0087] intermediate variable w tmp Consider it as equation R xx w tmp =L(L H w tmp The solution to ) = b; define LT = w tmp T is the forward solution, L and L H These are lower triangular matrices and upper triangular matrices, respectively;
[0088] The forward iterative solution process is described by the following formula:
[0089]
[0090] From the above equation, we obtain: Calculate the i-th element t of the forward solution i Only the elements of the lower triangular matrix L up to the i-th column need to be known; the process of solving each forward solution of the equation is as follows: First, subtract the i-th element b of the anti-interference criterion vector by subtracting the update element. i The update is performed, where b1 does not require a subtraction update; then, the i-th element t of the current forward solution is obtained through division. i Then use the lower triangular elements l of all lower triangular matrices after the i-th column and i-th row. ki With the i-th element t of the forward solution i Multiply, where k > i; if i is 1, store the multiplication result directly; otherwise, add the multiplication result to the element stored at the corresponding position before updating; solve each forward solution element in turn according to the above process until all elements are found;
[0091] Figure 5The process of forward iterative solution is demonstrated as follows: First, if the operations on the current column of the lower triangular matrix are completed, the elements of the current column are read in parallel from the storage RAM of the lower triangular matrix. Then, the forward iterative process begins, that is, the elements of the current forward solution are obtained by subtraction and division. The elements of the current forward solution and the subvectors of the lower triangular matrix are used to update the elements of the remaining forward solutions until the forward iteration ends. At this time, a floating-point complex subtraction module, a floating-point complex division module, and 63 floating-point complex multiplication and addition modules are required. During the forward iterative process, each column of the lower triangular matrix is read out and its elements are conjugated in parallel and serially stored into the storage RAM of the upper triangular matrix, that is, the i-th column.
[0092] The lower triangular matrix is serially stored into the i-th RAM after conjugation; after the forward iteration is completed, the backward iteration is performed in the same way using the same amount of hardware resources, and finally the weights are obtained by normalized division; 3) Cross-module parallel computing:
[0093] In step 2)a), the autocorrelation matrix solving module is computed in parallel using a block-based approach. Within each sub-block, multiple columns of the autocorrelation matrix elements are computed in parallel. In step 2)b), the Koleski decomposition module also computes the elements of the lower triangular matrix in a column-parallel manner, and the computation of each column only requires the lower triangular elements of the current column's autocorrelation matrix. In step 2)c), the forward iterative solving module, when solving for the i-th element of the forward solution, also only needs to use the lower triangular elements of the i-th column's lower triangular matrix. Thus, the three modules achieve cross-module parallel computation by marking the completion of each column's computation, without waiting for the previous module to finish its computation. The next module's operation begins; the autocorrelation matrix solving module outputs the column index currently written to RAM to the Kolesky decomposition module. If the column index of the lower triangular matrix to be calculated is less than the column index output by the autocorrelation matrix solving module, the Kolesky decomposition module begins to calculate the lower triangular elements of the current column lower triangular matrix; otherwise, it waits. The Kolesky decomposition module outputs the column index of the lower triangular matrix currently written to RAM to the forward iteration solving module. If the index of the element of the forward solution to be calculated is less than the output column index of the Kolesky decomposition module, the forward equation iteration solving module begins to solve the elements of the current forward solution; otherwise, it waits.
[0094] Under a clock period of 62MHz, this algorithm performs spatiotemporal analysis in a 16-element, 4-tap, 256-cycle configuration.
[0095] When solving for the anti-interference weights, it only takes about 0.6ms to complete the sequential iteration, which has high real-time performance; 4) The backward iterative solution module, such as Figure 6 As shown:
[0096] After the three modules have completed their parallel computations, the final iteration module is started to solve for the final spatiotemporal weight w.
[0097] The backward iterative solution process is described by the following formula:
[0098]
[0099] The coefficient matrix of the backward iteration equation is composed of the upper triangular matrix L H The backward iterative equation requires the M×P-th element w of the intermediate variable. tmpM×P Start by solving in reverse order until the first element w of the intermediate variable is reached. tmp1 And calculate the i-th element w of the intermediate variable. tmpi Only the upper triangular matrix L from the i-th row onwards needs to be known. H The elements of the matrix; therefore, to facilitate backward iteration, after reading the elements of each column of the upper triangular matrix L in parallel during forward iteration, the conjugate of the elements of that column is taken and serially stored into the upper triangular matrix storage RAM; after the forward iteration is completed, backward iteration is performed in the same parallel computing manner; and, for each intermediate variable, the i-th element w is calculated... tmpi That is, the i-th element of the normalization factor Perform a multiplication and accumulation. The conjugate of the i-th element of the anti-interference criterion vector; after the forward and backward equations are solved iteratively, the intermediate variable w is obtained simultaneously. tmp and normalization factor b H w tmp ,
[0100] b H The conjugate transpose of the anti-interference criterion vector is finally applied to the intermediate variable w. tmp Regarding the normalization factor b H w tmp Division normalization, through w = w tmp / b H w tmp The spatiotemporal weight w is obtained.
[0101] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art will understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection of the present invention is defined by the scope of the claims.
Claims
1. A high-speed, space-time anti-interference hardware algorithm with modular parallelism, characterized in that, The space-time anti-interference hardware algorithm includes the following steps: 1) Column-parallel storage and retrieval of matrices: Hardware computation is performed based on field-programmable gate arrays; the complex matrices involved in the computation are stored and read out according to the following rules: all complex matrices are stored in column-parallel using random access memory (RAM), and each column is read out in parallel when reading out the complex matrix; 2) Each parallel computing module: a) Autocorrelation matrix solving module: Satellite communication or satellite navigation receives communication signals and preprocesses the communication signals to obtain a complex sampling matrix; Each data vector is read in parallel from the complex sampling matrix, and then the outer product of the vectors is performed to obtain the autocorrelation matrix R. xx : In the process of vector outer product operation, the autocorrelation matrix is first divided into sub-blocks, and then parallel computation is performed in each sub-block; the reading and storage principle adopts column parallelism. b) Koleski decomposition module: The autocorrelation matrix is decomposed by Koleski decomposition, which decomposes the complex conjugate symmetric autocorrelation matrix into a lower triangular matrix L. Each element in the lower triangular matrix is solved iteratively. In the iterative process of solving the lower triangular matrix, one column of elements of the lower triangular matrix is solved each time, and then the other elements of the lower triangular matrix are updated by using the outer product of the column of elements. The solution and update process is repeated until all elements are iterated. c) Forward iterative solution module: The forward iterative equation LT = b is solved by the forward iterative solution module to obtain the forward solution T, where b is the anti-interference criterion vector; In the forward iterative solution module, each time an element of the forward solution is solved, this element is used to update the other elements of the forward solution. The solution and update process is repeated until all elements of the forward solution are iterated. 3) Cross-module parallel computing: In step 2)a), the autocorrelation matrix solving module is computed in parallel using a block-based approach. Within a single sub-block, the autocorrelation matrix elements of multiple columns are computed in parallel. In step 2)b), the Koleski decomposition module also computes the elements of the lower triangular matrix in a column-parallel manner, and the computation of each column only requires the lower triangular elements of the current column's autocorrelation matrix. In step 2)c), the forward iterative solving module, when solving for an element of a forward solution, also only requires the lower triangular elements of the current column's lower triangular matrix. Thus, the three modules achieve cross-module parallel computation by marking the completion of each column's computation, without waiting for the previous module to complete its computation before starting the computation of the next module. 4) Backward Iterative Solution Module: After the three modules have completed their parallel computations, the backward iterative equation L is solved by the backward iterative solution module. H w tmp =T obtains the intermediate variable w tmp L H It is an upper triangular matrix; In the process of solving the backward iterative equation in the backward iterative solution module, each time an element of the intermediate variable is solved, this element is used to update the other elements of the intermediate variable. The solution and update process is repeated until all elements are iterated. Then, according to the weighting formula, the intermediate variable w tmp Normalization yields the following spatiotemporal weights w: w=w tmp / b H w tmp Among them, b H This is the conjugate transpose of the anti-interference criterion vector.
2. The space-time anti-interference hardware algorithm as described in claim 1, characterized in that, In step 1), each memory stores one row of data for the complex matrix, and each complex element is stored using the same storage space in RAM, with the low-order bits representing the imaginary part and the high-order bits representing the real part. Furthermore, each column is read out in parallel when reading out complex matrices.
3. The space-time anti-interference hardware algorithm as described in claim 1, characterized in that, In step 2)a), the communication signal is preprocessed to obtain a complex sampling matrix, including the following steps: digital quadrature downconversion to obtain a complex signal, which is then sampled and stored in parallel to obtain a complex sampling matrix, and stored in the complex sampling matrix storage RAM.
4. The space-time anti-interference hardware algorithm as described in claim 1, characterized in that, In step 2)a), the lower triangular elements of the autocorrelation matrix are decomposed into K sub-blocks, and each sub-block element in the K sub-blocks is computed in parallel. In the process of vector outer product operation, the autocorrelation matrix is first divided into sub-blocks, where K is a natural number.
5. The space-time anti-interference hardware algorithm as described in claim 4, characterized in that, The specific partitioning method of the autocorrelation matrix is as follows: sort the lower triangular elements of the autocorrelation matrix in column-first traversal and stretch them into a row vector. Then, divide the row vector into K sub-blocks on an equal basis. The number of elements contained in each sub-block is C = (M×P×(M×P+1)) / (2×K). Then, perform parallel operations in each sub-block and obtain the elements in a sub-block by traversing all N sampling snapshots. Finally, the column-parallel fixed-to-float module built into the field-programmable gate array is called to perform column-parallel fixed-to-float operation on the calculated sub-block elements; M is the number of array elements, P is the number of taps, and N is the number of steps.
6. The space-time anti-interference hardware algorithm as described in claim 5, characterized in that, The principle of parallel reading and storage: When calculating each sub-block element, all sampling snapshots in the complex sampling matrix must be read in parallel. The reading operation and complex multiplication and accumulation operation are performed in parallel according to the pipeline to improve the calculation speed of each sub-block element. When the sampling snapshot is traversed, the autocorrelation matrix of the k-th sub-block is also calculated at the same time. Then, the calculated C elements are written into the first autocorrelation matrix storage RAM in the form of a two-dimensional matrix storage. After writing is complete, increment the number of sub-blocks by 1 and start parallel computation of the next sub-block until all sub-blocks have been calculated and stored.
7. The space-time anti-interference hardware algorithm as described in claim 1, characterized in that, In step 2)b), if the current calculation column is the first column, there is no need to update the autocorrelation matrix elements; otherwise, after each autocorrelation matrix element is read in, the autocorrelation matrix elements are first updated by parallel column subtraction, and then the diagonal elements are obtained by square root operation in a separate calculation cycle, and the off-diagonal elements are obtained by parallel division operation. This calculation order makes it possible to calculate only the elements of the autocorrelation matrix in the same column when calculating the lower triangular matrix of each column. After the calculation of a column of elements is completed, the elements of the column are written back to the lower triangular matrix storage RAM. After the calculation of a column is completed, the column elements are multiplied by the outer product and the previous update result is accumulated to obtain the current update result.
8. The space-time anti-interference hardware algorithm as described in claim 1, characterized in that, Step 3) specifically includes the following steps: The autocorrelation matrix solving module outputs the column index of the currently written autocorrelation matrix storage RAM to the Kolesky decomposition module. If the column index of the currently to be calculated lower triangular matrix is less than the column index output by the autocorrelation matrix solving module, the Kolesky decomposition module starts to calculate the lower triangular elements of the current column lower triangular matrix; otherwise, it waits. The Kolesky decomposition module outputs the column index of the currently written autocorrelation matrix storage RAM to the forward iteration solving module. If the index of the element of the currently to be calculated forward solution is less than the output column index of the Kolesky decomposition module, the forward equation iteration solving module starts to solve the elements of the current forward solution; otherwise, it waits.
9. The space-time anti-interference hardware algorithm as described in claim 1, characterized in that, In step 4), normalization is achieved through normalized inner product and normalized division.