Variable-order covariance matrix inversion method

The FPGA-based Cholesky decomposition method optimizes covariance matrix inversion in adaptive digital beamforming by parallelizing operations, reducing resource usage and computation time, addressing the inefficiencies of existing methods.

CN120296292APending Publication Date: 2025-07-11NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410039650.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When implementing covariance matrix inversion on FPGA, the existing technology has the problem of large amount of computing, high resource consumption and difficult to implement. Especially when the real-time requirements of array antenna adaptive anti-interference are high, how to shorten the computing time and save hardware resources is a challenge.

Method used

The covariance matrix is decomposed into the lower triangle matrix by using the Cholesky decomposition method, and calculated by the state machine into three steps. Parameterized design and time-sharing multiplexing technology are used to reduce hardware resource consumption and improve computing speed.

Benefits of technology

It realizes efficient calculation of covariance matrix inverse on FPGA, meets the real-time requirements of adaptive anti-interference in array antennas, and saves hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296292A_ABST
    Figure CN120296292A_ABST
Patent Text Reader

Abstract

The invention discloses a variable-order covariance matrix inversion method, and belongs to the technical field of array signal processing. Aiming at the inversion problem of the covariance matrix, according to the conjugate symmetry property of the covariance matrix, the covariance matrix inversion operation based on the FPGA is realized by utilizing a cholesky matrix decomposition method. The FPGA implementation architecture comprises a state control module, a state machine module, a storage module and a calculation module. According to the method, in a parameterization design mode, a user can customize the order number and the data bit width of an inversion matrix, the use of logic resources of an FPGA is saved in a time division multiplexing mode, and the operation speed is increased in an address index pipeline processing mode. According to the design, the self-definition of a user, the utilization rate of resources and the calculation precision and time are integrated, good compatibility is achieved, and the universal requirement for covariance matrix inversion is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of array signal processing, and specifically relates to a hardware structure for inverse covariance matrix based on FPGA. In array signal processing, the inverse of a conjugate symmetric covariance matrix is often required. The hardware structure proposed by the present invention uses a conjugate symmetric covariance matrix of any order and has universality. Background Art

[0002] Currently, the commonly used method for anti-interference of array antennas is the adaptive digital beamforming algorithm. The key to implementing the adaptive digital beamforming algorithm is to calculate the inverse matrix of the covariance matrix. General matrix inversion methods such as the elementary transformation method, the adjoint matrix method, and the Gaussian elimination method require a large number of determinant calculations. Their disadvantages of large computational amount, large storage space requirement, and high resource consumption make them not suitable for FPGA. The matrix decomposition method has a small amount of operations for matrix inversion and is easy to implement. Therefore, the matrix decomposition method is generally selected in FPGA to implement matrix inversion.

[0003] The matrix decomposition method decomposes a matrix into the product of some simple matrices, such as triangular matrices or unitary matrices, which all have certain characteristics and are relatively easy to obtain the inverse matrix. Commonly used matrix decomposition methods include: singular value decomposition, LU decomposition, QR decomposition, and Cholesky decomposition. Among them, the first three decomposition methods are used to calculate the inverse of all invertible complex matrices. The algorithms are relatively complex and difficult to implement, and more resources are used. The Cholesky decomposition is applicable to conjugate symmetric positive definite matrices. In the adaptive digital beamforming algorithm, the matrices to be inverted are all covariance matrices after autocorrelation and have conjugate symmetric characteristics. Therefore, the present invention selects the Cholesky decomposition method to solve the inverse matrix.

[0004] In the process of adaptive anti-interference of array antennas, there is a relatively high requirement for real-time performance. Therefore, how to shorten the calculation time, save hardware resources, and make the system modification more convenient using parametric design for matrix inversion on FPGA is a problem to be optimized. Summary of the Invention

[0005] Provide a positive definite matrix inversion structure on the FPGA chip that is easy to configure and modify, has low hardware complexity, and high resource utilization rate. While minimizing the consumption of hardware resources, ensure the operation speed by increasing the parallelism of matrix inversion operations.

[0006] An FPGA-based method for inverse covariance matrix includes six modules: a state control module, a matrix storage module, a multiplication and addition calculation and division calculation module, an LC matrix decomposition module, an inverse matrix C module, and an inverse matrix A module.

[0007] The state control module receives the order, data bit width, and system operating clock of the external input matrix, and controls the state changes, calculation, and storage addresses and timings of the other modules;

[0008] The matrix storage module stores the input covariance matrix A to be inverted, intermediate calculation matrices C, L, Y, and the inverse matrix W of the final matrix A;

[0009] The multiplication-addition calculation and division calculation module completes the calculation of the data input from the storage module according to the calculation instructions and enables of the state control module, and outputs the result to the matrix storage module for storage again.

[0010] The LC matrix decomposition module, matrix C inversion module, and matrix A inversion module are divided into three state machines by module. They are controlled by the state control module to change the state of the state machine and complete the commands of storage or calculation according to its instructions;

[0011] A method for inverting a covariance matrix based on FPGA, and its working steps are as follows:

[0012] Step 1: Pull up the enable for receiving external signals, perform system configuration according to the matrix order, data bit width, and system operating clock input externally, and store the input matrix A to be inverted row by row in the memory of the matrix storage module.

[0013] Step 2: Enter the LC matrix decomposition state machine, generate matrix index addresses, and according to the current calculated element index values (i, j) (i, j = 0, 1,... N-1), take out the i-th row from the RAM storing matrix C and the j-th row from the RAM storing matrix L; send the retrieved data and the current state to the multiplication-addition operation module; wait for the multiplication-addition operation module to output the operation result. Calculate c ij and store the result in the RAM; according to the current calculated element index values (i, j) (i, j = 0, 1,... N-1), send c ij and c jj to the division operation module; wait for the division operation to output the operation result; calculate l ij and store the result in the RAM; refresh the index values of the next round of calculated elements according to the element calculation order, and repeat until all elements are calculated.

[0014] Step 3: Enter the matrix C inversion state machine, and according to the current calculated element index values (i, j) (i, j = 0, 1,... N-1), take out the i-th row from the RAM storing matrix C and the j-th column from the RAM storing matrix Y. Send the retrieved data and the current state to the multiplication-addition operation module. Wait for the multiplication-addition operation module to output the operation result. According to the current calculated element index values (i, j) (i, j = 0, 1,... N-1), send the multiplication-addition operation result and c iiSend it to the division operation module. Wait for the division operation to output the operation result. Calculate y ij And store the result in the RAM. Refresh the index value of the next round of calculation elements according to the element calculation order, and repeat until all elements are calculated.

[0015] Step 4: Enter the inverse state machine of matrix A. According to the index values (i, j) (i, j = 0, 1,... N - 1) of the currently calculated elements, take out the i-th column from the RAM storing matrix L and the j-th column from the RAM storing matrix W. Send the retrieved data and the current state to the multiply-accumulate operation module. Wait for the multiply-accumulate operation module to output the operation result. Calculate w ij And store the result in the RAM. Refresh the index value of the next round of calculation elements according to the element calculation order, and repeat until all elements are calculated.

[0016] The beneficial effects of the present invention are as follows: In the implementation of the adaptive digital beamforming algorithm based on FPGA, it is necessary to invert the covariance matrix after autocorrelation. The general matrix inversion method structure will consume additional resources. The present invention proposes an inversion method and framework for any-order conjugate symmetric matrix. Through parametric design, the custom of the matrix order is realized. Through time-division multiplexing technology, the hardware resources of FPGA are saved. Through the element address index, the operation state pipeline is reduced, and the operation time is reduced, meeting the general requirements such as calculation accuracy and calculation time for covariance matrix inversion. Description of the Drawings

[0017] Figure 1 It is the flow chart of the FPGA design for covariance matrix inversion.

[0018] Figure 2 It is the flow chart of the covariance matrix inversion method.

[0019] Figure 3 It is the design block diagram of the covariance matrix inversion hardware module.

[0020] Figure 4 It is the block diagram of the calculation module structure.

[0021] Figure 5 It is the sequence diagram of the element address index of the 8-order matrix. Detailed Embodiment

[0022] The following details the specific embodiments of the embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0023] In the field of array signal processing, especially in the case of adaptive digital beamforming algorithms, the matrix to be inverted is often the covariance matrix after autocorrelation, which has the property of conjugate symmetry. Utilizing this property, in the FPGA implementation, the inverse matrix can be obtained by the Cholesky matrix decomposition method, and only the lower triangular matrix after decomposition needs to be calculated, greatly reducing the operation resources and time.

[0024] According to the Cholesky decomposition theorem, a symmetric positive definite matrix A can be represented as the product of a lower triangular matrix L and its conjugate transpose L H . By extracting the diagonal matrix D, we can get A = L H DL. It is expressed as

[0025]

[0026] To calculate the inverse matrix of A, first perform the Cholesky decomposition operation on A. Let C = LD, then the decomposition result can be expressed as A = CL H . According to the basic matrix multiplication operation, we can obtain

[0027]

[0028] Corresponding to c ij and l ij , for i, j = 1, 2,..., n, the calculation formula is

[0029]

[0030] After obtaining the lower triangular matrices C and L, calculate the inverse matrix Y of C. The calculation formula is

[0031]

[0032] Let the inverse matrix of A be W. According to AW = CY = I, the calculation formula of W can be obtained as

[0033]

[0034] The present invention designs the FPGA structure for the above formulas (3), (4), and (5). We divide the matrix inversion into three steps, each step is implemented by a state machine, and in each step, our division operation and multiply-accumulate operation modules are instantiated to complete the function of matrix inversion on the FPGA. Figure 1 It is the FPGA flow block diagram for the inverse of the covariance matrix.

[0035] The matrix storage module stores the input covariance matrix A to be inverted, intermediate calculation matrices C, L, Y, and the inverse matrix W of the final matrix A respectively. During the operation of the three state machines, the elements of matrices A, C, L, Y, and W need to be cross - used. Therefore, the RAM for storing the relevant matrices is placed at the top layer of the matrix inversion module. Each state machine fetches the elements of the corresponding matrix for operation when calculating the corresponding steps, realizing the time - sharing multiplexing of storage resources. The initial memory RAM_A of the matrix is used to store the input N - order conjugate - symmetric matrix A, and at the same time, the corresponding matrix order and data bit - width are input. After the matrix A data is stored in RAM_A, the state control module starts to control the change of the state process, performs the matrix decomposition shown in formula (3), and calculates the L matrix row - by - row after column - by - column.

[0036] The calculation module includes a multiply - add calculation and a division calculation module, and the structure diagram is as Figure 4 shown. The calculation module completes the calculation of the data input from the storage module according to the calculation instructions and enables of the state control module, and outputs the result to the matrix storage module for storage again. Due to the large dynamic range of the covariance matrix, the data bit - width is large during the calculation process. The designed bit - width of this system is 32 or 64 bits. There are a large number of multiply - add and division operations during the operation of the three state machines. To save FPGA resources, the basic multiply - add operation and division operation are abstracted into independent modules, and the state machine sends the data to be calculated to the corresponding module in a time - sharing manner during the operation.

[0037] For the generality of the basic multiply - add operation, the multiply - add operation is uniformly performed on N pairs of serially - input data. However, the range of multiply - add for calculating elements in different positions in different steps is inconsistent. Therefore, the module needs to generate corresponding masks during summation according to different states, and only sum the data within the range.

[0038] The division calculation module realizes the basic division operation. It should be noted that although complex division needs to be implemented, the divisors in the operation process are all the diagonal elements of the matrix, and the imaginary part of these elements is theoretically 0. Therefore, the basic division module of this design only needs to use two dividers to implement.

[0039] The state control module controls the state change, address indexing, matrix element calculation, element call, and storage enable of the three state machine modules.

[0040] The three state machine modules for matrix decomposition and inversion have the following functions:

[0041] State machine 1: Perform Cholesky decomposition on the input covariance matrix A (A = LDL H = CL H ), to obtain the lower triangular matrices C and L.

[0042] State machine 2: Calculate the inverse matrix Y of matrix C, Y = C-1 。

[0043] State machine 3: From AW = CL H W = E = CY to obtain L H W = Y, and then calculate the inverse matrix W = A of covariance matrix A -1 。

[0044] Since A is input serially by row, input control logic is required to parse the valid data input and write it into the corresponding RAM. Similarly, when outputting, in order to output serially by row, output logic is also required to read the RAM storing the inverse matrix according to the protocol and complete the data output.

[0045] The specific FPGA implementation process of the three state machine modules is as follows:

[0046] LC matrix decomposition module

[0047] To implement the matrix decomposition of Equation (3), we divide the state machine of this LC matrix decomposition module into the following states: state0, the initial state; statel, the state of generating the multiply-add calculation address; state2, the state of calculating c ij state; state3, the state of saving c jj state; state4, the state of calculating l ij state; state5, the state of saving l ij state; state6, the state of updating the address. The FPGA implementation process is as follows:

[0048] (1) According to the element index values (i, j) (i, j = 0, 1,... N-1) calculated currently, take the i-th row from the RAM storing matrix C and take the j-th column from the RAM storing matrix Y. Send the retrieved data and the current state to the multiply-add calculation module. Wait for the result of the multiply-add to output as c ij , and write c ij into the corresponding RAM. The read and write addresses of c ij and l ij are generated by the address generation state machine inside the cholesky decomposition module.

[0049] (2) According to the cholesky decomposition formula, send the result c ij output by the multiply-add module and c jj to the division calculation module, and then write the division output result l ij into the corresponding RAM. The write address of l ij is calculated by the state machine.

[0050] (3) Refresh the index values of the elements to be calculated in the next round according to the element calculation order. The element refresh order is as Figure 5As shown in the calculation order of elements C and L, that is, when i = j = N - 1, the state machine completes the calculation of all elements.

[0051] Matrix C Inversion Module

[0052] To implement the inverse matrix Y of C in Equation (4), we divide the state machine of the matrix C inversion module into the following states: state0, the initial state; statel, the state of generating the multiplication and addition calculation address; state2, the state of calculating the multiplication and addition result; state3, the state of calculating y ij state; state4, storing y ij state; state5, the state of updating the address. The FPGA implementation process is as follows:

[0053] (1) According to the element index values (i, j) (i, j = 0, 1,... N - 1) being calculated currently, the i-th row is fetched from the RAM storing matrix C, and the j-th column is fetched from the RAM storing matrix Y. The fetched data and the current state are sent to the multiplication and addition calculation module, waiting for the calculation result to be output.

[0054] (2) From Equation (4), the result output by the multiplication and addition calculation module is sent to the division calculation module together with c jj and waiting for the basic division operation to output the operation result, and y ij is stored in the corresponding RAM.

[0055] (3) Refresh the index values of the elements to be calculated in the next round according to the element calculation order. The element refresh order is as Figure 5 shown in the calculation order of elements Y, that is, when i = N - 1 and j = 0, the state machine completes the calculation of all elements.

[0056] Matrix A Inversion Module

[0057] To implement the inverse matrix W of A in Equation (4), we divide the state machine of the matrix A inversion module into the following states: state0, the initial state; state1, the state of generating the multiplication and addition calculation address; state2, the state of calculating the multiplication and addition result; state3, the state of calculating w ij state; state4, storing w ij state; state5, the state of updating the address. The FPGA implementation process is as follows:

[0058] (1) According to the element index values (i, j) (i, j = 0, 1,... N - 1) being calculated currently, the i-th column is fetched from the RAM storing matrix L, and the j-th column is fetched from the RAM storing matrix W. The fetched data and the current state are sent to the basic multiplication and addition operation module, waiting for the calculation result to be output.

[0059] (2) From equation (4), the basic multiplication and addition operation module outputs the operation result. Calculate w ij , and w ij The results are stored in RAM.

[0060] (3) Refresh the index value of the next round of calculation elements according to the element calculation order. The element refresh order is as follows: Figure 5 As shown in the calculation order of W elements, when i=j=0, the state machine completes the calculation of all elements.

[0061] Experimental verification

[0062] The present invention further verifies the variable-order covariance matrix inversion technology through the following experiments.

[0063] The hardware platform of this design is Xilinx's XC7VX690T chip with a main frequency of 250MHz; the algorithm programming software is Vivado 2018.3; and the auxiliary verification software is Matlab.

[0064] Experimental steps: Matlab simulates a set of 8-channel training data, and after autocorrelation calculation, generates an 8th-order covariance matrix. The matrix is ​​input into the matrix inversion module of FPGA and the order is configured. The simulation is performed through the simulation software provided by Vivado. The output of each node of FPGA is compared with the result of each node of Matlab matrix inversion, and the error is analyzed.

[0065] The results show that after normalization comparison with MATLAB calculation results, the FPGA calculation error is within 10 -5 At a frequency of 250M, the 8th-order covariance matrix inversion takes 38.21us. Table 1 shows the FPGA logic resource usage for the 8th-order covariance matrix inversion.

[0066] Table 18-order covariance matrix inversion logic resource occupancy table

[0067]

[0068] The above-mentioned embodiments only express the preferred implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A method for inverting a covariance matrix with variable order, characterized in that, It includes six modules: a status control module, a matrix storage module, a multiplication and addition calculation and division calculation module, an LC matrix decomposition module, an inverse matrix C module, and an inverse matrix A module. The status control module controls the state changes of the other modules and controls the address generation and refreshing of the calculation and storage of the other modules. The matrix storage module stores the input covariance matrix A to be inverted, the intermediate calculation matrices C, L, Y, and the inverse matrix W of the final matrix A. The multiplication and addition calculation and division calculation module completes the calculation of the data input from the storage module according to the calculation instructions and enables of the status control module, and outputs the result to the matrix storage module for storage again. The LC matrix decomposition module, the inverse matrix C module, and the inverse matrix A module are divided into three state machines by module. They are controlled by the status control module to change the state of the state machine and complete address generation, data calculation, and data storage according to its instructions. The method for inverting the covariance matrix specifically includes the following steps: Step 1: Receive configuration information: matrix order N, data bit width P, operating clock; arrange in parallel by row, and receive all data of the covariance matrix A to be inverted. When the number of columns is N, the matrix data reception is completed. Step 2: After the data reception is completed, the status control module controls the LC matrix decomposition module to start working. This module fetches the data from the storage unit according to the element index values (i, j) (i, j = 0, 1,... N-1) in the address order, and sends the data to the multiply-add calculation module and the division calculation module to calculate c ij and l ij . After the calculation is completed, the data results are sequentially stored in the RAM at the corresponding matrix element addresses. Repeat this step until i = j = N-1, at which point the LC matrix decomposition calculation is completed. Step 3: After completing the LC matrix decomposition, the state control module controls the matrix C inverse module to start working. This module retrieves data from the storage unit according to the element index values (i, j) (i, j = 0, 1,... N - 1) in the address order and sends the data to the multiply-add calculation module and the division calculation module to calculate y ij , after completing the calculation, the data results are sequentially stored in the RAM at the corresponding matrix element addresses, and this step is repeated until the inverse calculation of matrix C is completed when i = N - 1 and j = 0. Step 4: After completing the inversion of matrix C, the state control module controls the matrix A inversion module to start working. This module retrieves data from the storage unit according to the element index values (i, j) (i, j = 0, 1,... N - 1) in the address order and sends the data to the multiply-accumulate calculation module to calculate w ij , after the calculation is completed, the data results are sequentially stored in the RAM corresponding to the matrix elements. Repeat this step until i = j = N - 1, at which point the inversion calculation of matrix A is completed. Step 5: When the inverse matrix W of A is calculated, the status control module controls the data of W stored in the RAM to be output by row, and the covariance matrix inversion function is completed and output.

2. A method for inverting a covariance matrix with variable order, characterized in that, The matrix inversion method and framework proposed by the present invention realize the customization of the matrix order through parametric design, save the hardware resources of the FPGA through time-division multiplexing technology, reduce the operation time by making the operation status flow through element address indexing, and meet the general requirements such as the calculation accuracy and calculation time of covariance matrix inversion.

3. The LC matrix decomposition module according to claim 1, wherein the eigenvalue is that the input covariance matrix A is subjected to Cholesky decomposition (A = LDL H = CL H ), to obtain the lower triangular matrices C and L. The specific steps are as follows: Step 2.1: According to the currently calculated element index values (i, j) (i, j = 0, 1,... N - 1), fetch the i-th row from the RAM storing matrix C and the j-th column from the RAM storing matrix Y. Send the fetched data and the current state to the multiply-accumulate calculation module. Wait for the result of the multiply-accumulate operation to be output as c ij , and write c ij into the corresponding RAM. The addresses for reading and writing c ij and l ij are generated by the address generation state machine inside the Cholesky decomposition module. Step 2.2: According to the Cholesky decomposition formula, the result c output by the multiplication and addition module ij and c jj are sent to the division calculation module, and then the division output result l ij is written into the corresponding RAM. The write address of l ij is calculated and generated by the state machine. Step 2.3: Refresh the index value of the next round of calculated elements according to the element calculation order. The element refresh order is shown in the C and L element calculation orders in Figure 5, that is, when i = j = N - 1, the state machine completes the calculation of all elements.

4. The inverse matrix calculating module of matrix C according to claim 1, wherein the eigenvalue is to calculate the inverse matrix Y = C of matrix C -1 , and the specific steps are as follows: Step 3.1: According to the index value (i, j) (i, j = 0, 1,... N - 1) of the currently calculated element, take out the i-th row from the RAM storing matrix C and the j-th column from the RAM storing matrix Y. Send the taken data and the current state to the multiplication and addition calculation module and wait for the calculation result to be output. Step 3.2: According to the Cholesky decomposition formula, the result output by the multiplication and addition calculation module is combined with c jj and sent to the division calculation module to wait for the basic division operation to output the operation result, and y ij is stored in the corresponding RAM. Step 3.3: Refresh the index value of the next round of calculated elements according to the element calculation order. The element refresh order is shown in the Y element calculation order in Figure 5, that is, when i = N - 1 and j = 0, the state machine completes the calculation of all elements.

5. The inverse matrix module of matrix A according to claim 1, wherein the eigenvalue lies in AW = CL H From W = E = CY, L can be obtained H W = Y, and then the inverse matrix W = A of the covariance matrix A is calculated -1 , and the specific steps are as follows: Step 4.1: According to the index value (i, j) (i, j = 0, 1,... N - 1) of the currently calculated element, take out the i-th column from the RAM storing matrix L and the j-th column from the RAM storing matrix W. Send the taken data and the current state to the basic multiplication and addition operation module and wait for the calculation result to be output. Step 4.2: From Equation (4), wait for the basic multiply-accumulate operation module to output the operation result. Calculate w ij , and store the w ij result in the RAM. Step 4.3: Refresh the index value of the next round of calculated elements according to the element calculation order. The element refresh order is shown in the W element calculation order in Figure 5, that is, when i = j = 0, the state machine completes the calculation of all elements.