Calculation structure special for Kalman filtering for radar-guided target tracking

By designing a dedicated computing structure for Kalman filtering, using a modular parallel computing architecture and efficient data transmission mechanism, the problem of insufficient computing performance bottlenecks and real-time performance of the Kalman filtering algorithm in high-dimensional state space is solved, and efficient and real-time target tracking effect is achieved.

CN120216441APending Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277986.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems such as bottlenecks in the Kalman filtering algorithm, insufficient architecture optimization and insufficient real-time performance. Especially in the matrix inversion operation in the high-dimensional state space, the real-time and efficient algorithms are limited.

Method used

A dedicated Kalman filtering calculation structure for radar-guided target tracking is designed, including a state prediction module, a Kalman gain calculation module, a state update module, a covariance update module and a delay cache module. Through a modular parallel computing architecture and an efficient data transmission mechanism, matrix computing and data transmission are optimized to improve the calculation performance and real-timeness of the algorithm.

Benefits of technology

It significantly improves the computing performance and real-time performance of the Kalman filtering algorithm, can effectively track targets in high-precision and high-dynamic scenarios, meet real-time output requirements, and provides efficient, flexible and reliable hardware solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216441A_ABST
    Figure CN120216441A_ABST
Patent Text Reader

Abstract

The invention discloses a calculation structure special for Kalman filtering for radar guide target tracking, which comprises a state prediction module, a Kalman gain calculation module, a state updating module, a covariance updating module and a delay caching module which work cooperatively under the unified scheduling of a control unit designed based on a finite-state machine. Completing the calculation task of each stage of Kalman filtering; the radar collects the position information of the target at the current moment and stores the position information in a storage unit, the position information of the target forms a measurement vector of the moment, and the calculation structure is utilized to obtain an optimal prediction state vector containing target prediction information; according to the invention, an efficient matrix inversion unit and a configurable systolic array matrix multiplication unit are designed, and based on a parallel computing architecture and an optimized data transmission mechanism, the flexibility and expansibility of the system can be effectively enhanced, and the collaboration and stability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radar-guided target tracking, and particularly to a dedicated computing structure for Kalman filtering for radar-guided target tracking. Background Art

[0002] In the field of target tracking, visible light sensors are widely used in target detection and tracking due to their high image resolution and intuitive information acquisition ability. However, the field of view of visible light sensors is limited, making it difficult to efficiently capture and continuously track targets over a large range, resulting in poor actual application effects. Therefore, current technologies often use radar to guide visible light sensors for target tracking. Radar has high-precision ranging and large-range scanning capabilities, and can provide accurate target position estimates for visible light sensors, thereby achieving more effective target capture and tracking. The flowchart of the guiding process is shown in the appendix Figure 1 。

[0003] In this fusion scheme, the Kalman filtering algorithm is widely used for target state estimation and trajectory prediction. Especially when introducing the uniformly accelerated motion model, Kalman filtering can significantly improve the target tracking accuracy. However, the introduction of the uniformly accelerated motion model leads to an increase in the matrix dimension during the Kalman filtering process, especially the matrix inversion operation (HP k when calculating the Kalman gain K k,k-1 H T +R k ) -1 becomes the computational bottleneck, greatly limiting the real-time performance and efficiency of the algorithm. Specifically, the calculation formula for the Kalman gain is: K k =P k,k-1 H T (HP k,k-1 H T +R k ) -1 , where the matrix inversion operation (HP k,k-1 H T +R k ) -1 usually has a computational complexity of O(n 3 ) (n is the dimension of the matrix). In a twelve-dimensional state space, the computational amount of this operation is extremely large.

[0004] Currently, the hardware implementation for accelerating Kalman filtering is mainly based on reconfigurable hardware platforms such as FPGAs, but the existing technologies still have the following deficiencies:

[0005] 1. Limitations in optimizing general matrix operations:

[0006] Existing dedicated hardware computing architectures typically use LU decomposition or Cholesky decomposition to accelerate matrix inversion. These methods have high computational efficiency in general scenarios, but are not fully applicable in Kalman filter applications. Matrix inversion in Kalman filtering involves covariance matrices, which usually have specific symmetric positive definite structures and data correlations. General inversion algorithms incur additional computational overhead and resource waste when dealing with these characteristics.

[0007] 2. Insufficient utilization of parallelism:

[0008] Existing dedicated hardware computing architectures are designed without fully exploiting the parallelism of operations in matrix operations (such as multiplication, addition, and inversion). For example, mainstream solutions are mostly based on row-column block partitioning strategies, failing to fully utilize the sparsity or symmetry characteristics of matrices, resulting in unnecessary computational overhead. In addition, some hardware architectures adopt pipeline designs, but the depth and granularity of the pipelines are often insufficient. There are dependencies among multiple pipeline stages during matrix updates, leading to a decrease in throughput.

[0009] 3. Insufficient real-time performance:

[0010] In scenarios with high-precision requirements (such as target tracking under a uniformly accelerated model), Kalman filtering has extremely high real-time requirements. The filtering process needs to complete state prediction and update at each sensor data update and work in coordination with an external control system to ensure real-time output. Existing hardware architectures are difficult to meet the real-time requirements due to high computational latency and large data transmission overhead. For example, the execution time for a single filtering update of a 12×12 matrix may reach dozens of microseconds, unable to meet the real-time requirement of 1000 state updates per second.

[0011] In summary, existing technologies have problems such as prominent computational bottlenecks, unoptimized architectures, and insufficient real-time performance in accelerating target tracking based on Kalman filtering, and there is an urgent need for an efficient and dedicated hardware acceleration solution. Summary of the Invention

[0012] The purpose of the present invention is to provide a dedicated computing architecture for Kalman filtering in radar-guided target tracking, which solves the problems of computational performance bottlenecks, insufficient architecture optimization, and low real-time performance existing in Kalman filtering in radar-guided target tracking systems.

[0013] To achieve the above tasks, the present invention adopts the following technical solutions:

[0014] A dedicated computing architecture for Kalman filtering in radar-guided target tracking includes: a state prediction module, a Kalman gain calculation module, a state update module, a covariance update module, and a delay cache module. These modules work together under the unified scheduling of a control unit designed based on a finite state machine to complete the computational tasks of each stage of Kalman filtering;

[0015] The radar acquires the position information of the target at the current k-th moment and stores it in the storage unit, where the position information of the target constitutes the measurement vector Z at the k-th moment. k , and enters the dedicated computing structure for processing; where:

[0016] The state update module is used to receive the measurement vector Z at the k-th moment from the storage unit. k , the state vector X at the k-th moment predicted by using the state at the (k - 1)-th moment. k,k-1 and the Kalman gain matrix K provided by the Kalman gain calculation module. k , the measurement matrix H for predictive calculation, and outputs the optimal predicted state vector X of the target at the k-th moment. k,k ;

[0017] The Kalman gain calculation module includes a Kalman gain calculation A unit, a Kalman gain calculation B unit, and a Kalman gain multiplication unit.

[0018] The Kalman gain calculation A unit is a matrix inversion unit, which combines the process noise matrix Q from the storage unit. k , and calculates the gain K. A K = (HP k,k-1 H T + R k ) -1 ; R k represents the measurement noise matrix, P k,k-1 represents the covariance matrix of the state vector X. k,k-1 The superscript T represents the transpose of the matrix.

[0019] The Kalman gain calculation B unit is used to calculate K. B K = P k,k-1 H T ;

[0020] The Kalman gain multiplication unit is used to multiply the outputs of the Kalman gain calculation A unit and the Kalman gain calculation B unit, K A *K B , to obtain the Kalman gain K. k ;

[0021] The state prediction module is used to receive the optimal predicted state vector X from the state update module. k,k and the state transition matrix F from the storage unit, and completes the update of the state vector according to the following formula to predict the state vector X at the (k + 1)-th moment. k+1,k :

[0022] The covariance update module is used to complete the update of the covariance matrix P of the optimal predicted state vector X at the k-th moment. k,k of k,k ;

[0023] The delay cache module is used to temporarily store the intermediate calculation result X k,k and P k,k and provide inputs for the next iteration;

[0024] The optimal predicted state vector X at the target k moment obtained after being processed by the dedicated computing structure k,k is output to the visible light sensor for guiding the detection of the target in combination with the image data collected by the visible light sensor; wherein the optimal predicted state vector includes the position information and velocity information of the target at the k moment.

[0025] Furthermore, the dedicated computing structure further includes:

[0026] A configurable systolic array matrix multiplication unit for performing multiplication operations between matrices of different dimensions in the Kalman filtering process; wherein: the configurable systolic array matrix multiplication unit specifically includes:

[0027] An input register for receiving and temporarily storing input matrix data; a reconfigurable decoder for dynamically configuring the array computing scale and enabling different numbers of processing elements PE to adapt to different matrix dimensions; a control signal for coordinating the operations of the various components of the configurable systolic array matrix multiplication unit; a delay counter for controlling the transfer order of data between the processing elements PE to ensure synchronous transmission of data between the processing elements PE; a processing element PE for performing multiplication and accumulation operations; an output channel for outputting the calculation result to the target module or storage device.

[0028] Furthermore, for the input matrix C and input matrix D to be multiplied in the input register, the data of the input matrix C is loaded row by row into the row register and gradually input into the processing elements PE of each column through the delay counter; the data of the input matrix D is loaded column by column into the top register and synchronously input into the processing elements of each row; the number of processing elements PE is configured by the reconfigurable decoder, and under the coordination of the control signal, each processing element PE receives data from the left or above, and after completing the multiplication operation, transfers the intermediate result to the right or below, gradually completing the multiplication operation of the input matrix C and input matrix D; the final result is collected through the output channel to form a complete result matrix.

[0029] Further, the processing unit PE specifically includes: a PE input channel for receiving data from an adjacent processing unit PE or from the outside; a multiplier IP core for performing a multiplication operation on the input data to generate partial products; a register for storing the partial products for subsequent operations; an adder IP core for adding the partial products to an accumulated result to generate a new accumulated result; a PE output channel for passing the accumulated result to an adjacent processing unit PE; and a PE control unit designed based on a finite state machine for scheduling the operation process of the processing unit PE.

[0030] Further, the communication mechanism of the dedicated computing structure includes an on-chip bus OMBus and an inter-module bus MIBus; wherein the on-chip bus OMBus is used for data transmission between sub-units within each module of the dedicated computing structure, and the inter-module bus MIBus is responsible for multi-channel data interaction and cooperative operation between modules of the dedicated computing structure, and uses the AXI bus protocol as the core communication protocol.

[0031] Further, the Kalman gain K k , the measurement noise matrix R k and the covariance matrix P k,k-1 of the state vector X k-1,k-1 are transmitted through the inter-module bus and the intra-module bus, and a configurable systolic array multiplication unit is used to complete matrix multiplication calculations. The calculation results are stored in a FIFO queue, and data scheduling is performed through the FIFO to complete matrix addition to calculate the covariance matrix P k,k ; the configurable systolic array unit can dynamically adjust the operation scale to adapt to the calculation requirements of matrices with different dimensions.

[0032] Further, the Kalman gain calculation A unit adopts the following architecture:

[0033] A calculation execution unit: divided into multiple parallel sub-units for processing element-level blocks of the matrix input to this unit; a double-precision floating-point operation module: supporting multiplication, subtraction, division, and negation operations; an intermediate result register: for storing intermediate results of calculations; an intra-module data bus: for realizing efficient data transmission between sub-units.

[0034] Further, the calculation process in the Kalman gain calculation A unit is as follows:

[0035] The matrix (HP k,k-1 H T + R k ) -1 is expanded into the following form:

[0036]

[0037] The inverse matrix of the matrix whose inverse matrix is required is in the following form after Gaussian elimination inference:

[0038]

[0039] In the above formula, the mathematical calculation formulas for a, b, c, d, e, f, x, y, and z are as follows:

[0040]

[0041]

[0042] In the above expression, Δt represents the time step, and Θ (x,y) represents the optimal predicted state vector X of the target at the (k - 1)th moment k-1,k-1 covariance matrix P of k-1,k-1 the element in the xth row and yth column, Q (x,y) represents the process noise matrix Q k the element in the xth row and yth column, R (x,y) represents the measurement noise matrix R k the element in the xth row and yth column.

[0043] Furthermore, by deploying multiple identical computing modules to run in parallel to calculate a, b, c, d, e, f, x, y, and z, the computing efficiency is improved.

[0044] A target detection and tracking system, which adopts the target position information collected by radar and the image data collected by a visible light sensor, uses the dedicated Kalman filter computing structure for radar-guided target tracking to calculate the optimal predicted state vector of the target, and uses this vector combined with the image data to achieve target detection and tracking.

[0045] Compared with the prior art, the present invention has the following technical characteristics:

[0046] 1. Design of an efficient matrix inversion unit

[0047] Aiming at the high computational complexity problem of matrix inversion in the Kalman gain calculation process, a dedicated matrix inversion hardware unit is designed. This unit adopts an optimized method of block matrix decomposition and parallel operation, makes full use of the sparsity and symmetry of the matrix, significantly reduces the latency of the inversion operation, improves the computing efficiency, and thus breaks through the performance bottleneck of the traditional inversion algorithm for high-dimensional matrices.

[0048] 2. Modular parallel computing architecture

[0049] Adopt a modular design strategy to pipeline the core operations such as matrix multiplication, addition, and inversion in Kalman filtering, and combine it with a high-parallel processing architecture. Through multi-module parallel operation and pipelined design, achieve efficient hardware acceleration for the entire process of Kalman filtering, adapt to the large-scale matrix operation requirements in high-dynamic scenarios, and improve the overall system operation throughput.

[0050] 3. Optimized data transmission mechanism

[0051] Design an efficient data stream transmission mechanism, which combines the on-chip bus (OMBus) and the inter-module bus (MIBus) to optimize the data transmission paths inside and between modules. Introduce the AXI bus protocol, which supports multi-channel parallel communication and burst transmission mode, significantly reduces data transmission latency, improves the real-time performance of the overall system, and ensures that the high-frequency real-time computing requirements are fully met.

[0052] 4. Enhance system flexibility and scalability

[0053] Through the configurable systolic array matrix multiplication unit and the standardized design of module interfaces, achieve high flexibility and scalability of the system. According to different application scenarios and matrix dimension requirements, flexibly adjust the configuration and quantity of computing modules, which can not only achieve multi-module parallel acceleration when resources are sufficient, but also save hardware resources through pipelined design when resources are limited, ensuring the efficient operation of the system in various environments.

[0054] 5. Improve system coordination and stability

[0055] Based on the design of a control unit using a finite state machine (FSM), achieve efficient collaborative work among modules. Through precise state scheduling and signal control, ensure that the data dependency relationships between modules are correctly processed, avoid data competition and incorrect operation sequences, and improve the stability and reliability of the system. At the same time, support an error detection and feedback mechanism to enhance the fault tolerance of the system during complex computing processes.

[0056] In summary, through the dedicated hardware architecture design and optimized computing process, the present invention significantly improves the computing performance and real-time performance of the Kalman filtering algorithm in the radar-guided target tracking system. This dedicated computing structure for Kalman filtering not only effectively solves the computing bottleneck and real-time deficiency problems in the prior art, but also provides an efficient, flexible, and reliable hardware solution through modular design and optimized data transmission mechanism, providing strong support for high-precision target tracking applications in complex environments. Brief Description of the Drawings

[0057] Figure 1 It is a flowchart of a radar-guided visible sensor;

[0058] Figure 2 Schematic diagram of the overall architecture of the dedicated computing structure for Kalman filtering;

[0059] Figure 3 Flow chart of the operation of the dedicated computing structure for Kalman filtering;

[0060] Figure 4 Schematic diagram of the calculation method of the Kalman gain calculation module;

[0061] Figure 5 Schematic diagram of the hardware architecture of the Kalman gain calculation module;

[0062] Figure 6 Schematic diagram of the hardware architecture of Kalman gain calculation unit A;

[0063] Figure 7 Schematic diagram of the hardware architecture of the state update module;

[0064] Figure 8 Schematic diagram of the hardware architecture of the state prediction module;

[0065] Figure 9 Schematic diagram of the hardware architecture of the state prediction module;

[0066] Figure 10 Schematic diagram of the hardware architecture of the calculation execution unit in Kalman gain calculation unit A;

[0067] Figure 11 Schematic diagram of the architecture of the configurable systolic array multiplier unit;

[0068] Figure 12 Structure diagram of PE;

[0069] Figure 13 Finite state machine of the PE unit;

[0070] Figure 14 Finite state machine of the control unit of the dedicated computing structure for Kalman filtering. Detailed implementation manners

[0071] The present invention provides a dedicated computing structure for Kalman filtering used for radar-guided target tracking. Through innovative hardware architecture design and overall process optimization, the acceleration of the Kalman filtering algorithm is realized to meet the requirements of high-precision and high-real-time target tracking.

[0072] Referring to the accompanying drawings, a dedicated computing structure for Kalman filtering used for radar-guided target tracking provided by the present invention includes a state prediction module, a Kalman gain calculation module, a state update module, a covariance update module, and a delay buffer module. These modules cooperate under the unified scheduling of the control unit to complete the calculation tasks in each stage of Kalman filtering.

[0073] The radar acquires the position information of the target at the current k-th moment and stores it in the storage unit, where the position information of the target constitutes the measurement vector Z at the k-th moment k , and enters the dedicated computing structure for processing to obtain the optimal predicted state vector X of the target at the k-th moment k,k .

[0074] 1. State update module.

[0075] The state update module is used to receive the measurement vector Z at the k-th moment from the storage unit k , the state vector X at the k-th moment obtained by predicting with the state at the (k - 1)-th moment k,k-1 and the Kalman gain matrix K provided by the Kalman gain calculation module k , and performs prediction calculations according to the following formula to output the optimal predicted state vector X of the target at the k-th moment k,k ; providing a preliminary estimate for subsequent state updates.

[0076] X k,k = X k,k-1 + K k (Z k - HX k,k-1 )

[0077] where H represents the measurement matrix.

[0078]

[0079] In this module, Z k - HX k,k-1 is obtained by combining a multiplexer (Mux), an inverter (Neg), and an adder.

[0080] 2. Kalman gain calculation module.

[0081] The Kalman gain calculation module includes a Kalman gain calculation A unit (i.e., matrix inversion unit), a Kalman gain calculation B unit, and a Kalman gain multiplication unit, where:

[0082] The Kalman gain calculation A unit (i.e., matrix inversion unit), in combination with the process noise matrix Q from the storage unit k , calculates the gain K A = (HP k,k-1 H T + R k ) -1 ; R k represents the measurement noise matrix, P k,k-1 represents the covariance matrix of the state vector X k,k-1 , and the superscript T represents the transpose of the matrix.

[0083] Among them, the Kalman gain calculation A unit adopts the following architecture:

[0084] Calculation execution unit (CEU): Divided into multiple parallel sub-units, which process the element-level blocks of the matrix input to this unit; Double-precision floating-point operation module: Supports multiplication, subtraction, division, and negation operations; Intermediate result register: Used to store the intermediate results of the calculation, reduce redundancy in data transmission, and improve operation efficiency; Intra-module data bus: Enables efficient data transmission between sub-units.

[0085] Kalman gain calculation B unit, used to calculate K B = P k,k-1 H T ; The Kalman gain calculation B unit performs matrix operations based on the specific column elements of P k,k-1 and transmits data through the inter-module bus (MIBus) to support the subsequent operations of the covariance update module.

[0086] Kalman gain multiplication unit, used to multiply the outputs of the Kalman gain calculation A unit and the Kalman gain calculation B unit K A * K B to obtain the Kalman gain K k ; A configurable systolic array matrix multiplication unit is adopted during the calculation process.

[0087] The design concept of the Kalman gain calculation module is shown in Appendix Figure 4 , and separately calculating the two parts of the Kalman gain calculation A unit and the Kalman gain calculation B unit can make full use of the parallel potential and optimize the calculation efficiency. The hardware architecture of this module is shown in Appendix Figure 5 .

[0088] Existing dedicated calculation structures for Kalman filtering have a bottleneck in matrix inversion calculation. Especially under the high-precision uniformly accelerated motion model, the calculation complexity is further increased. To address this problem, the present invention designs an efficient matrix inversion unit as the Kalman gain calculation A unit, which is optimized specifically for the uniformly accelerated motion model. Through a strategy combining parallel calculation and serial scheduling, the latency of the inversion operation is significantly reduced, and the performance and efficiency of the gain matrix calculation are improved.

[0089] The present invention extracts the characteristics of the prediction equation and the update equation in Kalman filtering by utilizing the sparsity and symmetry of the matrix, and calculates the matrix (HP k,k-1 H T + R k ) -1 in advance, thereby optimizing the calculation efficiency. Specifically, this method transforms the high-dimensional matrix inversion into double-precision floating-point element four arithmetic operations with O(n) complexity, significantly reducing the calculation complexity and having the potential for parallel calculation, providing an effective way for the acceleration of Kalman filtering.

[0090] Matrix (HP k,k-1 H T +R k ) -1 is expanded into the following form:

[0091]

[0092] The inverse matrix of the required inverse matrix is in the following form after Gaussian elimination inference:

[0093]

[0094] In the above formula, the mathematical calculation formulas of a, b, c, d, e, f, x, y, z are as follows:

[0095]

[0096]

[0097] In the above expression, Δt represents the time step, Θ (x,y) represents the covariance matrix P k-1,k-1 of the optimal predicted state vector X k-1,k-1 at the (k - 1)th moment, the element in the x-th row and y-th column, Q (x,y) represents the process noise matrix Q k at the element in the x-th row and y-th column, R (x,y) represents the measurement noise matrix R k at the element in the x-th row and y-th column.

[0098] Through pre - calculation, the matrix inversion unit proposed by the present invention can complete the calculation of the Kalman gain with lower hardware resources and faster speed, significantly improving the real - time performance of Kalman filtering. Based on the above, and further combined with high - parallel hardware design, the schematic diagram of the hardware architecture is shown in the appendix Figure 6 .

[0099] In addition, the present invention supports flexible multi - module configuration. When there are sufficient hardware resources, especially in an environment rich in FPGA resources, multiple identical computing modules can be deployed to run in parallel; for example, 9 computing modules can be deployed to generate the required matrix elements a, b, c, d, e, f in parallel, thus significantly improving the computing efficiency and system throughput. For the case of resource constraints, a pipelining design can be adopted to gradually generate the calculation results through time slicing, thereby reducing the demand for hardware resources, especially in the use of double - precision floating - point multiplication IP, while still maintaining a high computing efficiency.

[0100] 3. State prediction module.

[0101] The state prediction module is used to receive the optimal predicted state vector X from the state update module k,k and the state transition matrix F from the storage unit, and complete the update of the state vector according to the following formula to predict the state vector X at time k+1 k+1,k :

[0102] X k+1,k = FX k,k

[0103] 4. Covariance update module.

[0104] The covariance update module is used to complete the update of the covariance matrix P of the optimal predicted state vector X at time k k,k The calculation formula is: k,k

[0105]

[0106] where I represents the identity matrix; this update process uses the matrix calculation unit for efficient parallel operations, reducing the system's computational load and ensuring that the system can respond to dynamic state changes in real time.

[0107] 5. Delay buffer module.

[0108] The delay buffer module is used to temporarily store the intermediate calculation results X k,k and P k,k and provide inputs for the next iteration to ensure the continuity of the data stream. The actual hardware architecture is implemented using FIFOs with a width of 64 bits and a depth of 12 and a width of 64 bits and a depth of 144 respectively.

[0109] 6. Control unit.

[0110] The control unit is designed based on a finite state machine (FSM) and is used to schedule the working order and data flow of each module. Through the precise control of the finite state machine, each module can process data in parallel at the appropriate time while ensuring the correctness of data dependencies. For example, while the state prediction module is performing calculations, the A unit of the Kalman gain calculation module can perform some preparatory work. Once the state prediction is completed, the Kalman gain calculation can be carried out quickly, reducing the overall calculation time. At the same time, the finite state machine ensures the state interlock between modules, avoiding data competition and incorrect operation sequences. Each module only starts working when it receives the correct control signal, ensuring data consistency and system stability. For example, the covariance update module will not start working when the Kalman gain has not been calculated, preventing the introduction of incorrect data

[0111] The optimal predicted state vector X at the target time k obtained after being processed by the dedicated calculation structure k,k and the corresponding covariance matrix P​k,k Output to a visible light sensor, which is used to guide the detection of a target by combining the image data collected by the visible light sensor, thereby ensuring the accuracy and stability of target detection; in addition, the optimal predicted state vector X k,k and the covariance matrix P k,k are also stored in the storage unit for external devices to use. The predicted state vector includes the position information and velocity information of the target at time k, and the covariance matrix P k,k includes the uncertainties of the position information and velocity information.

[0112] Among them, the control unit includes the following states:

[0113] Initialization state (Init): Clear all registers in the dedicated computing structure and empty the cache to prepare for the calculation process;

[0114] State prediction state (State_Prediction): Calculate the current predicted state vector according to the state vector and Kalman gain matrix of the previous moment;

[0115] Kalman gain calculation state (Cal_Kalman_Gain): Perform inverse matrix calculation, matrix multiplication and gain update operations;

[0116] State-covariance update state (State_Covariance_Update): Complete the update of the covariance matrix;

[0117] Measurement data input state (Measure_Data_Input): Receive a new measurement vector passed in from the outside and provide input for the next iteration;

[0118] State and covariance output state (State_Covariance_Output): Output the calculated predicted state vector and covariance matrix results to an external module;

[0119] End state (End): Enter the standby mode after all calculation tasks are completed.

[0120] Each state is triggered and transferred through valid signals (such as Init_Valid, SP_Done, CKG_Done, etc.), ensuring that the working order of each module is clear, the data transfer is accurate, and thus the orderly execution of tasks and the efficient operation of the system are realized.

[0121] 7. Configurable systolic array matrix multiplication unit.

[0122] Based on the above technical solutions, the present invention further provides a configurable systolic array matrix multiplication unit, which is used to perform multiplication operations between matrices of different dimensions during the Kalman filtering process. For example, the output multiplication K of the Kalman gain calculation A unit and the Kalman gain calculation B unit A *K B in the process can use this unit.

[0123] In the calculation of Kalman filtering, since a large number of fixed matrix multiplication operations are involved ([12,12]*[12,6], [12,6]*[6,6], [12,12]*[12,1], [12,6]*[6,12]), using a configurable systolic array matrix multiplication unit can effectively save hardware resources and improve calculation efficiency. The systolic array is an efficient parallel computing architecture, especially suitable for matrix operations. Data flows between the processing elements (PEs) of the systolic array along a fixed path and in a sequential order. Each processing element PE is responsible for simple calculations (such as multiplication and addition) and passes the intermediate results to the adjacent processing element PE to achieve pipelined data processing.

[0124] See Figure 11 , the configurable systolic array matrix multiplication unit specifically includes:

[0125] An input register, which is used to receive and temporarily store input matrix data; a reconfigurable decoder, which is used to dynamically configure the array calculation scale and enable different numbers of processing elements PE to adapt to different matrix dimensions; control signals (such as enb_1, enb_2_6, etc.), which are used to coordinate the operations of the various components of the configurable systolic array matrix multiplication unit; a delay counter, which controls the transfer order of data between the processing elements PE to ensure synchronous transmission of data between the processing elements PE; a processing element PE, which is used to perform multiplication and accumulation operations; an output channel, which is used to output the calculation result to the target module (the corresponding module in the dedicated computing structure) or a storage device.

[0126] For the input matrices C and D to be multiplied in the input register, the data of the input matrix C is loaded row by row into the row register and gradually input into the processing elements (PEs) of each column through a delay counter; the data of the input matrix D is loaded column by column into the top register and synchronously input into the processing elements of each row; the number of processing elements PEs is configured by a reconfigurable decoder. Under the coordination of control signals, each processing element PE receives data from the left or above, and after completing the multiplication operation, passes the intermediate result to the right or below, gradually completing the multiplication operation of the input matrices C and D; the final result is collected through the output channel to form a complete result matrix; through flexible configuration and efficient data flow, this architecture realizes low-latency and high-precision matrix multiplication calculation, adapting to the requirements of various dimensions and application scenarios.

[0127] The matrix multiplication unit adopts a streaming data transmission and synchronization mechanism, dynamically configures the number of enabled processing elements PEs according to the matrix scale to adapt to different matrix operation requirements; in the scenario of matrix operations with a fixed scale, redundant processing elements PEs can be turned off to save resources and reduce power consumption. This configurable systolic array matrix multiplication unit is applicable to fixed matrix multiplication operations in Kalman filtering and has efficient and real-time computing capabilities.

[0128] Among them, the processing element PE specifically includes: a PE input channel for receiving data from adjacent processing elements PEs or the outside; a multiplier IP core for performing the multiplication operation of the input data to generate partial products; a register for storing the partial products for subsequent operations; an adder IP core for adding the partial products to the accumulated result to generate a new accumulated result; a PE output channel for passing the accumulated result to adjacent processing elements PEs; a PE control unit designed based on a finite state machine for scheduling the operation process of the processing element PE.

[0129] The processing element PE realizes efficient operation by instantiating a double-precision floating-point arithmetic IP core (implemented by Vivado), supports pipelining operations and can be configured with bit widths to adapt to different precision requirements. The state transition of the control logic ensures the complete process of data from reception, operation to output, supports flexible configuration and dynamic adjustment, and is applicable to matrix operation scenarios of different scales.

[0130] The PE control unit controls the operation state through enable signals and reset signals to ensure the orderliness of the data stream and the correctness of the calculation results; the PE control unit includes the following states:

[0131] 1) Idle state (IDLE), used for standby and waiting for the start of a new task;

[0132] 2) Initialization state (INIT), used for loading initial data and configuring control signals;

[0133] 3) Multiplication state (Mul), which is used to perform multiplication operations and generate partial products;

[0134] 4) Addition state (Add), which is used to add the partial products to the accumulated result;

[0135] 5) Send data state (Send_data), which is used to send the calculation result to the next PE control unit or the PE output channel;

[0136] 6) Data Through state (Data_Through), which is used to directly transfer data under specific conditions without further processing;

[0137] 7) End state (End), which is used to complete the current operation task and prepare to return to the idle state.

[0138] The present invention proposes an efficient communication mechanism to meet the requirements of fast and reliable data interaction between modules of a dedicated computing structure for Kalman filtering.

[0139] The communication mechanism mainly consists of an On-Module Bus (OMBus) and a Module Interconnect Bus (MIBus). The OMBus is used for data transmission between sub-units within each module of the dedicated computing structure, and its design focus is to meet the requirements of short-distance and low-latency transmission. For example, within the Kalman gain calculation module, intermediate calculation results need to be frequently exchanged between sub-units. The OMBus ensures the smoothness and real-time performance of the calculations within the module through a high-speed and parallel data transmission mechanism. The MIBus is a data channel connecting each module of the dedicated computing structure, supporting efficient interaction of data between modules. It adopts a full-duplex communication mode, enabling two-way parallel transmission of data, thus significantly improving communication efficiency and avoiding bottleneck problems. For example, the result calculated by the state prediction module can be efficiently transmitted to the Kalman gain calculation module through the MIBus to provide input for subsequent operations.

[0140] The transfer of control signals is completed by the control unit, including the distribution of initialization signals, synchronization control, and error handling and feedback mechanisms. Through precise synchronization signal coordination, it ensures that the multi-module parallel calculations are in step. At the same time, when an exception occurs, error information is quickly fed back and the system state is adjusted to ensure system stability.

[0141] In addition, a multi-channel communication mechanism and a standardized design of module interfaces are adopted in the communication mechanism of the present invention. Multi-channel buffered transmission and hierarchical communication enable parallel processing of large-scale matrix data between modules, thus significantly improving the efficiency of matrix operations. For example, by dividing a large matrix into blocks and transmitting them to different modules for parallel operations, not only the transmission delay is reduced, but also the computing performance is optimized. The standardization of module interfaces ensures the compatibility and scalability between modules. Each module adopts standardized input interfaces, output interfaces, and control interfaces, and modules can be flexibly added or replaced without changing the communication architecture.

[0142] To support the above communication mechanism, the present invention adopts the AXI bus protocol as the core communication protocol. The AXI protocol supports burst transmission mode, which can transmit multiple data blocks at one time, greatly improving the transmission efficiency. Its pipeline architecture not only achieves high-bandwidth transmission, but also reduces communication latency. In addition, the master-slave device architecture design of the AXI protocol enables flexible expansion between modules. Whether it is the internal sub-units of a module or the communication between modules, they can be conveniently connected to the AXI bus, meeting the requirements of the modular architecture of the present invention.

[0143] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A Kalman filter dedicated computing structure for radar-guided target tracking, characterized in that: include: The state prediction module, Kalman gain calculation module, state update module, covariance update module and delay buffer module work together under the unified scheduling of the control unit designed based on the finite state machine to complete the calculation tasks of each stage of the Kalman filter; The radar collects the position information of the target at the current k moment and saves it to the storage unit, where the position information of the target constitutes the measurement vector Z at the k moment k , enter the dedicated computing structure for processing; wherein: The state update module is used to receive the measurement vector Z at time k from the storage unit k , the state vector X at time k obtained by using the state prediction at time k-1 k,k-1 And the Kalman gain matrix K provided by the Kalman gain calculation module k , the measurement matrix H is used for prediction calculation, and the optimal prediction state vector X of the target at time k is output k,k ; The Kalman gain calculation module includes a Kalman gain calculation A unit, a Kalman gain calculation B unit, and a Kalman gain multiplication unit; The Kalman gain calculation unit A is a matrix inversion unit, combined with the process noise matrix Q from the storage unit k , calculate the gain K A =(HP k,k-1 H T +R k ) -1 ; R k represents the measurement noise matrix, P k,k-1 Denotes the state vector X k,k-1 The covariance matrix of , the superscript T represents the transpose of the matrix; Kalman gain calculation unit B, used to calculate K B =P k,k-1 H T ; Kalman gain multiplication unit, used to multiply the output of Kalman gain calculation unit A and Kalman gain calculation unit B by K A *K B , get the Kalman gain K k ; The state prediction module is used to receive the optimal predicted state vector X from the state update module. k,k And the state transfer matrix F from the storage unit, and complete the update of the state vector according to the following formula, and predict the state vector X at time k+1 k+1,k : The covariance update module is used to complete the optimal prediction state vector X at time k k,k The covariance matrix P k,k Updates; The delay cache module is used to temporarily store intermediate calculation results X k,k and P k,k , providing input for the next round of iteration; The optimal predicted state vector X at the target k moment obtained after processing by a dedicated computing structure k,k The output is sent to the visible light sensor and used to guide the detection of the target in combination with the image data collected by the visible light sensor; wherein the optimal predicted state vector includes the position information and speed information of the target at time k.

2. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 1, characterized in that: The dedicated computing structure also includes: A configurable systolic array matrix multiplication unit is used to perform multiplication operations between matrices of different dimensions in a Kalman filtering process; wherein: the configurable systolic array matrix multiplication unit specifically includes: Input register, used to receive and temporarily store input matrix data; reconfigurable decoder, used to dynamically configure the array calculation scale and enable different numbers of processing units PE to adapt to different matrix dimensions; control signal, used to coordinate the operation of each component of the configurable systolic array matrix multiplication unit; delay counter, used to control the order of data transmission between processing units PE to ensure the synchronous transmission of data between processing units PE; processing unit PE, used to perform multiplication and accumulation operations; output channel, used to output the calculation results to the target module or storage device.

3. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 2, characterized in that: For the input matrix C and input matrix D to be multiplied in the input register, the data of the input matrix C is loaded into the row register by row, and is gradually input into the processing unit PE of each column through the delay counter; the data of the input matrix D is loaded into the top register by column, and is synchronously input into the processing unit of each row; The number of processing units PE is configured using the reconfigurable decoder. Under the coordination of the control signal, each processing unit PE receives data from the left or top, and after completing the multiplication operation, it passes the intermediate result to the right or bottom, gradually completing the multiplication operation of the input matrix C and the input matrix D; the final result is collected through the output channel to form a complete result matrix.

4. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 2, characterized in that: The processing unit PE specifically includes: a PE input channel for receiving data from an adjacent processing unit PE or the outside; a multiplier IP core for performing multiplication operations on input data to generate partial products; a register for storing partial products for subsequent operations; an adder IP core for adding partial products to accumulation results to generate new accumulation results; a PE output channel for transmitting the accumulation results to an adjacent processing unit PE; and a PE control unit designed based on a finite state machine for scheduling the operation flow of the processing unit PE.

5. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 1, characterized in that: The communication mechanism of the dedicated computing structure includes an intra-block bus OMBus and an inter-module bus MIBus; The intra-block bus OMBus is used for data transmission between sub-units within each module of the dedicated computing structure, and the inter-module bus MIBus is responsible for multi-channel data interaction and collaborative operation between modules of the dedicated computing structure, and uses the AXI bus protocol as the core communication protocol.

6. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 2, characterized in that: Kalman gain K k , measurement noise matrix R k and the state vector X k,k-1 The covariance matrix R k-1,k-1 Through the inter-module bus and intra-module bus transmission, the matrix multiplication calculation is completed using a configurable systolic array multiplication unit. The calculation results are stored in the FIFO queue. The data is scheduled through the FIFO to complete the matrix addition calculation of the covariance matrix P k,k The configurable systolic array unit can dynamically adjust the computing scale to adapt to the computing requirements of matrices of different dimensions.

7. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 1, characterized in that: The Kalman gain calculation A unit adopts the following architecture: Computation execution unit: divided into multiple parallel sub-units, which process the element-level blocks of the matrix input to the unit; double-precision floating-point operation module: supports multiplication, subtraction, division and negation operations; intermediate result register: used to store the intermediate results of the calculation; intra-module data bus: realizes efficient data transmission between sub-units.

8. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 1, characterized in that: The calculation process in the Kalman gain calculation unit A is: Matrix (HP k,k-1 H T +R k ) -1 Expands to the following form: The inverse matrix of the required inverse matrix is ​​in the following form after Gaussian elimination: In the above formula, the mathematical calculation formulas of a, b, c, d, e, f, x, y, and z are as follows: In the above expression, Δt represents the time step, Θ (x,y) Represents the optimal predicted state vector X of the target at time k-1 k-1,k-1 The covariance matrix P k-1,k-1 The element in row x and column y, Q (x,y) Denotes the process noise matrix Q k The element in row x and column y, R (x,y) Denotes the measurement noise matrix R k The element at row x, column y.

9. The Kalman filter dedicated computing structure for radar-guided target tracking according to claim 8, characterized in that: By deploying multiple identical computing modules to run in parallel, a, b, c, d, e, f, x, y, z can be calculated, thereby improving computing efficiency.

10. A target detection and tracking system, characterized in that: The system uses target position information collected by radar and image data collected by visible light sensors, and uses the Kalman filter dedicated computing structure for radar-guided target tracking according to any one of claims 1-9 to calculate the optimal predicted state vector of the target, and uses this vector in combination with image data to achieve target detection and tracking.