A method for constructing a parallel acceleration computing architecture for tracking faint targets in orbit

By adopting the Kalman filtering calculation method based on digital logic parallel acceleration architecture in orbit space dark weak target tracking, the calculation process is decomposed and the calculation period and parallelism are optimized, and the problems of low computing efficiency and strict noise covariance constraints in the existing technology are solved, and efficient real-time tracking calculation is achieved.

CN116579911BActive Publication Date: 2025-05-06CHINA ORDNANCE SCI INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310584899.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-05-06
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

In the real-time tracking of dark and weak targets in orbit space, it is difficult to complete the calculation at high frame rates in milliseconds, and the existing acceleration design has strict constraints on noise covariance, and it is impossible to ensure that all gain matrices can be inversely determined.

Method used

The Kalman filtering calculation method based on the digital logic parallel acceleration architecture is adopted to decompose the target state Kalman filtering calculation process into different VPE calculation units. Through flow division and parallel calculation, the calculation period and parallel degree are optimized to realize the parallel calculation of matrix row vectors and column vectors.

Benefits of technology

It significantly shortens the calculation time of a single filter iteration, improves the overall computing efficiency, reaches the computing speed of microseconds, and meets the real-time tracking needs at high frame rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116579911B_ABST
    Figure CN116579911B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing a parallel accelerated computing architecture for tracking dim targets in orbit, including: dividing the target state Kalman filter computing process into pipelines, obtaining multiple subprocesses one, determining the subprocess one for parallel computing, obtaining the merged target state Kalman filter computing subprocess two, wherein the subprocess two includes the subprocess one for determining parallel computing; taking the minimum computing cycle of the merged computing process as the objective function, taking the optimal overlap between the subprocesses two as the constraint condition, and determining the parallelism of the subprocess one in combination with the computing state and computing cycle of the VPE; determining the number of VPEs, which are used to perform the computing operations of the row vectors and column vectors of the matrix in the subprocess one; and the general CTM distributes the data in the subprocess one to each VPE for computing. The present invention is used for real-time computing of the dim target tracking process in orbit, has less single filtering iteration computing time, and improves the overall computing efficiency by one order of magnitude.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing for tracking faint targets in orbit, and relates to a construction method based on a digital logic parallel acceleration computing architecture, which is used for real-time computing of faint targets detection and tracking processes in orbit. Background Art

[0002] With the continuous launch of small satellites such as Star-Link in recent years, the number of spacecraft in orbit has increased, and collision incidents have occurred from time to time. Space traffic safety and the protection of key space assets (such as space stations) have become a hot topic of research at home and abroad. European and American countries have deployed a large number of ground-based telescopes for observing threats such as space debris, but debris often has a small effective reflection area and large attitude flip luminosity flicker. At the same time, the atmosphere affects target observation, making it difficult to observe a large number of dim targets in real time.

[0003] The space-based target observation system mainly relies on optical payloads to continuously detect the airspace, screen mobile targets that are different from the background under the star background, obtain the target's position and trajectory information, and make predictions based on the existing trajectory for applications such as collision warning. For the tracking of dim targets in space, especially the tracking and trajectory prediction of multiple targets, Kalman filtering is mainly used to complete the association of discontinuous trajectories and the distinction of adjacent trajectories. In particular, the attitude changes of close-range targets and the sudden changes in trajectory attitude during maneuvers can be effectively estimated and solved by using Kalman filtering.

[0004] Tracking of dim targets on orbit mainly combines the detection results of optical image data to form a target sequence trajectory, and obtains the final association and predicted trajectory after outlier removal and Kalman filter solution. The key to this process lies in the real-time calculation of target detection and tracking. After long-term development, inter-frame sequence detection and tracking has formed a mature framework. In the tracking process, how to determine whether the real-time Kalman filter calculation can be effectively completed on orbit is the key to building its computing architecture.

[0005] The real-time tracking process of faint space targets is mainly to achieve the association and prediction of point traces after the detection calculation process. The on-orbit calculation of the tracking part needs to be completed within milliseconds, and needs to reach the microsecond level at a high frame rate. For the real-time tracking of faint space targets, it is necessary to complete the general tracking calculation acceleration to cope with different sky areas and complex backgrounds, which puts high demands on the engineering implementation architecture. The existing design mainly relies on the combination of software and hardware acceleration design, using high-performance CPU to complete the control of complex processes, and the logic part to implement matrix multiplication and other operations, or using simplification for certain special applications to reduce resource and time consumption. Some literature also starts with algorithm optimization, combining sequential fusion and square root methods to reduce the matrix inversion in the gain calculation process to improve real-time performance, but this type of method requires the measurement noise covariance to be a constant matrix or a diagonal matrix, and the constraints are relatively strict. Some literature also uses block-based inversion methods for acceleration, but the block matrix in the gain matrix inversion process cannot be guaranteed to be inverted.

[0006] Therefore, how to provide a method for constructing a parallel accelerated computing architecture for tracking dim targets in orbit that can shorten the calculation time of a single filtering iteration and improve the overall computing efficiency is an urgent problem that technicians in this field need to solve. Summary of the invention

[0007] In view of this, the present invention proposes a Kalman filter calculation based on a digital logic parallel acceleration architecture, which is used for real-time solution of the on-orbit space dim target tracking process, has less single filter iteration calculation time, and the overall calculation efficiency is improved by an order of magnitude.

[0008] In order to achieve the above object, the present invention adopts the following technical solution:

[0009] The present invention discloses a method for constructing a parallel accelerated computing architecture for tracking dim targets in orbit. The steps of tracking dim targets in orbit sequentially include: inter-frame registration, target detection and inter-frame accumulation, trajectory association, target state Kalman filtering and prediction; decomposing the target state Kalman filtering calculation process into different VPE calculation units, including the following steps:

[0010] S1: Divide the target state Kalman filter calculation process into pipelines to obtain k sub-processes one, and determine the sub-process one for parallel calculation according to the correlation between the sub-processes, and obtain the merged target state Kalman filter calculation process, including n sub-processes two, n<k, and the sub-process two includes the sub-process one for parallel calculation;

[0011] S2: Taking the minimum calculation cycle of the merged calculation process as the objective function and the optimal overlap between sub-processes 2 as the constraint condition, the parallelism of sub-process 1 is determined in combination with the calculation state and calculation cycle of the VPE calculation unit;

[0012] S3: Determine the number of VPE computing units according to the parallelism, wherein the VPE computing units are used to perform computing operations on the row vectors and column vectors of the matrix in the first subprocess;

[0013] S4: The general configuration management module CTM distributes the data in the sub-process 1 to each VPE computing unit for calculation, and the calculation results of the VPE computing unit are stored through the selector MUX.

[0014] Preferably, in S1, sub-process one for parallel calculation is determined according to the calculation relationship between each matrix variable in the target state Kalman filter calculation process and the correlation between the data streams between sub-processes.

[0015] Preferably, the S2 specifically includes:

[0016] S21: Determine the calculation period function of the nth subprocess 2 using the calculation amount C(n), parallelism B(n) and overlapping period D(n) of the nth subprocess 2; calculate the optimal solution set of B(n) with the minimum sum of the calculation period functions of all subprocesses 2 as the objective function; n∈Z + , Z + represents a positive integer;

[0017] S22: according to the calculation frequency of different clock domains of the data reading and writing and calculation process of the second subprocess, combined with the calculation state and calculation cycle of the VPE calculation unit, obtain the optimal value set of B(n) when the data reading and writing and calculation of the VPE parallel acceleration matrix process in the output result of S21 are balanced;

[0018] S23: Taking the minimum variance of the sum of the parallelism of the VPE computing units working at any time as a constraint condition, find the single optimal solution of B(n) under the optimal overlap condition between each sub-process;

[0019] S24: Taking the parallelism between the parallel computing sub-processes 1 being proportional to the amount of computation as a constraint, the parallelism of the parallel sub-process 1 is obtained, thereby obtaining a complete parallel computing architecture.

[0020] Preferably, the S21 specifically includes:

[0021] The calculation period of the nth subprocess 2 is T(n)=f(C(n) / B(n))-D(n);

[0022] in, k is the proportionality constant;

[0023] The whole calculation period T is calculated according to the following formula A The solution of B(n) with the minimum value is:

[0024]

[0025] Preferably, the S22 specifically includes:

[0026] The data of subprocess 2 is stored in Buffer;

[0027] The constraints for access and computation balance are:

[0028] min|Θ1f1-max{Θ (2,i) f2}|

[0029] The data access and calculation process in the buffer select different clock domains. f1 is the calculation frequency of clock domain 1, and f2 is the calculation frequency of clock domain 2. Θ1 is the amount of data taken out of the buffer, Θ (2,i) is the i-th VPE computing unit v i calculation cycle.

[0030] Preferably, the S23 specifically includes:

[0031] The constraint condition on the variance of the sum of the parallelism of the VPE computing units working at any time is:

[0032]

[0033] Among them, B t (n) represents the parallelism of the VPE computing unit effective at time t, and n&t represents the effective computing subprocess two at time t.

[0034] Preferably, the S24 specifically includes:

[0035] The constraint that the degree of parallelism between the subprocesses of parallel computing is proportional to the amount of computing is:

[0036]

[0037] Among them, ε=1, 2 represents the subprocess number, B(n.ξ1) represents the parallelism of subprocess ξ1 in the n-th subprocess ξ2, and B(n.ξ2) represents the parallelism of subprocess ξ2 in the n-th subprocess ξ2 that is calculated in parallel with subprocess ξ1; C(n.ξ1) represents the computational amount of subprocess ξ1, and C(n.ξ2) represents the computational amount of subprocess ξ2.

[0038] Preferably, the architecture of the VPE computing unit in S3 executing the computing operation of the matrix row vector and column vector in subprocess 1 includes:

[0039] The input data calculation of the VPE computing unit includes two types of work: streaming data and register data;

[0040] When the input is stream data, the data in the row vector and column vector in the N-dimensional matrix multiplication operation and addition and subtraction operation are input to N cascaded D flip-flops to complete the serial-to-parallel conversion;

[0041] When the input is register data, the data is stored in parallel in the register;

[0042] The input stream data and register data are input into N calculators respectively according to the selected input data judged by the selector MUX; the N calculators select multiplication, addition or subtraction operation to complete the calculation according to the input of the selection signal, and the calculation results are added step by step through the adder to complete the accumulation operation and output the final result.

[0043] Preferably, the S4 specifically includes:

[0044] The input of the general configuration management module CTM is the row and column values ​​of the two input matrices of the matrix operation. When it is determined that there is data of the sub-process 2 that needs to be processed in the Buffer, the following steps are executed:

[0045] S41: Determine the status of each VPE computing unit. If all VPE computing units are busy, wait. If any VPE computing unit is idle, take out the matrix row vector or column vector data to be calculated from the Buffer, store it in the VPE computing unit and start the calculation.

[0046] S42: Return to S41, continue to loop to detect the status of the VPE computing unit and start the calculation.

[0047] Preferably, a fixed-point system is used to perform calculations of a single VPE calculation unit; and an iterative error calculation based on a target state Kalman filter is performed according to different decimal bit widths in the fixed-point calculation system to obtain a minimum bit width number that meets the accuracy requirements.

[0048] It can be seen from the above technical solution that compared with the prior art, the present invention has the following gain effects:

[0049] The present invention is used for real-time solution of the tracking process of dim targets in orbit. It establishes a computing unit structure VPE with a unified architecture for Kalman filter matrix calculation, and designs an array computing structure based on VPE to decompose the entire filtering process and form a parallel pipeline data flow. The present invention greatly reduces the number of VPE computing units, reduces the hardware resource occupation of FPGA, shortens the computing cycle, and greatly improves the maximum executable frequency of the entire FPGA program. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative work.

[0051] Figure 1 Schematic diagram of the data flow of the Kalman filter calculation process provided by an embodiment of the present invention;

[0052] Figure 2 It is a schematic diagram of the steps of detecting and tracking a dim target on orbit provided by an embodiment of the present invention;

[0053] (a) corresponds to the schematic diagram of inter-frame registration;

[0054] (b) Schematic diagram corresponding to target detection and inter-frame accumulation;

[0055] (c) corresponds to the trajectory association diagram;

[0056] (d) Schematic diagram of Kalman filtering and prediction corresponding to the target state;

[0057] Figure 3 It is a schematic diagram of the principle of the VPE parallel computing architecture provided by an embodiment of the present invention;

[0058] (a) Schematic diagram of the VPE parallel computing matrix multiplication architecture principle;

[0059] (b) Schematic diagram of the VPE parallel computing matrix addition / subtraction architecture principle;

[0060] Figure 4 It is a schematic diagram of the principle of the VPE vector computing architecture provided by an embodiment of the present invention;

[0061] Figure 5 It is a schematic diagram of the principle of a multi-VPE vector computing parallel acceleration architecture provided by an embodiment of the present invention;

[0062] Figure 6 It is a schematic diagram of the number of time cycles for each subprocess calculation under the parallel architecture provided by an embodiment of the present invention;

[0063] Figure 7 is a schematic diagram of the accumulation of effective parallelism at different computing moments provided by an embodiment of the present invention;

[0064] Figure 8 It is a schematic diagram comparing the target trajectory after Kalman filtering provided by an embodiment of the present invention with the theoretical trajectory and the observed trajectory;

[0065] Fig. 9is a schematic diagram of error comparison before and after Kalman filtering provided by an embodiment of the present invention;

[0066] Fig.10 It is a schematic diagram of the average error of filtering calculation results under different quantization bit widths provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0067] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0068] The parallel accelerated calculation of dim target tracking based on Kalman filtering provided by the embodiment of the present invention is implemented based on the VPE calculation unit (VPE, Vector Processing Element) of FPGA. The following first describes the principle of dim target detection and tracking on orbit in conjunction with the accompanying drawings.

[0069] See also Figure 2 , a schematic diagram of the calculation steps and results of each part of on-orbit faint target detection and tracking is given, which mainly includes three parts: alignment of the previous and next sequence image frames based on the stellar background, inter-frame moving target detection, and sequence target point trace filtering tracking.

[0070] Define the target position and velocity information on the camera imaging plane as X(t)=[p x (t),v x (t),p y (t),v y (t)], where p x (t) and p y (t) is the time t(t∈Z + ) is the coordinate value of the x and y directions in the camera plane coordinate system, v x (t) and v y (t) is the velocity in the x and y directions. The one-step transfer matrix F and the measurement matrix H are:

[0071]

[0072] Where S is the imaging interval.

[0073] Define Q and R as process noise variance and observation noise variance respectively, and the detection result Z(t)=[p x (t),p y (t)] T , then the t-th Kalman filter calculation process is:

[0074]

[0075] Among them, X pre (t) is the state prediction result, including the position and velocity information in the x and y directions;

[0076] X KF (t) is the t-th Kalman filtering result of the target position and velocity information on the imaging plane, and the initial value is the X(0) vector;

[0077] P(t) is the covariance matrix initialized to the identity matrix;

[0078] P pre (t) is the covariance prediction matrix;

[0079] K(t) is the Kalman gain matrix;

[0080] e(t) is the prediction residual vector, which refers to the predicted position error, that is, the error of the x and y coordinates of the camera plane target.

[0081] The variables in formula (2) are described in detail as shown in the following table.

[0082] Table 1. Variable definitions and descriptions

[0083] Serial number Logo name Dimensions Remark 1 F One-step transfer array 4×4 Preset Value 2 H Measurement Matrix 2×4 Preset Value 3 S Imaging interval 1×1 Preset Value 4 Q Process noise variance 4×4 Preset Value 5 R Observation noise variance 2×2 Preset Value 6 Z(t) Test results 2×1 Iterative Updates 7 <![CDATA[X pre (t)]]> Status prediction results 4×1 Iterative Updates 8 XKF(t) Kalman filter results 4×1 Iterative Updates 9 P(t) Covariance matrix 4×4 Iterative Updates 10 <![CDATA[P pre (t)]]> Covariance prediction matrix 4×4 Iterative Updates 11 K(t) Kalman gain matrix 4×4 Iterative Updates 12 e(t) Prediction residual vector 2×1 Iterative Updates

[0084] It can be seen from formula (2) that the calculation processes that take up the most time are mainly matrix multiplication, inversion and other processes. Therefore, how to decompose complex calculations into various VPE computing units and ensure that the busy-idle ratio is consistent so that the overall design can achieve a more optimized result is a technical problem to be solved by the embodiments of the present invention.

[0085] The method for constructing a parallel accelerated computing architecture for tracking dim targets in orbit disclosed in an embodiment of the present invention comprises the following steps:

[0086] S1: Divide the target state Kalman filter calculation process into pipelines to obtain k sub-processes one, and determine the sub-process one for parallel calculation according to the correlation between the sub-processes, and obtain the merged target state Kalman filter calculation process, including n sub-processes two, n<k, and the sub-process two includes the sub-process one for parallel calculation;

[0087] S2: Taking the minimum calculation cycle of the merged calculation process as the objective function and the optimal overlap between sub-processes 2 as the constraint condition, the parallelism of sub-process 1 is determined in combination with the calculation state and calculation cycle of the VPE calculation unit;

[0088] S3: Determine the number of VPE computing units according to the parallelism, wherein the VPE computing units are used to perform computing operations on the row vectors and column vectors of the matrix in the first subprocess;

[0089] S4: The general configuration management module CTM distributes the data in the sub-process 1 to each VPE computing unit for calculation, and the calculation results of the VPE computing unit are stored through the selector MUX.

[0090] The specific execution process of the above steps is described below with reference to the accompanying drawings:

[0091] See also Figure 1 , giving the relationship between each variable and the data flow, including the dimension information in each matrix calculation process. The data calculation is divided into pipelines according to the relevant information described in the figure. Among them, the transposition calculation in FPGA only converts between registers, and does not consider the occupation of storage space reading and writing time; the matrix inversion adopts the second-order matrix fast inversion method, and the division adopts fast calculation and only takes one cycle, and the efficiency of multiplication, addition and division is consistent.

[0092] In one embodiment, in S1, subprocess one for parallel computing is determined according to the computational relationship between matrix variables in the target state Kalman filter computation process and the correlation between data streams between subprocesses.

[0093] The entire calculation process is divided into 17 sub-processes according to relevance, and after merging the same type, it becomes 11 sub-processes. Figure 1 As shown, the same sub-process is considered to be able to be calculated in parallel.

[0094] Table 2 Process and corresponding formula description

[0095]

[0096]

[0097] Table 2 shows the subprocesses that can be run in parallel, including 1.1 and 1.2; 2.1 and 2.2; 3.1 and 3.2; 9.1 and 9.2; 10.1 and 10.2. Parallel subprocesses can be computed at the same time.

[0098] The number and computational complexity of each merged sub-process are shown in Table 3. It can be seen that the sub-processes are highly correlated during the iterative calculation of the Kalman filter and are difficult to operate in complete parallel. How to design the pipeline between the sub-processes in the engineering process is the key to balancing hardware resource usage and performance.

[0099] Table 3 Number and computational effort of subprocess 2

[0100] Subprocess 2n quantity Calculation Amount Subprocess 2 quantity Calculation Amount Subprocess 2 quantity Calculation Amount 1. 2 64 / 16 5. 1 16 9. 2 32 / 8 2. 2 64 / 32 6. 2 8 / 32 10. 2 4 / 16 3. 2 16 / 2 7. 1 7 11. 1 64 4. 1 32 8. 1 16

[0101] It can be seen that after the specific input and output calculation analysis, the iterative process of Kalman filter calculation becomes a simple multiplication and addition calculation process of the main variables.

[0102] In one embodiment, S2 specifically includes:

[0103] S21: Determine the calculation period function of the nth subprocess 2 using the calculation amount C(n), parallelism B(n) and overlapping period D(n) of the nth subprocess 2; calculate the optimal solution set of B(n) with the minimum sum of the calculation period functions of all subprocesses 2 as the objective function; n∈Z + , Z + represents a positive integer;

[0104] S22: according to the calculation frequency of different clock domains of the data reading and writing and calculation process of the second subprocess, combined with the calculation state and calculation cycle of the VPE calculation unit, obtain the optimal value set of B(n) when the data reading and writing and calculation of the VPE parallel acceleration matrix process in the output result of S21 are balanced;

[0105] S23: Taking the minimum variance of the sum of the parallelism of the VPE computing units working at any time as a constraint condition, find the single optimal solution of B(n) under the optimal overlap condition between each sub-process;

[0106] S24: Taking the parallelism between the parallel computing sub-processes 1 being proportional to the amount of computation as a constraint, the parallelism of the parallel sub-process 1 is obtained, thereby obtaining a complete parallel computing architecture.

[0107] The following is an explanation of the calculation steps of the parallelism of each process in S2 with the help of specific formulas:

[0108] In this embodiment, in S21, the calculation process n(n∈Z + ), as shown in Table 2, sub-trip 2, Z + represents a positive integer; the computational effort of the nth computational process is C(n), the degree of parallelism is B(n), and the overlap period between the two processes is k is a constant of proportional coefficient, and the calculation cycle of the nth process is T(n) = f(C(n) / B(n))-D(n). In the calculation process, the larger B(n) is, the smaller the calculation cycle T(n). When it increases to a certain extent, the calculation cycle tends to remain unchanged; at the same time, the larger B(n) is, the smaller D(n) is; the f(·) function represents the number of calculation cycles under different degrees of parallelism.

[0109]

[0110] Ultimately, we hope that the calculation cycle of all processes is minimized, that is, when the sum of the calculation cycles of n processes is TA To obtain the minimum, we need to obtain the B(n) values ​​of each level of the process, which is the entire optimal parallel computing architecture. Formula (3) is a nonlinear function, k is an empirical value or a test value, generally selected between 1 and 100, and the solution of formula (3) can be obtained by traversal search. Since the first formula in formula (3) is a conditional optimal solution, and the second formula in formula (3) is a nonlinear constraint, the settlement result of B(n) is the optimal solution set (more than one solution).

[0111] In this embodiment, in S22, VPE is defined as v i , i∈[0,…,Ψ-1], i∈Z + , Ψ is the number of VPEs. The data access and calculation process in the buffer select different clock domains. The calculation frequency of clock domain 1 is f1, and the calculation frequency of clock domain 2 is f2. The amount of data taken out from the buffer is defined as Θ1, v i The calculation period is Θ (2,i) , in order to ensure the balance between access and computing, it is necessary to satisfy

[0112] min|Θ1f1-max{Θ (2,i) f2}| (4)

[0113] Since the calculation frequencies f1 and f2 are determined according to the chip selection, the best Ψ selection result can be obtained according to formula (4), that is, the optimal numerical value set of B(n) (i.e., the value of Ψ) when the data reading and writing and calculation of the VPE parallel acceleration matrix process are balanced is obtained according to formula (5);

[0114] In this embodiment, in S23, considering that power is proportional to the maximum parallelism, the concept of cumulative sum of parallelism in any computing cycle is introduced, and the variance of the sum of parallelism of the VPEs working at any time should be minimized, that is,

[0115]

[0116] This ensures that there will not be too much parallelism at any time, where B t (n) represents the parallelism of the VPE effective at time t, and n&t represents the effective computing process at time t. The optimal value set of B(n) under the optimal overlap condition between each process is obtained for the result of S22 according to formula (5), and the single optimal solution of B(n) is obtained.

[0117] In this embodiment, in S24, in the process with parallel computing, for a single optimal solution of B(n) that has been obtained, the parallelism of the parallel computing process is proportional to the amount of computing:

[0118]

[0119] Where ε=1, 2 represents the subprocess number, B(n.ξ1) represents the parallelism of subprocess ξ1 in the nth subprocess ξ2, and B(n.ξ2) represents the parallelism of subprocess ξ2 in the nth subprocess ξ1; C(n.ξ1) represents the computational effort of subprocess ξ1, and C(n.ξ2) represents the computational effort of subprocess ξ2; ξ1,ξ2∈Z + , B(n.ξ1)∈Z + .

[0120] In one embodiment, a design architecture of a single vector computing unit VPE is provided for adapting matrix multiplication / addition / subtraction operations in a Kalman filter calculation process according to input parameters.

[0121] See also Figure 3 , the architecture of VPE is based on Figure 1 The matrix multiplication, addition, and subtraction operations in , which are sorted out, include the multiplication operation of 4×1 vector and 1×4 vector, and the addition and subtraction operations of 4×1 vector and 4×1 vector. According to the requirements of input parameters, they can satisfy the calculation of vectors of different lengths.

[0122] For example, in the calculation process 1.2, X pre (t) = FX KF (t-1), where F is a 4×4 matrix, X KF (t-1) is a 4×1 matrix. The process of matrix multiplication is to parallelize multiple VPEs and evenly decompose the calculation process to multiple VPEs, such as Figure 3 (a) shows that matrix multiplication is to assign the multiplication operation of each row vector and column vector of the matrix to a VPE computing unit. For the multiplication operation of d×d matrix, d parallel VPEs are constructed. Figure 3 (b) Taking the matrices Ta and Tb for matrix addition and subtraction as an example, the addition operation of the first row of matrix Ta and the first row of matrix Tb is assigned to a VPE computing unit. The addition operation of a d×d matrix constructs d parallel VPEs, and the output results are stored in Tc. B(n) is the number of VPEs (the parallelism value can be obtained according to formulas (3) to (5)).

[0123] In one embodiment, the architecture of the VPE computing unit in S3 performing the computing operation of the matrix row vector and column vector in subprocess 1 includes:

[0124] The input data calculation of the VPE computing unit includes two types of work: streaming data and register data;

[0125] When the input is stream data, the data in the row vector and column vector in the N-dimensional matrix multiplication operation and addition and subtraction operation are input to N cascaded D flip-flops to complete the serial-to-parallel conversion;

[0126] When the input is register data, the data is stored in parallel in the register;

[0127] The input stream data and register data are input into N calculators respectively according to the selected input data judged by the selector MUX; the N calculators select multiplication, addition or subtraction operation to complete the calculation according to the input of the selection signal, and the calculation results are added step by step through the adder to complete the accumulation operation and output the final result.

[0128] See also Figure 4 ,The input data computation of VPE includes two types of work: stream data and register data. Figure 4 x in m 、x m+1 、x m+2 、x m+3 、x m+4 、x m+5 ,y m ,y m+1 ,y m+2 ,y m+3 ,y m+4 ,y m+5 It is in stream data format, and the output of registers R0~R7 is in register data format.

[0129] When the input is stream data, the multiplicand / addend sequence x m and multiplier / addend y m Input to four cascaded D flip-flops (with D -1 Mark), complete the serial-to-parallel conversion;

[0130] When the input is register data, combined with the 4th order Kalman filter calculation process, the two operation vectors have a maximum of 8 values, so 8 registers are set, and the data is stored in parallel in registers R0 to R7. The input stream data and register data are selected according to the input data selected by the selector MUX and stored in the registers M of the four calculators respectively. (0,A) M (0,B) M (1,A) M (1,B) M (2,A) M (2,B) M (3,A) M (3,B) The four calculators choose to perform multiplication or addition / subtraction to complete the calculation according to the input of the selection signals e0 to e3. The calculation results are accumulated through three adders and the final results are output to register S. Under parallel pipeline operation, the final result can be obtained after 5 stages of pipeline.

[0131] In one embodiment, see Figure 5In the actual matrix computing application, the number of VPEs is set, and the general configuration management module (CTM) is designed to manage the scheduling of data between multiple VPEs, which is used to determine which VPE the data in the buffer is given to. The gating rule of CTM is the first-come-first-out principle, which realizes the general acceleration design of multiple types of matrix computing. CTM takes data from the buffer for calculation and distributes it to each VPE. i In, v i The calculated result is stored back into the Buffer through the MUX.

[0132] The input of CTM is the row and column values ​​of the two input matrices of the matrix operation. When it is determined that there is data to be processed in the Buffer, the status of each VPE is first determined, starting from v0 to v Ψ-1 If all VPEs are busy, wait. If any VPE is idle, take out the vector data to be calculated from the Buffer, store it in the VPE and start the calculation. Then continue to loop to check the status of the VPE. If there is an idle VPE, continue to read the stored data and calculate.

[0133] In one embodiment, the impact of the calculation base on the design architecture must also be considered. Single-precision floating-point data can ensure calculation accuracy, but a single floating-point calculation unit occupies more resources. As the bit width of fixed-point data decreases, it is difficult to ensure calculation accuracy. Therefore, it is necessary to reasonably design the fixed-point bit width to reduce resources while ensuring calculation accuracy to accelerate the design of parallel computing architecture.

[0134] In the calculation process of formula (2), due to the existence of iterative calculation, the error will form a certain form of accumulation, and the error of the final result will be affected by the length of the operation time and the single calculation error. The error of a single calculation is mainly caused by the calculation base. In the process of designing hardware acceleration, the selection of floating point, fixed point and the bit width of the fixed point will directly affect the occupation of hardware resources and calculation error.

[0135] In actual engineering applications, floating-point calculations occupy more than ten times the hardware resources of fixed-point calculations. Therefore, fixed-point calculations are the first choice for parallel calculations with limited resources. For a single result in a single matrix multiplication, four multiplications and three additions are required. Assuming that the decimal width in the fixed-point calculation system is σ, the single numerical calculation error is 0.5 σ+1 , the error of a single VPE calculation result can be simply recorded as 3.5 σ+1 In practical applications, σ is tested repeatedly, and the iterative error calculation based on formula (2) is performed according to different bit widths to obtain the minimum bit width σ that meets the accuracy requirements. optimal .

[0136] The technical effects of the present invention are described below in conjunction with specific experimental test results:

[0137] The software was designed based on Xilinx's XC7K325T FPGA chip. The filtering result error test was carried out for different bases, and the resource occupancy and computing performance were compared with the design results in previous literature.

[0138] First, calculate the parallelism parameter:

[0139] By using formulas (3) to (5), we can obtain Figure 1 The optimal parallel computing architecture in , we can get the optimal parallelism results of each sub-process as shown in Table 4. Figure 6 The number of overlapping cycles and the calculation time of each calculation process are given.

[0140] Table 4 Parallelism of each calculation process

[0141]

[0142] The design principle of overlapping cycles between subprocesses is to ensure that the input used by the next-level subprocess is the output of the previous-level subprocess after searching and screening D(n)∈[3,7]. Parallel computing subprocesses such as 1.1 and 1.2, due to simultaneous computing, the parallelism is based on the longest module time. The choice of the parallelism of the second-longest module is based on the computing time not exceeding the longest module computing time. Figure 6 As shown in , the entire architecture design uses a total of 52 VPEs, and the final calculation time is 67 clock cycles. If a 100MHz computing clock is used, the calculation time length is 0.67 microseconds, taking into account the data input and result output delays that will not exceed 1 microsecond.

[0143] Figure 7 The cumulative sum of effective parallelism at different computing moments is given, and the maximum cumulative sum of parallelism is 21 at the 58th clock cycle.

[0144] Then analyze the data test results:

[0145] Figure 8 The observation values ​​and filtering results of the trajectory movement process of the faint target in an approximately flat field of view are given, where the black line represents the true trajectory, the blue dotted line represents the observed trajectory, and the red plus line represents the trajectory after Kalman filtering. There are 10% missed targets in the observed trajectory. Fig. 9 The trajectory errors at different times are given. The average observation error before filtering reached 137 meters, and the average error after filtering was reduced to 85 meters.

[0146] In practical applications, fixed-point system is used for calculation. Figure 8The error of Kalman filter calculation results under different decimal bit widths is given. When the decimal bit width is not less than 17 bits (red arrow), the average error is about 85, which is consistent with the theoretical calculation results. When the decimal bit width is reduced to 10 bits, the average error increases to 285 meters; when it is reduced to less than 8 bits, the error increases to more than 500 meters. Therefore, in the actual design process, selecting a 17-bit decimal bit width can not only ensure the calculation accuracy, but also minimize resource usage.

[0147] Finally, compare resource usage and performance:

[0148] Vivado 2018.1 was used to develop VHDL programs on XC7K325T. The resource usage of the entire Kalman filter calculation is shown in Table 5. Specifically, the designed parallel computing architecture occupies 24.76% of the DSP, 2.7% of the BRAM, 1.78% of the LUTRAM, and 22.51% of the LUT in the FPGA. The maximum executable frequency of the entire FPGA program reaches 160MHz.

[0149] Table 5 Resource usage

[0150] resource Occupancy Proportion LUT 11470 / 50950 22.51% LUTRAM 71 / 4000(kb) 1.78% FF 20378 / 407600 5% BRAM 12 / 445(36kb) 2.7% DSP 208 / 840 24.76%

[0151] The efficiency index is selected, and the efficiency is defined as 1010 / LUTs / Perf. / Freq. / DSP. The normalized index is used to complete the comparison of different existing computing architecture designs.

[0152] Table 6 Comparison of efficiency indicators between the prior art and the embodiments of the present invention

[0153]

[0154] Table 5 gives the specific parameters and performance indicators of three FPGA-based filter designs in recent years. It can be seen that the embodiment of the present invention can run up to 160MHz, but considering the power consumption, stability and other factors under long-term working conditions, it runs at 100MHz; the embodiment of the present invention uses the largest number of DSPs, which is 4 times that of the other three designs, and the LUT and FF occupy the middle. BRAM is generally less used in each design. The final efficiency index of the embodiment of the present invention reached 59.87, which is an order of magnitude higher than the three references that do not exceed 5.8. At the same time, the calculation time does not exceed 0.7 microseconds, which can meet the current high frame rate target tracking needs.

[0155] The FPGA-based parallel acceleration architecture Kalman filter calculation proposed in the present invention is used for real-time solution of the on-orbit space dim target tracking process. A computing unit structure VPE with a unified architecture is established for Kalman filter matrix calculation, and an array computing structure based on VPE is designed to decompose the entire filtering process and solve the optimal parallel pipeline data flow. The prediction error under different fixed-point bit widths is analyzed and counted, and the optimal bit width under the statistical results is given, which is equivalent to the error obtained from the theoretical calculation results. The design and implementation is based on Xilinx's XC7K325T. Compared with previous designs, it has a relatively high maximum executable frequency, less single filter iteration calculation time, and the overall computing efficiency is improved by an order of magnitude.

[0156] The above is a detailed introduction to the method for constructing a parallel accelerated computing architecture for tracking dim targets in orbit provided by the present invention. In this embodiment, specific examples are used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

[0157] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the present embodiments may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown in the present embodiment, but will conform to the widest range consistent with the principles and novel features disclosed in the present embodiment.

Claims

1. A method for constructing a parallel acceleration computing architecture for tracking faint targets in orbit, wherein the steps of tracking faint targets in orbit include: Inter-frame registration, target detection and inter-frame accumulation, trajectory association, target state Kalman filtering and prediction; characterized in that the target state Kalman filtering calculation process is decomposed into different VPE calculation units, including the following steps: S1: Divide the target state Kalman filter calculation process into pipelines to obtain k sub-processes one, and determine the sub-process one for parallel calculation according to the correlation between the sub-processes, and obtain the merged target state Kalman filter calculation process, including n sub-processes two, n<k, and the sub-process two includes the sub-process one for parallel calculation; S2: Taking the minimum calculation cycle of the merged calculation process as the objective function and the optimal overlap between sub-processes 2 as the constraint condition, the parallelism of sub-process 1 is determined in combination with the calculation state and calculation cycle of the VPE calculation unit; S3: Determine the number of VPE computing units according to the parallelism, wherein the VPE computing units are used to perform computing operations on the row vectors and column vectors of the matrix in the first subprocess; S4: The general configuration management module CTM distributes the data in the sub-process 1 to each VPE computing unit for calculation, and the calculation results of the VPE computing unit are stored through the selector MUX.

2. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 1, characterized in that: In S1, the sub-process 1 for parallel calculation is determined according to the calculation relationship between the various matrix variables in the target state Kalman filter calculation process and the correlation between the data streams between the sub-processes.

3. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 1, characterized in that: The S2 specifically includes: S21: Determine the calculation period function of the nth subprocess 2 using the calculation amount C(n), parallelism B(n) and overlapping period D(n) of the nth subprocess 2; calculate the optimal solution set of B(n) with the minimum sum of the calculation period functions of all subprocesses 2 as the objective function; n∈Z + , Z + represents a positive integer; S22: according to the calculation frequency of different clock domains of the data reading and writing and calculation process of the second subprocess, combined with the calculation state and calculation cycle of the VPE calculation unit, obtain the optimal value set of B(n) when the data reading and writing and calculation of the VPE parallel acceleration matrix process in the output result of S21 are balanced; S23: Taking the minimum variance of the sum of the parallelism of the VPE computing units working at any time as a constraint condition, find the single optimal solution of B(n) under the optimal overlap condition between each sub-process; S24: Taking the parallelism between the parallel computing sub-processes 1 being proportional to the amount of computation as a constraint, the parallelism of the parallel sub-process 1 is obtained, thereby obtaining a complete parallel computing architecture.

4. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 3, characterized in that: The S21 specifically includes: The calculation period of the nth subprocess 2 is T(n)=f(C(n) / B(n))-D(n); in, k is the proportionality constant; The whole calculation period T is calculated according to the following formula A The solution of B(n) with the minimum value is:

5. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 3, characterized in that: The S22 specifically includes: The data of subprocess 2 is stored in Buffer; The constraints for access and computation balance are: min|Θ1f1-max{Θ (2,i) f2}| The data access and calculation process in the buffer select different clock domains. f1 is the calculation frequency of clock domain 1, and f2 is the calculation frequency of clock domain 2. Θ1 is the amount of data taken out of the buffer, Θ (2,i) is the i-th VPE computing unit v i calculation cycle.

6. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 3, characterized in that: The S23 specifically includes: The constraint condition on the variance of the sum of the parallelism of the VPE computing units working at any time is: Among them, B t (n) represents the parallelism of the VPE computing unit effective at time t, and n&t represents the effective computing subprocess two at time t.

7. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 3, characterized in that: The S24 specifically includes: The constraint that the degree of parallelism between the subprocesses of parallel computing is proportional to the amount of computing is: Among them, ε=1, 2 represents the subprocess number, B(n.ξ1) represents the parallelism of subprocess ξ1 in the n-th subprocess ξ2, and B(n.ξ2) represents the parallelism of subprocess ξ2 in the n-th subprocess ξ2 that is calculated in parallel with subprocess ξ1; C(n.ξ1) represents the computational amount of subprocess ξ1, and C(n.ξ2) represents the computational amount of subprocess ξ2.

8. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 1, characterized in that: The architecture of the VPE computing unit in S3 executing the computing operation of the matrix row vector and column vector in subprocess 1 includes: The input data calculation of the VPE computing unit includes two types of work: streaming data and register data; When the input is stream data, the data in the row vector and column vector in the N-dimensional matrix multiplication operation and addition and subtraction operation are input to N cascaded D flip-flops to complete the serial-to-parallel conversion; When the input is register data, the data is stored in parallel in the register; The input stream data and register data are input into N calculators respectively according to the selected input data judged by the selector MUX; the N calculators select multiplication, addition or subtraction operation to complete the calculation according to the input of the selection signal, and the calculation results are added step by step through the adder to complete the accumulation operation and output the final result.

9. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 1, characterized in that: The S4 specifically includes: The input of the general configuration management module CTM is the row and column values ​​of the two input matrices of the matrix operation. When it is determined that there is data of the sub-process 2 that needs to be processed in the Buffer, the following steps are executed: S41: Determine the status of each VPE computing unit. If all VPE computing units are busy, wait. If any VPE computing unit is idle, take out the matrix row vector or column vector data to be calculated from the Buffer, store it in the VPE computing unit and start the calculation. S42: Return to S41, continue to loop to detect the status of the VPE computing unit and start the calculation.

10. The method for constructing a parallel accelerated computing architecture for tracking faint targets in orbit according to claim 1, characterized in that: The calculation of a single VPE computing unit is performed using a fixed-point system. According to the different decimal bit widths in the fixed-point calculation system, iterative error calculation based on the target state Kalman filter is performed to obtain the minimum bit width that meets the accuracy requirements.

Citation Information

Patent Citations

  • Rapid implementation method for Kalman filter

    CN107506332A

  • Real-time parallel determination method of high-sampling-rate navigation satellite clock correction

    CN112462396A