ViSAR-based multi-target tracking method and device, storage medium and electronic equipment
By combining foreground target detection and motion background compensation for multi-target tracking, the accuracy problem of multi-vehicle target tracking in ViSAR images is solved, achieving higher tracking accuracy and continuity.
Patent Information
- Application Number
- CN202410728547.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-06-06
AI Technical Summary
Existing multi-vehicle target tracking methods for ViSAR have limited processing accuracy, are prone to missed detection or false detection, and find it difficult to effectively utilize the shadow imaging characteristics of moving targets, resulting in tracking failure.
A multi-target tracking method based on ViSAR is adopted, combining foreground target detection and motion background compensation. Through appearance feature extraction, motion clue extraction, foreground target detection, fast motion background compensation and inter-frame target association modules, multi-target loss function is used for training to improve tracking accuracy.
The accuracy of multi-target tracking in ViSAR images is improved, false associations and false detections are reduced, the recognition and tracking continuity of moving targets are enhanced, and the accuracy of trajectory prediction is improved.
Smart Images

Figure CN118628529B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a multi-target tracking method, device, storage medium and electronic device based on ViSAR. Background Art
[0002] Video Synthetic Aperture Radar (ViSAR) is a new type of synthetic aperture radar system that has emerged in recent years. It has the ability to observe the earth around the clock and in all weather conditions and monitor moving targets. It is of great application value in disaster relief, traffic management, and military reconnaissance. Therefore, high-value multi-vehicle target tracking on land for ViSAR has become a research hotspot.
[0003] Existing ViSAR-oriented visual multi-vehicle target tracking methods mainly eliminate clutter interference and improve small target detection performance through methods such as SAR image preprocessing, multi-scale design of target detection backbone network and feature optimization. These methods mainly follow the idea of target tracking by appearance modeling of moving targets, but are not closely integrated with the shadow imaging characteristics of ViSAR moving targets, resulting in limited overall processing accuracy and prone to missed detections or false detections. Summary of the Invention
[0004] The embodiments of the present application provide a ViSAR-based multi-target tracking method, device, storage medium, and electronic device, which can improve the multi-target tracking accuracy of ViSAR images.
[0005] The present invention provides a multi-target tracking method based on ViSAR, including:
[0006] Obtain the current frame ViSAR image, the previous frame ViSAR image and the previous frame target detection results;
[0007] Inputting the current frame ViSAR image, the previous frame ViSAR image, and the previous frame target detection result into a trained multi-target tracking model to obtain the current frame target tracking result; wherein the multi-target tracking model includes an appearance feature extraction module, a motion clue extraction module, a foreground target detection module, a motion background fast compensation module, and an inter-frame target association module;
[0008] Inputting the current frame ViSAR image, the previous frame ViSAR image, and the previous frame target detection result into a trained multi-target tracking model to obtain the current frame target tracking result includes:
[0009] Inputting the current frame ViSAR image and the previous frame ViSAR image into the appearance feature extraction module to perform feature extraction to obtain a previous frame appearance feature map and a current frame appearance feature map;
[0010] inputting the previous frame appearance feature map and the current frame appearance feature map into the motion clue extraction module to obtain an offset matrix;
[0011] inputting the previous frame appearance feature map, the current frame appearance feature map and the offset matrix into the foreground target detection module to obtain a current frame target detection result, and dividing the current frame target detection result into a first observation set and a second observation set;
[0012] inputting the previous frame target detection result and the offset matrix into the motion background fast compensation module to predict and correct a current frame position;
[0013] inputting the corrected current frame position into an inter-frame target association module respectively with the first observation set and the second observation set to obtain a current time multi-target tracking result.
[0014] Further, the above-mentioned ViSAR-based multi-target tracking method, wherein the motion clue extraction module comprises a linear self-attention layer, a linear cross-attention layer and a cost volume;
[0015] inputting the previous frame appearance feature map and the current frame appearance feature map into the motion clue extraction module to obtain an offset matrix, comprising:
[0016] flattening the previous frame appearance feature map and the current frame appearance feature map respectively to obtain a first flattened feature map and a second flattened feature map;
[0017] inputting the first flattened feature map and the second flattened feature map into the linear self-attention layer to obtain a first spatial semantic alignment feature map and a second spatial semantic alignment feature map;
[0018] inputting the first spatial semantic alignment feature map and the second spatial semantic alignment feature map into the linear cross-attention layer to obtain a previous frame feature embedding vector set and a current frame feature embedding vector set;
[0019] inputting the previous frame feature embedding vector set and the current frame feature embedding vector set into the cost volume to obtain an offset matrix.
[0020] Further, the above-mentioned ViSAR-based multi-target tracking method, wherein the inputting the previous frame appearance feature map, the current frame appearance feature map and the offset matrix into the foreground target detection module to obtain a current frame target detection result comprises:
[0021] fusing the previous frame appearance feature map and the offset matrix to obtain a current frame conversion feature;
[0022] fuse the current frame conversion feature and the current frame appearance feature map to obtain an enhanced current frame feature map;
[0023] obtain the current frame target detection result according to the enhanced current frame feature map.
[0024] Further, the above-mentioned ViSAR-based multi-target tracking method, wherein the dividing the current frame target detection result into a first observation set and a second observation set comprises:
[0025] calculating a confidence set of the current frame target detection result;
[0026] traversing the confidence set to determine an optimal confidence value;
[0027] determining the first observation set and the second observation set by taking the optimal confidence value as a threshold.
[0028] Further, the above-mentioned ViSAR-based multi-target tracking method, wherein the inputting the previous frame target detection result and the offset matrix into the motion background fast compensation module to predict and correct the current frame position comprises:
[0029] obtaining a background region homonym pair based on the current frame target detection result and the offset matrix;
[0030] calculating an affine transformation matrix based on the background region homonym pair, the affine transformation matrix comprising a linear transformation matrix and a translation transformation matrix;
[0031] inputting the previous frame target detection result and the offset matrix into the motion background fast compensation module to predict the current frame position through Kalman filtering;
[0032] correcting the current frame position based on the linear transformation matrix and the translation transformation matrix.
[0033] Further, the above-mentioned ViSAR-based multi-target tracking method, wherein the inputting the corrected current frame position into the inter-frame target association module to obtain the current frame target tracking result comprises:
[0034] pairing the current frame target detection result and the first observation set through the Hungarian algorithm to obtain a paired successful corrected current frame target detection result, a paired successful first observation set, and updating the position of the target at the current time based on the paired successful first observation set to obtain first position information;
[0035] pair the modified current frame target detection result which is not successfully paired with the second observation set to obtain a successfully paired second observation set, and update the position of the target at the current time based on the successfully paired second observation set to obtain second position information;
[0036] fuse the current frame target detection result which is successfully paired, the first position information and the second position information through Kalman filtering to obtain the multi-target tracking result at the current time.
[0037] Further, the above-mentioned multi-target tracking method based on ViSAR, wherein the training process of the multi-target tracking model is based on a multi-target loss function, and the multi-target loss function is:
[0038] L = λ FG L FG + λ BG L BG + λ Det L Det
[0039] wherein λ FG , λ BG and λ Det are weight values of different loss terms, L FG is a foreground motion loss function, L BG is a background motion self-supervised loss function, and L Det is a detection loss function.
[0040] The embodiment of the application also provides a multi-target tracking device based on ViSAR, comprising:
[0041] an acquisition module configured to acquire a current frame of ViSAR image, a previous frame of ViSAR image and a previous frame of target detection result;
[0042] an appearance feature extraction module configured to perform feature extraction based on the current frame of ViSAR image and the previous frame of ViSAR image to obtain a previous frame of appearance feature map and a current frame of appearance feature map;
[0043] a motion clue extraction module configured to obtain a displacement matrix based on the previous frame of appearance feature map and the current frame of appearance feature map;
[0044] a foreground target detection module configured to obtain a current frame of target detection result based on the previous frame of appearance feature map, the current frame of appearance feature map and the displacement matrix, and divide the current frame of target detection result into a first observation set and a second observation set;
[0045] a motion background fast compensation module configured to predict a current frame of position based on the previous frame of target detection result and perform correction;
[0046] an inter-frame target association module configured to obtain a multi-target tracking result at a current time based on the corrected current frame position and the second observation set.
[0047] The embodiment of the present application further provides a computer readable storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute any one of the above-mentioned ViSAR-based multi-target tracking methods.
[0048] The embodiment of the present application further provides an electronic device, which comprises a processor and a memory, the processor is electrically connected with the memory, the memory is used for storing instructions and data, and the processor is used for the steps in any one of the above-mentioned ViSAR-based multi-target tracking methods.
[0049] The embodiment of the present application provides a ViSAR-based multi-target tracking method, device, storage medium and electronic device, which combines foreground target detection and motion background compensation, predicts a current frame target detection result through a foreground target detection module, predicts a current frame position and corrects it through a motion background fast compensation module, obtains a multi-target tracking result at a current time based on the current frame target detection result and the corrected current frame position, and is closely combined with the imaging characteristics of a ViSAR moving target shadow, so that the multi-target tracking precision is improved. BRIEF DESCRIPTION OF DRAWINGS
[0050] The technical scheme and other beneficial effects of the present application will be apparent through the following detailed description of the specific embodiments of the present application combined with the drawings.
[0051] Figure 1 A flowchart of a ViSAR-based multi-target tracking method provided by the embodiment of the present application is shown.
[0052] Figure 2 Another flowchart of a ViSAR-based multi-target tracking method provided by the embodiment of the present application is shown.
[0053] Figure 3 A structure diagram of a ViSAR-based multi-target tracking device provided by the embodiment of the present application is shown.
[0054] Figure 4 A structure diagram of an electronic device provided by the embodiment of the present application is shown.
[0055] Figure 5 Another structure diagram of an electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0056] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort are within the scope of the present application.
[0057] The existing ViSAR-oriented visual multi-vehicle target tracking method mainly excludes clutter interference and improves the small target detection performance through SAR image preprocessing, multi-scale design of target detection backbone network and feature optimization and other methods. These methods mainly follow the idea of modeling the appearance of moving target shadow, and less enrich and perfect the feature description of the target from the perspective of motion information compensation. In addition, the existing method design fully absorbs the advantages of optical video MOT method, but the combination with the imaging characteristics of ViSAR moving target shadow is not close enough, which leads to limited overall processing precision improvement. Due to the huge difference between the image characteristics of moving targets in optical video and ViSAR, the existing vehicle multi-target tracking method applied to ViSAR may mainly face the following three challenges:
[0058] (1) Individual similarity leads to false association: The Doppler modulation of ViSAR moving target echo is very sensitive to its motion, which causes the shift and defocus of its backscattering body, forming a shadow at its real position, which cannot extract the distinctive texture features like vehicle targets in optical video. Therefore, there is appearance similarity between different vehicle individuals, moving target shadows and static target shadows on ViSAR, which makes the multi-target tracking method relying only on target appearance feature matching prone to false association.
[0059] (2) Appearance time-varying leads to tracking failure: The constantly changing imaging geometry of ViSAR makes the contrast between moving target shadow and background clutter change over time, and is easily affected by the occlusion of backscattering body, which cannot have stable appearance features like vehicle targets in optical video. Therefore, the detection confidence of the same vehicle target on ViSAR varies greatly at different times, and the real target that has appearance changes (such as occlusion and shadow defocus) and false alarm targets also have low detection confidence. This makes the multi-target tracking method using fixed hyperparameters to filter high-value detection boxes for association prone to extreme false detection or missed detection, causing moving target shadow tracking failure.
[0060] (3) Radar multi-angle imaging leads to increased trajectory prediction error: The constantly changing imaging geometry of ViSAR also causes random displacement and global defocusing of the background between frames. Therefore, the Kalman filter based on linear displacement assumption has a deviation in predicting the position of the next time of the vehicle target on ViSAR, and forms a cumulative error, leading to target loss or confusion.
[0061] To address the above issues, embodiments of the present application provide a ViSAR-based multi-target tracking method, device, storage medium, and electronic device. The ViSAR-based multi-target tracking device provided in embodiments of the present application can be integrated into an electronic device, such as a terminal or server. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other device.
[0062] See also Figure 1 and Figure 2 , Figure 1 A flowchart of a multi-target tracking method based on ViSAR provided in an embodiment of the present application is shown below. Figure 2 Another flowchart of a multi-target tracking method based on ViSAR provided in an embodiment of the present application, which is applied to an electronic device, includes the following steps:
[0063] S1, obtain the current frame ViSAR image, the previous frame ViSAR image and the previous frame target detection result.
[0064] Specifically, the current frame ViSAR image and the previous frame ViSAR image are divided into several frames of size h I ×w I Pixel slice image I t-1 and I t (The number of channels, height and width are 3, h I and w I ).
[0065] It should be noted that the target detection result refers to the size, position and category of the target.
[0066] In step S2, the current frame ViSAR image, the previous frame ViSAR image, and the previous frame target detection results are input into the trained multi-target tracking model to obtain the target tracking results for the current frame. The multi-target tracking model includes an appearance feature extraction module, a motion cue extraction module, a foreground target detection module, a motion background fast compensation module, and an inter-frame target association module.
[0067] The application combines foreground target detection and motion background compensation, predicts the target detection result of the current frame through the foreground target detection module, predicts the current frame position and corrects it through the motion background fast compensation module, obtains the multi-target tracking result at the current moment based on the target detection result of the current frame and the corrected current frame position, is closely combined with the ViSAR moving target shadow imaging characteristics, and improves the multi-target tracking precision. The method is a multi-task unified training and reasoning method for ViSAR vehicle moving target detection, tracking and implicit image stabilization, has the characteristics of high coupling and simple use.
[0068] Step S2 includes the following steps:
[0069] S21, input the current frame ViSAR image and the previous frame ViSAR image into the appearance feature extraction module for feature extraction to obtain the previous frame appearance feature map and the current frame appearance feature map.
[0070] Specifically, the appearance feature module is a DLA-34 backbone network, and the slice images I t-1 and I t are input into the DLA-34 backbone network for feature extraction to obtain the previous frame appearance feature map F t-1 and the current frame appearance feature map F t (channel number, height and width are 64, h F and w F respectively).
[0071] S22, input the previous frame appearance feature map and the current frame appearance feature map into the motion clue extraction module to obtain the offset matrix.
[0072] In an embodiment, the motion clue extraction module includes a linear self-attention layer, a linear cross-attention layer and a cost volume.
[0073] Step S22 specifically includes the following steps:
[0074] S221, flatten the previous frame appearance feature map and the current frame appearance feature map respectively to obtain the first flattened feature map and the second flattened feature map.
[0075] Specifically, the previous frame appearance feature map F t-1 and the current frame appearance feature map F t are respectively unfolded according to the channel dimension to obtain the unfolded first flattened feature map and the second flattened feature map (channel number and vector length are 64 and h F x w F respectively).
[0076] S222: Input the first flattened feature map and the second flattened feature map into a linear self-attention layer to obtain a first spatial semantic alignment feature map and a second spatial semantic alignment feature map.
[0077] Specifically, in the linear self-attention layer, the first flattened feature map and the second flattened feature map are processed by the first formula and the second formula respectively. The first formula and the second formula are:
[0078]
[0079] in, is the first spatial semantic alignment feature map in the flattened state, is the second spatial semantic alignment feature map in the flattened state, φ(·)=Elu(·)+1, Elu represents the Elu activation function, and ρ represents the sinusoidal position encoding adopted by the target detection method DETR.
[0080] S223, input the first spatial semantic alignment feature map and the second spatial semantic alignment feature map into the linear cross attention layer to obtain the previous frame feature embedding vector set and the current frame feature embedding vector set.
[0081] Specifically, the first spatial semantic alignment feature map and the second spatial semantic alignment feature map are processed respectively by the third formula and the fourth formula, and the third formula and the fourth formula are respectively:
[0082]
[0083] in is the feature embedding vector set of the previous frame, The feature embedding vector set for the current frame (the number of channels and vector length are 64 and h respectively) F ×w F ).
[0084] Will and Restore them to their original shapes and obtain a set of feature embedding vectors aligned with spatiotemporal semantics (the number of channels, height and width are 64, h F and w F ) is recorded as E t-1 and E t .
[0085] S224: Input the feature embedding vector set of the previous frame and the feature embedding vector set of the current frame into the cost volume to obtain an offset matrix.
[0086] E t-1 and E t Input into the cost volume, and calculate the offset matrix between the previous frame pixel and the current frame pixel through the Cost Volume (matching cost) method Specifically, for E t-1 and E t Performing Attention calculation (similar to cross-correlation, matrix multiplication of the feature maps of the previous frame and the current frame is performed to obtain the similarity of each pixel), we can obtain the pixel of the previous frame with the maximum similarity among each pixel of the current frame, thereby forming pixel pairs and their positions, and obtaining the offset matrix.
[0087] The motion clue extraction module provided in this application uses spatiotemporal context information to further improve the individual feature recognition, obtain more accurate dense optical flow, and enhance the association and detection of moving targets.
[0088] S23, input the appearance feature map of the previous frame, the appearance feature map of the current frame and the offset matrix into the foreground target detection module to obtain the target detection result of the current frame, and divide the target detection result of the current frame into a first observation set and a second observation set.
[0089] In one embodiment, the appearance feature map of the previous frame, the appearance feature map of the current frame, and the offset matrix are input into a foreground object detection module to obtain the object detection result of the current frame, including the following steps:
[0090] S231, fusing the appearance feature map of the previous frame with the offset matrix to obtain the conversion feature of the current frame;
[0091] S232, fusing the current frame conversion feature with the current frame appearance feature map to obtain an enhanced current frame feature map;
[0092] S233, obtaining the current frame target detection result according to the enhanced current frame feature map.
[0093] In one embodiment, dividing the current frame target detection result into a first observation set and a second observation set includes the following steps:
[0094] S234, calculating the confidence level set of the target detection result of the current frame.
[0095] Specifically, calculate the current frame target detection result D t The confidence level C t , C t Sort in descending order to get the confidence set C t_des .
[0096] S235, traverse the confidence set and determine the optimal confidence value.
[0097] Assume that there is a confidence classification threshold at the current moment The confidence set C t_des Divided into two categories, namely and The mean confidence levels of these two categories are and At the same time, the target detection results are divided into the first observation set D_high t and the second observation set D_low t The probabilities of and The between-class variance of confidence σ t 2 Expressed as:
[0098]
[0099] According to the above formula, traverse the confidence set C t_des All values of , when the between-class variance σ t 2 When it is maximum, the confidence classification threshold is obtained The optimal confidence value of
[0100] S236: Using the optimal confidence value as a threshold, determine the first observation set and the second observation set.
[0101] Specifically, according to the optimal value of confidence As the threshold, determine the first observation set D_high with good observation status t and the second observation set D_low with poor observation status t .
[0102] The foreground target detection module provided in this application performs spatiotemporal consistency judgment and secondary recall on potential targets with poor observation status based on the detection confidence, improves the tracking continuity of targets with changing appearance, and designs an adaptive mining method for observation status to alleviate the poor tracking effect caused by unreasonable hyperparameter settings.
[0103] S24, inputting the target detection result of the previous frame and the offset matrix into the moving background fast compensation module, predicting the current frame position and making corrections.
[0104] In one embodiment, step S24 includes the following steps:
[0105] S241, obtaining the same-name point pairs in the background area based on the current frame target detection result and the offset matrix.
[0106] Specifically, based on the current frame target detection result D t With the offset matrix (The number of channels, height and width are 2, h F and w F ), extract D t Offset outside the region (2 and N, respectively) and then the same point pairs of adjacent two frames of background region are obtained (N, 2 and 2, respectively).
[0107] It is to be noted that the D t The offset outside the region Specifically, the background pixel offset, and the same point pair refers to the same point on the ground in different images.
[0108] S242, the affine transformation matrix is calculated based on the same point pairs of the background region, and the affine transformation matrix includes a linear transformation matrix and a translation transformation matrix.
[0109] Specifically, the affine transformation matrix is calculated by the RanSAC algorithm The affine transformation matrix includes a linear transformation matrix M 2×2 and a translation transformation matrix T 2×1 , and the specific formula is as follows:
[0110]
[0111] S243, the target detection result of the previous frame and the offset matrix are input into the motion background fast compensation module, and the current frame position is predicted by Kalman filtering.
[0112] Specifically, the current frame position X(k│k-1) is predicted by Kalman filtering, and its covariance P(k│k-1) is calculated.
[0113] S244, the current frame position is corrected based on the linear transformation matrix and the translation transformation matrix.
[0114] Specifically, the linear transformation matrix M 2×2 and the translation transformation matrix T 2×1 are used to correct X(k│k-1) and P(k│k-1), and the corrected current frame prediction value and its covariance
[0115] Specifically, the correction is performed by the following formula:
[0116]
[0117] S25, the corrected current frame position is input into the inter-frame target association module with the first observation set and the second observation set respectively to obtain the multi-target tracking result at the current time.
[0118] In one embodiment, step S25 includes the following steps:
[0119] S251, pair the current frame target detection result with the first observation set through the Hungarian algorithm to obtain a successfully paired corrected current frame target detection result and a successfully paired first observation set, and update the target's current position based on the successfully paired first observation set to obtain the first position information.
[0120] Specifically, calculate the corrected current frame position With the first observation set D_high t The intersection-and-union ratio (the ratio of intersection and union) is used as the allocation cost to perform pairwise pairing through the Hungarian algorithm, where D_high is the first observation set without pairing. t Create a new trajectory, denoted as The first observation set D_high that is successfully paired t Used to update the position of the old track at the current moment, recorded as the first position information Keep the corrected current frame position without pairing Recorded as
[0121] S252, pairing the unpaired corrected current frame target detection result with the second observation set to obtain a successfully paired second observation set, and updating the target's current position based on the successfully paired second observation set to obtain second position information.
[0122] Specifically, calculate With the second observation set D_low t The intersection-over-union ratio of , and as the allocation cost, pairwise pairing is performed by the Hungarian method, where the second observation set D_low without pairing is t Delete the second observation set D_low that is successfully paired t Used to update the position of the old track at the current moment, recorded as the second position information Keep unpaired Recorded as
[0123] S253: The successfully paired current frame target detection result, the first position information, and the second position information are fused through Kalman filtering to obtain a multi-target tracking result at the current moment.
[0124] Specifically, the continuous Thr TMN Frames not paired Delete, where Thr TMN It is the threshold used to determine whether a track needs to be deleted. For example, if an existing track has not been associated with a target for N consecutive frames, the track can be deleted (indicating that the target may have disappeared).
[0125] The modified current frame target detection result of successful pairing The first position information as a prediction value And the second position information The optimal estimation value T1 is obtained by fusion calculation through Kalman filtering as an observation value t And T2 t , T1 t And T2 t The union T i is the final current time multi-target tracking result.
[0126] The background motion fast compensation module provided in the application fully utilizes the background optical flow in the cost volume compared with the traditional method, corrects the trajectory error through affine matrix parameter estimation and coordinate system conversion, and designs a self-supervised loss function based on mutual correlation pixels to constrain the background optical flow generation, thereby improving the multi-target tracking accuracy under the camera motion condition.
[0127] The above is the application process of the multi-target tracking model, and the training process of the multi-target tracking model is briefly introduced as follows:
[0128] (1) ViSAR image training samples and corresponding label information;
[0129] (2) input the ViSAR image sample into the multi-target tracking model to obtain the predicted current time multi-target tracking result;
[0130] (4) calculate the multi-target loss function according to the predicted current time multi-target tracking result and the label true value of the training sample, and iteratively train the multi-target tracking model based on the multi-target loss function to obtain the trained multi-target tracking model.
[0131] In an embodiment, the multi-target loss function is:
[0132] L=λ FG L FG +λ BG L BG +λ Det L Det
[0133] Wherein, λ FG , λ BG And λ Det are the weight values of different loss terms, L FG is the foreground motion loss function, L BG is the background motion self-supervised loss function, and L Det is the detection loss function.
[0134] In the foreground motion loss function L FGIn the current frame feature map, for the pixel point (i, j), the corresponding cost volume is Indicates the similarity between the pixel (i, j) of the feature map of the current frame and all the pixels of the feature map of the previous frame. Apply the two maximum pooling kernels MaxPool to CV_FG respectively. i,j , the pooling kernel sizes are h F ×1 and 1×w F , and after the softmax function σ is calculated, the cost vectors in the two directions can be obtained respectively and They respectively represent the cost vector formed by the maximum value of each column vector in the matrix and the cost vector formed by the maximum value of each horizontal quantity in the matrix.
[0135] Indicates the probability that the target of the pixel point (i, j) of the feature map of the current frame is at the column maximum vector position (0, l), Indicates the probability that the target of the pixel point (i, j) of the feature map of the current frame is at the row maximum vector position (k, 0).
[0136] Remember Y_FG ijkl is the supervision label of the foreground motion. If the target of the pixel point (i, j) of the feature map of the current frame appears at the position (k, l) of the previous frame, then Y_FG ijkl =1, otherwise Y_FG ijkl = 0. Therefore, L FG It can be expressed in the form of FocalLoss function:
[0137]
[0138] in, β is a hyperparameter of FocalLoss.
[0139] In the background motion self-supervision loss function L based on cross-correlated pixels BG In the example, similar to the foreground motion loss function L FG , which can be expressed in the form of FocalLoss function:
[0140]
[0141] in, γ is a hyperparameter of FocalLoss.
[0142] In the detection loss function L Det In the figure, the target position loss L of the current frame is included H , size loss L WH and classification loss L CLS ; Among them, the target position loss L HIn the form of FocalLoss function, β is a hyperparameter, and the target size loss L WH and the classification loss L CLS In the form of L1Los function:
[0143] L Det = λ H L H + λ WH L WH + λ CLS L CLS
[0144]
[0145] In the formula, H represents the output target position feature map, WH represents the output target size feature map, CLS represents the output target category feature map, Y WH and Y CLS are the training supervision labels of the size and classification output feature maps, respectively, Y ij is the training supervision label of the target position, Y ij = 1 if the current frame pixel point (i, j) is a target, otherwise Y ij = 0. Wherein, λ H , λ WH and λ CLS are the weight values of different loss terms.
[0146] Moreover, in step S24, based on the cross-correlation pixel background motion self-supervision loss function L BG Need background area homonymic point pair (i, j) and (k, l), its screening method needs to meet two conditions: one is that the correlation degree P of the two points is higher than the threshold θ, and the other is that the two points belong to the mutual neighbor relationship, that is, the most relevant points of the two points on the other frame are each other. Denote Y_BG ijkl as the supervision label of the background motion, if the current frame background pixel (i, j) and the (k, l) of the previous frame belong to the homonymic point, then Y_BG ijkl = 1, otherwise Y_BG ijkl = 0.
[0147] According to the method described in the above embodiment, this embodiment will be further described from the perspective of a ViSAR-based multi-target tracking device. The ViSAR-based multi-target tracking device can be implemented as an independent entity or integrated in an electronic device, which can be a terminal, a server, or the like. The terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro processing box, or other devices, etc.
[0148] Please refer to Figure 3 ,Figure 3 The application further provides a device for tracking multiple targets based on ViSAR, which is applied to an electronic device and can comprise:
[0149] a obtaining module, configured to obtain a current frame of ViSAR image, a previous frame of ViSAR image and a previous frame of target detection result;
[0150] an appearance feature extraction module, configured to perform feature extraction based on the current frame of ViSAR image and the previous frame of ViSAR image to obtain a previous frame of appearance feature map and a current frame of appearance feature map;
[0151] a motion clue extraction module, configured to obtain a displacement matrix based on the previous frame of appearance feature map and the current frame of appearance feature map;
[0152] a foreground target detection module, configured to obtain a current frame of target detection result based on the previous frame of appearance feature map, the current frame of appearance feature map and the displacement matrix, and divide the current frame of target detection result into a first observation set and a second observation set;
[0153] a motion background fast compensation module, configured to predict a current frame position based on the previous frame of target detection result and perform correction;
[0154] an inter-frame target association module, configured to obtain a current time multiple target tracking result based on the corrected current frame position and the second observation set.
[0155] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be combined as the same or several entities, and the specific implementation of each of the above modules and / or units can refer to the method embodiments above, and the beneficial effects that can be achieved can also refer to the beneficial effects in the method embodiments above, which will not be repeated here.
[0156] In addition, the application further provides an electronic device, which can be a computer, a tablet computer or the like. As shown in Figure 4 The electronic device 400 comprises a processor 401 and a memory 402. The processor 401 is electrically connected with the memory 402.
[0157] The processor 401 is the control center of the electronic device 400, and connects each part of the electronic device through various interfaces and lines. By running or loading the application program stored in the memory 402 and calling the data stored in the memory 402, the processor 401 can execute various functions of the electronic device and process data, thereby monitoring the whole electronic device.
[0158] In the embodiment, the processor 401 in the electronic device 400 loads the instructions corresponding to the processes of one or more application programs into the memory 402 and runs the application programs stored in the memory 402 by the processor 401 to implement various functions according to the following steps:
[0159] obtaining a current frame of ViSAR image, a previous frame of ViSAR image and a previous frame of target detection result;
[0160] inputting the current frame of ViSAR image, the previous frame of ViSAR image and the previous frame of target detection result into a trained multi-target tracking model to obtain a current frame of target tracking result; wherein the multi-target tracking model comprises an appearance feature extraction module, a motion clue extraction module, a foreground target detection module, a motion background fast compensation module and an inter-frame target association module;
[0161] inputting the current frame of ViSAR image, the previous frame of ViSAR image and the previous frame of target detection result into a trained multi-target tracking model to obtain a current frame of target tracking result comprises:
[0162] inputting the current frame of ViSAR image and the previous frame of ViSAR image into the appearance feature extraction module to extract features to obtain a previous frame of appearance feature map and a current frame of appearance feature map;
[0163] inputting the previous frame of appearance feature map and the current frame of appearance feature map into the motion clue extraction module to obtain a displacement matrix;
[0164] inputting the previous frame of appearance feature map, the current frame of appearance feature map and the displacement matrix into the foreground target detection module to obtain a current frame of target detection result, and dividing the current frame of target detection result into a first observation set and a second observation set;
[0165] inputting the previous frame of target detection result and the displacement matrix into the motion background fast compensation module to predict and correct a current frame position;
[0166] inputting the corrected current frame position into the first observation set and the second observation set into the inter-frame target association module to obtain a current time multi-target tracking result.
[0167] The electronic device can implement the steps in any embodiment of the multi-target tracking method based on ViSAR provided in the embodiments of the application, and thus can implement the beneficial effects of any multi-target tracking method based on ViSAR provided in the embodiments of the application. Details are described in the foregoing embodiments, which will not be described here.
[0168] Figure 5 A specific structural block diagram of an electronic device provided by an embodiment of the present application is shown, which can be used to implement the ViSAR-based multi-target tracking method provided in the above embodiments. The electronic device 500 can be a terminal, a server, or the like, wherein the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro processing box, or other devices, etc.
[0169] The RF circuit 510 is configured to receive and send electromagnetic waves, and to convert the electromagnetic waves and electrical signals to each other, so as to communicate with a communication network or other devices. The RF circuit 510 can include various existing circuit elements for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and the like. The RF circuit 510 can communicate with various networks, such as the Internet, an intranet, a wireless network, or other devices through the wireless network. The wireless network can include a cellular telephone network, a wireless local area network or metropolitan area network. The wireless network can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g and / or 802.11n standards), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short message service, and any other suitable communication protocol, even including those not yet developed as of the date of the application.
[0170] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above-described embodiments, and the processor 580 can execute various functions and data processing by running the software programs and modules stored in the memory 520, i.e., realize functions such as front camera shooting, processing of the captured image, and switching of display colors of the display content on the display screen. The memory 520 can include a high-speed random access memory, and can further include a non-volatile memory such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 520 can further include memories disposed remotely with respect to the processor 580, which can be connected to the electronic device 500 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0171] The input unit 530 can be used to receive inputted digital or character information, and generate a keyboard, a mouse, and the like related to user settings and function control.
[0172] The display unit 540 can be used to display information inputted by the user or provided to the user, and various graphical user interfaces which can be constituted by graphics, texts, icons, videos, and any combination thereof. The display unit 540 can include a display panel 541, which can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like.
[0173] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can convert received audio data into an electrical signal, transmit the electrical signal to the speaker 561, and convert the electrical signal into a sound signal outputted by the speaker 561; on the other hand, the microphone 562 can convert a sound signal collected into an electrical signal, and the audio circuit 560 can convert the electrical signal into audio data, output the audio data to the processor 580 for processing, and then transmit the audio data to another terminal through the RF circuit 510, or output the audio data to the memory 520 for further processing. The audio circuit 560 can further include an earphone jack to provide communication between an external earphone and the electronic device 500.
[0174] The electronic device 500 can help the user to receive requests, send information, and the like through the transmission module 570 (e.g., a Wi-Fi module), which provides the user with wireless broadband Internet access. Although the transmission module 570 is shown, it can be understood that it does not belong to the essential components of the electronic device 500, and can be omitted as needed without changing the essence of the application.
[0175] The processor 580 is a control center of the electronic device 500 that uses various interfaces and lines to connect the various parts of the entire mobile phone, performs various functions of the electronic device 500 and processes data by running or executing software programs and / or modules stored in the memory 520 and calling data stored in the memory 520, thereby overall monitoring the electronic device. Optionally, the processor 580 can include one or more processing cores; in some embodiments, the processor 580 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. Understandably, the above-mentioned modem processor can also not be integrated into the processor 580.
[0176] The electronic device 500 further includes a power supply 590 (such as a battery) for supplying power to various components, and in some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so that the power management system can realize functions such as management of charging, discharging, and power consumption management. The power supply 590 can also include one or more direct or alternating power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and any other components.
[0177] Although not shown, the electronic device 500 also includes a camera (such as a front camera or a rear camera), a Bluetooth module, and the like, which are not described here in detail. Specifically, in the present embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory and one or more programs stored in the memory and configured to be executed by one or more processors, wherein the one or more programs include instructions for:
[0178] obtaining a current frame of ViSAR image, a previous frame of ViSAR image and a previous frame of target detection result;
[0179] inputting the current frame of ViSAR image, the previous frame of ViSAR image and the previous frame of target detection result into a trained multi-target tracking model to obtain a current frame of target tracking result; wherein the multi-target tracking model includes an appearance feature extraction module, a motion clue extraction module, a foreground target detection module, a motion background fast compensation module and an inter-frame target association module;
[0180] inputting the current frame of ViSAR image, the previous frame of ViSAR image and the previous frame of target detection result into a trained multi-target tracking model to obtain a current frame of target tracking result includes:
[0181] input the current frame ViSAR image and the previous frame ViSAR image into the appearance feature extraction module to perform feature extraction to obtain a previous frame appearance feature map and a current frame appearance feature map;
[0182] input the previous frame appearance feature map and the current frame appearance feature map into the motion clue extraction module to obtain a displacement matrix;
[0183] input the previous frame appearance feature map, the current frame appearance feature map and the displacement matrix into the foreground target detection module to obtain a current frame target detection result, and divide the current frame target detection result into a first observation set and a second observation set;
[0184] input the previous frame target detection result and the displacement matrix into the motion background fast compensation module to predict and correct a current frame position;
[0185] input the corrected current frame position into the inter-frame target association module to obtain a current time multi-target tracking result.
[0186] In specific implementation, each of the above modules can be implemented as an independent entity, or can be combined as the same or several entities. The specific implementation of each of the above modules can be found in the method embodiments above, and will not be described here.
[0187] Those skilled in the art can understand that all or part of the steps in the above embodiments can be completed by instructions, or by related hardware controlled by the instructions. The instructions can be stored in a computer readable storage medium and loaded and executed by a processor. Therefore, the embodiments of the present application provide a storage medium storing a plurality of instructions, which can be loaded by a processor to execute the steps of any embodiment of the multi-target tracking method based on ViSAR provided by the embodiments of the present application.
[0188] The computer readable storage medium can include a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0189] Since the instructions stored in the storage medium can execute the steps in any embodiment of the multi-target tracking method based on ViSAR provided by the embodiments of the present application, the beneficial effects of any multi-target tracking method based on ViSAR provided by the embodiments of the present application can be achieved, which will be described in detail in the above embodiments and will not be described here.
[0190] The above describes in detail a multi-target tracking method and device based on ViSAR, a storage medium and an electronic device provided by the embodiments of the present application. The principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, the specific implementation manners and application ranges will be changed according to the idea of the present application. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A multi-target tracking method based on ViSAR, characterized in that: The method comprises: Obtain the current frame ViSAR image, the previous frame ViSAR image and the previous frame target detection results; Inputting the current frame ViSAR image, the previous frame ViSAR image, and the previous frame target detection result into a trained multi-target tracking model to obtain the current frame target tracking result; wherein the multi-target tracking model includes an appearance feature extraction module, a motion clue extraction module, a foreground target detection module, a motion background fast compensation module, and an inter-frame target association module; Inputting the current frame ViSAR image, the previous frame ViSAR image, and the previous frame target detection result into a trained multi-target tracking model to obtain the current frame target tracking result includes: Inputting the current frame ViSAR image and the previous frame ViSAR image into the appearance feature extraction module to perform feature extraction to obtain a previous frame appearance feature map and a current frame appearance feature map; Inputting the appearance feature map of the previous frame and the appearance feature map of the current frame into the motion clue extraction module to obtain an offset matrix; Inputting the appearance feature map of the previous frame, the appearance feature map of the current frame, and the offset matrix into the foreground object detection module to obtain a current frame object detection result, and dividing the current frame object detection result into a first observation set and a second observation set; Inputting the target detection result of the previous frame and the offset matrix into the moving background fast compensation module to predict the current frame position and perform correction; Inputting the corrected current frame position, the first observation set, and the second observation set into the inter-frame target association module to obtain the multi-target tracking result at the current moment; The step of inputting the corrected current frame position, the first observation set, and the second observation set into an inter-frame target association module to obtain a current frame target tracking result includes: The current frame target detection result is paired with the first observation set through the Hungarian algorithm to obtain a successfully paired corrected current frame target detection result and a successfully paired first observation set, and the position of the target at the current moment is updated based on the successfully paired first observation set to obtain the first position information; the unsuccessfully paired corrected current frame target detection result is paired with the second observation set to obtain a successfully paired second observation set, and the position of the target at the current moment is updated based on the successfully paired second observation set to obtain the second position information; the successfully paired current frame target detection result, the first position information and the second position information are fused through Kalman filtering to obtain the multi-target tracking result at the current moment.
2. The multi-target tracking method based on ViSAR according to claim 1, characterized in that: The motion clue extraction module includes a linear self-attention layer, a linear cross-attention layer and a cost body; Inputting the appearance feature map of the previous frame and the appearance feature map of the current frame into the motion clue extraction module to obtain an offset matrix, including: Flattening the appearance feature map of the previous frame and the appearance feature map of the current frame respectively to obtain a first flattened feature map and a second flattened feature map; Inputting the first flattened feature map and the second flattened feature map into the linear self-attention layer to obtain a first spatial semantic alignment feature map and a second spatial semantic alignment feature map; Inputting the first spatial semantic alignment feature map and the second spatial semantic alignment feature map into the linear cross attention layer to obtain a previous frame feature embedding vector set and a current frame feature embedding vector set; The previous frame feature embedding vector set and the current frame feature embedding vector set are input into the cost volume to obtain an offset matrix.
3. The multi-target tracking method based on ViSAR according to claim 1, characterized in that: The step of inputting the appearance feature map of the previous frame, the appearance feature map of the current frame, and the offset matrix into the foreground object detection module to obtain the current frame object detection result includes: Fusing the appearance feature map of the previous frame with the offset matrix to obtain the conversion feature of the current frame; fusing the current frame conversion feature with the current frame appearance feature map to obtain an enhanced current frame feature map; The current frame target detection result is obtained according to the enhanced current frame feature map.
4. The multi-target tracking method based on ViSAR according to claim 1, characterized in that: The dividing the current frame target detection result into a first observation set and a second observation set includes: Calculating a confidence level of the target detection result of the current frame; Traversing the confidence set to determine an optimal confidence value; The first observation set and the second observation set are determined by using the optimal confidence value as a threshold.
5. The multi-target tracking method based on ViSAR according to claim 1, characterized in that: The step of inputting the target detection result of the previous frame and the offset matrix into the moving background fast compensation module, predicting the current frame position and performing correction includes: Obtaining pairs of points with the same name in the background area based on the current frame target detection result and the offset matrix; Calculating an affine transformation matrix based on the same-name point pairs in the background area, wherein the affine transformation matrix includes a linear transformation matrix and a translation transformation matrix; Inputting the target detection result of the previous frame and the offset matrix into the motion background fast compensation module, and predicting the current frame position through Kalman filtering; The current frame position is corrected based on the linear transformation matrix and the translation transformation matrix.
6. The multi-target tracking method based on ViSAR according to claim 1, characterized in that: The training process of the multi-target tracking model is based on a multi-target loss function, which is: in, 、 and are the weights of different loss terms, is the foreground motion loss function, is the background motion self-supervision loss function, is the detection loss function.
7. A multi-target tracking device based on ViSAR, wherein the multi-target tracking device based on ViSAR is used to implement the multi-target tracking method based on ViSAR according to claim 1, characterized in that: include: The acquisition module is used to obtain the current frame ViSAR image, the previous frame ViSAR image and the previous frame target detection result; An appearance feature extraction module is used to extract features based on the current frame ViSAR image and the previous frame ViSAR image to obtain a previous frame appearance feature map and a current frame appearance feature map; A motion clue extraction module, configured to obtain an offset matrix based on the appearance feature map of the previous frame and the appearance feature map of the current frame; a foreground object detection module, configured to obtain a current frame object detection result based on the previous frame appearance feature map, the current frame appearance feature map, and the offset matrix, and divide the current frame object detection result into a first observation set and a second observation set; A moving background fast compensation module, configured to predict the current frame position and perform correction based on the target detection result of the previous frame and the offset matrix; The inter-frame target association module is used to obtain the multi-target tracking result at the current moment based on the corrected current frame position and the first observation set and the second observation set.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the ViSAR-based multi-target tracking method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the ViSAR-based multi-target tracking method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Micro-robot high-precision attitude measurement correlation filtering method based on Kalman filtering
CN115900702A
Video SAR moving target shadow detection method based on time sequence appearance feature aggregation
CN116152213A