Video SAR target tracking method and system based on spatiotemporal expansion-accumulation strategy
By batch processing video SAR images and constructing a multi-frame tree structure, combined with adaptive template updating, the multi-peak and interference problems of shadow targets in video SAR target tracking are solved, and high-precision target tracking effect is achieved.
Patent Information
- Application Number
- CN202411568004.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-11-05
AI Technical Summary
In the existing video SAR target tracking technology, the shadow target is small in size and its complex texture features are masked by speckle noise. The complex background makes the tracker susceptible to interference, resulting in multi-peak problems, causing the target tracking frame to drift and be lost.
A method based on spatiotemporal expansion-accumulation strategy is adopted to divide video SAR images into multiple batch packages. Multi-frame prediction-accumulation and screening are performed on each frame, and a multi-frame tree structure is constructed. The target trajectory is found through the trajectory backtracing algorithm, and the adaptive template update method is used to improve the tracking accuracy.
It effectively reduces the probability of target loss, improves the calculation speed and tracking success rate, enhances the ability to recognize shadow targets, reduces the risk of template contamination, and achieves high-precision target tracking.
Smart Images

Figure CN119515918B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video SAR target tracking, and in particular relates to a video SAR target tracking method and system based on a spatiotemporal expansion-accumulation strategy. Background Art
[0002] According to SAR imaging principles, in high-frame-rate video SAR images, the energy of a moving target will experience significant shifts and defocus relative to its true position, casting a shadow at its true location. Therefore, the state information of the moving target's shadow is consistent with the target's true state information. In video SAR images, because there is no phase information from the target's shadow region, the target shadow imaging is unaffected by any distortion effects. Furthermore, the contrast between the target shadow and clutter is independent of the target's radar cross section, meaning that shadows cannot be eliminated or reduced by stealth technology. Due to these advantages, analyzing the characteristics of moving target shadows and developing a series of detection and tracking methods has become a research hotspot in the field of video SAR target tracking.
[0003] Algorithms based on correlation filtering have become popular in video SAR target tracking in recent years. Zhong et al. proposed a joint kernel correlation filter (JKCF) framework for video SAR moving target tracking, combining shadows in video SAR images with the energy in the corresponding range-Doppler spectrum to track small targets. Zhao et al. equipped the tracker with a detection mechanism based on spatiotemporal information and saliency for interference and background clutter. Tian et al. proposed a new BACF algorithm based on appearance-range information-assisted probabilistic data association (AD-PDA), which uses video SAR images to estimate target state.
[0004] TBD technology is a detection and tracking technique for faint targets. Its basic steps are: no threshold or a very low threshold is set for a single frame; then, based on the correlation of the target's motion across frames, data from multiple frames is accumulated. Finally, the target is determined and its estimated track is determined by comparing the data with the set threshold. The dynamic programming-based track-before-detection (DP-TBD) and particle filter-based track-before-detection (PF-TBD) algorithms are two typical algorithms. Their excellent overall performance has made them popular among researchers for weak target tracking. Using the DP-TBD algorithm to process video SAR image sequences allows for the design of a value function based on target amplitude, leveraging target track correlation to accumulate target amplitudes and improve the signal-to-noise ratio of faint targets. Furthermore, it avoids the value function accumulation errors caused by the model's assumed distribution when the video SAR measurement model is unknown. Tian Xiaoqing et al. proposed an improved ES-TBD algorithm to detect and track a fixed number of slow-moving targets. Based on the DP-TBD algorithm, Qin et al. proposed a JP-DP-TBD algorithm based on a joint processing strategy. By analyzing the target's position and radial velocity information, this method can screen the state candidate areas and retain more states that conform to the target's motion laws, thereby tracking multiple maneuvering targets.
[0005] While the aforementioned trackers achieve excellent performance by striking a balance between tracking accuracy and speed, the following factors still limit their application in video SAR target tracking. First, the shadow target is small, and its complex texture features are masked by speckle noise. Second, the surrounding background is complex, and the tracker is likely to be distracted by man-made structures or similar objects in the background. All of these situations result in multiple peaks in the correlation filter response graph, with the actual target not being at the highest peak. This is a major challenge in tracking small shadow targets in video SAR. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to address the deficiencies in the above-mentioned prior art and provide a video SAR target tracking method and system based on a spatiotemporal expansion-accumulation strategy, which is used to solve the technical problem of target tracking frame drift in video SAR target tracking and achieve high-precision tracking of weak shadow moving targets in video SAR.
[0007] The present invention adopts the following technical solutions:
[0008] The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy includes the following steps:
[0009] S1、Contain M The video SAR image dataset of the frame is divided into N Batch packages, each containing K Frame video SAR image;
[0010] S2: Manually select the tracking target in the initial frame of the first batch package obtained in step S1, and extract the features of the initial position using the background-aware correlation filter. F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial target candidate box ;
[0011] S3, read the second frame image in the first batch package obtained in step S1, based on the candidate frame of the first frame obtained in step S2 Define the target search area for the second frame , The background-aware correlation filter is divided into multiple test samples, each test sample and Calculate similarity and obtain a response matrix;
[0012] S4. Select the response matrix with the maximum response value from the response matrix obtained in step S3 and from Select S The maximum response value, each response value corresponds to a predicted position of the target; extract features for each predicted position and update the corresponding template, and finally generate S target candidate boxes;
[0013] S5, for the subsequent k Frame No. j Candidate boxes , repeat steps S3 to S4 to obtain target candidate boxes; K Frame, repeat step S3 to step S4 to get target candidate frames, at this time all target candidate frames are generated to form a multi-frame tree structure;
[0014] S6, using the trajectory backtracking algorithm to reversely search for the estimated position of each frame in the tree structure obtained in step S5, K Frame No. Candidate frames are backtracked, and a total of candidate paths of target trajectories;
[0015] S7. Calculate the value function corresponding to each candidate path obtained in step S6, select the candidate path with the largest value function as the estimated path of the target in a batch package, and output the estimated state of the target in each frame.
[0016] Preferably, in step S1, the last frame of the previous batch package and the first frame of the next batch package are the same frame, the target initial state of the first frame of the next batch package is the target estimated state of the previous batch package after screening by the trajectory backtracking module, and the corresponding candidate box information including the position-size state vector, features, templates, and response values all come from the last frame of the previous batch package.
[0017] Preferably, in step S2, the feature F 1 is a 32-dimensional feature, including 31-dimensional directional histogram features and grayscale features. The target tracking frame is , The superscript indicates the number of the current candidate box, and the subscript indicates the frame number of the image in the current batch package.
[0018] Preferably, in step S3, the target search area The location is based on Position generation, target search area The size is right The width and height of the position-size state vector are expanded Double generation, is an adjustable parameter ranging from [4,8]; the resulting response matrix is the matrix with the maximum value among all response matrices.
[0019] Preferably, in step S4, each response value corresponds to the predicted position of the target, which satisfies the following relationship:
[0020]
[0021] in,( ) is the response value The corresponding s The coordinates of the predicted target position.
[0022] Preferably, in step S5, the tree structure constituting the multiple frames is specifically:
[0023] S501, initial target candidate frame and the position-size state vector of the candidate box ;
[0024] S502: Generate a search area in the second frame , The background-aware correlation filter is used to crop the sample into multiple test samples, each with the same size and dimension. Same, each test sample and Perform cross-correlation similarity calculation and obtain a response matrix, and select the response matrix with the largest response value , It is a two-dimensional matrix. The value in each grid is between 0 and 1, and each grid corresponds to a position coordinate of the current frame.
[0025] S503. Generate response matrix from search area Select the front S The maximum response value is obtained S target candidate boxes, the size of each candidate box is determined by the scale adjustment factor of the background perception correlation filter;
[0026] S504: For any frame in the first batch package k -1, from k -Select one from all candidate frames in 1 frame , After repeating step S502 and step S503, k Frame Generation candidate boxes, represented as , ; And so on, in the last frame of the first batch ( k=K )Total The candidate box of a target is represented as , ; All candidate box information in the first batch package is stored in the buffer for subsequent calculation. n When >1, the candidate box information of the initial frame is directly retrieved from the target estimation result of the last frame of the previous batch package, and the target candidate box generation process of the remaining frames is the same as that of the first batch package.
[0027] Preferably, an adaptive template updating method is used to calculate the first batch of initial frames. n The grayscale mean difference between the first frame of the batch package and the initial frame of the first batch package is used to adjust the online learning rate variables.
[0028] Preferably, the online learning rate for:
[0029]
[0030] in, is the grayscale difference index, is the threshold, and is the value of the online learning rate.
[0031] Preferably, in step S7, Select the function with the largest value from the candidate paths The path is:
[0032]
[0033] in, Represents the set of all candidate paths in a batch package, is the estimated position of the target for all frames estimated by the algorithm, Estimate the target position for the Kth frame, All predicted positions for all frames of the target.
[0034] In a second aspect, an embodiment of the present invention provides a video SAR target tracking system based on a spatiotemporal expansion-accumulation strategy, comprising:
[0035] Divide the module into a M The dataset of frame video SAR images is divided into N Batch packages, each containing K Frame image;
[0036] The selection module manually selects the tracking target in the initial frame of the first batch package obtained by the division module, and the background perception correlation filter extracts the features of the initial position F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial candidate box ;
[0037] The reading module reads the second frame image in the first batch package obtained by the segmentation module, based on the candidate frame of the first frame obtained by the selection module Define the target search area for the second frame , The background-aware correlation filter is divided into multiple test samples, each test sample and Calculate similarity and obtain a response matrix;
[0038] Update module, select the response matrix with the maximum response value from the response matrix obtained by the reading module and from Select S The maximum response value, each response value corresponds to a target prediction position; extract features for each prediction position and update the corresponding template, and finally generate S target candidate boxes;
[0039] Repeat module, for the subsequent k Frame No. j Candidate boxes , after repeatedly reading the module and updating the module, we can get candidate boxes;K Frame, repeatedly read the module and update the module to get candidate frames, at this time all candidate frames are generated to form a multi-frame tree structure;
[0040] The backtracking module uses the trajectory backtracking module to reversely search for the target's estimated position in each frame of the tree structure obtained by the repeating module, starting from the first K Frame No. The target candidate frames are backtracked, and a total of candidate paths of target trajectories;
[0041] The tracking module calculates the value function corresponding to each candidate path obtained by the backtracking module, selects the candidate path with the largest value function as the estimated path of the target in a batch package, and outputs the estimated state of the target in each frame.
[0042] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy are implemented.
[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy.
[0044] Compared with the prior art, the present invention has at least the following beneficial effects:
[0045] A video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy aims to solve the multi-peak problem caused by target occlusion or interference from similar targets in correlation filtering. Inspired by the TBD idea, a new tracking mode is designed, which changes the original single-frame prediction-single-frame update method to a multi-frame prediction-accumulation and screening-multi-frame update method, successfully reducing the probability of target loss.
[0046] Furthermore, dividing the video SAR images into multiple batch packages has the following main advantages: first, it reduces the computational complexity of the algorithm and improves the calculation speed of the algorithm; second, by accumulating multiple frames of information in each batch package according to the batch images, it can effectively avoid the problems of missed tracking and wrong tracking caused by the loss of single-frame target information, thereby improving the success rate of target tracking.
[0047] Furthermore, the use of 32-dimensional features can more accurately and comprehensively characterize the target, making it easier for the tracker to identify the target.
[0048] Furthermore, within a batch, searching for the target position in the current frame based on the predicted target position in the previous frame allows for more efficient target finding, as the target moves less distance between two consecutive frames. Furthermore, the maximum response value indicates the highest correlation between the test sample and the target template, and therefore the highest probability that the test sample is the target.
[0049] Furthermore, the method of selecting the location corresponding to the maximum response value from a response matrix as the target prediction location is abandoned. This is because the area where similar targets are located may also have the maximum response value. If the tracker mistakenly tracks similar targets, the template will be contaminated, and the target will eventually be lost. To reduce the probability of template contamination, S response values are selected and their corresponding locations are obtained as candidate target locations. As much as possible, the candidate locations contain the target's true location.
[0050] Further, such as Figure 2 As shown, the target prediction positions of all frames within a batch together form a multi-frame tree. For each target prediction position at the bottom layer of the multi-frame tree, the backtracking module can find a target prediction trajectory. The target's true motion trajectory is hidden in these predicted trajectories. This multi-frame tree structure covers the target's true trajectory as comprehensively as possible, reducing the probability of target loss. Properly setting the total number of frames in the tree structure and the number of next-level branches for each target prediction position can both control the algorithm's computational complexity and improve the success rate of target tracking.
[0051] Furthermore, in correlation filtering, the response value represents the similarity between the predicted target and the real target. The value function of the candidate path in the present invention refers to the sum of the response values of the target predicted position under the path. The maximum value function indicates that the target predicted position corresponding to the candidate path is closest to the real target position. If there is interference from a similar target in a certain frame, resulting in the maximum response value at that position, the traditional correlation filtering algorithm will update the template according to the predicted position here, causing the template to be contaminated, and eventually the tracker will lose the target. However, in the present invention, the possible positions of all targets in a batch package are first predicted, and then the corresponding response values are summed, which can effectively avoid the occurrence of the above situation.
[0052] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0053] In summary, the algorithm proposed in the present invention uses a method of accumulating multi-frame information to track weak shadow targets in video SAR. After accumulating multi-frame image information, it obtains enough real target information and can successfully distinguish the target from clutter. A multi-frame tree structure is built to deal with the multi-peak problem in the correlation filter response graph. By generating multiple candidate frames, the real target is covered as much as possible. Even if the candidate frame of the real target in a certain frame does not have the maximum response value, the real trajectory of the target can still be successfully found after the tree structure is expanded in the time domain and space domain. The online learning rate of the correlation filter is flexibly adjusted based on the change of the target appearance, and the appropriate online learning rate is adaptively selected to update the template, which can improve the algorithm's ability to adapt to different tracking scenarios. It has the ability to continuously and stably track targets, and has achieved excellent tracking accuracy and success rate, as well as very small positioning error, and has achieved excellent performance compared with other algorithms based on correlation filtering.
[0054] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Schematic diagram of batch division;
[0056] Figure 2 This is a schematic diagram of the tree structure.
[0057] Figure 3 It is a framework diagram of the method of the present invention;
[0058] Figure 4 A schematic diagram of a computer device provided in accordance with an embodiment of the present invention;
[0059] Figure 5 The present invention is a block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0061] In the description of the present invention, it is to be understood that the terms “include” and “comprise” indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0062] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0063] It should be further understood that the term "and / or" as used in the present specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects are in an "or" relationship.
[0064] It should be understood that although the terms "first," "second," and "third" may be used to describe preset ranges in embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are merely used to distinguish one preset range from another. For example, without departing from the scope of embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0065] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0066] The accompanying drawings illustrate various schematic diagrams of structures according to embodiments disclosed herein. These figures are not drawn to scale; for clarity, some details are exaggerated and some details may be omitted. The shapes of the various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art may design regions / layers with different shapes, sizes, and relative positions as needed.
[0067] The present invention provides a video SAR target tracking method based on a spatiotemporal expansion-accumulation strategy, which includes training a target detector; using the target detector to detect whether the predicted position of the next frame is a target; and using the new detection result to update the training set and then update the target detector.
[0068] The BACF target tracking process includes two key stages: filter training and target detection.
[0069] Filter training
[0070] BACF applies a cyclic shift operator and a cropping operator to the image, and uses the cyclic shifted and then twice cropped samples as the real training samples. Then BACF obtains an optimal filter by minimizing the following objectives: h :
[0071]
[0072] in, is an ideal two-dimensional Gaussian response centered on the target, and denote the samples and the correlation filter, respectively. channels, is the number of feature channels, the superscript represents the transpose operation, Indicates the sample Conduct a The circular shift operation of the step, Is the size of The binary matrix is used to extract the middle of the candidate area elements to select appropriate training samples to alleviate boundary effects. In general, ,because is equal to the size of the target object, T is the size of the training sample (including both the target object and the fill area), is the regularization coefficient.
[0073] Object Detection
[0074] In the detection phase, the calculation starts from Samples cropped from the frame search area With the Correlation filter obtained by frame training Correlation operation response ,Right now , according to the response At the same time, a new frame of samples is obtained at the target position to retrain the filter and prepare for tracking in the next frame.
[0075] See also Figure 3 The present invention provides a video SAR target tracking method based on a spatiotemporal expansion-accumulation strategy, comprising the following steps:
[0076] S1、Contain M The dataset of frame video SAR images is divided into NBatch packages (batch), each batch package (batch) contains K Frame video SAR image;
[0077] The last frame of the previous batch and the first frame of the next batch are the same frame. The target estimation state of the first frame of the next batch is the state of the previous batch after being filtered by the trajectory backtracking module. The corresponding candidate box information, including position-size state vector, features, template, response value, etc., all comes from the last frame of the previous batch.
[0078] S2, in the initial frame of the first batch ( k =1) Manually select the tracking target, and its state vector is , Background Aware Correlation Filter (BACF) extracts features at the initial position F 1. Obtain the initialization parameters and initial filter template of BACF , generate an initial candidate box , is the response value corresponding to the candidate box, and its value is 1;
[0079] The BACF here can also be replaced by other correlation filters, such as kernel correlation filter (KCF), etc. The present invention selects BACF based on the effect of the algorithm.
[0080] feature F 1 is a 32-dimensional feature, including 31-dimensional directional histogram features and grayscale features. The target tracking frame can be further expressed as ,in, The superscript indicates the number of the current candidate box, and the subscript indicates the frame number of the image in the current batch.
[0081] S3, read the second frame image in the first batch, based on the candidate box of the first frame Define the target search area for the second frame , It is an adjustable parameter; Divided into multiple test samples by BACF, each test sample and Calculate similarity and obtain a response matrix ;
[0082] Search Area , whose location is based on The position is generated, and its size is The width and height are expanded Double generation, is an adjustable parameter, and its range is [4,8]. The final response matrix is the matrix with the maximum value among all response matrices.
[0083] S4, from the matrix Select S The maximum response value, each response value corresponds to the predicted position of the target, that is, it satisfies the relationship ; Extract features for each predicted position and update the corresponding template, and finally generate S candidate boxes, i.e. ;
[0084] Each value of the response matrix corresponds to a target prediction position, and the BACF algorithm applies a subgrid interpolation strategy to calculate the response results on the test sample. R , use the same method to calculate the response of each scale in the scale pool, and then use the scale with the largest response value as the size of the current frame target, and the maximum response value is the position of the target. S The maximum response value is calculated and its corresponding position is finally generated. S target candidate boxes.
[0085] S5. For the subsequent k Frame No. j Candidate boxes , repeat steps S3 to S4 to obtain candidate boxes; K Frame, repeat steps S3 to S4 to get At the end of this step, all candidate frames are generated, and the generation process forms a multi-frame tree structure;
[0086] Building a tree structure is as follows:
[0087] During the BACF tracking process, the response matrix The position corresponding to the maximum value of represents the estimated position of the target in the current frame. However, due to the existence of multi-peak problems, the position with the maximum response value is not necessarily the target, but may also be clutter or similar structures. Therefore, it is necessary to select as many positions as possible as candidate positions of the target in the current frame. Take the first batch as an example to introduce the process of building the tree structure:
[0088] (1) Filter initialization
[0089] Manually input the target tracking frame information of the first frame (initial frame), and its position-size state vector is , Represents the center position of the initial target tracking frame, and Represent the width and height of the tracking frame respectively. Extract features at the tracking frame Including 31-dimensional directional histogram features and grayscale features, using the extracted features Generate an initial template for the target And store it in the buffer, and finally generate a candidate frame of the target in the initial frame In order to more simply introduce the algorithm idea of the present invention, Has a double meaning:
[0090] 1) The name of the initial candidate box;
[0091] 2) The state vector of the initial candidate box.
[0092] Correspondingly, its state vector can be expressed as In the initial frame of the first batch, the filter mainly performs autocorrelation calculation. The value of is 1.
[0093] (2) Search area generation
[0094] See also Figure 2 , based on the initial candidate box The position size state vector , generate the search area in the second frame ,and The definition of is similar to that of , and the state vector of the search area is expressed as , is an adjustable parameter. In the present invention, The adjustment range is [4,8]. It is cut into multiple test samples by BACF, and the size and dimension of each test sample are the same as Same, each test sample and Perform cross-correlation similarity calculation and obtain a response matrix, and select the response matrix with the largest response value ,in, It is a two-dimensional matrix. The value in each grid is between 0 and 1, and each grid corresponds to a position coordinate of the current frame.
[0095] In traditional BACF algorithms, the position corresponding to the maximum value of the response matrix represents the estimated target position. However, in video SAR shadow target tracking, due to the small size of the target and the lack of effective texture features, the target is easily interfered with by clutter and similar structures. The response matrix may contain multiple high peaks, but the actual target is not located at the highest peak, causing the tracker to lose the target. Therefore, the present invention does not directly select the position corresponding to the maximum value of the response matrix as the estimated target position in the current frame.
[0096] (3) Selection of multiple candidate boxes
[0097] Based on the problem of multiple peaks described in the search region generation part and inspired by the idea of DP-TBD algorithm, the present invention obtains the maximum response matrix Select the front S The maximum response value, S The response values satisfy the following relationship:
[0098]
[0099] in, Represent this S The position corresponding to the response value is obtained. S The size of each candidate box is determined by the scale adjustment factor of BACF. The state vector of the target candidate box is defined as ,in , the candidate box includes the position-size state vector ,feature ,template and response values .
[0100] (4) Repetition and iteration
[0101] For any frame in the first batch k -1, from k -Select one from all candidate frames in 1 frame , , repeat the process (2) and (3) and then k Frames can be generated candidate boxes, represented as , . And so on, in the last frame of the first batch ( k=K )Total The candidate box of a target is represented as , All candidate box information in the first batch is stored in the buffer area for calculation of the trajectory backtracking module.
[0102] The process of building the multi-frame tree structure in all subsequent batches is the same as that of the first batch, but one thing needs to be explained: when the batch number n When >1, the candidate box information of the initial frame is directly retrieved from the target estimation result of the last frame of the previous batch instead of manual delineation.
[0103] In order to balance computational complexity, speed and accuracy, the present invention K The value of is 4 but not limited to 4. The specific situation is analyzed specifically. The schematic diagram of the construction of the multi-frame tree structure is as follows Figure 2 shown.
[0104] The template of the extension module needs to be updated in the template update module, and the updated template is retransmitted back to the extension module. When the template update module updates the template, the features at the candidate box of the extension module need to be updated, and the two cannot be separated.
[0105] The update method is the same as the original BACF template update, that is,
[0106]
[0107] in, Represents a multidimensional feature read from an extension module. Represents the template of the previous frame, represents the online learning rate, which represents the learning ability of the filter to the changes in the target appearance.
[0108] During the long-term tracking process of video SAR, factors such as illumination changes, target deformation, and noise interference will affect the appearance characteristics of shadow targets. At this time, a fixed learning rate is likely to cause target loss.
[0109] In order to improve the algorithm's adaptability to changes in target appearance features, this paper proposes an adaptive template updating method, which takes the initial frame of the first batch as the benchmark and calculates the first batch's initial frame. n The grayscale mean difference between the first frame of a batch and the initial frame of the first batch is used to adjust the online learning rate. variables;
[0110] First calculate the grayscale difference index :
[0111]
[0112] Among them, the grayscale mean of the target area of the initial frame of the first batch is , No. n The grayscale mean of the target area of the initial frame of the batch is .
[0113] Secondly, the algorithm follows Choose a suitable online learning rate :
[0114]
[0115] in, , is the threshold.
[0116] In the data sample used in this invention, , , =17.
[0117] S6, use the trajectory backtracking algorithm to reversely find the estimated position of the target in each frame, for the tree structure in step S5, from K Frame No. Candidate frames are backtracked, and a total of The candidate paths of the target trajectory, namely , ;
[0118] S7. Calculate the value function corresponding to each candidate path, that is, , select the candidate path with the largest value function as the estimated path of the target in a batch, and output the estimated state of the target in each frame.
[0119] For each subsequent batch, the target state of its initial frame is extracted from the last frame of the previous batch, and steps S3 to S7 are repeated to complete the target tracking of the entire dataset.
[0120] The multi-frame tree structure of each batch is built, and the information of all candidate boxes is stored in the data buffer. The backtracking process is performed to filter out the estimated state of the target from the numerous target candidate boxes.
[0121] Inspired by the DP-TBD concept, this paper uses the response value of the candidate box as the accumulated value of the value function. The detailed ideas are as follows:
[0122] Taking the first batch as an example, the last frame has target candidate frames, and the corresponding target candidate paths have Any candidate path consisting of multiple target candidate boxes from different frames Expressed as:
[0123]
[0124] in, ;
[0125] Candidate Path The accumulated value function is expressed as:
[0126]
[0127] from Select the function with the largest value from the candidate paths The path, that is
[0128]
[0129] in, Represents the set of all candidate paths in a batch.
[0130] Finally, the state vector of the filtered path is output, which is the estimated state of the target in each frame within a batch.
[0131] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "platforms."
[0132] In one embodiment of the present invention, a video SAR target tracking system based on a spatiotemporal expansion-accumulation strategy is provided. The system can be used to implement the above-mentioned video SAR target tracking method based on a spatiotemporal expansion-accumulation strategy. Specifically, the video SAR target tracking system based on a spatiotemporal expansion-accumulation strategy includes a division module, a selection module, a reading module, an update module, a repetition module, a backtracking module, and a tracking module.
[0133] Among them, the module is divided into a M The dataset of frame video SAR images is divided into N Batch packages, each containing K Frame image;
[0134] The selection module manually selects the tracking target in the initial frame of the first batch package obtained by the division module, and the background perception correlation filter extracts the features of the initial position F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial candidate box ;
[0135] The reading module reads the second frame image in the first batch package obtained by the segmentation module, based on the candidate frame of the first frame obtained by the selection module Define the target search area for the second frame , The background-aware correlation filter is divided into multiple test samples, each test sample and Calculate similarity and obtain a response matrix;
[0136] Update module, select the response matrix with the maximum response value from the response matrix obtained by the reading module and from Select SThe maximum response value, each response value corresponds to a predicted position of the target, extract features for each predicted position and update the corresponding template, and finally generate S target candidate boxes;
[0137] Repeat module, for the subsequent k Frame No. j Candidate boxes , after repeatedly reading the module and updating the module, we can get candidate boxes; K Frame, repeatedly read the module and update the module to get candidate frames, at this time all candidate frames are generated to form a multi-frame tree structure;
[0138] The backtracking module uses the trajectory backtracking module to reversely find the estimated position of the target in each frame in the tree structure obtained by the repeating module, starting from the K Frame No. Candidate frames are backtracked, and a total of candidate paths of target trajectories;
[0139] The tracking module calculates the value function corresponding to each candidate path obtained by the backtracking module, selects the candidate path with the largest value function as the estimated path of the target in a batch, and outputs the estimated state of the target in each frame.
[0140] In another embodiment of the present invention, a terminal device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions; the processor described in the embodiment of the present invention can be used for the operation of the video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy, including:
[0141] Include a MThe dataset of frame video SAR images is divided into N Batch packages, each containing K Frame image; the tracking target is manually selected in the initial frame of the first batch package, and the background-aware correlation filter extracts the features of the initial position F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial target candidate box ; Read the second frame image in the first batch, based on the candidate box of the first frame Define the target search area for the second frame , Divided into multiple test samples by BACF, each test sample and Calculate the similarity and obtain a response matrix, select the response matrix with the largest response value ; and from Select S The maximum response value, each response value corresponds to a predicted position of the target; extract features for each predicted position and update the corresponding template, and finally generate S target candidate box; for the subsequent k Frame No. j Candidate boxes , repeat the above steps to obtain candidate boxes; K Frame, repeat the above steps to get At the end, all candidate frames are generated to form a multi-frame tree structure; the trajectory backtracking algorithm is used to reversely find the estimated position of the target in each frame in the tree structure, starting from the first frame. K Frame No. Candidate frames are backtracked, and a total of candidate paths for the target trajectory; calculate the value function corresponding to each candidate path, select the candidate path with the largest value function as the estimated path of the target in a batch package, and output the estimated state of the target in each frame.
[0142] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device, used to store programs and data. It is understood that the computer-readable storage medium herein may include both built-in storage media in the terminal device and, of course, extended storage media supported by the terminal device. It may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for being loaded and executed by a processor. These instructions may be one or more computer programs (including program code). It should be noted that more specific examples (a non-exhaustive list) of computer-readable storage media herein include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk-read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0143] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, which carry readable program code. Such propagated data signals can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0144] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0145] The processor may load and execute one or more instructions stored in a computer-readable storage medium to implement the corresponding steps of the video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy in the above embodiment. The processor may load and execute the following steps:
[0146] Include a M The dataset of frame video SAR images is divided into N Batch packages, each containing K Frame image; the tracking target is manually selected in the initial frame of the first batch package, and the background-aware correlation filter extracts the features of the initial position F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial target candidate box ; Read the second frame image in the first batch, based on the candidate box of the first frame Define the target search area for the second frame , Divided into multiple test samples by BACF, each test sample and Calculate the similarity and obtain a response matrix, select the response matrix with the largest response value ; and from Select S The maximum response value, each response value corresponds to a predicted position of the target; extract features for each predicted position and update the corresponding template, and finally generate S target candidate box; for the subsequent k Frame No. j Candidate boxes , repeat the above steps to obtain candidate boxes; K Frame, repeat the above steps to get At the end, all candidate frames are generated to form a multi-frame tree structure; the trajectory backtracking algorithm is used to reversely find the estimated position of the target in each frame in the tree structure, starting from the first frame. K Frame No. Candidate frames are backtracked, and a total of candidate paths for the target trajectory; calculate the value function corresponding to each candidate path, select the candidate path with the largest value function as the estimated path of the target in a batch package, and output the estimated state of the target in each frame.
[0147] See also Figure 4The terminal device is a computer device. The computer device 60 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable by the processor 61. When executed by the processor 61, the computer program 63 implements the method for calculating the fluid composition in a reservoir-stimulated wellbore according to the embodiment. To avoid repetition, a detailed description thereof is omitted here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the system for calculating the fluid composition in a reservoir-stimulated wellbore according to the embodiment. To avoid repetition, a detailed description thereof is omitted here.
[0148] The computer device 60 may be a desktop computer, a notebook computer, a PDA, a cloud server, or other computing devices. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. It will be understood by those skilled in the art that Figure 4 This is merely an example of the computer device 60 and does not constitute a limitation of the computer device 60 . The computer device 60 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.
[0149] The processor 61 may be a central processing unit (CPU), other general-purpose processors, central processing units (CPUs), graphics processors (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), other programmable logic devices, discrete gate or transistor logic devices, quantum computing-based data processing logic, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0150] The memory 62 may be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 60.
[0151] Furthermore, the memory 62 may include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is about to be output.
[0152] Any reference to memory, database, or other media used in the various embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0153] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0154] See also Figure 5 The terminal device 600 is an electronic device that is implemented as a general-purpose computing device. The components of the electronic device may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), and a display unit 640.
[0155] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above method section of this specification. For example, the processing unit 610 can perform the following steps: Figure 3 Follow the steps shown in .
[0156] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0157] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0158] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0159] The electronic device 600 can also communicate with one or more external devices 700 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 650. Furthermore, the electronic device 600 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 via the bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0160] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Specific embodiments
[0162] Input Data
[0163] The method of the present invention was tested on video SAR datasets published by Sandia National Laboratories in the United States, including the Victr and Highway datasets. The Victr dataset contains 398 images with a single target traveling on a highway and an image resolution of 324*267. The Highway dataset contains 300 frames of images with a resolution of 660*720 showing multiple targets turning on the highway. In the implementation example (Highway) of the present invention, a total of five targets were selected for tracking.
[0164] Parameter initialization
[0165] The parameters of the entire algorithm are all taken from the optimal values during the experiment, where the number of frames in each batch is , the number of maximum response values selected , a multiplication factor of the search area , online learning rate , , the intensity difference threshold .
[0166] Evaluation indicators
[0167] The evaluation indicators used in the present invention are:
[0168] success rate
[0169] Indicates the The target box tracked by the frame algorithm, Indicates the true target box, overlap rate Defined as:
[0170]
[0171] in, 、 represents the intersection and union of two target boxes, and |∙| represents the number of pixels in the region. If the overlap ratio is greater than a given threshold (0.5), tracking for the current frame is considered successful. The success rate is defined as the percentage of successfully tracked frames to the total number of frames.
[0172] CLE and precision
[0173] Center Location Error (CLE) is defined as the center coordinate of the predicted target box The center coordinates of the real target frame The average Euclidean distance between them is expressed as:
[0174]
[0175] The accuracy is defined as the ratio of the number of accurately tracked frames to the total number of frames in the dataset. When the center position error (CLE) is less than a certain threshold, the current frame can be considered to be accurately tracked.
[0176] FPS (Frames Per Second, FPS)
[0177] The processing time of a tracking algorithm is one of the important indicators for evaluating the real-time performance of the algorithm. FPS represents the number of images processed by the tracking algorithm in one second.
[0178] Tracking results display
[0179] In the table, red represents the best, green represents the second, and blue represents the third.
[0180] Table 1 Results and comparison of algorithms on the Victr dataset
[0181]
[0182] In summary, the present invention provides a video SAR target tracking method and system based on a spatiotemporal expansion-accumulation strategy, which has the following characteristics:
[0183] (1) Tracking weak shadow targets in video SAR using the method of accumulating multiple frames of information. Because video SAR targets exist continuously in the surveillance area for a long time, while clutter is randomly distributed and does not exist for a long time, after accumulating multiple frames of image information, sufficient real target information can be obtained, and the target can be successfully distinguished from the clutter.
[0184] (2) Build a multi-frame tree structure to handle the multi-peak problem in the correlation filter response graph. By generating multiple candidate boxes, the real target is covered as much as possible. Even if the candidate box of the real target in a certain frame does not have the maximum response value, the real trajectory of the target can still be successfully found after expanding the tree structure in the time domain and spatial domain.
[0185] (3) The online learning rate of the correlation filter is flexibly adjusted based on changes in target appearance. By comparing the grayscale difference between the first frame of the subsequent batch and the initial frame of the first batch, the algorithm adaptively selects an appropriate online learning rate to update the template, improving the algorithm's ability to adapt to different tracking scenarios. During long-term video SAR target tracking, the target's appearance characteristics may change significantly due to factors such as deformation and illumination changes. In this case, the template update method proposed in this invention is more adaptable to these changes and can achieve long-term target tracking.
[0186] (4) It has the ability to continuously and stably track the target, and has achieved excellent tracking accuracy and success rate, as well as very small positioning error, and has achieved excellent performance compared with other algorithms based on correlation filtering.
[0187] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0188] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0189] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0190] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms.
[0191] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0192] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0193] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0194] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0195] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0196] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0197] The above content is only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A video SAR target tracking method based on a spatiotemporal expansion-accumulation strategy, characterized in that: The following steps are involved: S1、Contain M The video SAR image dataset of the frame is divided into N Batch packages, each of which contains K Frame video SAR image; S2: Manually select the tracking target in the initial frame of the first batch package obtained in step S1, and use the background-aware correlation filter to extract the features of the initial position. F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial candidate box ; S3, read the second frame image in the first batch package obtained in step S1, based on the candidate frame of the first frame obtained in step S2 Define the target search area for the second frame , The background-aware correlation filter is divided into multiple test samples, each test sample and Calculate similarity and obtain a response matrix; S4. Select the response matrix with the maximum response value from the response matrix obtained in step S3 , from the response matrix Select S The maximum response value, each response value corresponds to a predicted position of the target; extract features for each predicted position and update the corresponding template, and finally generate S candidate boxes; S5, for the subsequent k Frame No. j Candidate boxes , repeat steps S3 to S4 to obtain candidate boxes; K Frame, repeat step S3 to step S4 to get candidate frames, and at the end all candidate frames are generated to form a multi-frame tree structure; S6, use the trajectory backtracking to reversely search for the estimated position of the target in each frame in the tree structure obtained in step S5, K Frame No. Candidate frames are backtracked, and a total of candidate paths of target trajectories; S7. Calculate the value function corresponding to each candidate path obtained in step S6, select the candidate path with the largest value function as the predicted path of the target in a batch package, and output the estimated state of the target in each frame.
2. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 1 is characterized in that: In step S1, the last frame of the previous batch package and the first frame of the next batch package are the same frame. The initial target state of the first frame of the next batch package is the target estimated state of the previous batch package after screening by the trajectory backtracking module. The corresponding candidate box information including position-size state vector, features, template, and response value all comes from the last frame of the previous batch package.
3. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 1, characterized in that: In step S2, the feature F 1 is a 32-dimensional feature, including 31-dimensional directional histogram features and grayscale features; The target tracking box is , The superscript indicates the number of the current candidate box, and the subscript indicates the frame number of the image in the current batch package.
4. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 1, characterized in that: In step S3, the target search area The location is based on Position generation, target search area The size is right The position-size state vector is expanded in width and height Double generation, is an adjustable parameter ranging from [4,8]; the resulting response matrix is the matrix with the maximum value among all response matrices.
5. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 1, characterized in that: In step S4, each response value corresponds to the predicted position of the target and satisfies the following relationship: in,( ) are the response values The corresponding s The coordinates of the predicted target position.
6. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 1, characterized in that: In step S5, the tree structure constituting the multiple frames is specifically as follows: S501, initial target candidate frame and the position-size state vector of the candidate box ; S502: Generate a search area in the second frame , The background-aware correlation filter is used to crop the sample into multiple test samples, each with the same size and dimension. Same, each test sample and Perform cross-correlation similarity calculation and obtain a response matrix , select the response matrix with the maximum response value , It is a two-dimensional matrix. The value in each grid is between 0 and 1, and each grid corresponds to a position coordinate of the current frame. S503. Generate response matrix from search area Select the front S The maximum response value is obtained S target candidate boxes, the size of each candidate box is determined by the scale adjustment factor of the background perception correlation filter; S504: For any frame in the first batch package k -1, from k -Select one from all candidate frames in 1 frame , After repeating step S502 and step S503, k Frame Generation candidate boxes, represented as , ; And so on, in the last frame of the first batch ( k=K )Total The candidate box of a target is represented as , ; All candidate box information in the first batch package is stored in the buffer area, among which, when the batch package number n When >1, the candidate box information of the initial frame is directly retrieved from the target estimation result of the last frame of the previous batch package, and the target candidate box generation process of the remaining frames is the same as that of the first batch package.
7. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 6, characterized in that: Adopting the adaptive template updating method, taking the initial frame of the first batch package as the benchmark, calculate the grayscale mean difference between the first frame of the nth batch package and the initial frame of the first batch package, and use the difference as the basis for adjusting the online learning rate. variables.
8. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 7, characterized in that: Online learning rate for: in, is the grayscale difference index, is the threshold, and is the value of the online learning rate.
9. The video SAR target tracking method based on the spatiotemporal expansion-accumulation strategy according to claim 1, characterized in that: In step S7, Select the function with the largest value from the candidate paths The path is: in, Represents the set of all candidate paths in a batch package, is the estimated position of the target for all frames estimated by the algorithm, Estimate the target position for the Kth frame, All predicted locations for the target within a batch.
10. A video SAR target tracking system based on spatiotemporal expansion-accumulation strategy, characterized in that: include: Divide the module into a M The dataset of frame video SAR images is divided into N Batch packages, each containing K Frame video SAR image; Select the module, manually select the tracking target in the initial frame of the first batch package obtained by the division module, and extract the features of the initial position through the background perception correlation filter F 1. Obtain the initialization parameters and initial filter template of the background perception correlation filter , generate an initial candidate box ; The reading module reads the second frame image in the first batch package obtained by the segmentation module, based on the candidate frame of the first frame obtained by the selection module Define the target search area for the second frame , The background-aware correlation filter is divided into multiple test samples, each test sample and Calculate similarity and obtain a response matrix; Update module, select the response matrix with the maximum response value from the response matrix obtained by the reading module , from the response matrix Select S The maximum response value, each response value corresponds to the predicted position of the target; extract features for each predicted position and update the corresponding template, and finally generate S candidate boxes; Repeat module, for the subsequent k Frame No. j Candidate boxes , after repeatedly reading the module and updating the module, we can get candidate boxes; K Frame, repeatedly read the module and update the module to get candidate frames, and at the end all candidate frames are generated to form a multi-frame tree structure; The backtracking module uses the trajectory backtracking algorithm to reversely find the estimated position of the target in each frame in the tree structure obtained by the repeated module. K Frame No. Candidate frames are backtracked, and a total of candidate paths of target trajectories; The tracking module calculates the value function corresponding to each candidate path obtained by the backtracking module, selects the candidate path with the largest value function as the estimated path of the target in a batch package, and outputs the estimated state of the target in each frame.
Citation Information
Patent Citations
Target tracking method and system based on fusion twin network and Kalman filtering
CN115471525A
Efficient autofocus method for swath SAR
US20060109164A1