Target capture method, system and medium in a video sequence scenario

Through the method of object detection on video sequence images, imaging size range filtering, multi-object tracking and cumulative confidence calculation, single-object tracking and cross-verification method cross-verification, the problems of insufficient anti-interference ability and poor stability of object detection tracking in video sequence scenarios are solved, and efficient and accurate target capture and tracking are achieved.

CN119228843BActive Publication Date: 2025-06-17HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411249939.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-06-17
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient anti-interference ability, poor tracking stability and accuracy, and many false alarms and missed detection phenomena in target detection and tracking in video sequence scenarios, making it difficult to adapt to the appearance, motion mode and environmental changes of the target.

Method used

Target capture and tracking are achieved through object detection, imaging size range filtering, multi-objective tracking and cumulative confidence calculation, single-objective tracking and cross-validation of cross-validation of the video sequence images.

Benefits of technology

It improves the anti-interference ability, tracking stability and accuracy of the target capture method in video sequence scenarios, reduces false alarms and missed detection, and ensures effective detection and tracking in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119228843B_ABST
    Figure CN119228843B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and medium for target capture in a video sequence scenario, belonging to the field of single-object tracking in images. The method includes: performing target detection on video sequence images to obtain the position and size of the target, and deleting the targets whose sizes are not within the specified range; initializing a multi-object tracker using the positions of the undeleted targets to perform multi-object tracking and cumulative confidence calculation; for each object tracker, when its tracking hit count reaches the first set value and the cumulative confidence reaches the set confidence threshold, output the target it tracks and the cumulative confidence, and when its consecutive tracking miss count reaches the second set value, delete the object tracker; initializing a single-object tracker using the target with the highest cumulative confidence to perform single-object tracking and tracking confidence calculation and output, and cross-verifying the results output by the two types of object tracking to determine whether to output and the output results. This method realizes stable capture and recognition of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of single-object image tracking, and more specifically, relates to a target capture method, system, and medium in a video sequence scenario. Background Art

[0002] As an information-intensive carrier, video plays an increasingly important role in modern society. Whether it is for security monitoring in daily life or complex military surveillance tasks, video data provides a rich source of information. Among them, accurate detection and stable tracking of targets in video are one of the key technologies for applications such as behavior recognition analysis, intelligent monitoring analysis, abnormal event detection, and even military surveillance. This technology not only has broad application prospects in the civilian field, such as intelligent transportation and smart cities, but also has significant strategic importance in the military field, such as target recognition, tracking, and battlefield situation awareness. It is a hot topic that countries around the world focus on developing and competing to research.

[0003] However, there are still some problems in the existing technical means for target detection and tracking in a video sequence scenario. Traditional tracking methods adopt a method of first detecting and then directly tracking. This method highly depends on the results of the first detection and the performance of subsequent tracking methods. Once a false alarm occurs in the initial detection of the target, it will lead to an incorrect tracking initialization, thereby affecting the accuracy of subsequent tracking. If the target area is redetected every once in a while during the tracking process to verify the accuracy of the tracking result, although this method avoids the influence of the initial capture, if a false alarm occurs during the verification, it may lead to the failure of tracking, and ultimately still affect the performance of tracking and capture. In addition, factors such as the appearance of the target, the motion pattern, and the shooting environment will change with time and space, and simple tracking algorithms are difficult to adapt to these changes, which will also lead to tracking failure.

[0004] Therefore, there are still bottleneck problems in the existing target detection and tracking technologies for video scenarios, and further research and development are urgently needed. How to improve the anti-interference ability of algorithm detection, the stability and accuracy of tracking, reduce false alarms and missed detections, and ensure effective detection and tracking of various targets in various complex environments has important research significance. Summary of the Invention

[0005] In view of the defects and improvement requirements of the existing technology, the present invention provides a target capture method, system, and medium in a video sequence scenario, aiming to improve the anti-interference ability, tracking stability and accuracy of the target capture method in a video sequence scenario, reduce false alarms and missed detections, and ensure effective detection and tracking of various targets in various complex environments.

[0006] To achieve the above object, according to one aspect of the present invention, there is provided a target capture method in a video sequence scenario, including: S1, performing target detection on the input video sequence images to obtain the positions and sizes of all targets in each frame of the image; S2, designing the imaging size range of the desired target according to its theoretical size in the image, deleting the targets whose sizes in S1 are not within the imaging size range, and outputting the undeleted targets; S3, initializing a multi-target tracker with the positions of all the targets output in S2 to perform multi-target tracking and cumulative confidence calculation on the video sequence images; during the multi-target tracking process, for each target tracker, when its tracking hit count reaches the first set value and the cumulative confidence reaches the set confidence threshold, output the target it tracks and the cumulative confidence, and when its consecutive tracking miss count reaches the second set value, delete the target tracker; S4, initializing a single-target tracker with the target having the highest cumulative confidence to perform single-target tracking and tracking confidence calculation on the video sequence images and output; S5, using the intersection over union matching method to cross-verify the targets and cumulative confidence output in S3 and the targets and tracking confidence output in S4 to determine whether to output the target tracked in S4, or output the target with the highest cumulative confidence tracked in S3, or output the fusion result of the targets matched in S4 and S3, or not output.

[0007] Further, the method further includes: judging whether each frame of the image is captured according to the result of the cross-validation in S5; when the number of frames of un-captured images within consecutive several frames of images reaches the third set value, exit the tracking, if there are targets output in S3, re-execute S4 and S5, and if there are no output targets in S3, do not output.

[0008] Further, S1 specifically includes: preprocessing the input video sequence images; performing target detection on the preprocessed images to obtain horizontal detection frames; using border contraction to find the minimum bounding rectangle of the target within the horizontal detection frames to obtain rotated detection frames, and the positions and sizes of the rotated detection frames are the positions and sizes of the targets therein.

[0009] Further, the using border contraction to find the minimum bounding rectangle of the target within the horizontal detection frames specifically includes: sequentially performing saliency detection, image segmentation, filtering, and edge detection on the images within the horizontal detection frames to achieve border contraction, and obtaining the minimum bounding rectangle of the frame after border contraction as the rotated detection frame.

[0010] Further, the imaging size range is:

[0011] Th wmax =target w ×rate w

[0012] Th hmax = target h × rate w

[0013] Th wmin = target w ÷ rate h

[0014] Th hmin = target h ÷ rate h

[0015]

[0016] Among them, Th wmax 、Th hmax 、Th wmin 、Th hmin are respectively the upper limit of the imaging length, the upper limit of the imaging width, the lower limit of the imaging length, and the lower limit of the imaging width of the expected target. target w 、target h are respectively the theoretical length and the theoretical width of the expected target in the image. The theoretical dimensions include target w and target h , rate w is the length scaling factor, rate h is the width scaling factor, base is the first adjustable parameter, α is the second adjustable parameter, β is the third adjustable parameter. The imaging size range includes Th wmax 、Th hmax 、Th wmin and Th hmin .

[0017] Furthermore, the multi-target tracking of the video sequence images specifically includes: for each target tracker, using the target tracker to perform target tracking on the video sequence images. If the target it tracks matches the corresponding target obtained by target detection, the tracking hits, and the position of the corresponding tracking target that is matched is corrected using the target position obtained by target detection. If the target it tracks does not match any of the targets obtained by target detection, the tracking misses; for each target output in S2, if the target does not match any of the targets obtained by the multi-target tracker, a new target tracker is initialized for multi-target tracking.

[0018] Furthermore, when the number of tracking hits is greater than the fourth set value, the cumulative confidence is:

[0019] cumConf ← cumConf × (1 - cum_factor) + conf × cum_factor

[0020] When the number of tracking hits is not greater than the fourth set value, the cumulative confidence is as follows:

[0021] cumConf ← 1 - (1 - cumConf) × (1 - conf × cum_factor)

[0022] Where cumConf is the cumulative confidence, the left and right sides of ← are the current cumulative confidence and the previous cumulative confidence respectively, cum_factor is an adjustable weight, and conf is the object detection confidence of the current frame.

[0023] Furthermore, the tracking confidence is: the correlation score of the position of the object tracked by single-object tracking.

[0024] According to another aspect of the present invention, there is provided an object capture system in a video sequence scenario, including: a processor; a memory storing computer-executable programs, which when executed by the processor, cause the processor to execute the object capture method in the video sequence scenario as described above.

[0025] According to another aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the object capture method in the video sequence scenario as described above.

[0026] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0027] (1) It provides an object capture method in a video sequence scenario. After performing preliminary object detection on an image, some detected objects are initially filtered out according to the imaging size of the object in the image to achieve spatial filtering. The undeleted objects are passed into the multi-object tracking module, where temporal filtering is achieved at this stage to exclude accidental misdetections in the detection and achieve multi-frame accumulation of confidence. The output is controlled by setting a cumulative count threshold (the first set value) and a confidence threshold. The output result is sent to the single-object tracking module, and the object with the highest confidence is initially selected for initialization of single-object tracking to achieve capture. Subsequently, each time the single-object tracking module independently completes the tracking prediction of the captured object and cross-verifies with the object output by the multi-object tracking to determine whether the object is correctly captured and tracked, and perform capture replacement in a timely manner to ensure correct capture and tracking of the object;

[0028] (2) provides a preferred object detection method. After obtaining the horizontal detection frame, the minimum circumscribed rectangle of the object is searched within the horizontal detection frame to obtain the rotated detection frame, thereby determining the position and size of the object, and having higher object detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of the object capture method in the video sequence scenario provided by the embodiment of the present invention;

[0030] Figure 2 is an example diagram of controllable generation of cloud interference provided by the embodiment of the present invention;

[0031] Figure 3 is a flowchart of the implementation of the border contraction algorithm provided by the embodiment of the present invention;

[0032] Figure 4 is an example diagram of imaging size calculation provided by the embodiment of the present invention;

[0033] Figure 5 is a flowchart of SORT multi-object tracking time series information filtering provided by the embodiment of the present invention

[0034] Figure 6 is a flowchart of DSST single-object tracking prediction provided by the embodiment of the present invention

[0035] Figure 7 is a flowchart of cross-validation between DSST and SORT provided by an embodiment of the present invention;

[0036] Figure 8 is a flowchart of cross-validation between DSST and SORT provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0038] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0039] Embodiment 1

[0040] An object capture method in a video sequence scenario. Refer to Figure 1 , in combination with Figures 2 - 8, the target capture method in the video sequence scenario of this embodiment will be described in detail. The method includes operations S1 - S5.

[0041] Operation S1: Detect the targets in the input video sequence images to obtain the positions and sizes of all targets in each frame of the image.

[0042] Preferably, operation S1 specifically includes the following sub - operations S11 - S13.

[0043] In sub - operation S11, pre - process the input video sequence images. Specifically, for example, input the video sequence frame by frame. Before input, perform image pre - processing such as quantization denoising on the video sequence; it is also possible to directly input the video sequence without pre - processing.

[0044] In sub - operation S12, detect the targets in the pre - processed images to obtain horizontal detection boxes.

[0045] Specifically, traditional target detection methods can be directly used for target detection; intelligent methods can also be used for target detection, and network training needs to be carried out before use. When using intelligent methods for target detection, direct use of rotated - box target detection can be selected to better represent the position and size of the target; horizontal - box target detection can also be used to improve the detection speed and accuracy. After detecting the horizontal box, find the rotated box of the target within the local area represented by the horizontal detection box.

[0046] Taking the use of the yolov5 target detection algorithm for target detection as an example, the specific implementation process of sub - operation S12 is described. Other target detection methods can also be used to implement sub - operation S12.

[0047] Before using yolov5 for target detection, network training needs to be carried out first. To improve the anti - interference ability, target images with interference such as occlusion and cloud and fog are added for training during training. Therefore, a method for controllable generation of cloud - layer occlusion data based on target texture mapping and a method for controllable simulation of foggy scenes are used to generate target images with cloud and fog occlusion. As Figure 2 shown, cloud and fog simulation and controllable generation are carried out on the ship targets in DOTAv1.0. The cloud image is extracted by manual image cropping or intelligent cloud segmentation method for subsequent image texture mapping, and the original image is directly selected from the DOTA dataset; fog images of various degrees are generated using a foggy - day image simulator based on the atmospheric scattering model to simulate a more realistic haze effect at any altitude.

[0048] The core architecture of YOLOv5 consists of four parts: the input end, Backbone, Neck, and Head. At the input end, Mosaic data augmentation and adaptive anchor box calculation are used to enrich the detection dataset. In particular, random scaling adds many small targets, making the network more robust. In Backbone, two important structures are Focus and CSP. Focus is a special downsampling method that can speed up the process, and CSP can solve the problem of large computational complexity in inference and also play a role in downsampling. Neck adopts the structure of FPN+PAN. FPN conveys strong semantic features from top to bottom, while PAN conveys strong localization features from bottom to top, enabling parameter aggregation for different detection layers from different backbone layers. Head uses the neighborhood positive and negative sample assignment strategy, which can significantly accelerate convergence. As an intelligent object detection method, YOLOv5 can learn the features of the object through training. At the same time, due to the addition of interference such as cloud and fog occlusion during training, it has anti-interference ability during inference testing and can accurately detect the object in the image. In addition, during the object detection stage, more attention should be paid to the recall rate to ensure that the object is detected as much as possible without worrying about false alarms, because subsequent operations will filter the detection area.

[0049] In sub-operation S13, the minimum bounding rectangle of the target is searched within the horizontal detection box using border contraction to obtain the rotated detection box, and the position and size of the rotated detection box are the position and size of the target therein.

[0050] Preferably, searching for the minimum bounding rectangle of the target within the horizontal detection box using border contraction specifically includes: successively performing saliency detection, image segmentation, filtering, and edge detection on the image within the horizontal detection box to achieve border contraction, and obtaining the minimum bounding rectangle of the box after border contraction as the rotated detection box, as Figure 3 shown. It should be noted that Figure 3 the method shown is only the preferred method provided in this embodiment, and other border contraction methods can also be used to obtain the rotated detection box.

[0051] Taking Figure 3 the border contraction method shown as an example, the specific implementation process of sub-operation S13 is described, including the following steps (1)-(5).

[0052] Step (1), first, saliency features of the input image are extracted using saliency detection to obtain a saliency map. The saliency detection is, for example, FT saliency detection.

[0053] Specifically, first perform Gaussian smoothing on the image, and then convert its color space from the RGB color space to the CIELAB color space. Calculate the average values of L, a, and b for the entire image, and calculate the Euclidean distance between the L, a, and b values of each pixel after Gaussian smoothing and the average values of L, a, and b of the image according to the following formula:

[0054]

[0055] where I μ =[L μ a μ b μ T represents the matrix of the average values of L, a, and b of the image, represents the matrix of the L, a, and b values of each pixel after Gaussian smoothing. Then perform a normalization operation to obtain the final saliency map.

[0056] Step (2), then obtain different regions in the image through image segmentation. The image segmentation is, for example, OTSU image segmentation. OTSU is an algorithm for determining the binarization segmentation threshold of an image, and its essential idea is to maximize the between-class variance.

[0057] Taking a grayscale image as an example, for the image, it can be regarded as a matrix of size M×N, that is, the pixels in the image, and each value is the pixel value, where the pixel value is between (0, 255). The segmentation threshold between the foreground (i.e., the target) and the background is denoted as optimal threshold , the proportion of the number of foreground pixels in the entire image is denoted as w0, the average grayscale of the foreground is denoted as μ0; the proportion of the number of background pixels in the entire image is denoted as w1, and its average grayscale is denoted as μ1. The total average grayscale of the image is denoted as μ, and the between-class variance is denoted as maximum. Assume that the background of the image is darker, and the size of the image is M×N, and the number of pixels in the image with a grayscale value less than the threshold optimal threshold is denoted as N0, and the number of pixels with a grayscale value greater than or equal to the threshold optimal threshold is denoted as N1, then there are:

[0058]

[0059] N0 + N1 = M×N

[0060] w0 + w1 = 1

[0061] μ = w0×μ0 + w1×μ1

[0062] maximum = w0×(μ0 - μ) 2 + w1×(μ1 - μ) 2

[0063] ​Substituting the expression of μ into the expression of maximum, we get:

[0064] maximum = w0 × w1 × (μ0 - μ1) 2

[0065] The threshold maximum that maximizes the between-class variance is obtained by using a traversal method. After solving for the optimal threshold using the OSTU algorithm, binary segmentation is performed.

[0066] Step (3), the remaining image noise is suppressed through filtering operations, such as morphological filtering. Morphological filtering includes erosion operations and dilation operations. Erosion operations can eliminate noise points, filter sea clutter, suppress clutter interference, and at the same time eliminate some boundary values, which may cause the overall reduction of the target image. Dilation operations can increase the target feature values, resulting in the overall enlargement of the target image. Using erosion operations and dilation operations in combination can achieve a better purpose of segmenting independent graphic elements.

[0067] Step (4), the target edge information is extracted through an edge detection algorithm, such as the LSD edge detection method. The LSD edge detection method is used to detect local straight contours in the image. First, the horizontal line angles of each pixel point are calculated, thus forming a horizontal line field, that is, a unit vector field. Here, the horizontal line angle of the pixel point is the vertical angle of the gradient direction of the point, and the horizontal line field is a matrix corresponding one-to-one with the points in the image, and the value of the element in the matrix is the horizontal line angle corresponding to the point in the image. For all horizontal lines in the horizontal field, if the directions are approximately within the tolerance range, these regions are divided into the same support domain, and then the pixel range of the support domain is obtained, thereby obtaining the minimum bounding rectangle. Compared with traditional edge detection operators such as Canny, the LSD edge detection algorithm has stronger robustness in the case of image interference and occlusion.

[0068] Step (5), a directed box after border contraction is obtained through the minimum bounding rectangle, and a rotated target detection box is obtained.

[0069] Operation S2, design the imaging size range of the expected target according to its theoretical size in the image, delete the targets in operation S1 whose sizes are not within the imaging size range, and output the undeleted targets.

[0070] The theoretical size of the expected target in the image can be calculated according to the pose position of the camera and the actual size of the expected target. The specific calculation method is as follows.

[0071] In the real scene, the pose and position of the camera during shooting are known. Taking high-altitude shooting as an example, as Figure 4As shown in the figure, taking the ground as the reference system, the position of the camera is (x, y, z), and the angular attitude is (α, β, γ), where α is the pitch angle, and the field of view angles in the length and width directions are 40° and 30° respectively. The target is on the ground, and the true length and width of the target are real w 、real h , then:

[0072]

[0073] Therefore, the actual imaging range is:

[0074]

[0075] The imaging size is 1920×1280. Then, by scaling proportionally, the long side (length) and short side (width) of the target in the image are respectively:

[0076]

[0077] Thus, the theoretical size of the expected target in the image is obtained (target w and target h ).

[0078] According to the calculated imaging size and aspect ratio, a certain range of thresholds is set to determine whether the detection frame exceeds the range. For the target frames that exceed the range, they are filtered out to achieve the purpose of filtering false alarms.

[0079] Preferably, the imaging size range set for the expected target is:

[0080] Th wmax = target w × rate w

[0081] Th hmax = target h × rate w

[0082] Th wmin = target w ÷ rate h

[0083] Th hmin = target h ÷ rate h

[0084]

[0085] Among them, Th wmax 、Th hmax 、Th wmin 、Thhmin They are respectively the upper limit of the imaging length, the upper limit of the imaging width, the lower limit of the imaging length, the lower limit of the imaging width of the expected target, target w 、target h They are respectively the theoretical length and the theoretical width of the expected target in the image. The theoretical size includes target w and target h , rate w is the length scaling coefficient, rate h is the width scaling coefficient, base is the first adjustable parameter, α is the second adjustable parameter, β is the third adjustable parameter. The imaging size range includes Th wmax 、Th hmax 、Th wmin and Th hmin .

[0086] Three parameters, base, α and β, are introduced in the imaging size range. This is because when the distance is different, the size difference of the target in the field of view changes greatly. Therefore, different scaling ranges should be set for targets of different sizes. By adjusting the base value, the tolerable scaling range for large targets can be determined, and by adjusting the α value, the scaling range for small targets can be changed. Preferably, base = 1.2, α = 2.5, β = 0.9.

[0087] By judging whether Th wmin <det w <Th wmax 、Th hmin <det h <Th hmax , the rotation detection boxes outside this range are deleted to complete the filtering of the prior knowledge space information. Among them, det w 、det h are respectively the length and width of the target detected by calculating the rotation detection box obtained in operation S1.

[0088] Operation S3: Initialize the multi-object tracker with the positions of all the targets output in operation S2 to perform multi-object tracking and cumulative confidence calculation on the video sequence images; during the multi-object tracking process, for each object tracker, when its tracking hit count reaches the first set value and the cumulative confidence reaches the set confidence threshold, output the tracked object and the cumulative confidence, and when its consecutive tracking miss count reaches the second set value, delete the object tracker.

[0089] The positions of the targets output in operation S2 are the results after filtering by prior knowledge and meet the requirements in terms of spatial scale, so they are used to initialize the multi-object tracker. The multi-object tracker adopts, for example, the SORT multi-object tracking algorithm, or other multi-object tracking algorithms can also be used.

[0090] The SORT multi-object tracking algorithm consists of three parts: object detection, Kalman filtering, and Hungarian matching. In this embodiment, the object detection result comes from the output of operation S2, so the focus is on Kalman filtering and Hungarian matching. The specific process is as Figure 5 shown.

[0091] After obtaining the output result of the object detection module, the accurate position of the object cannot be directly obtained. Because the detection result may have errors, the detected bounding box will be noisy. Kalman filtering can fuse the object detection result and the prediction result of the tracking algorithm to further improve the accuracy of the detected bounding box obtained by the detection algorithm.

[0092] Kalman filtering uses a signal and noise state space model, and uses the estimated value of the previous moment and the observed value of the current moment to estimate the state variable and obtain the estimated value of the current moment. Kalman filtering believes that the observation value credibility model is a normal distribution, and the control model is a normal distribution. All observable signals are linearly variable.

[0093] In the SORT multi-object tracking algorithm, the observation variable where u represents the horizontal pixel position of the object center, v represents the vertical pixel position of the object center, s and r respectively represent the ratio (area) and aspect ratio of the object bounding box. respectively represent the horizontal movement speed, vertical movement speed and the change speed of the area of the object. After associating and detecting an object, the detected bounding box is used to update the object state, and the speed component is optimized and solved through the Kalman filter framework. If there is no detection related to the object, the linear speed model is used to simply predict its state without correction.

[0094] Hungarian matching is mainly used to solve the assignment problem. There are n different tasks, and n people are required to complete 1 of them respectively, and the time for each person to complete the task is different. So there is a problem of how to assign tasks to minimize the time spent. The optimal solution problem mainly relies on the property of the optimal solution of the matrix: if the minimum element of each row (column) of the matrix is subtracted from the elements of that row (column) respectively to obtain a reduced matrix, its optimal solution is the same as that of the original matrix. In the SORT multi-object tracking algorithm, the distance between the detected object and the tracked object is determined by calculating the IOU distance between the two objects, and then the corresponding cost matrix is generated. Then, the Hungarian matching algorithm is used to match the detection result of the current frame and the prediction result of the tracker according to the cost matrix.

[0095] Since object detection can only output the confidence of the current frame each time, but the confidence of the object bounding box in the tracking result after multiple frames of accumulation should also be cumulative, so the calculated multi-frame cumulative confidence is used as the confidence of the tracking result of the current frame.

[0096] In one embodiment of the present invention, the cumulative confidence is calculated as follows:

[0097] cumConf←cumConf×(1 - cum_factor)+conf×cum_factor

[0098] In another embodiment of the present invention, the cumulative confidence is calculated as follows:

[0099] cumConf←1-(1 - cumConf)×(1 - conf×cum_factor)

[0100] In yet another embodiment of the present invention, the cumulative confidence is calculated as follows:

[0101]

[0102] Wherein, cumConf is the cumulative confidence, the left and right sides of ← are the current cumulative confidence and the previous cumulative confidence respectively, cum_factor is an adjustable weight, conf is the target detection confidence of the current frame, N is the number of tracking hits, and N4 is the fourth set value. The initial value of cumConf is 0; cum_factor is used to prevent cumConf from growing too fast; N4 is, for example, 20. It should be noted that other calculation methods of cumulative confidence are also applicable to this method.

[0103] According to the embodiments of the present invention, during the multi - target tracking process: for each target tracker, the target tracker is used to perform target tracking on the video sequence image. If the target tracked by it matches the corresponding target obtained by target detection, the tracking is successful, and the position of the corresponding tracked target that is matched is corrected using the target position obtained by target detection. If the target tracked by it does not match any of the targets obtained by target detection, the tracking fails; for each target output in operation S2, if the target does not match any of the targets obtained by the multi - target tracker, a new initialized target tracker is added for multi - target tracking.

[0104] When the number of consecutive misses of a tracker reaches the second set value (such as 10), it is considered to reach the survival period, and the tracker will be deleted to eliminate occasional false alarms. When the cumulative number of hits of a tracker reaches the first set value (such as 5) and the cumulative confidence reaches the set confidence threshold (such as 0.5), the target in the current tracker is considered to be a real and effective target, and it will be output to the subsequent single - target tracking module for initialization or cross - verification.

[0105] In operation S4, the single - target tracker is initialized with the target having the highest cumulative confidence to perform single - target tracking on the video sequence image, calculate the tracking confidence, and output it.

[0106] Specifically, first, select the target with the highest confidence in the results after spatio-temporal information filtering for single-target tracker initialization to complete the initial capture; then, use the initialized single-target tracker to predict the target position and tracking confidence.

[0107] Taking the DSST single-target tracking algorithm used by the single-target tracker as an example, the specific process of operation S4 is described. Refer to Figure 6 , after obtaining the region center in the initialization stage, perform feature extraction f to train the filter, extract feature z after determining the candidate region, calculate the correlation score y, so as to determine the target center, and continuously iterate to achieve tracking.

[0108] DSST introduces a multi-feature fusion mechanism. The features used are HOG (Histogram of Oriented Gradients) + CN (Color Feature) + Grayscale Feature, and a scale filter is added to handle the scale change of the target in the image. Therefore, DSST mainly has two important modules: a position filter and a scale filter. The two filters are independent of each other and complementary to each other.

[0109] The position filter can complete the position estimation during single-target tracking. Assume that the size of the image patch P where the target is located is M×N, extract the features of P (such as Fhog features) from it to obtain a feature f of size M×N×d, where the dimension of the feature is d-dimensional. Then, construct a response g through a Gaussian function, with a size of M×N, the maximum value in the middle, and decreasing gradually towards the surroundings. Perform a two-dimensional discrete Fourier transform (DFT) on each dimension of the feature f to obtain Fl, perform a two-dimensional DFT on g to obtain G, and obtain the filter template H.

[0110] The feature of the new frame of the picture is z. Perform a two-dimensional DFT on each dimension of the feature to obtain Zl, and let each dimension of Z pass through the filter template to obtain the final response matrix y. The position of the maximum value in y is the new center position of the target, completing the target positioning.

[0111] The scale filter can complete the scale estimation during single-target tracking. Assume that the size of the image patch P where the target is located is M×N. Taking the center of the image patch as the center, intercept pictures of different scales according to different scale factors to obtain a series of image patches of different scales. For each image patch, find its feature descriptor (with a dimension of d-dimensional, where d here has no relation to the dimension d in position estimation). g is the output response constructed by a Gaussian function, with a size of 1×S, the maximum value in the middle, and decreasing gradually towards both ends. If 33 scales are selected, then S = 33. Each dimension of the feature fl is a 1×S vector. Perform a one-dimensional DFT on each dimension of the feature f to obtain Fl, perform a one-dimensional DFT on g to obtain G. The calculation process is the same as that of position estimation, except that the dimension becomes one-dimensional, and finally the filter template H is obtained.

[0112] For a new frame of image, centered at the position obtained by the position filter, S image patches of different scales are intercepted, and their feature descriptors are respectively calculated to form a new feature z. The one-dimensional DFT of each dimension feature is obtained to get Zl, and then through the regularization loss, a matrix y containing the responses of each scale is obtained. y is a 1×S vector, and the scale corresponding to the maximum value in y is the result of the final scale estimation.

[0113] The position and size of the target are respectively obtained by using the position filter and the scale filter for tracking, and the y value of the position filter is used as the confidence of the single-object tracking itself. For the single-object tracking result, there are also outputs of the position size and confidence. Preferably, in this embodiment, the tracking confidence of the single-object tracker is: the correlation score of the position of the target tracked by the single-object.

[0114] Operation S5: Using the intersection over union matching method, cross-validation is performed on the target and cumulative confidence output in operation S3 and the target and tracking confidence output in operation S4 to determine whether to output the target tracked in operation S4, or the target with the maximum cumulative confidence output in operation S3, or the fusion result of the targets matched in operations S4 and S3, or not to output.

[0115] In this embodiment, several (such as 3) results with the highest cumulative confidence of the multi-object tracker are cross-validated with the prediction results of the single-object tracking by the intersection over union matching, and the following two specific cross-validation methods are provided as Figure 7 and Figure 8 shown.

[0116] Referring to Figure 7 the cross-validation method shown, the specific process is as follows.

[0117] When the tracking confidence of the single-object tracking is greater than the set value A1 (such as 0.4): If there is no output in operation S3 at this time, only the result of the single-object tracking in operation S4 is output, and the number of uncaught frames remains unchanged; if there is an output in operation S3 at this time, select the top three target detection frames with the highest cumulative confidence to perform intersection over union matching with the single-object tracking frame output in operation S4. If the intersection over union is greater than the set value B1 (such as 0.2), the number of uncaught frames remains unchanged, and the result of the single-object tracking in operation S4 is output. If the intersection over union is not greater than the set value B1, the number of uncaught frames +1, and the result of the single-object tracking in operation S4 is output. At this time, the tracking frame is not updated with the detection result.

[0118] When the tracking confidence of single-object tracking is not greater than the set value A1: If there is no output in operation S3 at this time, then no output is made, and the number of uncaught frames is incremented by 1; If there is an output in operation S3 at this time, select the top three target detection boxes with the highest cumulative confidence and perform an intersection-over-union ratio match with the single-object tracking box output in operation S4. If the intersection-over-union ratio is greater than the set value B2 (such as 0.5), then the number of uncaught frames remains unchanged, the result of single-object tracking in operation S4 is output, and the tracking confidence is output by weighted fusion according to the confidence, and the detection result updates the tracking box. If the intersection-over-union ratio is not greater than the set value B2, then the number of uncaught frames is incremented by 1, no output is made, and the tracking box is not updated with the detection result.

[0119] Refer to Figure 8 the cross-validation method shown, and its specific process is as follows.

[0120] When the tracking confidence of single-object tracking is greater than the set value A2 (such as 0.6): If there is no output in operation S3 at this time, then only the result of single-object tracking in operation S4 is output, and the number of uncaught frames remains unchanged; If there is an output in operation S3 at this time, select the top three target detection boxes with the highest cumulative confidence and perform an intersection-over-union ratio match with the single-object tracking box output in operation S4. If the intersection-over-union ratio is greater than the set value B3 (such as 0.2), then the number of uncaught frames remains unchanged, the result of single-object tracking in operation S4 is output, and the tracking box is not updated with the detection result. If the intersection-over-union ratio is not greater than the set value B3, then the number of uncaught frames is incremented by 1, the result of single-object tracking in operation S4 is output, and the tracking box is not updated with the detection result.

[0121] When the tracking confidence of single-object tracking is between the set value A2 and the set value A3 (such as 0.1), where A3 < A2: If there is no output in operation S3 at this time, then only the result of single-object tracking in operation S4 is output, and the number of uncaught frames is incremented by 1; If there is an output in operation S3 at this time, select the top three target detection boxes with the highest cumulative confidence and perform an intersection-over-union ratio match with the single-object tracking box output in operation S4. If the intersection-over-union ratio is greater than the set value B4 (such as 0.5), then the number of uncaught frames remains unchanged, the result of single-object tracking in operation S4 and the weighted fusion confidence are output, and the tracking box is updated with the detection result. If the intersection-over-union ratio is between the set value B4 and the set value B3, the number of uncaught frames remains unchanged, the weighted fusion result of the confidences of operations S4 and S3 is output, and the tracking box is updated with the detection result. If the intersection-over-union ratio is less than the set value B3, then the number of uncaught frames is incremented by 1, no output is made, and the tracking box is not updated with the detection result.

[0122] When the tracking confidence of single-object tracking is less than the set value A3: If there is no output in operation S3 at this time, no output is made; if there is an output in operation S3 at this time, select the top three target detection frames with the highest cumulative confidence and perform an alternating ratio matching with the single-object tracking frame output in operation S4. If the alternating ratio is greater than the set value B5 (such as 0.8), update the tracking frame with the detection result, increment the number of uncaught frames by 1, and output the weighted fusion result according to a given value (such as 0.5). If the alternating ratio is not greater than the set value B5, increment the number of uncaught frames by 1, do not output, and do not update the tracking frame with the detection result.

[0123] According to an embodiment of the present invention, the method further includes: judging whether each frame of image is captured according to the result of cross-validation in operation S5; when the number of frames of uncaught images within a continuous number of frames reaches the third set value, exit the tracking. If there is a target output in operation S3, re-execute operation S4 and operation S5. If there is no target output in operation S3, no output is made.

[0124] The target capture method in the video sequence scenario provided by the embodiment of the present invention first performs a preliminary detection on the target in the image; then performs a preliminary filtering according to prior knowledge such as the imaging size and aspect ratio of the target in the image; then transmits the result after target detection to the multi-object tracking module, where temporal filtering is realized at this stage, accidental misdetections in the detection are excluded and multi-frame accumulation of confidence is realized, and the output is controlled by setting a cumulative number threshold and a confidence threshold; the output result is sent to the single-object tracking module, and initially the target with the highest confidence is selected for the initialization of single-object tracking to achieve capture. Subsequently, each time the single-object tracking module independently completes the tracking prediction of the captured target and performs cross-validation with the target output by the multi-object tracking to judge whether the target capture and tracking are correct, and perform replacement capture in a timely manner to ensure the correct capture and tracking of the target. This method can solve the following technical problems existing in the prior art: the detection and recognition ability of target detection is weak under interference; more false alarms will occur when the target is small, resulting in inaccurate capture; the single-object tracking ability is limited, and when the target features become weak due to too long a time sequence, the tracking and capture will fail.

[0125] Embodiment Two

[0126] A target capture system in a video sequence scenario includes: a processor; a memory that stores a computer-executable program, and when the program is executed by the processor, the processor executes the above-mentioned target capture method in the video sequence scenario. The related technical solutions are the same as those in Embodiment One and will not be elaborated here.

[0127] Embodiment Three

[0128] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the above-mentioned target capture method in the video sequence scenario. The related technical solutions are the same as those in Embodiment One and will not be elaborated here.

[0129] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for capturing a target in a video sequence scene, characterized in that: include: S1, performs target detection on the input video sequence images to obtain the position and size of all targets in each frame image; S2, designing an imaging size range of the desired target according to its theoretical size in the image, deleting the targets whose sizes in S1 are not within the imaging size range, and outputting the targets that are not deleted; S3, initializing a multi-target tracker using the positions of all targets output in S2 to perform multi-target tracking and cumulative confidence calculation on the video sequence images; During multi-target tracking, for each target tracker, when the number of tracking hits reaches a first set value and the cumulative confidence reaches a set confidence threshold, the tracked target and the cumulative confidence are output; when the number of consecutive tracking misses reaches a second set value, the target tracker is deleted; S4, initializing a single target tracker using the target with the highest cumulative confidence, so as to perform single target tracking and tracking confidence calculation on the video sequence images and output the result; S5, cross-validates the target and cumulative confidence output in S3 and the target and tracking confidence output in S4 using the intersection-over-union matching method to determine whether to output the target tracked in S4, or to output the target with the largest cumulative confidence tracked in S3, or to output the fusion result of the targets matched in S4 and S3, or not to output it.

2. The target capture method in a video sequence scene according to claim 1, characterized in that: The method further comprises: Determine whether each frame image is captured according to the result of the cross-validation in S5; When the number of frames in which the image is not captured in a continuous number of frames reaches a third set value, the tracking is exited. If a target is output in S3, S4 and S5 are re-executed. If no target is output in S3, no output is performed.

3. The target capture method in a video sequence scene according to claim 1, characterized in that: The S1 specifically includes: Preprocess the input video sequence images; Perform object detection on the preprocessed image to obtain a horizontal detection frame; The minimum circumscribed rectangle of the target is found in the horizontal detection frame by using the frame shrinkage to obtain a rotation detection frame, and the position and size of the rotation detection frame are the position and size of the target in the rotation detection frame.

4. The target capture method in a video sequence scene as claimed in claim 3, characterized in that: The step of searching for the minimum bounding rectangle of the target in the horizontal detection frame by shrinking the frame specifically includes: The image in the horizontal detection frame is sequentially subjected to saliency detection, image segmentation, filtering and edge detection to achieve frame shrinkage, and the minimum circumscribed rectangular frame of the frame after frame shrinkage is obtained as the rotation detection frame.

5. The target capture method in a video sequence scene according to claim 1, characterized in that: The imaging size range is: Th wmax =target w ×rate w Th hmax =target h ×rate w Th wmin =target w ÷rate h Th hmin =target h ÷rate h Among them, Th wmax ,Th hmax ,Th wmin ,Th hmin are the upper limit of imaging length, upper limit of imaging width, lower limit of imaging length, and lower limit of imaging width of the desired target, respectively. w 、target h are the theoretical length and width of the target in the image, respectively. The theoretical size includes target w and target h , rate w is the length scaling factor, rate h is the width scaling factor, base is the first adjustable parameter, α is the second adjustable parameter, β is the third adjustable parameter, and the imaging size range includes Th wmax ,Th hmax ,Th wmin and Th hmin .

6. The target capture method in a video sequence scene according to any one of claims 1 to 5, characterized in that: The performing multi-target tracking on the video sequence images specifically includes: For each target tracker, the target tracker is used to track the target of the video sequence image. If the tracked target matches the corresponding target obtained by target detection, the tracking is successful. The position of the corresponding tracked target is corrected by using the target position obtained by target detection. If the tracked target does not match any target obtained by target detection, the tracking is unsuccessful. For each target output in S2, if the target does not match any target obtained by the multi-target tracker, a new target tracker is initialized for multi-target tracking.

7. The target capture method in a video sequence scene according to claim 6, characterized in that: When the tracking hit count is greater than the fourth set value, the cumulative confidence is: cumConf←cumConf×(1-cum_factor)+conf×cum_factor When the tracking hit count is not greater than the fourth set value, the cumulative confidence is: cumConf←1-(1-cumConf)×(1-conf×cum_factor) Among them, cumConf is the cumulative confidence, the left and right sides are the current cumulative confidence and the previous cumulative confidence respectively, cum_factor is the adjustable weight, and conf is the target detection confidence of the current frame.

8. The target capture method in a video sequence scene according to any one of claims 1 to 5, characterized in that: The tracking confidence is: a correlation score of the position of a target tracked by a single target.

9. A target capture system in a video sequence scene, characterized in that: include: processor; A memory storing a computer executable program, wherein when the program is executed by the processor, the processor executes the target capturing method in the video sequence scene according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a target capturing method in a video sequence scenario as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Target tracking method and device, electronic equipment and storage medium

    CN113177968A

  • System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform

    US20200184278A1