Method and apparatus for tracking target

KR103016227B1Active Publication Date: 2026-09-09SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020200179773
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-16
Filing Date
2020-12-21
Publication Date
2026-09-09
Estimated Expiration
2040-12-21

Smart Images

  • Figure 112020138888237-PAT00022_ABST
    Figure 112020138888237-PAT00022_ABST
Patent Text Reader

Abstract

A target tracking method and apparatus are disclosed. According to one embodiment, the target tracking apparatus includes at least one processor. The processor may obtain a first depth feature from a target area image and obtain a second depth feature from a search area image. The processor may obtain a global response diagram between the first depth feature and the second depth feature. The processor may obtain temporary bounding box information based on the global response diagram. The processor may update the second depth feature based on the temporary bounding box information to obtain the updated second depth feature. The processor may obtain a plurality of local feature blocks based on the first depth feature. The processor may obtain a local response diagram based on the plurality of local feature blocks and the updated second depth feature. The processor may obtain output bounding box information based on the local response diagram.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This relates to a technology for tracking targets in images, specifically a technology for tracking targets step-by-step. Background Technology

[0002] Visual object tracking is one of the important fields in computer vision. Visual object tracking involves continuously predicting the bounding box of a target in subsequent frames of a video sequence based on the first frame image and a given bounding box. The target can be an object or a part of an object. means of solving the problem

[0003] According to one embodiment, a target tracking method may include the steps of: obtaining a first depth feature from a target area image and obtaining a second depth feature from a search area image; obtaining a global response diagram between the first depth feature and the second depth feature; obtaining temporary bounding box information based on the global response diagram; updating the second depth feature based on the temporary bounding box information to obtain an updated second depth feature; obtaining a plurality of local feature blocks based on the first depth feature; obtaining a local response diagram based on the plurality of local feature blocks and the updated second depth feature; and obtaining output bounding box information based on the local response diagram.

[0004] The step of acquiring the plurality of local feature blocks includes acquiring the plurality of local feature blocks by additionally extracting a third depth feature from the first depth feature or by dividing the first depth feature, and the step of acquiring the local response diagram may include acquiring the local response diagram based on a fourth depth feature additionally extracted from the updated second depth feature or the correlation between the second depth feature and the plurality of local feature blocks.

[0005] The step of obtaining the local response diagram based on the above correlation may include: obtaining a plurality of local sub-response diagrams based on the correlation between each of the plurality of local feature blocks and the second depth feature or the fourth depth feature; and synthesizing the plurality of local sub-response diagrams to obtain the local response diagram.

[0006] The step of obtaining the local response diagram by synthesizing the plurality of local sub-response diagrams may include: a step of classifying the plurality of local feature blocks into target feature blocks or background feature blocks; and a step of obtaining the local response diagram by synthesizing the local sub-response diagrams based on the classification results.

[0007] The above classification step may include a step of classifying the plurality of local feature blocks into target feature blocks or background feature blocks based on the overlap ratio of each of the temporary bounding box and the plurality of local feature blocks.

[0008] The output bounding box information includes a coordinate offset between the center position coordinates of the temporary bounding box included in the temporary bounding box information and the center position coordinates of the output bounding box, and a size offset between the size of the output bounding box and a preset size. The step of obtaining the output bounding box information may output the temporary bounding box information as the output bounding box information when the sum of the absolute values ​​of the coordinate offsets is greater than a threshold value, and output the result of summing the center position of the temporary bounding box and the coordinate offsets and the result of summing the size of the temporary bounding box and the size offsets as the output bounding box information when the sum of the absolute values ​​of the coordinate offsets is less than or equal to the threshold value.

[0009] The step of obtaining the above temporary bounding box information may output the coordinates having the greatest correlation in the global response diagram of the current frame as the center position coordinates of the above temporary bounding box of the current frame, and output the size of the output bounding box estimated in the previous frame as the size of the above temporary bounding box of the current frame.

[0010] The step of obtaining a plurality of local feature blocks by additionally extracting a third depth feature from the first depth feature or by dividing the first depth feature may involve dividing the first depth feature or the third depth feature so that the plurality of local feature blocks do not overlap, dividing the plurality of local feature blocks so that they overlap, or dividing based on a preset block distribution.

[0011] A computer program according to one embodiment may be stored in a computer-readable recording medium in combination with hardware to execute the method.

[0012] According to one embodiment, a target tracking device may include at least one processor. The processor may acquire a first depth feature from a target area image and acquire a second depth feature from a search area image. The processor may acquire a global response diagram between the first depth feature and the second depth feature. The processor may acquire temporary bounding box information based on the global response diagram. The processor may update the second depth feature based on the temporary bounding box information to acquire the updated second depth feature. The processor may acquire a plurality of local feature blocks based on the first depth feature. The processor may acquire a local response diagram based on the plurality of local feature blocks and the updated second depth feature. The processor may acquire output bounding box information based on the local response diagram. Brief explanation of the drawing

[0013] FIG. 1 is a flowchart illustrating the overall operation of a target tracking method according to one embodiment. FIG. 2 is a sequence illustrating the operation of a target tracking method according to one embodiment. FIG. 3 is a flowchart illustrating a target tracking method according to one embodiment. FIG. 4 is an example of a global correlation operation according to one embodiment, and FIG. 6 is an example of a global correlation operation according to the present invention. FIG. 5 is an exemplary diagram of the first stage of a target tracking method according to one embodiment. FIG. 6 is an exemplary diagram of a block division method according to one embodiment. FIG. 7 is an example diagram of a block correlation operation according to one embodiment. FIG. 8 is an example diagram of a fused interference suppression and response diagram according to one embodiment. FIG. 9 is an example diagram of self-adaptation prediction according to one embodiment. FIG. 10 is an example diagram of a second stage operation of a target tracking method according to one embodiment. FIG. 11 is an example diagram of network training according to one embodiment. FIG. 12 illustrates a block correlation method combined with interference suppression according to one embodiment and a comparison of the difference in effects between the global correlation method and the block correlation method. FIG. 13 is a diagram illustrating the configuration of a target estimation device according to one embodiment. Specific details for implementing the invention

[0014] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, actual implementations are not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or substitutions included in the technical concept described by the embodiments.

[0015] Terms such as "first" or "second" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may be named the first component.

[0016] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or coupled with that other component, or that there may be other components in between.

[0017] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0018] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.

[0019] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are given the same reference numeral regardless of the drawing number, and redundant descriptions thereof will be omitted.

[0021] FIG. 1 is a flowchart illustrating the overall operation of a target tracking method according to one embodiment.

[0022] According to one embodiment, a target tracking device can track a target on a sequence of images using two stages. The target tracking device can determine a target area image (101) and a search area image (103). In the first stage (105), the target tracking device can perform an approximate prediction (107) for an area matching the target area image (101) from the search area image (103). As a result of the approximate prediction (107), provisional bounding box information can be obtained. In the second stage (109), the target tracking device can perform an accurate prediction (111) using the provisional bounding box information. As a result of the accurate prediction (111), output bounding box information can be output.

[0023] The target tracking device performs target tracking by dividing it into two stages and can track targets using block correlation and global correlation. By applying block correlation, the target tracking device can extract information to adjust temporary bounding box information. Through this, the target tracking device can achieve high-accuracy target tracking using fewer resources. Because it uses fewer resources, the target tracking device can perform high-accuracy and stable real-time tracking even in mobile environments.

[0024] In the following description, the input image refers to an image input to a target tracking device. The input image may include, but is not limited to, consecutive frames. The target refers to an object being tracked in the input image. The target area image refers to a reference image representing the target and may be referred to as a template image. The search area image refers to an image serving as the area where the target is searched.

[0025] The first depth feature refers to the features extracted from the target region image. The second depth feature refers to the features extracted from the search region image. Global correlation refers to the correlation operation between all first depth features and the second depth features. The global response diagram refers to the result of the global correlation. The third depth feature refers to the depth feature extracted through additional convolution operations on the first depth feature, and the fourth depth feature refers to the depth feature extracted through additional convolution operations on the second depth feature.

[0026] Local feature blocks refer to the results of segmenting the features of the target region image. Here, the features of the target region may include first depth features or third depth features derived from the first depth features. Local feature blocks do not necessarily have to be rectangular and may include blocks of various shapes. Local correlation refers to the correlation operation between each local feature block and the second depth feature or the updated second depth feature. Local response diagrams refer to the results of local correlation.

[0027] Temporary bounding box information refers to information regarding a temporary bounding box as a result of the first stage, and output bounding box information refers to information regarding an output bounding box as a result of the second stage. Each bounding box information may include location information and size information of a target within the search area.

[0029] FIG. 2 is a sequence illustrating the operation of a target tracking method according to one embodiment.

[0030] In FIG. 2, steps (201) to (205) correspond to the first stage, and steps (207) to (213) correspond to the second stage. Here, the distinction between the first stage and the second stage is for convenience of explanation, and it is not necessary to distinguish between the stages.

[0031] According to one embodiment, in step (201), the target tracking device can obtain a first depth feature from a target area image and obtain a second depth feature from a search area image.

[0032] For example, the target tracking device may acquire consecutive frames. The target tracking device may use a first neural network to set a portion of the first frame image among the consecutive frames as the target area image. For example, the first neural network may be a Siamese twin convolution network, but is not limited thereto. The target tracking device may extract a first depth feature from the target area image. The target tracking device sets the current frame as the search area image and extracts a second depth feature from the search area image. Here, the target area image may be acquired by clipping the first frame image according to an initial bounding box set manually or an output bounding box for the previous frame, but is not limited thereto. Additionally, the first depth feature is a global feature of the target area image, and the second depth feature is a global feature of the search area image.

[0033] In step (203), the target tracking device can obtain a global response diagram between the first depth feature and the second depth feature. For example, the target tracking device can obtain the global response diagram by calculating the global correlation between the first depth feature and the second depth feature.

[0034] When a correlation operation is applied to two images, a response diagram Y representing the similarity between the two images can be obtained. A higher similarity value indicates that the corresponding area of ​​the search area image Z and the target area image X have a higher degree of similarity. For example, the correlation operation can be performed by Equation 1.

[0035]

[0036]

[0037] In mathematical formula 1, h and w represent the size of image X, and i, j, u, and v represent the coordinates of the image, respectively.

[0038] In step (205), the target tracking device can obtain temporary bounding box information based on the global response diagram. The target tracking device can output the coordinates having the greatest correlation in the global response diagram of the current frame as the center position coordinates of the temporary bounding box of the current frame, and output the size of the output bounding box estimated in the previous frame as the size of the temporary bounding box of the current frame.

[0039] In step (207), the target tracking device can obtain the updated second depth feature by updating the second depth feature based on temporary bounding box information. For example, the target tracking device can obtain a search area image of a reduced area by clipping the search area image according to the temporary bounding box. The target tracking device can update the second depth feature by extracting the depth feature from the search area image of the reduced area.

[0040] In step (209), the target tracking device can acquire a plurality of local feature blocks based on the first depth feature. The target tracking device can acquire a plurality of local feature blocks by additionally extracting a third depth feature from the first depth feature or by splitting the first depth feature. The target tracking device can additionally extract a third depth feature by inputting the first depth feature into a second neural network.

[0041] The target tracking device can divide a first depth feature or a third depth feature into a plurality of local feature blocks. The target tracking device can divide the first depth feature or the third depth feature such that the plurality of local feature blocks do not overlap, divide such that the plurality of local feature blocks overlap, or divide based on a preset block distribution.

[0042] Here, the pre-set block distribution may be an artificially determined distribution or a distribution derived by a trained neural network. An artificially determined distribution may include a Gaussian distribution. The neural network outputting the distribution can be trained to find optimized parameters targeting specific distribution parameters, such as the Gaussian distribution (e.g., the mean and variance of the Gaussian distribution).

[0043] In step (211), the target tracking device can obtain a local response diagram based on a plurality of local feature blocks and an updated second depth feature. The target tracking device can obtain a local response diagram based on a fourth depth feature additionally extracted from the updated second depth feature or a correlation between the second depth feature and the plurality of local feature blocks. The target tracking device can input the second depth feature into a second neural network to additionally extract a fourth depth feature.

[0044] The target tracking device can acquire multiple local sub-response diagrams based on the correlation between each of the multiple local feature blocks and the second depth feature or the fourth depth feature.

[0045] The target tracking device can obtain a local response diagram by synthesizing multiple local sub-response diagrams. The target tracking device can classify multiple local feature blocks into target feature blocks or background feature blocks. Multiple local feature blocks can be classified into target feature blocks or background feature blocks based on the overlap ratio between a temporary bounding box and each of the multiple local feature blocks. The target tracking device can obtain a local response diagram by synthesizing local sub-response diagrams based on the classification results.

[0046] In step (213), the target tracking device can obtain output bounding box information based on a local response diagram. The output bounding box information may include a coordinate offset between the center position coordinates of the temporary bounding box included in the temporary bounding box information and the center position coordinates of the output bounding box, and a size offset between the size of the output bounding box and a preset size. If the sum of the absolute values ​​of the coordinate offsets is greater than a threshold, the target tracking device can output the temporary bounding box information as output bounding box information. If the sum of the absolute values ​​of the coordinate offsets is less than or equal to the threshold, the target tracking device can output the result of summing the center position of the temporary bounding box and the coordinate offsets, and the result of summing the size of the temporary bounding box and the size offsets as output bounding box information.

[0048] FIG. 3 is a flowchart illustrating a target tracking method according to one embodiment.

[0049] As illustrated in FIG. 3, the target tracking method comprises two stages. In the first stage, an approximate prediction (107) is performed. In the first stage (105), a target image (101) and a search area image (103) may be determined. The target tracking device may extract global features (311, 312) of the target area image (101) and the search area image (103), respectively, and calculate a global correlation (313) for the extracted global features (311, 312) to obtain a global response diagram. The target tracking device may obtain temporary bounding box information based on the global response diagram.

[0050] In the second stage, an accurate prediction (111) is made. The target tracking device can obtain a plurality of local feature blocks (321, 322) by dividing the features of the target area image, and obtain a local response diagram by calculating the block correlation (323) between the plurality of local feature blocks (321, 322), which are the divided features of the target area image, and the updated search area image features. The target tracking device can output output bounding box information according to the local response diagram.

[0052] FIG. 4 is an example diagram of a global correlation operation according to one embodiment.

[0053] Referring to FIG. 4, the target tracking device has a first depth feature F T (401) and second depth feature F St (402) Correlation operations between can be performed. The target tracking device can perform the first depth feature F T (401) 2nd depth feature F St The correlation at each position can be calculated by sliding on (402). The target tracking device can obtain a global response diagram (403) through global correlation.

[0055] FIG. 5 is an exemplary diagram of the first stage of a target tracking method according to one embodiment.

[0056] In the first stage (105), the target tracking device may output a temporary bounding box. To do this, the target tracking device may perform feature extraction, global correlation (313), and feature clipping (505).

[0057] The target tracking device can extract features for the target area image (101) and the search area image (103), respectively, using a first neural network. The first neural network is a convolutional neural network It may include. The target tracking device uses a convolutional neural network to target area image Z (101). Input into to first depth feature (Z) can be output. The target tracking device uses a convolutional neural network to search area image X (103). Input into to 2nd depth features It can output (X). For example, a convolutional neural network It may be a Siamese twin convolution network, and the parameters of the two branches of Fig. 5 may be shared so that the input images are mapped to the same feature space.

[0058] The target tracking device is the extracted first depth feature (Z) and second depth features A global response diagram f (501) can be output by performing global correlation (313) on (X). Based on the global response diagram f (501), the target tracking device can output temporary bounding box information P1 (503), which is the result of the first stage prediction. The target tracking device uses Equation 2 to [provide] a first depth feature (Z) and second depth features A global response diagram between (X) can be obtained.

[0059]

[0060] The target tracking device can output the location with the largest value in the global response diagram as location information for a temporary bounding box. The target tracking device can output the size of the output bounding box of the previous frame as size information for a temporary bounding box.

[0061] The target tracking device can perform feature clipping (505) on the first stage prediction result P1 (503). The target tracking device can update the second depth feature by extracting depth features from the search area image clipped by feature clipping (505). Consequently, the target tracking device updates the second depth feature updated in the first stage (105). Section 1 Depth Feature (Z) can be output. Temporary bounding box information P1 (503) can be expressed as P1 = (x1, y1, w1, h1). Here, x1 and y1 are the horizontal and vertical coordinates of the center position of the temporary bounding box of the first stage, respectively, and w1 and h1 represent the width and height of the temporary bounding box, respectively.

[0062] The target tracking device can obtain a smaller search area image X' by clipping the search area image X according to the center position and size of the temporary bounding box. The target tracking device extracts depth features of the search area image X' to obtain updated second depth features You can obtain.

[0064] FIG. 6 is an exemplary diagram of a block division method according to one embodiment.

[0065] Referring to FIG. 6, various types of segmentation methods (600) performed by a target tracking device are disclosed. The target tracking device can segment a first depth feature (401) or a third depth feature (not shown) into a plurality of local feature blocks. According to a non-overlapping image segmentation method (601), the target tracking device can segment the first depth feature (401) or the third depth feature such that the plurality of local feature blocks do not overlap. According to an overlapping image segmentation method (603), the target tracking device can segment the first depth feature (401) or the third depth feature such that the plurality of local feature blocks overlap. According to a segmentation method based on a predetermined block distribution (605), the target tracking device can segment the first depth feature (401) or the third depth feature based on a predetermined block distribution.

[0067] FIG. 7 is an example diagram of a block correlation operation according to one embodiment.

[0068] FIG. 7 is illustrated assuming a first depth feature (401) for convenience of explanation, but block correlation operations may be performed on a third depth feature. Referring to FIG. 7, region feature partitioning may be performed on the first depth feature (401). The first depth feature (401) may be partitioned into a plurality of local feature blocks.

[0069] The target tracking device can calculate the correlation between each local feature and the second depth feature (402). As a result of the calculation, the target tracking device may output a plurality of local sub-diagrams (701, 702, 703, 704, 705, 706, 707, 708, 709). The target tracking device can obtain a local response diagram (711) by synthesizing the plurality of local sub-diagrams (701, 702, 703, 704, 705, 706, 707, 708, 709).

[0070] For example, the target tracking device can classify each local feature block among a plurality of local feature blocks into a target feature block or a background feature block. The target tracking device can obtain a local sub-response diagram corresponding to the target feature block and a local sub-response diagram corresponding to the background feature block. The target tracking device can synthesize all local sub-response diagrams to output a local response diagram (711).

[0071] In a target area image, a background area exists in addition to the target area, and the characteristics of the background area can affect the stability and accuracy of target tracking. The target tracking device can improve the stability and accuracy of target tracking through a method such as that shown in Fig. 7. The target tracking device can effectively reduce interference caused by the background by classifying multiple local feature blocks into target feature blocks and background feature blocks and then synthesizing them.

[0073] FIG. 8 is an example diagram of a fused interference suppression and response diagram according to one embodiment.

[0074] The target tracking device can obtain a local response diagram by synthesizing a plurality of local sub-response diagrams. The target tracking device can classify a plurality of local feature blocks into target feature blocks or background feature blocks. Based on the overlap ratio of each of the temporary bounding box (801) and the plurality of local feature blocks (801, 802, 803, 804, 805, 806, 807, 808, 809), the plurality of local feature blocks (801, 802, 803, 804, 805, 806, 807, 808, 809) can be classified into target feature blocks or background feature blocks.

[0075] The target tracking device can classify each local feature block (801, 802, 803, 804, 805, 806, 807, 808, 809) as a target feature block or a background feature block depending on the proportion that each local feature block (801, 802, 803, 804, 805, 806, 807, 808, 809) occupies in the overlapping area between each local feature block (801, 802, 803, 804, 805, 806, 807, 808, 809) and the temporary bounding box (801). For example, based on a corrected temporary bounding box (801) on a target area image, if a local feature block has an area of ​​p% or more within the temporary bounding box, it can be classified as a target feature block, and if the overlap between the local feature block (801, 802, 803, 804, 805, 806, 807, 808, 809) and the temporary bounding box (801) is less than p%, it can be classified as a background feature block. Here, p may be a predetermined threshold value.

[0076] The target tracking device can obtain a local response diagram by synthesizing the local sub-response diagram corresponding to the target feature block and the local sub-response diagram corresponding to the background feature block using mathematical formula 3.

[0077]

[0078] In mathematical formula 3, S is a local response diagram, So is a local sub-response diagram corresponding to a target feature block, Sb is a sub-response diagram corresponding to a background feature block, no is the number of target feature blocks, and nb is the number of background feature blocks.

[0079] The target tracking device can output an output bounding box based on a local response diagram. The target tracking device can predict the position offset and size offset of the temporary bounding box (801) according to the local response diagram. The target tracking device can output an output bounding box according to the predicted position offset and size offset.

[0080] For example, the target tracking device can process the local response diagram using a third neural network and predict the position offset and size offset of the output bounding box. The third neural network may be different from the first and second neural networks mentioned above. Here, the prediction result of the output bounding box may include position information and size information of the target bounding box. Hereinafter, the process of outputting the output bounding box based on the local response diagram may be referred to as a self-adaptive prediction process.

[0082] FIG. 9 is an example diagram of self-adaptation prediction according to one embodiment.

[0083] The target tracking device can process the local response diagram (S) (901) using a convolutional neural network (902, 903). The target tracking device can output an offset D (905) through an offset prediction (904). The target tracking device can perform a self-adaptive prediction (907) based on the first stage prediction result P1 (906) and the offset D (905). As a result of the self-adaptive prediction (907), a second stage prediction result P2 (908) can be output.

[0084] The target tracking device can predict an offset D=(dx, dy, dw, dh), and the offset includes a position offset and a size offset. For example, the position offset may be a coordinate offset between the center position coordinates of the output bounding box of the second stage and the center position coordinates of the temporary bounding box of the first stage, and the size offset may be a size offset between the output bounding box of the second stage and a predetermined bounding box.

[0085] The target tracking device can obtain a prediction result of the output bounding box of the second stage based on the predicted position offset and size offset. If the sum of the absolute values ​​of the coordinate offsets is greater than a preset threshold, the target tracking device can output the prediction result of the temporary bounding box of the first stage as the prediction result of the output bounding box of the second stage.

[0086] If the sum of the absolute values ​​of the coordinate offsets is less than or equal to a preset threshold, the target tracking device can obtain a predicted result of the output bounding box of the second stage by adding the center position of the temporary bounding box of the first stage and the predicted position offset, and adding the size of the predetermined bounding box and the predicted size offset.

[0087] For example, if the prediction result of the temporary bounding box of the first stage is P1=(x1, y1, w1, h1) and the size of the pre-specified bounding box is (w0, h0)(width w0, height h0), the prediction result of the output bounding box of the second stage may be P2=(x1+dx, y1+dy, w0+dw, h0+dh).

[0089] FIG. 10 is an example diagram of a second stage operation of a target tracking method according to one embodiment.

[0090] The target tracking device has a first depth feature through the first stage. and updated second depth features It can acquire. The target tracking device has a first depth feature. and updated second depth features It can be input into a convolution network (1003). The target tracking device can perform block equilibrium (1004). The target tracking device can perform interference suppression and fuse local sub-response diagrams. The target tracking device can obtain local sub-response diagrams (1006). The target tracking device can perform self-adaptive prediction (1007). The target tracking device can output the prediction result P2 of the second stage.

[0092] FIG. 11 is an example diagram of network training according to one embodiment.

[0093] A target tracking device can track a target using a cascade network (including a first neural network, a second neural network, and a third neural network). The cascade network can be trained using multiple supervised signals. Here, the multiple supervised signals include a global response diagram, a local response diagram, and a target bounding box.

[0094] The steps performed in the training process below can be applied in a similar manner to some extent in the inference process. Multiple supervised signals can be used to optimize the loss value of the loss function. Multiple supervised signals can be used to learn the network parameters through iterative recurrent learning.

[0095] In a learning process using multiple supervised signals, the learning device may obtain a global response diagram (1103) by performing a first stage tracing (1101) on a template image (1101) and a search area image (1102). In the global response diagram, if the distance from the center position is less than a specific threshold, it may be set to +1, and if the distance from the center position is greater than a specific threshold, it may be set to -1. The learning device may output an approximate prediction box (1105) based on the global response diagram (1103). The learning device may output a shared feature (1104) based on the result of the first stage tracing (1102) and the clipped approximate prediction box (1105).

[0096] The learning device can perform a second stage tracking (1108). The learning device can obtain a segmentation result (obtained by a segmentation algorithm or manually) for a target on the search area image. The learning device can perform a distance transformation on the segmentation result and numerically normalize the distance transformation map to obtain a supervised signal of a local response diagram (1109).

[0097] The learning device can obtain accurate prediction results through self-adaptive prediction (1106). During the learning process, the global response diagram, the local response diagram (1109), and the target bounding box can be used as supervised signals, and the parameters of the cascade network can be learned by optimizing the loss function through iterative recurrent learning.

[0098] For example, the learning device can obtain temporary bounding box information by inputting image pairs (including a template image and a search area image) extracted from the same video sequence into the neural network of the first stage. The learning device can calculate the loss (Loss0) between the predicted value and the actual value of the global response diagram using Binary Cross Entropy.

[0099] The learning device can obtain a local response diagram (1109) by inputting the shared features (1104) of the first stage into the neural network of the second stage based on the approximate prediction results. The learning device can obtain the accurate prediction results of the second stage based on the local response diagram. The learning device can measure the loss (Loss1) between the predicted value and the actual value of the local response diagram using KL divergence (Kullback-Leibler Divergence). The learning device can measure the loss (Loss2) between the accurately predicted bounding box and the actual box using L1 distance. The learning device can learn the parameters of the neural network by optimizing the loss (Loss = Loss0 + (a1) * Loss1 + (a2) * Loss2). Here, a1 and a2 represent the weights of the loss, respectively.

[0101] FIG. 12 illustrates a block correlation method combined with interference suppression according to one embodiment and a comparison of the difference in effects between the global correlation method and the block correlation method.

[0102] Referring to FIG. 12, a global correlation response diagram (1201), a block correlation response diagram (1202), and the result (1203) of interference suppression performed on the block correlation response diagram are shown. By performing interference suppression based on block correlation, the target tracking device can effectively extract detailed information about the target and further improve the accuracy of tracking.

[0104] FIG. 13 is a diagram illustrating the configuration of a target estimation device according to one embodiment.

[0105] According to one embodiment, the target tracking device includes at least one processor.

[0106] The processor can acquire a first depth feature from a target region image and a second depth feature from a search region image. The processor can acquire a global response diagram between the first depth feature and the second depth feature. The processor can acquire temporary bounding box information based on the global response diagram. The processor can update the second depth feature based on the temporary bounding box information to acquire the updated second depth feature. The processor can acquire multiple local feature blocks based on the first depth feature. The processor can acquire a local response diagram based on the multiple local feature blocks and the updated second depth feature. The processor can acquire output bounding box information based on the local response diagram.

[0108] For the time being, a target tracking method and a target tracking device according to one embodiment have been described with reference to FIGS. 1 through 13. It should be understood that each operation of the device illustrated in FIG. 13 may be implemented by software, hardware, firmware, or any combination thereof to perform a specific function. For example, the operation of the target estimation device may be implemented by a specialized integrated circuit, pure software code, or a module combining software and hardware. For example, the target tracking device may be, but is not limited to, a PC computer, a tablet device, a personal information terminal, a smartphone, a web application, or other device capable of executing program commands.

[0109] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0110] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.

[0111] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination, and the program instructions recorded on the medium may be those specifically designed and configured for the embodiment or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0112] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0113] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based thereon. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0114] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

Claim 1 A target tracking method comprising: a step of obtaining a first depth feature from a target area image and obtaining a second depth feature from a search area image; a step of obtaining a global response diagram between the first depth feature and the second depth feature; a step of obtaining temporary bounding box information based on the global response diagram; a step of updating the second depth feature based on the temporary bounding box information to obtain an updated second depth feature; a step of obtaining a plurality of local feature blocks based on the first depth feature; a step of obtaining a local response diagram based on the plurality of local feature blocks and the updated second depth feature; and a step of obtaining output bounding box information based on the local response diagram. Claim 2 A target tracking method according to claim 1, wherein the step of acquiring the plurality of local feature blocks includes acquiring the plurality of local feature blocks by dividing the first depth feature or a third depth feature additionally extracted from the first depth feature, and the step of acquiring the local response diagram includes acquiring the local response diagram based on the correlation between the second depth feature and the plurality of local feature blocks or a fourth depth feature additionally extracted from the updated second depth feature. Claim 3 A target tracking method according to claim 2, wherein the step of obtaining the local response diagram based on the correlation comprises: obtaining a plurality of local sub-response diagrams based on the correlation between each of the plurality of local feature blocks and the second depth feature or the fourth depth feature; and synthesizing the plurality of local sub-response diagrams to obtain the local response diagram. Claim 4 A target tracking method according to claim 3, wherein the step of obtaining a local response diagram by synthesizing the plurality of local sub-response diagrams comprises: a step of classifying the plurality of local feature blocks into target feature blocks or background feature blocks; and a step of obtaining a local response diagram by synthesizing the local sub-response diagrams based on the classification results. Claim 5 A target tracking method according to claim 4, wherein the classifying step comprises classifying the plurality of local feature blocks into target feature blocks or background feature blocks based on the overlap ratio of each of the temporary bounding box and the plurality of local feature blocks. Claim 6 In paragraph 3, the output bounding box information includes a coordinate offset between the center position coordinates of the temporary bounding box included in the temporary bounding box information and the center position coordinates of the output bounding box, and a size offset between the size of the output bounding box and a preset size, and the step of obtaining the output bounding box information is to output the temporary bounding box information as the output bounding box information when the sum of the absolute values ​​of the coordinate offsets is greater than a threshold value, and when the sum of the absolute values ​​of the coordinate offsets is less than or equal to the threshold value, output the result of summing the center position of the temporary bounding box and the coordinate offsets and the result of summing the size of the temporary bounding box and the size offsets as the output bounding box information, a target tracking method. Claim 7 In claim 6, the step of obtaining the temporary bounding box information comprises outputting the coordinate having the greatest correlation in the global response diagram of the current frame as the center position coordinate of the temporary bounding box of the current frame, and outputting the size of the output bounding box estimated in the previous frame as the size of the temporary bounding box of the current frame, a target tracking method. Claim 8 In paragraph 2, the step of obtaining a plurality of local feature blocks by additionally extracting a third depth feature from the first depth feature or by dividing the first depth feature is a target tracking method in which the first depth feature or the third depth feature is divided such that the plurality of local feature blocks do not overlap, divided such that the plurality of local feature blocks overlap, or divided based on a preset block distribution. Claim 9 A computer program stored on a computer-readable recording medium in combination with hardware to execute the method of any one of claims 1 through 8. Claim 10 A target tracking device comprising at least one processor, wherein the processor obtains a first depth feature from a target area image, obtains a second depth feature from a search area image, obtains a global response diagram between the first depth feature and the second depth feature, obtains temporary bounding box information based on the global response diagram, updates the second depth feature based on the temporary bounding box information to obtain an updated second depth feature, obtains a plurality of local feature blocks based on the first depth feature, obtains a local response diagram based on the plurality of local feature blocks and the updated second depth feature, and obtains output bounding box information based on the local response diagram.

Citation Information

Patent Citations

  • Enhanced siamese trackers

    US20180129934A1

  • Hybrid and self-aware long-term object tracking

    US20190147602A1

  • System and method for siamese instance search tracker with a recurrent neural network

    US20190332935A1