A target tracking method based on deep optical flow

Through the deep neural network model, deep optical flow technology is integrated into the target tracking technology, end-to-end object detection and tracking is achieved, solving the problem of poor robustness of the existing technology in complex scenarios, and improving the real-time and robustness of tracking.

CN113888604BActive Publication Date: 2025-06-03ANHUI TSINGLINK INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111138039.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-06-03
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

The existing target tracking technology is poorly robust in complex scenarios, making it difficult to achieve long-term tracking, and at the same time, the calculation is expensive and slow.

Method used

A deep optical flow-based target tracking method is adopted to process the object detection and tracking process end-to-end through the deep neural network model, and the continuous tracking of the target is achieved using feature extraction, detection and optical flow modules.

Benefits of technology

Without increasing the calculation cost, end-to-end object detection and tracking is achieved, with high versatility, real-time and robustness, and can be tracked for a long time, with excellent tracking results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113888604B_ABST
    Figure CN113888604B_ABST
Patent Text Reader

Abstract

The present invention discloses a target tracking method based on deep optical flow, belonging to the technical field of target tracking, which includes: S31, selecting an initial tracking image as the previous frame image; S32, forming a motion image pair with the previous frame image and the current frame image and inputting it into a deep neural network model to predict the positions of all targets in the current frame image and the corresponding motion optical flow field; S33, updating the positions of the targets to be tracked in the previous frame image according to the positions of all targets in the current frame image and the corresponding motion optical flow field to obtain an updated frame image; S34, taking the updated frame image as the previous frame image, and repeating steps S32 to S33 to achieve continuous tracking of the target. The present invention realizes end-to-end target detection and tracking without increasing the computational cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and particularly relates to a target tracking method based on deep optical flow. Background Art

[0002] Target tracking refers to determining the boundary position of a target of interest in the current frame image based on its boundary position in the previous frame image according to spatio-temporal correlation. It is a core technology in the field of computer vision, with a very wide range of application fields and is a necessary technology for many downstream applications, such as action analysis, behavior recognition, monitoring, and human-computer interaction, etc.

[0003] Currently, target tracking technologies are mainly divided into two categories, specifically as follows:

[0004] (1) Target tracking technologies based on traditional technologies, and the representative technologies mainly include Kalman filter tracking, optical flow method tracking, template matching tracking, TLD tracking, CT tracking, KCF tracking, etc. The advantages of this type of technology are simple principles and relatively fast operating speeds, and good results can be obtained in relatively simple scenarios, suitable for short-term tracking; its drawback is poor robustness, and it is easy to lose the target and mis-track the target in slightly more complex scenarios and cannot adapt to long-term tracking.

[0005] (2) Target tracking technologies based on deep learning technologies, and this type of technology mainly adopts the strategy of object detection plus object matching to complete the target tracking process. The process is to first locate the target position in each frame of image with the help of a powerful deep learning object detection framework (such as: faster-rcnn, ssd, yolo), and then associate the same target in the front and rear frame images with the help of the nearest neighbor matching algorithm or feature vector matching algorithm, thereby completing the target tracking process. The advantages of this type of technology are strong robustness and the ability to perform long-term tracking; its disadvantage is over-reliance on the object detection framework, the target running speed cannot be too fast, and the superposition of the two-step algorithm is time-consuming. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects existing in the prior art and achieve end-to-end object detection and tracking without increasing the computational cost.

[0007] To achieve the above purpose, the present invention adopts a target tracking method based on deep optical flow, including:

[0008] S31. Select the initial tracking image as the previous frame image;

[0009] S32. Form a motion image pair with the previous frame image and the current frame image and input it into a deep neural network model to predict the positions of all targets in the current frame image and the corresponding motion optical flow field;

[0010] S33. Update the positions of the targets to be tracked in the previous frame image according to the positions of all the targets in the current frame image and the corresponding motion optical flow fields, and obtain the updated frame image;

[0011] S34. Take the updated frame image as the previous frame image, and repeat steps S32 - S33 to achieve continuous tracking of the targets.

[0012] Furthermore, the deep neural network model includes a feature extraction module, a detection module, and an optical flow module;

[0013] The feature extraction module is used to obtain the high - level feature maps of the motion image pair;

[0014] The detection module is used to predict whether there are targets in the current frame image and the positions of all the targets according to the high - level feature maps;

[0015] The optical flow module is used to predict the motion optical flow field of the motion image pair based on the high - level feature maps.

[0016] Furthermore, the feature extraction module includes a concatenation layer concat, a backbone network backbone, a feature pyramid network FPN, and output feature layers out_feature1 and out_feature2. The motion image pair is used as the input of the concatenation layer concat. The concatenation layer concat concatenates the motion image pair along the channel dimension and outputs a concatenated image. The output of the concatenation layer concat is connected to the backbone network backbone and the feature pyramid network FPN. The feature pyramid network FPN outputs feature maps fused with features of different scales, and outputs them through the output feature layers out_feature1 and out_feature2.

[0017] Furthermore, the detection module includes convolutional layers dconv1_0, dconv2_0, dconv1_1, dconv2_1, and a target information parsing layer yolo. The outputs of the output feature layer out_feature1 and the output feature layer out_feature2 are respectively connected to the convolutional layers dconv1_0 and dconv2_0. The convolutional layers dconv1_0 and dconv2_0 are respectively connected to the target information parsing layer yolo through the convolutional layers dconv1_1 and dconv2_1. The target information parsing layer yolo is used to extract effective target position information.

[0018] Further, the optical flow detection module outputs the forward optical flow field and the backward optical flow field of the motion image pair. The optical flow detection module includes a splicing layer concat1, upsampling layers upsample0, upsample1, upsample2, and convolutional layers lconv0, lconv1, lconv2, lconv3. The output feature layers out_feature1 and out_feature2 are respectively connected to the splicing layer concat1 and the upsampling layer upsample0. The upsampling layer upsample0 is connected to the splicing layer concat1. The splicing layer concat1, the upsampling layer upsample1, the convolutional layer lconv1, the upsampling layer upsample2, the convolutional layer lconv2, and the convolutional layer lconv3 are connected in sequence.

[0019] Further, updating the position of the target to be tracked in the previous frame image according to the positions of all targets in the current frame image and the corresponding motion optical flow field to obtain the updated frame image includes:

[0020] Obtaining the rough position of the target to be tracked in the current frame image according to the position of the target to be tracked in the previous frame image and the motion optical flow field;

[0021] Performing association matching according to the rough position of the target to be tracked in the current frame image and the positions of all targets in the current frame image to obtain the accurate position of the target to be tracked in the current frame image;

[0022] Generating the updated frame image according to the accurate position of the target to be tracked in the current frame image.

[0023] Further, obtaining the rough position of the target to be tracked in the current frame image according to the position of the target to be tracked in the previous frame image and the motion optical flow field includes:

[0024] Reducing the positions of the detected targets in the current frame image and intercepting the corresponding optical flow regions in all motion optical flow fields;

[0025] Comparing the motion displacement amounts and directions of the same sub-pixels in the forward motion optical flow field and the backward motion optical flow field of the intercepted corresponding optical flow regions to determine the correctly tracked pixels;

[0026] Obtaining the motion displacement amount of the target to be tracked by a statistical method according to the correctly tracked pixels;

[0027] Accumulating the motion displacement amount of the target to be tracked in the previous frame image to obtain the rough position of the target to be tracked in the current frame image.

[0028] Further, before inputting the previous frame image and the current frame image as a motion image pair into the deep neural network model to predict the positions of all targets in the current frame image and the corresponding motion optical flow field, the following steps are also included:

[0029] Collect pedestrian videos;

[0030] Annotate the pedestrian motion position information for each frame image in the pedestrian video, and construct a set of motion image pairs;

[0031] Use the set of motion image pairs to train the deep neural model and learn the model parameters.

[0032] Further, the step of annotating the pedestrian motion position information for each frame image in the pedestrian video and constructing a set of motion image pairs includes:

[0033] Obtain and annotate the target positions in each frame image of the pedestrian video;

[0034] Randomly select a frame image containing a target as the previous frame image, and randomly select an image frame after the previous frame image as the current frame image, and form a motion image pair with the previous frame image;

[0035] Downsample each generated motion image pair, and based on the optical flow field generation tool, obtain the forward motion optical flow field and the backward motion optical flow field of the image pair and annotate them, and construct the set of motion image pairs.

[0036] Further, the loss function L used during the training of the deep neural network model is:

[0037] L = αL loc + βL offset

[0038] where L loc represents the detection loss function, L offset represents the loss function of the optical flow module, and α and β represent the weighting coefficients.

[0039] Compared with the prior art, the present invention has the following technical effects: By means of a deep neural network model, in the object detection framework based on deep learning, an object matching strategy is incorporated. Without almost any increase in computational cost, end-to-end object detection and tracking can be achieved, with strong generality, high real-time performance, fewer error sources, the ability to track for a long time, and strong robustness of the tracking effect. Description of the Drawings

[0040] The following combines the drawings to describe the specific embodiments of the present invention in detail:

[0041] Figure 1 is the overall structure diagram of the deep neural network model;

[0042] Figure 2 It is the network structure diagram of the feature extraction module;

[0043] Figure 3 It is the network structure diagram of the detection and tracking module;

[0044] Figure 4 This is the network structure diagram of the optical flow tracking module;

[0045] Figure 5 It is the target tracking flow chart.

[0046] Among them, the logo on the left side of each neural network structure layer graph indicates the output feature map size of the network structure: feature map width × feature map height × number of feature map channels. DETAILED DESCRIPTION

[0047] In order to further illustrate the features of the present invention, please refer to the following detailed description and drawings of the present invention. The drawings are for reference and illustration only and are not intended to limit the scope of protection of the present invention.

[0048] This embodiment is applicable to all multi-target tracking scenarios. For the convenience of description, the present invention takes pedestrian multi-target tracking as an example. This embodiment discloses a target tracking method based on deep optical flow, including the following steps:

[0049] S1. Design of deep neural network model:

[0050] The deep neural network model designed by the present invention has the main function of directly completing the detection and tracking of pedestrian targets in each frame of the image with the help of a deep neural network model with a fusion mechanism. Since the steps of pedestrian detection and positioning and pedestrian association matching are no longer deliberately distinguished, the entire pedestrian tracking system has a faster operation speed, fewer sources of error, and a more robust tracking effect. The present invention adopts a convolutional neural network (CNN). In order to facilitate the description of the present invention, some terms are defined: feature map resolution refers to feature map height × feature map width, feature map size refers to feature map width × feature map height × feature map channel number, kernel size refers to kernel width × kernel height, span refers to width direction span × height direction span, and each convolution layer is followed by a batch normalization layer and a nonlinear activation layer.

[0051] like Figure 1As shown in the figure, the present invention is optimized and improved based on the single-stage object detection framework yolov4-tiny with excellent performance. The designed deep neural network model includes four modules: the feature extraction module backbone module, the detection module detect module, the optical flow module flow module, and the update module update module. Among them, the update module update module does not participate in training and only works during testing. The specific design steps are as follows:

[0052] S11. Design the feature extraction module backbone module:

[0053] The feature extraction module is mainly used to obtain high-level features of the input moving image pair with high abstraction and rich expression ability. The quality of high-level feature extraction directly affects the performance of subsequent pedestrian target tracking. The present invention adopts the same feature extraction module as yolov4-tiny, as Figure 2 shown. The input of this feature extraction network is a moving image pair, which is composed of 2 three-channel RGB images with a resolution of 320×320. Among them, one image is the current frame image and the other is the previous frame image. Concat is a splicing layer, and its main function is to splice the input 2 three-channel RGB images into a 6-channel image with the same resolution according to the channel dimension. Backbone is the backbone network of yolov4-tiny, and FPN is the feature pyramid network, which is mainly used to fuse features of different scales. The specific network structure is the same as that of yolov4-tiny. Out_feature1 and out_feature2 are the output feature layers of the feature extraction module, which are used for subsequent detection and tracking of pedestrian targets. Among them, the resolution of the feature map of out_feature1 is 20x20x384, and the resolution of the feature map of out_feature2 is 10x10x256.

[0054] S12. Design the detection module detect module:

[0055] The detection module is mainly based on the feature map output by the feature extraction module to predict whether there is a pedestrian target in the current frame image and the position of the pedestrian target. The present invention improves on the detection module of yolov4-tiny, and the specific network structure is as Figure 3As shown in the figure, both dconv1_0 and dconv2_0 are convolutional layers with a kernel size of 3x3 and a stride of 1x1, and both dconv1_1 and dconv2_1 are convolutional layers with a kernel size of 1x1 and a stride of 1x1. The yolo layer is a pedestrian target information parsing layer used to extract effective pedestrian target information and only works during testing. The yolo layer is the same as this layer in yolov4-tiny. The resolution of the feature map of the yolo layer is Nx5, where N represents the number of detected pedestrian targets.

[0056] S13. Design the optical flow module flow module:

[0057] The optical flow module mainly predicts the motion displacement amount and direction of each pixel between the front and back frames of the input image pair based on the feature map output by the feature extraction module, that is, predicts the motion optical flow field of the input image pair. The present invention adopts a forward and backward bidirectional optical flow field design, and the optical flow prediction is more accurate. The specific network structure is as Figure 4 shown in the figure. Concat1 is a splicing layer, and its main function is to splice multiple input feature maps into an output feature map according to the channel dimension. Upsample0, upsample1, and upsample2 are all 2-fold upsampling layers, and the specific principle is the same as that in the yolov4-tiny structure. Lconv0, lconv1, and lconv2 are all convolutional layers with a kernel size of 3x3 and a stride of 1x1. Lconv3 is a convolutional layer with a kernel size of 1x1 and a stride of 1x1, and its output feature map represents the corresponding motion optical flow field of the input image pair. Among them, the first output feature map represents the motion optical flow field in the x coordinate direction from the previous frame image to the current frame image in the input image pair, the second output feature map represents the motion optical flow field in the y coordinate direction from the previous frame image to the current frame image in the input image pair, the third output feature map represents the motion optical flow field in the x coordinate direction from the current frame image to the previous frame image in the input image pair, and the fourth output feature map represents the motion optical flow field in the y coordinate direction from the current frame image to the previous frame image in the input image pair. The first 2 output feature maps form the forward motion optical flow field, and the last 2 output feature maps form the backward motion optical flow field.

[0058] S13. Design the update module update module:

[0059] The update module mainly calculates the accurate position of the pedestrian target to be tracked in the previous frame image in the current frame image according to the output information of the detection module and the optical flow module, and updates the information of the previous frame image accordingly. The specific steps are as follows:

[0060] S131. Obtain the rough new position of the pedestrian target to be tracked. Mainly based on the position of the pedestrian target to be tracked in the previous frame image, add the corresponding motion offset obtained by the optical flow module, and then the rough position of the pedestrian target to be tracked in the current frame image can be obtained. To obtain the corresponding motion offset according to the output information of the optical flow module, the specific method is as follows:

[0061] S1311. Obtain the motion optical flow field of the pedestrian target to be tracked. The main method is to reduce the position of the pedestrian target detected in the current frame image by 4 times, and then intercept the corresponding optical flow area in all the output motion optical flow fields of the optical flow module.

[0062] S1312. Obtain the correct motion pixels. Mainly in the forward motion optical flow field and the backward motion optical flow field, compare the motion displacement and direction of the pixels at the same position. When both the motion displacement and the motion direction have small differences, it is considered that the pixel is correctly tracked.

[0063] S1313. Obtain the running displacement of the pedestrian target. The main method is to obtain the accurate running pixels according to step S1312, and obtain the motion displacement of the pedestrian target through the statistical method.

[0064] S132. Obtain the accurate new position of the pedestrian target to be tracked. The main method is to perform associative matching between the rough new position of the pedestrian target to be tracked obtained in step S131 and each pedestrian target position obtained by the detection module, and select the best-matched pedestrian target position as the accurate new position of the pedestrian target to be tracked in the current frame image. Among them, the IOU (Intersection over Union) matching degree adopted by the present invention is used as the associative matching function, and any other similarity metric function can also be used as the associative matching function. When the matching degree of the best-matched pedestrian target is greater than a certain threshold, it is considered that the pedestrian target in the current frame image has a corresponding historical record in the previous frame image, that is, the pedestrian target in the current frame image is a trackable pedestrian target. When the matching degree of the best-matched pedestrian target is lower than a certain threshold, it is considered that the pedestrian target in the current frame image has no corresponding historical record in the previous frame image, that is, the pedestrian target in the current frame image is a newly emerged pedestrian target.

[0065] S133. Update the pedestrian target to be tracked. Mainly generate a new previous frame image and the existing pedestrian targets in the new previous frame according to the pedestrian target position in the current frame image. First, turn the current frame image into the previous frame image, and then update the known pedestrian target position information in the previous frame image according to the trackable pedestrian targets and the newly emerged pedestrian targets obtained in step S132; when a certain pedestrian target in the previous frame image is not associated with the pedestrian target in the current frame image in step S132, it is considered that the pedestrian target has disappeared from the video screen, and the corresponding tracking record should be deleted.

[0066] S2. Train the deep neural network model:

[0067] After the deep neural network model is designed, the next step is to collect pedestrian video images in various scenarios, input them into the deep neural network model, and learn the relevant model parameters. The specific steps are as follows:

[0068] S21. Collect pedestrian videos, mainly collect pedestrian videos in various scenarios, various lights, and various angles.

[0069] S22. Label the pedestrian movement position information, mainly label the pedestrian position information in each frame of the video and the movement information between different frames of moving images. The specific steps are as follows:

[0070] S221. Label the pedestrian target position information. The main method is to use the existing pedestrian detection framework based on deep learning to obtain the pedestrian position in each frame of the video as the pedestrian position information.

[0071] S222. Form moving image pairs. Mainly turn the video into an image sequence. Arbitrarily select a frame of image containing the pedestrian target as the previous frame image, and then within the next 120 frames, arbitrarily select an image as the current frame image, and form a moving image pair with the previous frame image.

[0072] S223. Obtain the movement information of the pedestrian target. The main method is to first perform 4-fold downsampling on each moving image pair, and then based on the existing optical flow field generation tool, obtain the forward movement optical flow field and the backward movement optical flow field of the image pair. Among them, each movement optical flow field is represented by 2 grayscale images with the same image resolution, representing the movement optical flow field in the x coordinate direction and the movement optical flow field in the y coordinate direction of the movement optical flow field respectively.

[0073] S23. Train the deep neural network model, input the sorted set of moving image pairs into the defined deep neural network model, and learn the relevant model parameters. The loss function L during network model training is shown in the following formula:

[0074] L = αL loc + βL offset

[0075] Where L loc represents the detection loss function, and the meaning of this loss function is the same as that of yolov4-tiny. L offset represents the loss function of the optical flow module. This loss function uses the mean square error loss function. α and β represent the weighting coefficients.

[0076] S3. After training the deep neural network model, the next step is to use the model in a real environment for pedestrian tracking. For any given frame of pedestrian image, form a pair of motion images and feed them into the trained deep neural network model to directly output the new positions of pedestrian targets in the current frame image. Repeat this process to achieve continuous tracking of multiple pedestrian targets. As Figure 5 shown, the specific steps are as follows:

[0077] S31. Select the initial tracking image, mainly by arbitrarily selecting a frame of pedestrian image as the previous frame image.

[0078] S32. Predict the motion information of pedestrian positions in the current frame image. The main method is to form a pair of images consisting of the previous frame image and the current frame image, and feed them into the deep neural network model to directly predict the positions of all pedestrian targets in the current frame image and the forward and backward motion optical flow fields of this pair of images.

[0079] S33. Update the positions of the pedestrian targets to be tracked. Mainly according to the pedestrian target positions and the corresponding motion optical flow fields predicted in step S32, with the help of the update module, obtain the new previous frame image and the new existing pedestrian targets.

[0080] S34. Continuously track. Repeat steps S32 - S33 to achieve continuous tracking of pedestrian targets.

[0081] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A target tracking method based on deep optical flow, characterized in that, it includes: S31. Select the initial tracking image as the previous frame image; S32. Compose the previous frame image and the current frame image into a motion image pair and input it into the deep neural network model to predict the positions of all targets and the corresponding motion optical flow field in the current frame image; S33. Update the position of the target to be tracked in the previous frame image according to the positions of all targets and the corresponding motion optical flow field in the current frame image to obtain the updated frame image; S34. Take the updated frame image as the previous frame image and repeat steps S32 - S33 to achieve continuous tracking of the target; In step S33, it specifically includes: Obtain the rough position of the target to be tracked in the current frame image according to the position of the target to be tracked in the previous frame image and the motion optical flow field; Perform association matching according to the rough position of the target to be tracked in the current frame image and the positions of all targets in the current frame image to obtain the accurate position of the target to be tracked in the current frame image; Generate the updated frame image according to the accurate position of the target to be tracked in the current frame image; The obtaining the rough position of the target to be tracked in the current frame image according to the position of the target to be tracked in the previous frame image and the motion optical flow field includes: Reduce the positions of the detected targets in the current frame image and intercept the corresponding optical flow regions in all motion optical flow fields; In the forward motion optical flow field and the backward optical flow field of the intercepted corresponding optical flow regions, compare the motion displacement amounts and directions of the same sub - pixels to determine the correctly tracked pixels; Obtain the motion displacement amount of the target to be tracked by the statistical method according to the correctly tracked pixels; Accumulate the motion displacement amount of the target to be tracked to the position of the target to be tracked in the previous frame image to obtain the rough position of the target to be tracked in the current frame image.

2. The target tracking method based on deep optical flow according to claim 1, characterized in that, the deep neural network model includes a feature extraction module, a detection module and an optical flow module; The feature extraction module is used to obtain the high - level feature map of the motion image pair; The detection module is used to predict whether there are targets in the current frame image and the positions of all targets according to the high - level feature map; The optical flow module is used to predict the motion optical flow field of the motion image pair based on the high - level feature map.

3. The target tracking method based on deep optical flow according to claim 2, characterized in that, The feature extraction module includes a concatenation layer "concat", a backbone network "backbone", a Feature Pyramid Network "FPN", and output feature layers "out_feature1" and "out_feature2". The pair of motion images serves as the input to the concatenation layer "concat". The concatenation layer "concat" concatenates the pair of motion images along the channel dimension and outputs a concatenated image. The output of the concatenation layer "concat" is connected to the backbone network "backbone" and then to the Feature Pyramid Network "FPN". The Feature Pyramid Network "FPN" outputs a feature map that fuses features of different scales and is output through the output feature layers "out_feature1" and "out_feature2".

4. The method for object tracking based on deep optical flow according to claim 3, wherein, the detection module includes convolutional layers "dconv1_0", "dconv2_0", "dconv1_1", "dconv2_1" and an object information parsing layer "yolo". The outputs of the output feature layer "out_feature1" and the output feature layer "out_feature2" are respectively connected to the convolutional layers "dconv1_0" and "dconv2_0". The convolutional layers "dconv1_0" and "dconv2_0" are respectively connected to the object information parsing layer "yolo" through the convolutional layers "dconv1_1" and "dconv2_1". The object information parsing layer "yolo" is used to extract effective object position information.

5. The method for object tracking based on deep optical flow according to claim 3, wherein, the output of the optical flow module is the forward optical flow field and the backward optical flow field of the pair of motion images. The optical flow module includes a concatenation layer "concat1", upsampling layers "upsample0", "upsample1", "upsample2" and convolutional layers "lconv0", "lconv1", "lconv2", "lconv3". The output feature layers "out_feature1" and "out_feature2" are respectively connected to the concatenation layer "concat1" and the upsampling layer "upsample0". The upsampling layer "upsample0" is connected to the concatenation layer "concat1". The concatenation layer "concat1", the upsampling layer "upsample1", the convolutional layer "lconv1", the upsampling layer "upsample2", the convolutional layer "lconv2", and the convolutional layer "lconv3" are connected in sequence.

6. The method for object tracking based on deep optical flow according to claim 1, wherein, before inputting the previous frame image and the current frame image as a pair of motion images into the deep neural network model to predict the positions of all objects in the current frame image and the corresponding motion optical flow field, it further includes: collecting pedestrian videos; annotating the pedestrian motion position information for each frame image in the pedestrian videos to construct a set of pairs of motion images; using the set of pairs of motion images to train the deep neural model to learn the model parameters.

7. The method for object tracking based on deep optical flow according to claim 6, wherein, Annotating the pedestrian motion position information for each frame image in the pedestrian video, and constructing a set of motion image pairs, including: Obtaining and annotating the target position in each frame image of the pedestrian video; Randomly selecting a frame image containing the target as the previous frame image, and randomly selecting an image frame after the previous frame image as the current frame image, and forming a motion image pair with the previous frame image; Downsampling each generated motion image pair, and based on an optical flow field generation tool, obtaining the forward motion optical flow field and the backward motion optical flow field of the image pair and performing annotation, and constructing the set of motion image pairs.

8. The target tracking method based on depth optical flow according to claim 6, wherein, the loss function L used during the training of the deep neural network model is: Among them, represents the loss function of detection, represents the loss function of the optical flow module, , represents the weighting coefficient.

Citation Information

Patent Citations

  • Object tracking method and system based on depth characteristic flow, terminal and medium

    CN108242062A

  • Unmanned aerial vehicle aerial video moving small target real-time detection and tracking method

    CN109785363A