Target Tracking Method, Device, Electronic Device, and Storage Medium

Through inter-sampling and intra-multi-scale sampling, inter-frame correlation information and multi-scale distribution information with high correlation are obtained, which solves the problems of low accuracy and target loss in target tracking, and achieves higher tracking accuracy and accuracy.

CN114202559BActive Publication Date: 2025-05-27BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010986849.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-18
Publication Date
2025-05-27
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

The existing method of target tracking through convolutional neural networks has problems with low accuracy or accuracy, and it is prone to loss of tracking targets.

Method used

By inter-sampling the image sequence, inter-correlation information with high correlation degree is obtained, and multi-scale sampling of the sampled inter-frame images is performed to obtain multi-scale distribution information of the target image in the frame, and target positioning is combined with this information.

Benefits of technology

Improve the accuracy and accuracy of the target tracking task, and avoid the occurrence of tracking target loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202559B_ABST
    Figure CN114202559B_ABST
Patent Text Reader

Abstract

The present application relates to a target tracking method, device, electronic device and storage medium. When performing target tracking on the current frame image in an image sequence, the present application performs inter-frame sampling on the image sequence with an inverse correlation between the sampling density at different positions and the time interval from different positions to the current frame image, and extracts inter-frame correlation information with a high degree of correlation with the current frame image. The inter-frame correlation information with a high degree of correlation can guide or assist the current frame image to perform more accurate / precise positioning of the target; and further perform intra-frame multi-scale sampling on each of the inter-frame images sampled in the above manner, which can obtain the multi-scale distribution information of the target in the intra-frame image, improve the understanding of a single image frame, and overcome the situation such as the loss of the tracked target easily caused by the fixed receptive field of the convolution kernel in the convolutional neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision, and particularly relates to an object tracking method, device, electronic device, and storage medium. Background Art

[0002] Object tracking is a hot topic in the field of computer vision. By processing an image sequence to track a moving object within the camera's field of view, it provides semantic or non-semantic information support for researching the motion law of the object, understanding behavior, event detection, etc., and can include single-object tracking, multi-object tracking, cross-camera multi-object tracking, and so on.

[0003] Currently, generally, traditional methods such as optical flow analysis are used to achieve object tracking, or convolutional neural network methods are used to achieve object tracking. Among them, when using traditional methods such as optical flow analysis for object tracking, it has the advantage of simple calculation, but the extracted feature information (such as gray-scale features) is insufficient, and the accuracy of object tracking is limited; when using convolutional neural network methods for object tracking, complex feature information can be extracted, which to a certain extent solves the defect of insufficient feature information existing in traditional methods such as optical flow analysis, but there are still problems such as low accuracy or precision in object tracking tasks, and situations such as the loss of the tracked object are likely to occur. Summary of the Invention

[0004] In view of this, this application provides an object tracking method, device, electronic device, and storage medium, which are used to solve the problems existing in object tracking by convolutional neural network methods, improve the accuracy / precision of object tracking tasks, and avoid situations such as the loss of the tracked object.

[0005] The specific technical solutions are as follows:

[0006] An object tracking method, comprising:

[0007] Determine the current frame image to be processed in the image sequence; the image sequence includes multiple frames of images arranged in the order of imaging time, and the image contents of at least some of the multiple frames of images are related;

[0008] Perform a first sampling process on the image sequence to obtain a set of first sampling images corresponding to the current frame image; the sampling density at different positions in the image sequence is inversely related to the time interval from the different positions to the current frame image;

[0009] Perform a second sampling process on each of the first sampling images to obtain a set of second sampling images corresponding to each of the first sampling images, each of which includes multiple scales of second sampling images;

[0010] In the scale space composed of the set of first sampled images and the sets of second sampled images corresponding to each first sampled image, based on a single set of inter-frame sampled images with the same scale, determine the positioning result of the current frame image for the target to be tracked at the same scale, so as to obtain the respective positioning results of the target to be tracked at different scales corresponding to the current frame image;

[0011] Based on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image, locate the target to be tracked in the current frame image.

[0012] Optionally, the first sampling process on the image sequence to obtain a set of first sampled images corresponding to the current frame image includes:

[0013] Perform inter-frame sampling on the image sequence in a Gaussian distribution manner;

[0014] Among them, in the inter-frame sampling in the Gaussian distribution manner, the current frame image corresponds to the center of the Gaussian distribution, and the sampling density at different positions in the image sequence presents a Gaussian distribution state centered on the current frame image.

[0015] Optionally, the second sampling process on each first sampled image respectively includes:

[0016] Perform multiple downsampling operations on each first sampled image to obtain multiple second sampled images with different scales corresponding to each first sampled image;

[0017] Among them, the second sampled images with different scales respectively correspond to different resolutions.

[0018] Optionally, the step of determining the positioning result of the current frame image for the target to be tracked at the same scale in the scale space composed of the set of first sampled images and the sets of second sampled images corresponding to each first sampled image, so as to obtain the respective positioning results of the target to be tracked at different scales corresponding to the current frame image, includes:

[0019] Input the set of first sampled images and the sets of second sampled images corresponding to each first sampled image into a pre-constructed convolutional neural network model including multiple convolutional layers; different convolutional layers correspond one by one to different scales in the scale space; each convolutional layer processes at least a set of inter-frame sampled images at the corresponding scale, so as to use the set of inter-frame sampled images to guide the positioning of the current frame image for the target to be tracked at the corresponding scale;

[0020] Obtain the output information of the convolutional neural network model;

[0021] Among them, the output information includes the respective positioning results of the target to be tracked at different scales corresponding to the current frame image; the positioning result of the target to be tracked at one scale corresponding to the current frame image includes: the positioning position information of at least one target to be tracked corresponding to the one scale in the current frame image and / or the class labels of at least one target to be tracked respectively.

[0022] Optionally, in the method:

[0023] The input of the first convolutional layer among the multiple convolutional layers is the set of first sampled images;

[0024] The input of the other convolutional layers except the first convolutional layer among the multiple convolutional layers is: the stacked result of the set of sampled images corresponding to the corresponding scale and the feature image output by the previous convolutional layer.

[0025] Optionally, based on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image, positioning the target to be tracked in the current frame image includes:

[0026] Performing non-maximum suppression processing on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image to locate the target to be tracked in the current frame image.

[0027] A target tracking device includes:

[0028] A first determination unit for determining the current frame image to be processed in the image sequence; the image sequence includes multiple frames of images arranged in the order of imaging time, and the image contents of at least some of the multiple frames of images are related;

[0029] A first sampling unit for performing first sampling processing on the image sequence to obtain a set of first sampled images corresponding to the current frame image; the sampling density at different positions in the image sequence is inversely correlated with the time interval from the different positions to the current frame image;

[0030] A second sampling unit for performing second sampling processing on each of the first sampled images respectively to obtain a set of second sampled images including multiple scales corresponding to each of the first sampled images respectively;

[0031] A second determination unit for determining the positioning result of the current frame image for the target to be tracked at the same scale in the scale space composed of the set of first sampled images and the sets of second sampled images corresponding to each of the first sampled images, so as to obtain the respective positioning results of the target to be tracked at different scales corresponding to the current frame image;

[0032] A positioning unit, configured to locate a target to be tracked in the current frame image based on respective positioning results of the target to be tracked at different scales corresponding to the current frame image.

[0033] Optionally, the first sampling unit is specifically configured to:

[0034] Perform inter-frame sampling on the image sequence in a Gaussian distribution manner;

[0035] Wherein, in the inter-frame sampling in the Gaussian distribution manner, the current frame image corresponds to the center of the Gaussian distribution, and the sampling density at different positions in the image sequence presents a Gaussian distribution state centered on the current frame image.

[0036] Optionally, the second sampling unit is specifically configured to:

[0037] Perform multiple downsampling operations on each first sampled image to obtain multiple second sampled images with different scales corresponding to each first sampled image;

[0038] Wherein, the second sampled images with different scales respectively correspond to different resolutions.

[0039] Optionally, the second determination unit is specifically configured to:

[0040] Input the group of first sampled images and the groups of second sampled images corresponding to each first sampled image into a pre-constructed convolutional neural network model including multiple convolutional layers; different convolutional layers correspond to different scales in the scale space one by one; each convolutional layer processes at least a group of inter-frame sampled images corresponding to the corresponding scale, so as to use the group of inter-frame sampled images to guide the positioning of the target to be tracked in the current frame image at the corresponding scale;

[0041] Obtain the output information of the convolutional neural network model;

[0042] Wherein, the output information includes respective positioning results of the target to be tracked at different scales corresponding to the current frame image; the positioning result of the target to be tracked at one scale corresponding to the current frame image includes: positioning position information of at least one target to be tracked corresponding to the one scale in the current frame image and / or class labels of at least one target to be tracked respectively.

[0043] Optionally, in the device:

[0044] The input of the first convolutional layer of the convolutional neural network model is the group of first sampled images;

[0045] The input of other convolutional layers except the first convolutional layer of the convolutional neural network model is: the stacked result of a set of sampled images at the corresponding scale and the feature image output by the previous convolutional layer.

[0046] An electronic device, comprising:

[0047] A memory for storing a computer instruction set;

[0048] A processor for implementing the target tracking method as described in any one of the above by executing the instruction set stored on the memory.

[0049] A computer-readable storage medium storing a computer instruction set, and when the computer instruction set is executed by a processor, the target tracking method as described in any one of the above is implemented.

[0050] When performing target tracking on the current frame image in the image sequence, the target tracking method, device, electronic device, and storage medium provided in the embodiments of the present application perform inter-frame sampling on the image sequence with an "inverse correlation between the sampling density at different positions and the time interval from different positions to the current frame image", and extract highly correlated inter-frame correlation information for the current frame image. The highly correlated inter-frame correlation information can guide or assist the current frame image to perform more accurate / precise positioning of the target; and further perform intra-frame multi-scale sampling on each inter-frame image sampled in the above manner, and can obtain the multi-scale distribution information of the target in the intra-frame image, improving the understanding of a single image frame, and overcoming the situation such as the loss of the tracked target easily caused by the fixed receptive field of the convolutional kernel in the convolutional neural network. Therefore, by combining the highly correlated inter-frame correlation information and the multi-scale distribution information of the intra-frame image, the present application can at least improve the accuracy / precision of the target tracking task and avoid the occurrence of situations such as the loss of the tracked target. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0052] Figure 1 It is a schematic flowchart of a target tracking method provided by an embodiment of the present application;

[0053] Figure 2 It is a schematic distribution diagram of the sampling density at different positions in the image sequence during Gaussian inter-frame sampling provided by an embodiment of the present application;

[0054] Figure 3 It is a schematic diagram for comparing the effects of a multi-scale second sampled image obtained by repeatedly downsampling a first sampled image with the first sampled image provided by an embodiment of the present application;

[0055] Figure 4 It is a schematic diagram of a single set of inter-frame sampled images at the same scale provided by an embodiment of the present application;

[0056] Figure 5 It is another schematic flow diagram of a target tracking method provided by an embodiment of the present application;

[0057] Figure 6 It is a schematic diagram of input information of different convolutional layers in a convolutional neural network model provided by an embodiment of the present application;

[0058] Figure 7 It is a schematic structural diagram of a target tracking device provided by an embodiment of the present application;

[0059] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0060] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0061] By using the method of convolutional neural network for target tracking, complex feature information (deep features) can be extracted. Commonly used methods mainly include processing methods such as SiamFC and FlowTrack. However, the inventor has found through research that the existing methods for target tracking by convolutional neural network still have the following defects:

[0062] In the existing method, object tracking of video frame images is achieved by inputting the original frame images in the video into a convolutional neural network. When using the convolutional neural network for feature extraction, the receptive field of the convolutional kernel remains fixed. However, in object tracking of image sequences such as videos, the size of the region where the object is located in each frame image often varies, and there are also often differences in the change intensity and the scale of the changed region of the object in different frame images. For example, when an athlete is playing football in a video, since the shooting camera is fixed, the distance between the athlete and the camera will change over time, which will cause the size of the athlete in different frame images of the video to be different, and there are also differences in the change intensity and the scale of the changed region of the athlete in different frame images. When using the convolutional neural network for feature extraction, the receptive field of the convolutional kernel is fixed. It is difficult for the fixed receptive field of the convolutional kernel to adapt to accurately extract features of an object with continuously changing size, different change intensities, and different changed regions in each frame image of the video, which easily leads to the loss of object tracking during the tracking process and the omission of object positioning in some images;

[0063] At the same time, in the existing method of object tracking through a convolutional neural network, an equal-interval extraction method is used to sample the inter-frame images of the video, which cannot reflect the differences in the correlation of inter-frame information, resulting in problems of low accuracy or precision in the object tracking task.

[0064] To at least solve the above defects existing in the existing method of object tracking through a convolutional neural network, the embodiments of the present application provide an object tracking method, device, electronic device, and storage medium.

[0065] Refer to Figure 1 , which shows a schematic flowchart of an object tracking method provided by an embodiment of the present application. The object tracking method of the embodiment of the present application can be applied to, but is not limited to, terminal devices such as mobile phones, tablets, and personal PCs (such as laptops, all-in-ones, and desktops) with computer vision processing functions, or corresponding physical machines such as private cloud / public cloud platforms and servers (computer vision processing servers) with computer vision processing functions.

[0066] As Figure 1 shown, in this embodiment, the object tracking method includes the following processing steps:

[0067] Step 101: Determine the current frame image to be processed in the image sequence; the image sequence includes multiple frame images arranged in the order of imaging time, and the image contents of at least some of the multiple frame images are related.

[0068] The image sequence to be subjected to target tracking may be, but is not limited to: a complete video downloaded or captured by a device camera, or a video clip intercepted from a complete video, or a set of image sequences extracted from a complete video / video clip, or a set of image sequences captured by a device camera in chronological order.

[0069] It is easy to understand that the image contents of at least some of the multiple frames of images included in the image sequence to be subjected to target tracking are related. Essentially, the image sequence includes multiple frames of images corresponding to different time points in the time sequence. The image contents of the multiple frames of images included in the image sequence change over time and are related to a certain extent. Specifically, it is like the video example of the athlete kicking a football as described above.

[0070] When performing target tracking on an image sequence such as a video, usually, multiple frames of images in the image sequence are processed frame by frame, such as processing frame by frame in a serial or parallel manner based on the time order, etc. By positioning the target to be tracked for each frame of image, the target tracking of the image sequence is realized.

[0071] The target to be tracked can be any object or one or more components of an object within the camera's field of view, such as objects like people, animals, and objects, or one or more components of these objects (such as, a human face), etc.

[0072] Step 102: Perform a first sampling process on the image sequence to obtain a set of first sampled images corresponding to the current frame image.

[0073] In a set of image sequences (such as, a video), at different time periods, the target usually has different states (which may include the position of the target in the image and the shape / pose of the target), and the states of the target in different frame images that are closer in time are closer. Similarly, there are greater differences in the states of the target in different frame images with a larger time interval, and even the target has disappeared or does not exist. That is, the greater the correlation between the image contents of the images with a smaller time interval in the image sequence, and the smaller the correlation between the image contents of the images with a larger time interval.

[0074] In view of this, when performing target tracking, for the current frame image to be processed (to be subjected to target positioning), the present application proposes to perform the following first sampling process on the image sequence:

[0075] In the image sequence, the sampling is denser at positions closer to the current frame image and sparser at positions farther from the current frame image, such that the sampling density at different positions in the image sequence is inversely correlated with the time interval from the different positions to the current frame image.

[0076] That is, it is considered that the images with a smaller time interval from the current frame image in the image sequence have a greater correlation with the current frame image, and the images with a larger time interval from the current frame image have a smaller correlation with the current frame image.

[0077] In implementation, optionally, the image sequence can be sampled between frames in a Gaussian distribution manner.

[0078] Among them, as Figure 2 shown, in the inter-frame sampling in this Gaussian distribution manner, the current frame image corresponds to the center of the Gaussian distribution. The sampling density at different positions in the image sequence presents a Gaussian distribution state centered on the current frame image. The sampling interval is smaller and the sampling is denser at positions closer to the current frame image, and the sampling interval is larger and the sampling is sparser at positions farther from the current frame image.

[0079] As another implementation manner, multiple different time periods can also be pre-divided for the image sequence, and a relatively small sampling time interval can be set for the time period close to the current frame image, and a relatively large sampling time interval can be set for the time period far from the current frame image. As the time interval between different time periods and the current frame image becomes larger (or smaller), the sampling time interval corresponding to different time periods is continuously increased (or decreased), so as to extract as much as possible the inter-frame correlation information with a high degree of correlation with the current frame image in the image sequence.

[0080] Based on the above sampling process, a group of first sampling images corresponding to the current frame image can be obtained. This group of first sampling images includes the current frame image to be target-located currently.

[0081] Step 103: Perform a second sampling process on each of the first sampling images to obtain a group of second sampling images corresponding to each of the first sampling images, where each group of second sampling images includes multiple scales of second sampling images.

[0082] In order to eliminate the adverse effects of the fixed receptive field of the convolution kernel in the convolutional neural network on target tracking, after the image sequence is inter-frame first sampled based on a Gaussian distribution or other methods, the embodiments of the present application continue to perform intra-frame multi-scale sampling on each image in the group of first sampling images obtained after the first sampling, that is, the above-mentioned second sampling process.

[0083] By performing intra-frame multi-scale sampling processing on each of the first sampling images, the multi-scale distribution information of the target in the intra-frame image of each first sampling image can be obtained. Through the intra-frame multi-scale distribution information of the target, the adaptability of the convolution kernel to the size change of the target between different frames during the feature extraction process based on the fixed receptive field can be improved.

[0084] Optionally, in this embodiment, the intra-frame multi-scale sampling process (second sampling process) is performed on each first sampled image by using the image pyramid method.

[0085] An image pyramid is a form of multi-scale representation of an image, and is an effective but conceptually simple structure for interpreting an image at multiple resolutions. The pyramid of an image is a set of images arranged in a pyramid shape with gradually decreasing resolutions and derived from the same original image. It is obtained by successive downsampling until a certain termination condition is reached (for example, reaching a predetermined number of sampling times, or the scale of the sampled image reaching a predetermined scale), and the sampled images at each layer are likened to a pyramid. The higher the level, the smaller the image scale and the lower the resolution.

[0086] Therefore, the second sampling process for the first sampled image may include:

[0087] Performing multiple downsampling operations on each first sampled image to obtain multiple second sampled images with different scales corresponding to each first sampled image; wherein, the second sampled images with different scales respectively correspond to different resolutions.

[0088] In this downsampling operation, specifically, multiple downsamplings may be respectively performed on the pixel points in the first sampled image based on different pixel point intervals / scale factors to obtain multiple sampled images with different scales; however, this is not limited thereto, and downsampling may also be performed on the original image, and multiple iterative downsamplings may be performed on the obtained sampled image to obtain multiple sampled images with different scales. Herein, the scale may be understood as a scale factor used to control the scaling ratio in the X and Y (horizontal and vertical) directions of the image. The sampled images with different scales respectively correspond to different resolutions. The larger the pixel point interval (the smaller the scale factor) based on which the downsampling is performed, the smaller the scale and the lower the resolution of the obtained sampled image.

[0089] After the above downsampling operation, a set of second sampled images in a pyramid shape corresponding to each first sampled image can be obtained. Refer to Figure 3 the provided effect comparison diagram of the original image of a certain first sampled image and the sampled images obtained after multiple downsamplings. Among them, image (a) is the original image of the first sampled image, and images (b), (c), and (d) are second sampled images with different scales obtained by respectively performing different downsamplings on the first sampled image. The scales of the second sampled images (b), (c), and (d) obtained by downsampling gradually decrease, and the resolutions gradually decrease.

[0090] Step 104. In the scale space composed of the set of first sampled images and the sets of second sampled images corresponding to each first sampled image, based on a single set of inter-frame sampled images with the same scale, determine the positioning result of the target to be tracked in the current frame image at the same scale, so as to obtain the respective positioning results of the target to be tracked at different scales corresponding to the current frame image.

[0091] As described above, in a set of image sequences (such as a video), the states of the target in different images with closer time are more similar, while there are greater differences in the states of the target in images with a larger time interval, and even the target has disappeared or does not exist. That is to say, the greater the correlation between the image contents of images with a smaller time interval in the image sequence, and the smaller the correlation between the image contents of images with a larger time interval. Combining the correlation characteristics between images at different positions in the image sequence, inter-frame images, especially the front and rear frame images of the current frame image (multiple images that are closer to the current frame image in time before and after the current frame image), can play a guiding role in the positioning of the target in the current frame image.

[0092] On this basis, in this embodiment, specifically, each single set of inter-frame sampled images with the same scale in the scale space composed of the set of first sampled images and the sets of second sampled images corresponding to each first sampled image is used to guide the positioning of the target in the current frame image at this scale. The specific forms of each single set of inter-frame sampled images with the same scale in this scale space are as Figure 4 shown. Corresponding to each scale in this scale space, the respective positioning results of the current frame image for the target at each scale can be obtained accordingly.

[0093] Step 105. Based on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image, locate the target to be tracked in the current frame image.

[0094] It is easy to understand that for the positioning results of the target in the current frame image at each scale, there may be a situation where the positioning accuracy / precision of the target is poor at one or several scales, or the target tracking is lost, resulting in the inability to locate the target at this scale. For example, this situation may occur due to the influence of the convolution kernel with a fixed receptive field described above.

[0095] However, since during target tracking, further multi-scale sampling is performed on each first sampled image obtained by inter-frame sampling, covering the multi-scale distribution information of the target within the frame. Therefore, among the positioning results of the target in the current frame image at each scale, there will also be a situation where the positioning accuracy / precision of the target is high at one or several scales. By comprehensively considering the positioning results of the target in the current frame image at each scale, it is finally possible to achieve high-accuracy / high-precision positioning of the target to be tracked in the current frame image.

[0096] Specifically, the positioning results of the target to be tracked at different scales corresponding to the current frame image can be processed by non-maximum suppression to determine the target to be tracked in the current frame image.

[0097] In this embodiment, when performing target tracking on the current frame image in the image sequence, by performing inter-frame sampling on the image sequence where "the sampling density at different positions is inversely correlated with the time interval from different positions to the current frame image", inter-frame correlation information with a high degree of correlation with the current frame image is extracted for the current frame image. The inter-frame correlation information with a high degree of correlation can guide or assist the current frame image to perform more accurate / precise positioning of the target; and intra-frame multi-scale sampling is further performed on each of the inter-frame images sampled in the above manner, and the multi-scale distribution information of the target in the intra-frame image can be obtained, improving the understanding of a single image frame and overcoming the situation where the tracked target is easily lost due to the fixed receptive field of the convolutional kernel in the convolutional neural network. Therefore, by combining the inter-frame correlation information with a high degree of correlation and the multi-scale distribution information of the intra-frame image, this application can at least improve the accuracy / precision of the target tracking task and avoid the occurrence of situations such as the loss of the tracked target.

[0098] The following provides Figure 1 An optional implementation manner of step 104 in the shown target tracking method. Specifically, through a convolutional neural network, each single-group inter-frame sampling image with the same scale is used to guide the current frame image to perform target positioning at this scale.

[0099] Refer to Figure 5 Another process schematic diagram of the shown target tracking method. In this implementation manner based on the convolutional neural network, step 104 can be implemented as the following processing process:

[0100] Step 501: Input the group of first sampling images and each group of second sampling images corresponding to each first sampling image into a pre-constructed convolutional neural network model including multiple convolutional layers;

[0101] Step 502: Obtain the output information of the convolutional neural network model.

[0102] In implementation, a convolutional neural network model can be pre-constructed based on the training process.

[0103] When constructing the model, an image sequence used as a training dataset, such as a series of videos, can be processed as shown in the previous embodiment. For the current video frame to be processed in the image sequence, perform inter-frame first sampling in the image sequence, and further perform intra-frame second sampling on each of the first sampled images obtained from the inter-frame sampling. On this basis, input a set of first sampled images corresponding to the current frame image obtained from the two types of sampling and each group of second sampled images corresponding to each first sampled image into the model. The output of the model is the localization result of the target at each scale corresponding to the current frame image. By matching the model output result with the pre-calibrated target localization result and based on the matching situation, feedback to the model and adjust the weights of each dimension of the convolution kernels in each convolutional layer of the model to achieve one training of the model. Loop through the above processing process for each frame image in the training dataset such as a series of videos, so that the training process of the model is iteratively repeated until the set target is reached (for example, the matching degree between the model output result and the pre-calibrated target localization result reaches the expectation), and the model construction is completed.

[0104] For the implementation where the first sampling is inter-frame Gaussian sampling and the second sampling is pyramidal multi-scale sampling, the specific input of the model for one time is the image set obtained by inter-frame Gaussian sampling and intra-frame pyramidal sampling of the current frame image.

[0105] Based on the constructed convolutional neural network model, when it is necessary to track the target in the image sequence, for the current frame image to be tracked, input a set of first sampled images (inter-frame Gaussian sampled images) corresponding to it and each group of second sampled images (intra-frame pyramidal multi-scale sampled images) corresponding to each first sampled image into the model, and the localization results of the target at different scales corresponding to the current frame image output by the model can be obtained.

[0106] Among them, the constructed convolutional neural network model includes multiple convolutional layers, and different convolutional layers correspond one by one to different scales in the scale space (the scale space composed of the set of first sampled images and each group of second sampled images corresponding to each first sampled image); each convolutional layer processes at least a set of inter-frame sampled images at the corresponding scale, so that the set of inter-frame sampled images is used to guide the localization of the target to be tracked in the current frame image at the corresponding scale.

[0107] Further, refer to Figure 6 , in one model processing process for the current frame image:

[0108] The input of the first convolutional layer of the convolutional neural network model is:

[0109] The set of first sampled images corresponding to the current frame image;

[0110] The input to the convolutional layers of the convolutional neural network model, except for the first convolutional layer, is as follows:

[0111] The concatenation result of a set of sampled images at the corresponding scale and the feature image output by the previous convolutional layer.

[0112] In the model output information, the localization result of the target at a scale corresponding to the current frame image may include: the localization position information of at least one target to be tracked corresponding to the one scale in the current frame image and / or the class labels of at least one target to be tracked.

[0113] In implementation, optionally, the model can specifically output the localization windows of the target corresponding to each scale in the current frame image and the class labels associated with each localization window respectively. The class labels, such as "0", "1", etc., each class label indicates a corresponding class of the target. For example, label "0" indicates "person", and label "1" indicates "football", etc.

[0114] In the embodiments of the present application, for the current frame image to be target-tracked in the image sequence, by performing inter-frame first sampling (e.g., inter-frame Gaussian sampling) on the image sequence and intra-frame second sampling (e.g., intra-frame pyramidal multi-scale sampling) on each of the first sampled images obtained by the sampling, the highly correlated correlation information of the current frame image in the video sequence can be extracted, which can play a better guiding role in the target localization of the current frame image, improve the localization accuracy / precision of the target in the current frame image. At the same time, it avoids sampling and extracting the correlation information with low correlation, reduces the computational amount during the analysis of the entire image sequence in the target tracking process, reduces the use of computing resources. In addition, it can also obtain the multi-scale distribution information of the target in the intra-frame image, and can avoid situations such as the loss of the tracked target easily caused by the fixed receptive field of the convolutional kernel in the convolutional neural network.

[0115] Corresponding to the above image denoising method, the embodiments of the present application also disclose a target tracking device. Refer to Figure 7 the structural schematic diagram of the target tracking device shown, the device may include:

[0116] The first determination unit 701 is used to determine the current frame image to be processed in the image sequence; the image sequence includes multiple frames of images arranged in the order of imaging time, and the image contents of at least some of the multiple frames of images are related;

[0117] The first sampling unit 702 is used to perform first sampling processing on the image sequence to obtain a set of first sampled images corresponding to the current frame image; the sampling density at different positions in the image sequence is inversely correlated with the time interval from the different positions to the current frame image;

[0118] A second sampling unit 703, configured to perform second sampling processing on each first sampled image respectively, so as to obtain a group of second sampled images corresponding to each first sampled image respectively, where each second sampled image includes multiple scales.

[0119] A second determination unit 704, configured to determine, in a scale space formed by the group of first sampled images and each group of second sampled images corresponding to each first sampled image, a positioning result of the current frame image for the target to be tracked at the same scale based on a single group of inter-frame sampled images having the same scale, so as to obtain each positioning result of the target to be tracked at different scales corresponding to the current frame image.

[0120] A positioning unit 705, configured to locate the target to be tracked in the current frame image based on each positioning result of the target to be tracked at different scales corresponding to the current frame image.

[0121] In an optional implementation manner of the embodiment of the present application, the first sampling unit 702 is specifically configured to:

[0122] Perform inter-frame sampling on the image sequence in a Gaussian distribution manner.

[0123] Wherein, in the inter-frame sampling in the Gaussian distribution manner, the current frame image corresponds to the center of the Gaussian distribution, and the sampling density at different positions in the image sequence presents a Gaussian distribution state centered on the current frame image.

[0124] In an optional implementation manner of the embodiment of the present application, the second sampling unit 703 is specifically configured to:

[0125] Perform multiple downsampling operations on each first sampled image to obtain multiple second sampled images with different scales corresponding to each first sampled image.

[0126] Wherein, the second sampled images with different scales respectively correspond to different resolutions.

[0127] In an optional implementation manner of the embodiment of the present application, the second determination unit 704 is specifically configured to:

[0128] Input the group of first sampled images and each group of second sampled images corresponding to each first sampled image into a pre-constructed convolutional neural network model including multiple convolutional layers; different convolutional layers correspond to different scales in the scale space one by one; each convolutional layer processes at least a group of inter-frame sampled images at the corresponding scale, so as to use the group of inter-frame sampled images to guide the positioning of the current frame image for the target to be tracked at the corresponding scale.

[0129] Obtain the output information of the convolutional neural network model;

[0130] Among them, the output information includes the respective localization results of the target to be tracked at different scales corresponding to the current frame image; the localization result of the target to be tracked at one scale corresponding to the current frame image includes: the localization position information of at least one target to be tracked corresponding to the one scale in the current frame image and / or the class labels of at least one target to be tracked respectively.

[0131] In an alternative embodiment of the embodiment of the present application:

[0132] The input of the first convolutional layer in the convolutional neural network model is the set of first sampled images;

[0133] The input of the other convolutional layers except the first convolutional layer in the convolutional neural network model is: the stacked result of a set of sampled images corresponding to the scale and the feature image output by the previous convolutional layer.

[0134] In an alternative embodiment of the embodiment of the present application, the positioning unit 705 is specifically configured to:

[0135] Perform non-maximum suppression processing on the respective localization results of the target to be tracked at different scales corresponding to the current frame image to locate the target to be tracked in the current frame image.

[0136] For the target tracking device disclosed in the embodiment of the present application, since it corresponds to the target tracking method disclosed in the corresponding method embodiment above, the description is relatively simple. For relevant similarities, please refer to the description of the target tracking method part in the method embodiment above, and details are not described here.

[0137] The embodiment of the present application also discloses an electronic device, which may be, but is not limited to, a mobile phone, a tablet computer, a personal PC (such as a notebook, an all-in-one machine, a desktop computer, etc.) with computer vision processing functions, or a physical machine corresponding to a private cloud / public cloud platform with computer vision processing functions, a server (image processing server), etc.

[0138] The composition structure of the electronic device is as Figure 8 shown, and at least includes:

[0139] A memory 801, used to store a set of computer instructions;

[0140] The set of computer instructions can be implemented in the form of a computer program.

[0141] The memory 801 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0142] A processor 802, configured to implement the target tracking method disclosed in the above method embodiments by executing the instruction set stored in the memory.

[0143] The processor 802 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, etc.

[0144] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor, and the communication interface complete communication with each other through the communication bus.

[0145] The communication interface is used for communication between the electronic device and other devices. The communication bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc.

[0146] In this embodiment, when the processor in the electronic device performs target tracking on the current frame image in the image sequence, by performing inter-frame sampling on the image sequence where "the sampling density at different positions is inversely correlated with the time interval from different positions to the current frame image", inter-frame correlation information with a high degree of correlation with the current frame image is extracted for the current frame image. The inter-frame correlation information with a high degree of correlation can guide or assist the current frame image to perform more accurate / precise positioning of the target; and further perform intra-frame multi-scale sampling on each of the inter-frame images sampled in the above manner, so as to obtain the multi-scale distribution information of the target in the intra-frame image, improve the understanding of a single image frame, and overcome the situation such as the loss of the tracked target easily caused by the fixed receptive field of the convolutional kernel in the convolutional neural network. Thus, by combining the inter-frame correlation information with a high degree of correlation and the multi-scale distribution information of the intra-frame image, the present application can at least improve the accuracy / precision of the target tracking task and avoid the occurrence of situations such as the loss of the tracked target.

[0147] In addition, the embodiment of the present application also discloses a computer-readable storage medium, in which a computer instruction set is stored, and when the computer instruction set is executed by a processor, the target tracking method disclosed in the above method embodiments is implemented.

[0148] When the instructions stored in the computer-readable storage medium are running, for the current frame image to be tracked and located for the target in the image sequence, through the inter-frame sampling of the image sequence with the relationship that "the sampling density at different positions is inversely correlated with the time interval from different positions to the current frame image", the inter-frame correlation information with high correlation degree with the current frame image is extracted for the current frame image. The inter-frame correlation information with high correlation degree can guide or assist the current frame image to more accurately / precisely locate the target; and the intra-frame multi-scale sampling is further respectively performed on each of the inter-frame images sampled in the above manner, and the multi-scale distribution information of the target in the intra-frame image can be obtained, improving the understanding of a single image frame, and overcoming the situation such as the loss of the tracked target easily caused by the fixed receptive field of the convolution kernel in the convolutional neural network. Thus, by combining the inter-frame correlation information with high correlation degree and the multi-scale distribution information of the intra-frame image, the present application can at least improve the accuracy / precision of the target tracking task and avoid the occurrence of situations such as the loss of the tracked target.

[0149] It should be noted that each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0150] For the convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0151] From the description of the above implementation manners, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.

[0152] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0153] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present application.

Claims

1. A target tracking method, characterized in that, it includes: Determine the current frame image to be processed in the image sequence; The image sequence includes multiple frames of images arranged in chronological order of imaging, and the image contents of at least some of the multiple frames of images are related; Perform a first sampling process on the image sequence to obtain a set of first sampling images corresponding to the current frame image; the sampling density at different positions in the image sequence is inversely correlated with the time interval from the different positions to the current frame image; Perform a second sampling process on each of the first sampling images to obtain a set of second sampling images including multiple scales of second sampling images corresponding to each of the first sampling images; In the scale space composed of the set of first sampling images and the sets of second sampling images corresponding to each of the first sampling images, based on a single set of inter-frame sampling images with the same scale, determine the positioning result of the current frame image for the target to be tracked at the same scale, so as to obtain the positioning results of the target to be tracked at different scales corresponding to the current frame image; Based on the positioning results of the target to be tracked at different scales corresponding to the current frame image, locate the target to be tracked in the current frame image; Wherein, in the scale space composed of the set of first sampling images and the sets of second sampling images corresponding to each of the first sampling images, based on a single set of inter-frame sampling images with the same scale, determine the positioning result of the current frame image for the target to be tracked at the same scale, so as to obtain the positioning results of the target to be tracked at different scales corresponding to the current frame image, including: Input the set of first sampling images and the sets of second sampling images corresponding to each of the first sampling images into a pre-constructed convolutional neural network model including multiple convolutional layers; different convolutional layers correspond one by one to different scales in the scale space; each convolutional layer processes at least a set of inter-frame sampling images at the corresponding scale, so that the set of inter-frame sampling images is used to guide the positioning of the current frame image for the target to be tracked at the corresponding scale; Obtain the output information of the convolutional neural network model; Wherein, the output information includes the positioning results of the target to be tracked at different scales corresponding to the current frame image; the positioning result of the target to be tracked at one scale corresponding to the current frame image includes: the positioning position information of at least one target to be tracked corresponding to the one scale in the current frame image and / or the class labels of at least one target to be tracked respectively.

2. The method according to claim 1, characterized in that, The performing a first sampling process on the image sequence to obtain a set of first sampling images corresponding to the current frame image includes: Performing inter-frame sampling on the image sequence in a Gaussian distribution manner; Wherein, in the inter-frame sampling in the Gaussian distribution manner, the current frame image corresponds to the center of the Gaussian distribution, and the sampling density at different positions in the image sequence presents a Gaussian distribution state centered on the current frame image.

3. The method according to claim 1, characterized in that, Performing second sampling processing on each of the first sampled images respectively includes: Subjecting each first sampled image to multiple downsampling operations to obtain a plurality of second sampled images with different scales corresponding to each first sampled image; Among them, the second sampled images with different scales respectively correspond to different resolutions.

4. The method according to claim 1, characterized in that: The input of the first convolutional layer among the multiple convolutional layers is the set of first sampled images; The input of the other convolutional layers except the first convolutional layer among the multiple convolutional layers is: the stacked result of a set of sampled images with the corresponding scale and the feature image output by the previous convolutional layer.

5. The method according to claim 1, characterized in that, Locating the target to be tracked in the current frame image based on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image includes: Performing non-maximum suppression processing on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image to locate the target to be tracked in the current frame image.

6. An object tracking device, characterized in that, including: A first determination unit for determining the current frame image to be processed in the image sequence; The image sequence includes multiple frames of images arranged in the order of imaging time, and at least part of the images in the multiple frames of images have associations in image content; A first sampling unit for performing first sampling processing on the image sequence to obtain a set of first sampled images corresponding to the current frame image; the sampling density at different positions in the image sequence is inversely correlated with the time interval from the different positions to the current frame image; A second sampling unit for performing second sampling processing on each of the first sampled images respectively to obtain a set of second sampled images including second sampled images with multiple scales corresponding to each of the first sampled images respectively; A second determination unit for determining the positioning result of the current frame image for the target to be tracked at the same scale in the scale space composed of the set of first sampled images and the sets of second sampled images corresponding to each of the first sampled images, so as to obtain the respective positioning results of the target to be tracked at different scales corresponding to the current frame image; A positioning unit for locating the target to be tracked in the current frame image based on the respective positioning results of the target to be tracked at different scales corresponding to the current frame image; Among them, the second determination unit is specifically used for: Inputting the set of first sampled images and the sets of second sampled images corresponding to each of the first sampled images into a pre-constructed convolutional neural network model including multiple convolutional layers; different convolutional layers correspond one-to-one with different scales in the scale space; each convolutional layer processes at least a set of inter-frame sampled images with the corresponding scale, so that the set of inter-frame sampled images is used to guide the positioning of the current frame image for the target to be tracked at the corresponding scale; Obtaining the output information of the convolutional neural network model; Among them, the output information includes the respective localization results of the target to be tracked at different scales corresponding to the current frame image; the localization result of the target to be tracked at one scale corresponding to the current frame image includes: the localization position information corresponding to the one scale of at least one target to be tracked in the current frame image and / or the class labels of at least one target to be tracked respectively.

7. The apparatus according to claim 6, wherein, the first sampling unit is specifically configured to: perform inter-frame sampling on the image sequence in a Gaussian distribution manner; wherein, in the inter-frame sampling in the Gaussian distribution manner, the current frame image corresponds to the center of the Gaussian distribution, and the sampling density at different positions in the image sequence presents a Gaussian distribution state centered on the current frame image.

8. The apparatus according to claim 6, wherein, the second sampling unit is specifically configured to: perform multiple downsampling operations on each first sampled image to obtain a plurality of second sampled images corresponding to each first sampled image at different scales; wherein the second sampled images at different scales correspond to different resolutions respectively.

9. The apparatus according to claim 6, wherein: the input of the first convolutional layer in the plurality of convolutional layers is the set of first sampled images; the input of the other convolutional layers except the first convolutional layer in the plurality of convolutional layers is: the stacked result of a set of sampled images corresponding to the corresponding scale and the feature image output by the previous convolutional layer.

10. An electronic device, wherein, comprising: a memory for storing a computer instruction set; a processor for implementing the target tracking method according to any one of claims 1-5 by executing the instruction set stored on the memory.

11. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer instruction set, and when the computer instruction set is executed by a processor, the target tracking method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Pedestrian target tracking method based on convolution association network in automatic driving scene

    CN111652903A

  • KR20190056161A