Target tracking method and device, electronic equipment and storage medium

By extracting and utilizing the motion displacement field information in the video frame, combining encoding and decoding operations for target prediction and tracking, the problem of implicitness and high computing resource consumption in target tracking is solved, and a more accurate and interpretable target tracking effect is achieved.

CN120107308APending Publication Date: 2025-06-06IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510043973.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, Transformer can only provide implicit target tracking internally, resulting in errors in the results and is not interpretable. Relying solely on Transformer to achieve target tracking requires a large amount of computing resources, which has the problem of excessive training cost.

Method used

By extracting the motion displacement field information in the current video frame, combining the encoder and the decoder for encoding and decoding operations, target prediction is performed based on the motion displacement field information, and target tracking is performed using the identification sequence, and the motion information is explicitly used to improve the accuracy and interpretability of the tracking results.

Benefits of technology

It improves the accuracy and interpretability of the target tracking results, reduces the computing resource requirements and training costs, and avoids the problems of excessive computing resources and high training costs caused by a single network structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107308A_ABST
    Figure CN120107308A_ABST
Patent Text Reader

Abstract

The invention provides a target tracking method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: extracting motion displacement field information in a current video frame, determining the motion information of a target in a display mode, further achieving the tracking of the displayed target, improving the accuracy of a target tracking result, and improving the user experience. And the target tracking result can be explained through the motion displacement field information, and the interpretability is achieved. Moreover, according to the method, the motion displacement field information is used as priori knowledge to be applied to target tracking, so that the problems of large computing resource demand and too high training cost caused by target tracking realized by only depending on a single network structure are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a target tracking method, device, electronic equipment and storage medium. Background Art

[0002] Target tracking is a hot topic in the field of computer vision. It uses the contextual information of video or image sequences to model the appearance and motion information of the target, thereby predicting the target's motion state and calibrating the target's position, and then studying the laws of moving targets, or providing semantic and non-semantic information support for the system's decision-making alarms, including motion detection, target classification, target tracking, behavior understanding, event detection, etc.

[0003] The core idea of ​​the Transformer-based multi-target tracking solution is that the output target set of an ideal tracking model should be complete and ordered. Therefore, the solution uses learnable target queries and target feature queries of the previous frame as input query objects of the Transformer. The former becomes the detection box of the current frame after passing through the decoder, which is used to detect the position of the target in the current frame, and the latter becomes the tracking box after passing through the decoder, which is used to detect the position of the target in the previous frame in the current frame, and can complete the data association step in the current frame.

[0004] However, the Transformer used in the above solution can only provide implicit target tracking internally, which may lead to errors in the target tracking results and lack of interpretability. Moreover, relying solely on the Transformer to achieve target tracking requires a large amount of computing resources, resulting in the problem of high training costs. Summary of the invention

[0005] The present invention provides a target tracking method, device, electronic equipment and storage medium to solve the defects existing in the related technology.

[0006] The present invention provides a target tracking method, comprising: Get the current video frame in the video to be processed; Extracting motion displacement field information in the current video frame; Encode the current video frame to obtain a current encoding result, decode the current encoding result and the motion displacement field information to obtain a current decoding result, and perform target prediction on the current video frame based on the current decoding result; Based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame, an identification sequence is applied to perform target tracking on the current video frame.

[0007] According to a target tracking method provided by the present invention, decoding the current encoding result and the motion displacement field information to obtain the current decoding result includes: Based on the motion displacement field information, re-encoding the current encoding result to obtain a re-encoding result; The re-encoding result is decoded to obtain the current decoding result.

[0008] According to a target tracking method provided by the present invention, the step of extracting the motion displacement field information in the current video frame includes: Inputting the current video frame into a motion displacement field prediction network to obtain the motion displacement field information output by the motion displacement field prediction network; The motion displacement field prediction network is trained based on video samples carrying motion displacement field labels of moving objects.

[0009] According to a target tracking method provided by the present invention, the extracting of the motion displacement field information in the current video frame specifically includes: Inputting the current video frame and the historical decoding results into a tracking identification prediction network, wherein the tracking identification prediction network includes an encoder, a decoder, a prediction module and an identification decoder; The encoder is used to encode the current video frame to obtain the current encoding result; The decoder is used to decode the current encoding result and the motion displacement field information to obtain the current decoding result; The prediction module is used to perform target prediction on the current video frame based on the current decoding result; The identification decoder is used to apply the identification sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

[0010] According to a target tracking method provided by the present invention, the identification sequence includes known identifications and unknown identifications; The known identifier is used to mark the prediction results in the historical video frame whose confidence is greater than a threshold; The unknown identifier is used to mark the prediction results in the current video frame whose confidence is greater than the threshold and are not assigned the known identifier.

[0011] According to a target tracking method provided by the present invention, based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame, the identification sequence is applied to perform target tracking on the current video frame, and further includes: If there is a designated identifier in the known identifiers, and the designated identifier has not been allocated in a continuous preset number of video frames, the designated identifier in the identifier sequence is deleted.

[0012] The present invention also provides a target tracking device, comprising: A video frame acquisition module is used to acquire the current video frame in the video to be processed; A motion displacement field extraction module, used to extract the motion displacement field information in the current video frame; A tracking mark prediction module is used to encode the current video frame to obtain a current encoding result, decode the current encoding result and the motion displacement field information to obtain a current decoding result, and perform target prediction on the current video frame based on the current decoding result; The tracking mark prediction module is further used to apply the mark sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the target tracking method as described above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the target tracking method as described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the target tracking method as described above is implemented.

[0016] The target tracking method, device, electronic device and storage medium provided by the present invention can realize displayed target tracking by extracting motion displacement field information in the current video frame, and determine the motion information of the target in a display manner, thereby improving the accuracy of the target tracking result, and the target tracking result can be explained by the motion displacement field information, and has interpretability. Moreover, the method applies the motion displacement field information as prior knowledge to target tracking, avoiding the problem of large computing resource requirements and high training costs caused by relying on a single network structure to achieve target tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 This is one of the flow charts of the target tracking method provided by the present invention.

[0019] Figure 2 This is the second flow chart of the target tracking method provided by the present invention.

[0020] Figure 3 This is the third flow chart of the target tracking method provided by the present invention.

[0021] Figure 4 It is a schematic diagram of the application process of the identification decoder in the target tracking method provided by the present invention.

[0022] Figure 5 It is a structural schematic diagram of the target tracking device provided by the present invention.

[0023] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] Computer vision is a highly interdisciplinary subject that integrates image processing, computer graphics, pattern recognition, artificial intelligence, artificial neural networks, computers, psychology, physics and mathematics. The purpose of computer vision research is to enable computers to perceive the geometric information of objects in the environment, including its shape, position, posture, movement, etc., and to describe, store, identify and understand it, so it has become one of the hottest topics today.

[0026] Target tracking is a hot topic in the field of computer vision. It can be divided into single-target tracking and multi-target tracking based on the number of targets to be tracked. Single-target tracking can handle problems such as illumination, deformation, and occlusion by modeling the target's appearance or motion. Multi-target tracking is much more complicated. In addition to the problems encountered in single-target tracking, it also requires correlation matching between targets. In addition, in multi-target tracking tasks, problems such as frequent occlusion of targets, unknown start and end times of trajectories, targets that are too small, similar appearances, interactions between targets, and low frame rates are often encountered.

[0027] Existing multi-target tracking solutions can be divided into two categories: traditional feature solutions represented by kernel correlation filtering and deep learning solutions represented by neural networks. At present, deep learning solutions are far ahead of traditional feature solutions and can be divided into four categories: detection-first-then-tracking solutions, multi-target tracking based on a joint detection and tracking framework, sequence modeling methods represented by Transformer, and tracking methods that introduce diffusion models as guidance.

[0028] The core of the detection-then-tracking solution is to first use the detection network to detect the target in each frame, then crop the target according to the detection box to obtain all the targets in the image, and then convert the tracking operation into a data association problem between the previous and next frames. In the data association part, the target of the previous frame is usually predicted, and then the intersection and union ratio between the prediction result and the detection result of the current frame is used to form a matching cost matrix. Finally, the Hungarian matching algorithm is used to solve the association information between the targets to form the target trajectory. In this solution, the target's apparent feature similarity information is sometimes used, so an additional apparent feature extraction network is trained separately.

[0029] The multi-target tracking solution based on the joint detection and tracking framework is an improved version of the above-mentioned detection-first-then-tracking solution. In order to solve the conflict between the detection network and the appearance feature extraction network, this solution integrates target detection and appearance feature extraction into one network.

[0030] However, the existing multi-target tracking solutions based on detection-first-tracking or joint detection-tracking frameworks require the maximum distance between classes for detection and the maximum distance within classes for appearance features, which cannot completely solve the contradiction between detection and appearance feature extraction; Hungarian matching in data association may lead to false matches. Moreover, if the appearance feature extraction network is introduced, the overall tracking efficiency will be reduced, especially when there are many targets, this shortcoming becomes more serious.

[0031] The Transformer-based multi-target tracking solution can use the learnable target query and the target feature query of the previous frame as the input query object of the Transformer, and complete the data association step in the current frame. However, the Transformer used in this solution can only provide implicit target tracking internally, cannot explicitly use motion information, and is not interpretable. Moreover, relying solely on Transformer to achieve target tracking requires a lot of computing resources and the training cost is too high.

[0032] With the rapid development of diffusion models, a multi-target tracking method that uses diffusion models as a guide has also emerged in the field of target tracking. The core of this solution is to add Gaussian noise to the coordinate frame of the t-1 frame image, and then use the denoising ability of the diffusion model to obtain the coordinate frame on the t frame image, and then use strategies such as Hungarian matching to perform relationship matching to achieve target tracking. However, this solution only predicts the four points of the coordinate frame and is easily interfered; moreover, this solution only considers the spatial position information between the target frames and ignores the motion information, which leads to inaccurate target tracking results.

[0033] Based on this, an embodiment of the present invention provides a target tracking method to solve the defects existing in the prior art.

[0034] Figure 1 A flow chart of a target tracking method provided in an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes: S1, obtaining the current video frame in the video to be processed; S2, extracting motion displacement field information in the current video frame; S3, encoding the current video frame to obtain a current encoding result, decoding the current encoding result and the motion displacement field information to obtain a current decoding result, and performing target prediction on the current video frame based on the current decoding result; S4, based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame, apply the identification sequence to perform target tracking on the current video frame.

[0035] Specifically, the target tracking method provided in the embodiment of the present invention is executed by a target tracking device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0036] First, step S1 is performed to obtain a current video frame in the video to be processed. The video to be processed may include multiple video frames, each of which may be used as the current video frame to perform operations in subsequent steps.

[0037] Then, step S2 is performed to extract the motion displacement field information in the current video frame. The motion displacement field information may include the motion direction and motion speed of the target in the current video frame, and may be explicitly obtained through Kalman filtering or neural networks. The motion displacement field has physical meaning, can well describe the motion of the target, thereby playing an accurate guiding role, and is applied to target tracking.

[0038] Then, step S3 is performed to encode the current video frame to obtain a current encoding result, which may be a high-dimensional feature.

[0039] Thereafter, the current encoding result can be fused with the motion displacement field information, and the fusion result can be decoded to obtain the current decoding result. The current decoding result is a decoding result that contains the prior knowledge of the motion displacement field information. Furthermore, based on the current decoding result, the target prediction can be performed on the current video frame to obtain a prediction result. Here, the target prediction can include category prediction and position prediction, and the prediction result can include the category prediction result (Class Predict) and the position prediction result (Box Predict) of the target in the current video frame, and the position prediction result can be the coordinate box of the target in the current video frame.

[0040] Finally, step S4 is performed to introduce an identification sequence, which can be updated according to the prediction result. The identification sequence can include a preset number of identification documents (IDs), which refers to the number of targets that can be tracked and can be set as needed, and is not specifically limited here.

[0041] Furthermore, the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame are used, and the identification sequence is applied to track the target in the current video frame to obtain the tracking result, that is, the identification of the target in the current video frame, and then the targets in different video frames are associated through the identification to achieve target tracking.

[0042] For example, if the current video frame is the first frame in the video to be processed, when the confidence of the prediction result is greater than the threshold, the target in the prediction result will be recorded as a tracking object and a unique identifier will be assigned to each target, and the identifier will be updated as a known identifier to the identifier sequence. In each subsequent frame, the tracking result is filtered using the threshold, and the identifiers in the tracking result with a confidence greater than the threshold that are not known identifiers are updated to the identifier sequence.

[0043] The target tracking method provided in the embodiment of the present invention first obtains the current video frame in the video to be processed; then extracts the motion displacement field information in the current video frame; thereafter encodes the current video frame to obtain the current encoding result, and decodes the current encoding result and the motion displacement field information to obtain the current decoding result, and performs target prediction on the current video frame based on the current decoding result; finally, based on the current decoding result and the historical decoding results of the historical video frames located before the current video frame in the video to be processed, the identification sequence is applied to perform target tracking on the current video frame. The method extracts the motion displacement field information in the current video frame to determine the motion information of the target in a display manner, thereby realizing displayed target tracking, improving the accuracy of the target tracking result, and the target tracking result can be explained by the motion displacement field information, which is interpretable. Moreover, the method applies the motion displacement field information as a priori knowledge to target tracking, avoiding the problems of large computing resource requirements and high training costs caused by relying only on a single network structure to achieve target tracking.

[0044] On the basis of the above embodiment, the current encoding result and the motion displacement field information are decoded to obtain the current decoding result, including: Based on the motion displacement field information, re-encoding the current encoding result to obtain a re-encoding result; The re-encoding result is decoded to obtain the current decoding result.

[0045] Specifically, in the process of decoding the current encoding result and the motion displacement field information to obtain the current decoding result, the motion displacement field information can be used to re-encode the current encoding result to obtain the re-encoded result. The re-encoding process can be a process of fusing the motion displacement field information with the current encoding result, and the obtained re-encoding result is the fused result.

[0046] Thereafter, the re-encoded result may be decoded to obtain the current decoding result.

[0047] In the embodiment of the present invention, by re-encoding the current encoding result using the motion displacement field information, the re-encoded result can include the motion displacement field information, thereby improving the accuracy of subsequent target tracking results.

[0048] Based on the above embodiment, the step of extracting the motion displacement field information in the current video frame includes: Inputting the current video frame into a motion displacement field prediction network to obtain the motion displacement field information output by the motion displacement field prediction network; The motion displacement field prediction network is trained based on video samples carrying motion field displacement labels of moving objects.

[0049] Specifically, when extracting the motion displacement field information in the current video frame, a motion displacement field prediction network can be introduced. The motion displacement field prediction network can use the video sample to extract the motion displacement field result of the moving object in advance and use it as a motion displacement field label. Then, each video frame sample in the video sample is input into the initial motion displacement field prediction network, so that the initial motion displacement field prediction network predicts the motion displacement field information of the moving object in each video frame in the offline video, and uses the prediction result and the training label to iteratively train the initial motion displacement field prediction network to obtain an applicable motion displacement field prediction network. The initial motion displacement field prediction network can be constructed using the Encoder-Decoder paradigm, for example, it can be a UNet network.

[0050] Thereafter, the current video frame is input into the motion displacement field prediction network, and the motion displacement field information of the target in the current video frame can be output through the motion displacement field prediction network.

[0051] In the embodiment of the present invention, by introducing a motion displacement field prediction network, the extraction efficiency of the motion displacement field information can be improved.

[0052] Based on the above embodiment, the step of extracting the motion displacement field information in the current video frame specifically includes: Inputting the current video frame and the historical decoding results into a tracking identification prediction network, wherein the tracking identification prediction network includes an encoder, a decoder, a prediction module and an identification decoder; The encoder is used to encode the current video frame to obtain the current encoding result; The decoder is used to decode the current encoding result and the motion displacement field information to obtain the current decoding result; The prediction module is used to perform target prediction on the current video frame based on the current decoding result; The identification decoder is used to apply the identification sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

[0053] Specifically, in an embodiment of the present invention, steps S3 and S4 can be implemented by a tracking identification prediction network, which can use Transformer as a basic structure, including an encoder (Encoder), a decoder (Dncoder), a prediction module and an identification decoder (ID Dncoder).

[0054] The current video frame and the historical decoding results are input into the tracking identification prediction network. The current video frame can be encoded by the encoder to obtain the current encoding result. The current encoding result and the motion displacement field information are decoded by the decoder to obtain the current decoding result. The prediction module predicts the target of the current video frame based on the current decoding result. The identification decoder uses the identification sequence to track the target of the current video frame based on the current decoding result and the historical decoding results of the historical video frames before the current video frame in the video to be processed.

[0055] like Figure 2 As shown, the tracking mark prediction network may also include a re-encoding module, which is connected between the encoder and the decoder, and the input of the re-encoding module is the motion displacement field information and the current encoding result, and the output is the re-encoding result. Furthermore, the input of the decoder is the re-encoding result.

[0056] In the embodiment of the present invention, target prediction and target tracking are implemented through a tracking mark prediction network, which can improve the efficiency and accuracy of target prediction and target tracking.

[0057] Based on the above embodiment, the identification sequence includes known identification bits and unknown identification bits; The known identifier is used to mark the prediction results in the historical video frame whose confidence is greater than a threshold; The unknown identification bit is used to mark the prediction results in the current video frame whose confidence is greater than the threshold and are not assigned the known identification.

[0058] Specifically, the identification sequence includes known identifications and unknown identification bits. The known identifications may include one or more, whose initial number is the same as the number of targets in the first frame of the video to be processed. The number of known identifications may increase according to the targets newly added in the subsequent video frames, and the number of known identifications may decrease according to the targets reduced in the subsequent video frames.

[0059] The known identifier can mark the prediction results in the historical video frame whose confidence is greater than a threshold, that is, only the prediction results whose confidence is greater than the threshold will be regarded as a target. The threshold can be set as needed and is not specifically limited here.

[0060] The unknown identification bit can mark the prediction results in the current video frame whose confidence is greater than the threshold and no known identification is assigned. In other words, the unknown identification bit can be used to record the identification of the newly added target. After recording, the unknown identification bit disappears, and the identification recorded by it becomes a known identification.

[0061] In the embodiment of the present invention, known identifiers and identifiers that may appear in the future are recorded in an identifier sequence, which can provide a basis for target tracking.

[0062] On the basis of the above embodiment, the target tracking of the current video frame is performed by applying an identification sequence based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame, further comprising: If there is a designated identifier in the known identifiers, and the designated identifier has not been allocated in a continuous preset number of video frames, the designated identifier in the identifier sequence is deleted.

[0063] Specifically, when tracking a target in the current video frame, if there is a specified identifier among the known identifiers, and the specified identifier has not been assigned to the target in a consecutive preset number of video frames, that is, there is no target corresponding to the specified identifier in a consecutive preset number of video frames, then it is considered that the target has been lost a preset number of times, and the specified identifier in the identifier sequence can be deleted, thereby improving the allocation efficiency of known identifiers during target tracking.

[0064] In summary, if Figure 3 As shown, the target tracking method provided in the embodiment of the present invention includes: Get the current video frame in the video to be processed; Extract the motion displacement field information in the current video frame through the motion displacement field prediction network; By tracking the encoder in the identification prediction network, the current video frame is encoded to obtain the current encoding result; By tracking the re-encoding module in the identification prediction network, the current encoding result is re-encoded based on the motion displacement field information to obtain a re-encoding result; By tracking the decoder in the identification prediction network, the re-encoded result is decoded to obtain the current decoding result; By tracking the prediction module in the identification prediction network, the target prediction of the current video frame is performed based on the current decoding result; By tracking the identifier decoder in the identifier prediction network, based on the current decoding result and the historical decoding results of the historical video frames before the current video frame in the video to be processed, the identifier sequence is applied to track the target on the current video frame.

[0065] The application process of the identification decoder is as follows Figure 4 As shown, assume that the current video frame is T frame, and the historical video frame is the previous N frames, namely TN frames, ..., T-1 frames respectively. The historical decoding results of the historical video frames are spliced ​​with the known identifiers therein, and the current decoding results of the T frames are spliced ​​with the unknown identifiers, and are input into the identifier decoder together. The unknown identifier is identified by the identifier decoder to achieve target tracking.

[0066] The target tracking method provided in the embodiment of the present invention is an end-to-end target tracking method based on motion displacement field information prediction. The core of the method is to represent the prior knowledge required for tracking the target by motion displacement field information with physical meaning, and at the same time use the Encoder-Decoder paradigm to model and seamlessly insert it into the tracking identification prediction network to achieve end-to-end target tracking. The method explicitly combines the motion displacement field information, can obtain more accurate tracking results, and does not require complex post-processing operations such as Hungarian matching.

[0067] like Figure 5 As shown, based on the above embodiment, an embodiment of the present invention provides a target tracking device, including: The video frame acquisition module 51 is used to acquire the current video frame in the video to be processed; A motion displacement field extraction module 52, used to extract the motion displacement field information in the current video frame; A tracking mark prediction module 53 is used to encode the current video frame to obtain a current encoding result, decode the current encoding result and the motion displacement field information to obtain a current decoding result, and perform target prediction on the current video frame based on the current decoding result; The tracking mark prediction module 53 is further used to apply the mark sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

[0068] Based on the above embodiment, the tracking mark prediction module is specifically used for: Based on the motion displacement field information, re-encoding the current encoding result to obtain a re-encoding result; The re-encoding result is decoded to obtain the current decoding result.

[0069] Based on the above embodiment, the motion displacement field extraction module is specifically used for: Inputting the current video frame into a motion displacement field prediction network to obtain the motion displacement field information output by the motion displacement field prediction network; The motion displacement field prediction network is trained based on video samples carrying motion displacement field labels of moving objects.

[0070] Based on the above embodiment, the tracking mark prediction module is specifically used for: Inputting the current video frame and the historical decoding results into a tracking identification prediction network, wherein the tracking identification prediction network includes an encoder, a decoder, a prediction module and an identification decoder; The encoder is used to encode the current video frame to obtain the current encoding result; The decoder is used to decode the current encoding result and the motion displacement field information to obtain the current decoding result; The prediction module is used to perform target prediction on the current video frame based on the current decoding result; The identification decoder is used to apply the identification sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

[0071] Based on the above embodiment, the identification sequence includes known identifications and unknown identifications; The known identifier is used to mark the prediction results in the historical video frame whose confidence is greater than a threshold; The unknown identifier is used to mark the prediction results in the current video frame whose confidence is greater than the threshold and are not assigned the known identifier.

[0072] Based on the above embodiment, the tracking mark prediction module is further used for: If there is a designated identifier in the known identifiers, and the designated identifier has not been allocated in a continuous preset number of video frames, the designated identifier in the identifier sequence is deleted.

[0073] Specifically, the functions of each module in the target tracking device provided in the embodiment of the present invention correspond one-to-one to the operation procedures of each step in the above-mentioned method embodiment, and the effects achieved are also consistent. Please refer to the above-mentioned embodiment for details, and no further details will be given in the embodiment of the present invention.

[0074] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the target tracking method provided in the above embodiments.

[0075] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0076] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target tracking method provided in the above embodiments.

[0077] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the target tracking method provided in the above embodiments.

[0078] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0079] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target tracking method, characterized in that: include: Get the current video frame in the video to be processed; Extracting motion displacement field information in the current video frame; Encode the current video frame to obtain a current encoding result, decode the current encoding result and the motion displacement field information to obtain a current decoding result, and perform target prediction on the current video frame based on the current decoding result; Based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame, an identification sequence is applied to perform target tracking on the current video frame.

2. The target tracking method according to claim 1, characterized in that: The decoding of the current encoding result and the motion displacement field information to obtain the current decoding result includes: Based on the motion displacement field information, re-encoding the current encoding result to obtain a re-encoding result; The re-encoding result is decoded to obtain the current decoding result.

3. The target tracking method according to claim 1, characterized in that: The extracting the motion displacement field information in the current video frame includes: Inputting the current video frame into a motion displacement field prediction network to obtain the motion displacement field information output by the motion displacement field prediction network; The motion displacement field prediction network is trained based on video samples carrying motion displacement field labels of moving objects.

4. The target tracking method according to claim 1, characterized in that: The step of extracting the motion displacement field information in the current video frame specifically includes: Inputting the current video frame and the historical decoding results into a tracking identification prediction network, wherein the tracking identification prediction network includes an encoder, a decoder, a prediction module and an identification decoder; The encoder is used to encode the current video frame to obtain the current encoding result; The decoder is used to decode the current encoding result and the motion displacement field information to obtain the current decoding result; The prediction module is used to perform target prediction on the current video frame based on the current decoding result; The identification decoder is used to apply the identification sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

5. The target tracking method according to any one of claims 1 to 4, characterized in that: The identification sequence includes known identifications and unknown identifications; The known identifier is used to mark the prediction results in the historical video frame whose confidence is greater than a threshold; The unknown identifier is used to mark the prediction results in the current video frame whose confidence is greater than the threshold and are not assigned the known identifier.

6. The target tracking method according to claim 5, characterized in that: The method of applying an identification sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame, further includes: If there is a designated identifier in the known identifiers, and the designated identifier has not been assigned in a continuous preset number of video frames, the designated identifier in the identifier sequence is deleted.

7. A target tracking device, characterized in that: include: A video frame acquisition module is used to acquire the current video frame in the video to be processed; A motion displacement field extraction module, used to extract the motion displacement field information in the current video frame; A tracking mark prediction module is used to encode the current video frame to obtain a current encoding result, decode the current encoding result and the motion displacement field information to obtain a current decoding result, and perform target prediction on the current video frame based on the current decoding result; The tracking mark prediction module is further used to apply the mark sequence to perform target tracking on the current video frame based on the current decoding result and the historical decoding results of the historical video frames in the video to be processed that are located before the current video frame.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the target tracking method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target tracking method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the target tracking method according to any one of claims 1 to 6 is implemented.