Target tracking method, device, apparatus and storage medium

By extracting time domain images and optical flow feature maps from the target video, determining the three-dimensional reference position of the target object and utilizing the motion bias map, the problem of target tracking in dynamic scenes being difficult to accurately restore motion trajectories in existing technologies is solved, achieving efficient and accurate target tracking effects.

CN117152204BActive Publication Date: 2025-10-21BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310967443.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2025-10-21
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

Existing target tracking technology has difficulty in accurately restoring the motion trajectory of the target object in three-dimensional space when faced with problems such as occlusion, lighting changes, and camera motion, especially in dynamic scenes.

Method used

By extracting the temporal image feature map and optical flow feature map between multiple frames of the target video, the three-dimensional reference position of the target object is determined, and the motion trajectory of the target object is tracked in three-dimensional space based on the motion bias map. This is achieved using a single-step, end-to-end network.

Benefits of technology

It achieves efficient and accurate tracking of target objects in dynamic scenes, reduces algorithm complexity, and is suitable for a variety of practical application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152204B_ABST
    Figure CN117152204B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a target tracking method, device, equipment and medium. The method described herein includes: extracting a plurality of time domain image feature maps and a plurality of optical flow feature maps between a plurality of frames of a target video; determining a three-dimensional reference position of at least one object in each frame of the plurality of frames based on the plurality of time domain image feature maps, the at least one object including a target object; determining at least a plurality of motion bias maps for the plurality of frames based on the plurality of time domain image feature maps and the plurality of optical flow feature maps, each motion bias map indicating a motion bias of the at least one object in a corresponding frame relative to a previous frame in a three-dimensional space; and determining a motion trajectory of the target object in the three-dimensional space based on at least the plurality of motion bias maps by referring to the three-dimensional reference position of the target object in each frame of the plurality of frames. Thus, accurate and efficient target tracking can be achieved based on the motion bias maps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure relate generally to the field of computer vision, and more particularly, to object tracking methods, apparatuses, devices, and computer-readable storage media. Background Art

[0002] With the development of computer technology, image processing technology has also progressed. Computer vision and image processing methods enable the automatic identification, tracking, and location of target objects in video or image sequences. For example, one or more objects of interest can be tracked across consecutive frames, providing information such as their 3D shape, pose, position, velocity, or trajectory. However, object tracking technology still faces challenges in practical applications, such as occlusion, illumination variations, and scale changes. Therefore, selecting the appropriate object tracking technology based on specific scenarios and requirements is a critical and pressing issue. Summary of the Invention

[0003] In a first aspect of the present disclosure, a target tracking method is provided. The method includes: extracting multiple temporal image feature maps and multiple optical flow feature maps between multiple frames of a target video; determining a three-dimensional reference position of at least one object in each of the multiple frames based on the multiple temporal image feature maps, the at least one object including a target object; determining at least a plurality of motion bias maps for the multiple frames based on the multiple temporal image feature maps and the multiple optical flow feature maps, each motion bias map indicating a motion bias of the at least one object in a corresponding frame relative to a previous frame in three-dimensional space; and determining a motion trajectory of the target object in three-dimensional space based on at least the multiple motion bias maps by referring to the three-dimensional reference position of the target object in each of the multiple frames.

[0004] In a second aspect of the present disclosure, a target tracking device is provided. The device includes: an extraction module configured to extract multiple temporal image feature maps and multiple optical flow feature maps between multiple frames of a target video; a three-dimensional reference position determination module configured to determine the three-dimensional reference position of at least one object in each of the multiple frames based on the multiple temporal image feature maps, the at least one object including the target object; a motion bias map determination module configured to determine at least multiple motion bias maps for the multiple frames based on the multiple temporal image feature maps and the multiple optical flow feature maps, each motion bias map indicating a motion bias of the at least one object in the corresponding frame relative to the previous frame in the three-dimensional space; and a motion trajectory determination module configured to determine the motion trajectory of the target object in the three-dimensional space based on at least the multiple motion bias maps by referring to the three-dimensional reference position of the target object in each of the multiple frames.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium and can be executed by a processor to perform the method according to the first aspect of the present disclosure.

[0007] It should be understood that the contents described in the summary of the present invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 shows a flow chart for a target tracking process according to some embodiments of the present disclosure;

[0011] Figure 3A A schematic diagram illustrating an example model architecture for target tracking according to some embodiments of the present disclosure is shown;

[0012] Figure 3B A schematic diagram illustrating an example architecture of a detection model according to some embodiments of the present disclosure;

[0013] Figure 3C A schematic diagram illustrating an example architecture of a first tracking model according to some embodiments of the present disclosure;

[0014] Figure 3D A schematic diagram illustrating an example architecture of a second tracking model according to some embodiments of the present disclosure;

[0015] Figure 3E A schematic diagram illustrating an example architecture of a morphology determination model according to some embodiments of the present disclosure;

[0016] Figure 4 A block diagram illustrating an apparatus for target tracking according to some embodiments of the present disclosure is shown; and

[0017] Figure 5 A block diagram of an electronic device is shown in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] It should be noted that the acquisition, storage and application of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0020] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit and implicit definitions may be included below.

[0021] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, "model" may also be referred to as "machine learning model", "machine learning network" or "network", and these terms are used interchangeably in this article.

[0022] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also known as the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also known as input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values ​​obtained through training to determine the corresponding output.

[0023] As briefly mentioned above, target tracking technology can be used to track and recover the 3D shape, pose, or trajectory of a target object from a monocular video captured by a camera. However, target tracking technology faces some challenges. For example, in some scenarios, the target object captured by the camera interacts with and is blocked by other objects during its motion, making tracking the target object in complex scenes very difficult. In other scenarios, the camera itself is also moving while capturing the target object. Because monocular video has depth ambiguity, it is necessary to accurately analyze the 3D motion of the camera and target object from the blurred image information to calculate the 3D motion trajectory of the target object in the world coordinate system.

[0024] In some scenarios, the target object is, for example, a human body. By inputting an image including the human body into a trained single-step network, relatively complete relative relationship information of the human body can be obtained, thereby realizing the 3D shape and posture estimation of the human body in multi-person scenarios.

[0025] However, such solutions do not intuitively track human motion in the time domain, nor do they model the camera's 3D motion. Instead, they estimate each person's 3D pose and shape solely through image features. These methods are essentially based on the representation of a single 2D image and fail to model temporal information. When faced with problems like occlusion, these representations cannot leverage temporal motion consistency to achieve accurate and stable estimation. Furthermore, due to the lack of camera motion modeling, the target object's motion trajectory in the world coordinate system cannot be recovered.

[0026] In other solutions, the local motion of the target object is modeled, and the motion trajectory of the target object in the world coordinate system is inferred based on prior indications. For example, using a multi-step network, each human body in the image is first detected, and the human body tracking model is used to estimate the 2D trajectory of the target human body in the image. On this basis, the 3D shape and posture of each human body in each detection frame are estimated, and then the motion of the local body parts of the human body is analyzed to infer the motion trajectory of each person in the world coordinate system. For example, the running trajectory of the target human body in the world coordinate system can be inferred by analyzing the 3D running action.

[0027] However, in real-world scenarios, this approach doesn't just rely on the body's movements in the world coordinate system; it's also affected by external objects and forces. For example, in scenarios like rowing, cycling, and skating, the body may only experience minor or even no movement, yet its 3D motion in the world coordinate system is often significant. Therefore, this approach isn't universally applicable.

[0028] Some solutions use a multi-step network to first estimate the motion trajectory of the target object in the camera space, and then use the Structure From Motion (SFM) algorithm to estimate the motion trajectory of the camera from the monocular video, thereby inferring the motion trajectory of the target object in the world coordinate system.

[0029] However, most SFM-based camera motion estimation methods are designed for static scenes without moving objects, relying on statically correlated keypoints across multiple frames to calculate camera motion. For dynamic scenes tracking moving objects, finding stable statically correlated keypoints is difficult, resulting in suboptimal camera motion estimation results in practice.

[0030] In order to at least partially solve the above problems and other potential problems that may exist in traditional solutions, the present disclosure provides an improved solution for target tracking. According to various embodiments of the present disclosure, multiple time-domain image feature maps and multiple optical flow feature maps are extracted between multiple frames of a target video. Based on the multiple time-domain image feature maps, the three-dimensional reference position of at least one object in each of the multiple frames is determined. The at least one object includes a target object. Based on the multiple time-domain image feature maps and the multiple optical flow feature maps, at least multiple motion bias maps for the multiple frames are determined. Each motion bias map indicates the motion bias of at least one object in the corresponding frame relative to the previous frame in the three-dimensional space. By referring to the three-dimensional reference position of the target object in each of the multiple frames, the motion trajectory of the target object in the three-dimensional space is determined based on at least multiple motion bias maps. Thus, accurate and efficient target tracking can be achieved based on the motion bias map. In this way, no additional model design such as post-processing is required, the algorithm has low complexity, and is easy to apply in practice.

[0031] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0032] Figure 1 A schematic diagram illustrates an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, various client applications, such as image processing applications and 3D modeling applications, may be installed on a terminal device 110. A model 130 is configured to process a target video, such as performing object tracking on the target video.

[0033] In some embodiments of the present disclosure, the model 130 used can be deployed on the remote device 120. The terminal device 110 can communicate with the remote device 130 (for example, via network communication) to use the model 130 stored thereon to perform a model reasoning task, that is, a target tracking task for a target video. The target video can be stored directly on the remote device 120, or collected by the terminal device 110 and sent to the remote device 120. In some embodiments, the model 130 can also be partially or completely deployed locally on the terminal device 110 and run by the terminal device 110 to perform the target tracking task for the target video. The embodiments of the present disclosure do not impose specific limitations on this.

[0034] The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The remote device 120 can, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, a virtual machine, and the like. Although a single device is shown, the remote device 130 can include multiple physical devices. In addition, although only a single terminal device 110 is shown, the remote device 120 or the model 130 deployed therein can be accessed by multiple terminal devices 110 to provide reasoning capabilities of the model 130.

[0035] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0036] To achieve target tracking, the present disclosure uses a motion bias map to characterize the temporal motion of a target object in camera space and / or three-dimensional space (e.g., the space corresponding to the world coordinate system). Based on the target object's reference position, the target object's three-dimensional motion trajectory is tracked and recovered. In some embodiments, three-dimensional reconstruction can also be used to recover the target object's morphology.

[0037] According to various embodiments of the present disclosure, a single-step, end-to-end target tracking network can be implemented. In some embodiments, the remote device 120 can implement such a target tracking network using the model 130, and the training, testing, and application of the model 130 are all end-to-end.

[0038] Figure 2Flowchart of target tracking process 200 according to some embodiments of the present disclosure is shown. Process 200 can be implemented in environment 100. Process 200 can be implemented at terminal device 110. In some embodiments, terminal device 110 can implement target tracking using model 130 running locally or at remote device 120.

[0039] The terminal device 110 obtains a target video containing a target object. Such a target video may include a video or image sequence. The target object may include a human body, an animal, a vehicle, etc. For example, the terminal device 110 may obtain a video containing a target object from a local or other storage device (e.g., another terminal device 110 or a third-party data platform). In some embodiments, the target video includes a monocular video captured by a dynamic camera.

[0040] The number of target objects to be tracked in the target video may be single or multiple. However, in some practical scenarios, such a video may include multiple objects. The terminal device 110 may receive a designation regarding a target object (e.g., a user selection is received) and determine one or more of the multiple objects appearing in the target video as the target object. In some embodiments, the terminal device 110 may also automatically determine all objects in a reference frame (e.g., the first frame of the target video or the first frame in which the object appears) as the target object.

[0041] In box 210, the terminal device 110 extracts feature maps between multiple frames of the target video. Such feature maps include time-domain image feature maps and optical flow feature maps. The time-domain image feature map can be a feature representation of the analysis and extraction of consecutive frames in the time dimension. Based on such a time-domain image feature map, dynamic information of pixel points or local areas changing over time can be obtained. For example, the frame difference method can be used to express the difference between the current frame and the previous frame as a feature to facilitate the detection of changes in moving objects or scenes. The optical flow feature map can be used to calculate the optical flow vector and map it into a color or grayscale image to more intuitively understand and analyze the motion pattern of the target object between consecutive frames.

[0042] In some embodiments, the remote device 130 may utilize the model 130 for feature extraction. Figure 3A FIG2 is a schematic diagram showing the architecture of an exemplary model 130 for target tracking according to some embodiments of the present disclosure. The model 130 is generally divided into a feature extraction module 310 and a target tracking model 320 .

[0043] In some embodiments, the feature extraction module 310 includes an image feature extractor 302. The terminal device 110 uses the image feature extractor 302 to extract a temporal image feature map between multiple frames of the target video.

[0044] As an example, the image feature extractor 302 may include an image backbone network and a temporal feature transfer module. For example, it is desired to extract the image feature from the first frame f1 to the Nth frame f1 in the target video 301. N Perform object tracking. The terminal device 110 inputs the target video 301 into the image backbone network to extract the image features of a single frame. Such image backbone networks include, for example, high-performance deep neural networks such as ResNet and HRNet. Such image backbone networks can have multi-resolution feature fusion performance, the ability to handle multi-scale changes, and network depth. It should be understood that any appropriate image feature extractor 302 can be selected based on the overall requirements of the model 130, and this disclosure does not limit this.

[0045] The time domain feature transfer module is used to receive image features and convert them into time domain image features 303, so as to achieve modeling of time domain information. Such a time domain feature transfer module is composed of, for example, a two-layer convolutional gated recurrent unit (GRU). Convolutional GRU is a model that combines convolutional neural networks (CNN) and gated recurrent units (GRU) and can process time domain features of multiple frames. For example, the image backbone network receives the i-1th frame and the ith frame of the target video 301, extracts the image features of each frame, and outputs the image feature map F i-1 and image feature map F i The time domain feature transfer module receives the image feature map F i-1 and image feature map F i , and output the time domain image feature map F i `.

[0046] In some embodiments, the feature extraction module 310 further includes an optical flow network 304. The terminal device 110 uses the optical flow network 304 to extract an optical flow feature map 305 between adjacent frames (e.g., adjacent frames) in the target video. Such an optical flow network 304 uses, for example, a RAFT algorithm to extract an optical flow map from the previous frame to the current frame. For example, the optical flow network 304 receives the i-1th frame and the ith frame of the target video 301 and outputs an optical flow feature O i .

[0047] Continue to refer Figure 2 In block 220 , the terminal device 110 determines a three-dimensional reference position of an object in each of the plurality of frames based on the plurality of temporal image feature maps.

[0048] In some embodiments, target tracking module 320 includes detection model 330 . Figure 3BA schematic diagram of an example architecture 300B of a detection model according to some embodiments of the present disclosure is shown. The detection model 330 receives the temporal image feature map 303, and outputs a three-dimensional reference position 336 after processing. The three-dimensional reference position 336 can be the center position of the target object, for example, represented by the coordinates of the center point. As an example, when the target object is a human body, the three-dimensional reference position 336 can be, for example, the center of gravity position, shoulder position, torso center position, etc. of the human body. The three-dimensional reference position 336 can also be a bounding box to represent the spatial position occupied by the target object, for example, represented by the coordinates of the upper left corner and lower right corner of the bounding box, or represented by the coordinates of the center point, width, and height. Additionally or alternatively, the detection model 330 can also output a confidence level 333 corresponding to the three-dimensional reference position 336. The higher the confidence level, the higher the probability that each pixel belongs to the three-dimensional reference position of the object.

[0049] In some embodiments, the terminal device 110 may determine, based on the temporal image feature map 303, a plurality of reference position probability maps 332 and a plurality of reference position offset vector maps 335 for a plurality of frames through a three-dimensional reconstruction process. Each reference position probability map 332 indicates a confidence level 333 that each pixel in the corresponding frame belongs to the reference position of the object. Each reference position offset vector map 335 indicates an offset of a pixel in the corresponding frame relative to the reference position of the object. Further, the terminal device 110 determines a candidate reference position of the object in each frame based on the confidence levels 333 in the plurality of reference position probability maps, and determines a three-dimensional reference position 336 of the object based on the candidate reference position and its corresponding offset.

[0050] As an example, the target video 301 includes consecutive frames f1 to f N The target object is a human body. The detection model 330 receives the temporal image feature map F output by the feature extraction module 310 for the i-1th frame and the ith frame of the target video 301. i `, and output the reference position probability map 332 for the i-th frame after the three-dimensional reconstruction process (for example, using tensor ) and a reference position offset vector map (e.g., using a tensor Such a reference position probability map 332 includes, for example, the rough position of the center of gravity of the human body. Such a reference position offset vector map 335 includes, for example, the precise positioning offset vector Δt of the human body. i Combined with the rough position of the center of gravity of the human body and precise positioning offset vector Δt i , the detection model 330 outputs the three-dimensional reference position t of the human body for the i-th frame i and confidence c i .

[0051] In order to solve the problem of mutual occlusion between objects, the terminal device 110 can determine the three-dimensional reference position 336 based on the reference position probability map 332 and the reference position offset vector map 335 under two perspectives.

[0052] In some embodiments, terminal device 110 may determine, based on temporal image feature map 303, a first reference position probability map and a first reference position offset vector map at a first perspective, and a second position probability map and a second reference position offset vector map at a second perspective, respectively, through a three-dimensional reconstruction process. Furthermore, terminal device 110 may determine a reference position probability map 332 for the frame by combining the first reference position probability map and the second reference position probability map for the same frame, and determine a reference position offset vector map 335 for the frame by combining the first reference position offset vector map and the second reference position offset vector map.

[0053] In some embodiments, the terminal device 110 may determine the reference position probability map 332 and the reference position offset vector map 335 based on the temporal image feature map 303 through a three-dimensional reconstruction process based on a bird's eye view (BEV). Thus, the first perspective includes the main perspective, and the second perspective includes the bird's eye view.

[0054] Continuing with the above example, the detection model 330 receives the temporal image feature map F output by the feature extraction module 310 for the i-1th frame and the i-th frame of the target video 301. i `, and outputs a reference position probability map 332 and a reference position offset vector map 336 after the BEV-based 3D reconstruction process. Such a reference position probability map 332 includes, for example, the rough position of the center of gravity of the human body in the main view and the rough position of the center of gravity of the human body in the bird's-eye view. Such a reference position offset vector map 336 includes, for example, the precise positioning offset vector of the center of gravity of the human body in the main view and the precise positioning offset vector of the center of gravity of the human body in the bird's-eye view. By combining these positions and positioning offset vectors in the main view and the bird's-eye view, the detection model 330 outputs the 3D reference position t of the target object in the i-th frame. i and confidence c i .

[0055] Continue to refer Figure 2 At block 230, the terminal device 110 may determine at least a plurality of motion bias maps for the plurality of frames based on the plurality of temporal image feature maps 303 and the plurality of optical flow feature maps 305. At block 240, the terminal device 110 may determine a motion trajectory of the target object in the three-dimensional space based at least on the plurality of motion bias maps by referring to the three-dimensional reference position of the target object in each of the plurality of frames.

[0056] exist Figure 3AIn the example of , the target tracking module 320 can determine a motion bias map for all objects in each frame based on the image feature map 303 and the optical flow feature map 305, and then sample the motion bias map based on the three-dimensional reference position 336 of the target object to determine the motion trajectory 306 of the target object. In some embodiments, the target tracking module 320 also includes a first tracking model. The terminal device 110 can use this first tracking model to track the motion trajectory 306 of the target object. In some embodiments, the target tracking module 320 also includes a second tracking model. The terminal device 110 can use this second tracking model to determine the target object from all objects. In some embodiments, the target tracking module 320 also includes a morphology determination model 350. The terminal device 110 can use this morphology determination model 350 to perform three-dimensional reconstruction based on the shape and posture of the target object, and then combine the motion trajectory of the target object with the three-dimensional reconstruction result to output the motion trajectory 306 of the target object in three-dimensional space.

[0057] It should be understood that the functional division of the detection model 330, the first tracking model 340, the second tracking model 350, and the morphology determination model 360 included in the target tracking module 320 is merely exemplary. These models can be further split or combined, and the extracted feature maps can be processed in parallel or at least partially in parallel. The target tracking module 320 can also include any appropriate model to implement an end-to-end network from receiving the target video to outputting the motion trajectory of the target object.

[0058] In some embodiments, the first tracking model 340 determines multiple second motion bias maps based on the temporal image feature map 303 and the optical flow feature map 305. The first tracking model 340 determines the motion trajectory of the target object in three-dimensional space based at least on such second motion bias maps. Such a three-dimensional space is, for example, a space in a world coordinate system with the initial position of the target object in the reference frame as the origin. If there are multiple target objects, such a three-dimensional space can be, for example, a space in a world coordinate system with the center position of one of the target objects in the reference frame as the origin.

[0059] In some embodiments, the second tracking model 350 determines multiple first motion bias maps based on the temporal image feature map 303 and the optical flow feature map 305. The second tracking model 350 determines the motion trajectory of the target object in three-dimensional space based at least on these first motion bias maps. Such a three-dimensional space is, for example, a three-dimensional space based on a camera coordinate system. Therefore, the first motion bias map output by the second tracking model 350 and the second motion bias map output by the first tracking model 340 include motion biases in different coordinate systems.

[0060] The following will refer to Figures 3C to 3EA first tracking model 340 , a second tracking model 350 , and a morphology determination model 360 are described separately.

[0061] Figure 3C A schematic diagram of an example architecture 300C of a first tracking model according to some embodiments of the present disclosure is shown. The first tracking model 340 receives a temporal image feature map 303 and an optical flow feature map 305 of each frame, and outputs a motion trajectory 346 of a target object in a world coordinate system.

[0062] In some embodiments, the first tracking model 340 can determine a second motion offset map 344 for each frame based on the temporal image feature map 303 and the optical flow feature map 305. The second motion offset map 344 indicates a three-dimensional motion offset 345 of all objects in the corresponding frame relative to the previous frame in the world coordinate system. Furthermore, the first tracking model 340 determines a three-dimensional orientation map 342 for each frame based on the multiple temporal image feature maps 303 and the multiple optical flow feature maps 305. The three-dimensional orientation map 342 indicates the three-dimensional orientation 343 of all objects in the corresponding frame in the world coordinate system.

[0063] In some embodiments, the first tracking model 340 may determine the second motion bias map 344 through a BEV-based 3D reconstruction process.

[0064] In some embodiments, the first tracking model 340 may determine a motion trajectory 346 of the target object in the world coordinate system based on the second motion offset map 344 and the 3D orientation map 342 with reference to the 3D reference position of the target object.

[0065] As an example, the first tracking model 340 includes two ResNet deep convolutional neural networks. The first ResNet receives the temporal image feature map F output by the feature extraction module 310 for the i-1th frame and the i-th frame of the target video 301. i `And the optical flow feature map O i , to estimate the first motion bias map 344 in the world coordinate system (e.g., using tensor The second ResNet receives the time domain image feature map F i `And the optical flow feature map O i , to estimate the three-dimensional orientation map 342 in the world coordinate system. The first tracking model 340 samples the three-dimensional orientation map 342 based on the three-dimensional reference position corresponding to the confidence exceeding the confidence threshold, and outputs the three-dimensional orientation τ of the target object i The first tracking model 340 samples the second motion offset map 344 based on the three-dimensional reference position corresponding to the confidence exceeding the confidence threshold, and outputs the three-dimensional motion offset ΔT of the target object. iThus, the first tracking model 340 is based on the three-dimensional orientation τ of the target object. i and 3D motion offset ΔT i To determine the motion trajectory in the world coordinate system 346. For example, the motion trajectory of the k-th target object in the i-th frame can be represented as a three-dimensional motion offset and three-dimensional orientation

[0066] Figure 3D A schematic diagram of an example architecture 300D of a second tracking model according to some embodiments of the present disclosure is shown. The second tracking model 350 receives the temporal image feature map 303 and the optical flow feature map 305 of each frame, and outputs a motion trajectory 346 of the target object in the camera coordinate system.

[0067] In some embodiments, the second tracking model 350 can determine a first motion bias map for each frame based on the temporal image feature map and the optical flow feature map. The first motion bias map indicates the motion bias of all objects in the corresponding frame relative to the previous frame in the camera coordinate system. Based on this first motion bias map 351, the second tracking model 350 can determine the motion trajectory 355 of the target object in the camera coordinate system. This motion trajectory can be used as intermediate data during the model training process or as the final output data during the model application process.

[0068] In some embodiments, the second tracking model 340 may determine the first motion bias map 351 through a BEV-based 3D reconstruction process.

[0069] In some embodiments, the second tracking model 350 identifies the target object from all objects based on the object's tracking identifier. Furthermore, the second tracking model 340 can determine the target object's motion trajectory 355 in the camera coordinate system using the target object's 3D reference position in each frame and the first motion offset map 351. Correspondingly, the first tracking model 340 can also determine the target object's motion trajectory 346 in the world coordinate system using the target object's 3D reference position in each frame and the second motion offset map 344.

[0070] As an example, the second tracking model 350 includes a deep convolutional neural network such as a ResNet. Such a ResNet receives the temporal image feature map F output by the feature extraction module 310 for the i-1th frame and the i-th frame of the target video 301. i `And the optical flow feature map O i , to estimate the second motion bias map 351 in the camera coordinate system (e.g., using tensor denoted by ). Second motion offset map 351 includes 3D motion offsets 352 for all objects. Second tracking model 350 receives 3D reference positions 336 and corresponding confidence levels 333 for all objects from detection model 330 and inputs them into memory unit 353 along with 3D motion offsets 352. Based on tracking identifiers 354 output by memory unit 350, second tracking model 350 can distinguish the target object from all other objects.

[0071] In some embodiments, the second tracking model 350 determines an initial reference position of a target object in a reference frame of the target video 301 based on a user's selection of the target object. Such target objects are assigned a target tracking identifier. The second tracking model 350 determines a first motion offset map 351 based on the temporal image feature map 303 and the optical flow feature map 305, and extracts a 3D motion offset 352 for all objects from the first motion offset map 351. The 3D motion offset 352 may indicate the change in the 3D position of an object across adjacent frames. Furthermore, the second tracking model 350 may determine tracking identifiers 354 for each of the remaining objects, excluding the target object, based on the matching results of the 3D reference positions 336, confidence levels 333, and 3D motion offsets 352 of all objects in each frame with the initial position of the target object. Furthermore, the second tracking model 350 may determine the target object from among all objects based on the tracking identifiers 354 for each of these objects.

[0072] As an example, the terminal device 110 can receive a user's designation and determine a target object from the first frame of the target video 301. The target object is assigned a target tracking identifier 0. The memory unit 353 stores the initial reference position of the target object and the target tracking identifier 0. The second tracking model 350 matches the three-dimensional reference positions and three-dimensional motion offsets corresponding to the confidence levels exceeding the threshold confidence level of all objects in each frame with the initial reference position of the target object. For matched objects, the second tracking model 350 determines that they are the target objects, and their target tracking identifier is 0. For unmatched objects, such as objects that appear newly relative to the first frame, the second tracking model 350 assigns them a new tracking identifier, such as tracking identifier 1.

[0073] Furthermore, the second tracking model 350 samples the first motion offset map 351 based on the three-dimensional reference position corresponding to the determined confidence level of the target object exceeding the confidence threshold, and outputs the three-dimensional motion offset Δm of the target object. i The second tracking model 350 is based on the three-dimensional motion offset Δm of the target object in each frame. i , to determine the motion trajectory of the target object in the camera coordinate system 355. For example, the motion trajectory of the k-th target object in the i-1 frame corresponds to Combined 3D motion bias After that, its motion trajectory in the i-th frame corresponds to

[0074] Figure 3E FIG3 is a schematic diagram illustrating an example architecture of a morphology determination model 360 according to some embodiments of the present disclosure. The morphology determination model 360 receives a temporal image feature map 303 and an optical flow feature map 305 of each frame, and outputs a shape and pose 363 of a target object.

[0075] In some embodiments, the morphology determination model 360 determines a morphology feature map 361 for all objects in each frame based on the temporal image feature map 304 and the optical flow feature map 305. Such a morphology feature map 361 indicates the shape and pose of all objects in the corresponding frame. Furthermore, the morphology determination model 360 receives the three-dimensional reference position and confidence 341 of the target object for the corresponding frame determined by the second tracking model 350, and determines the shape and pose 363 of the target object in the frame from the morphology feature map 361.

[0076] As an example, the morphology determination model 360 is based on a Skinned Multi-Person Linear (SMPL) model for modeling and generating multiple person poses. Such a model can represent each person's body shape, joint angles, and limb deformations by parameterization. Specifically, the SMPL model is based on the temporal image feature map F of the i-th frame. i `And optical flow feature map O i , determine the grid parameters (e.g. using tensors denoted), and the feature vector of the human body is sampled according to the center of gravity position of the human body with the highest confidence corresponding to the target human body. Such feature vector is input to the fully connected layer 362 to estimate the SMPL parameters of the human body, such as the shape parameter θ i and attitude parameter β i .

[0077] The above describes the various models included in the target tracking module 320. In some embodiments, the target tracking module 320 is further configured to perform a three-dimensional reconstruction of the target object based on the motion trajectory of the target object in the three-dimensional space and the shape and posture of each frame. Based on the three-dimensional reconstruction result, the target tracking module 320 can output a motion trajectory with a three-dimensional form of the target object, such as Figure 3A The motion trajectory 306 of the target object in the example is shown.

[0078] In some embodiments, reference Figure 1, the terminal device 110 uses the model 130 to achieve target tracking. The application of the model 130 is described above. In some embodiments, the model 130 can be trained by the remote device 120, trained by the terminal device 110, or can be trained by other devices or in a cloud environment. In the model training stage, the model 130 receives the input of the sample video and supervises the three-dimensional reference position of the sample object in the sample video in the camera coordinate system and the world coordinate system, the two-dimensional shape and the two-dimensional position projected into the image. The loss function includes the true value and the predicted value of these parameters. In the model application stage, the model 130 receives the target video and the specified target object, and can obtain the shape of the target object in each frame and the motion trajectory in the world coordinate system and the camera coordinate system in an end-to-end and single-step manner.

[0079] In summary, this paper proposes a target tracking network based on motion bias graph representation, which can estimate the 3D shape, pose, and motion trajectory of a target object in the world coordinate system from monocular video captured by a dynamic camera. This target tracking network enables a single-step, end-to-end algorithm implementation.

[0080] Example device

[0081] Figure 4 FIG4 shows a block diagram of an image verification apparatus 400 according to some embodiments of the present disclosure. The apparatus 400 may be implemented in the remote device 120 and / or the terminal device 110. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0082] Apparatus 400 includes an extraction module 410 configured to extract multiple temporal image feature maps and multiple optical flow feature maps between multiple frames of a target video. Apparatus 400 also includes a three-dimensional reference position determination module 420 configured to determine a three-dimensional reference position of at least one object in each of the multiple frames based on the multiple temporal image feature maps. The at least one object includes a target object. Apparatus 400 also includes a motion bias map determination module 430 configured to determine at least a plurality of motion bias maps for the multiple frames based on the multiple temporal image feature maps and the multiple optical flow feature maps. Each motion bias map indicates a motion bias of the at least one object in a corresponding frame relative to a previous frame in three-dimensional space. Apparatus 400 also includes a motion trajectory determination module 440 configured to determine a motion trajectory of the target object in three-dimensional space based on at least the multiple motion bias maps by referencing the three-dimensional reference position of the target object in each of the multiple frames.

[0083] In some embodiments, the three-dimensional reference position determination module 420 is further configured to determine multiple reference position probability maps and multiple reference position offset vector maps for multiple frames through a three-dimensional reconstruction process based on multiple time-domain image feature maps, each reference position probability map indicating the confidence that each pixel in the corresponding frame belongs to the reference position of the object, and each reference position offset vector map indicating the offset of the pixel in the corresponding frame relative to the reference position of the object; based on the confidence in the multiple reference position probability maps, determine the candidate reference position of at least one object in each of the multiple frames; and based on the candidate reference position of at least one object in each of the multiple frames and the corresponding offset of the candidate reference position, determine the three-dimensional reference position of at least one object in each of the multiple frames.

[0084] In some embodiments, the three-dimensional reference position determination module 420 is further configured to determine multiple first reference position probability maps and multiple first reference position offset vector maps for multiple frames at a first perspective based on multiple time domain image feature maps; determine multiple second position probability maps and multiple second reference position offset vector maps for multiple frames at a second perspective based on multiple time domain image feature maps; for each of the multiple frames, determine the reference position probability map of the frame by combining the first reference position probability map and the second reference position probability map of the frame; and for each of the multiple frames, determine the reference position offset vector map of the frame by combining the first reference position offset vector map and the second reference position offset vector map of the frame.

[0085] In some embodiments, the first perspective comprises a main perspective, and the second perspective comprises a bird's-eye perspective.

[0086] In some embodiments, the motion bias map determination module 430 is further configured to determine a plurality of first motion bias maps for a plurality of frames based on a plurality of temporal image feature maps and a plurality of optical flow feature maps, each first motion bias map indicating a motion bias of at least one object in a corresponding frame relative to a previous frame in a camera coordinate system; and wherein the motion trajectory of the target object includes the motion trajectory of the target object in the camera coordinate system.

[0087] In some embodiments, the motion bias map determination module 430 is further configured to determine a plurality of second motion bias maps for a plurality of frames based on a plurality of temporal image feature maps and a plurality of optical flow feature maps, each second motion bias map indicating a three-dimensional motion bias of at least one object in a corresponding frame relative to a previous frame in a world coordinate system; and to determine a plurality of three-dimensional orientation maps for a plurality of frames based on a plurality of temporal image feature maps and a plurality of optical flow feature maps, each three-dimensional orientation map indicating a three-dimensional orientation of at least one object in a corresponding frame in a world coordinate system.

[0088] In some embodiments, the motion trajectory determination module 440 is further configured to determine the motion trajectory of the target object in the world coordinate system based on multiple second motion offset maps and multiple three-dimensional orientation maps by referring to the three-dimensional reference position of the target object.

[0089] In some embodiments, the motion trajectory determination module 440 is further configured to identify a target object from at least one object based on the respective tracking identifiers of the at least one object; extract the three-dimensional motion bias of the target object from at least a plurality of motion bias maps respectively by the three-dimensional reference position of the target object in each of the plurality of frames; and determine the motion trajectory of the target object in the three-dimensional space based at least on the extracted three-dimensional motion bias of the target object.

[0090] In some embodiments, the motion trajectory determination module 440 is further configured to determine the initial reference position of the target object in the reference frame of multiple frames based on the user's selection of the target object in the reference frame, and the target object is assigned a target tracking identifier; determine the three-dimensional motion bias of at least one object in each frame based on multiple time domain image feature maps and multiple optical flow feature maps, and the three-dimensional motion bias indicates the three-dimensional position change of at least one object across adjacent frames; determine the tracking identifiers of the remaining objects in at least one object except the target object based on the matching results of the three-dimensional reference position, confidence, and three-dimensional motion bias of at least one object in each frame and the initial position of the target object; and determine the target object from at least one object based on the tracking identifier of the at least one object.

[0091] In some embodiments, the device 400 also includes a morphological feature map determination module, which is configured to determine multiple object morphological feature maps for multiple frames based on multiple time domain image feature maps and multiple optical flow feature maps, each object morphological feature map indicating the shape and posture of at least one object in the corresponding frame; and determine the shape and posture of the target object in the multiple frames from the multiple morphological feature maps based on the three-dimensional reference position of the target object in the multiple frames.

[0092] In some embodiments, the apparatus 400 further includes a three-dimensional reconstruction module configured to perform three-dimensional reconstruction of the target object in multiple frames based on the motion trajectory of the target object in the three-dimensional space and the shape and posture of the target object in the multiple frames.

[0093] The units included in the device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 400 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0094] Figure 5 1 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. Figure 5 The illustrated electronic device 500 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to implement Figure 1 remote device 120 and / or terminal device 110.

[0095] like Figure 5 As shown, electronic device 500 is in the form of a general electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 500.

[0096] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 500.

[0097] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0098] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0099] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0100] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0101] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0103] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0104] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0105] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

1. A target tracking method, comprising: Extracting multiple temporal image feature maps and multiple optical flow feature maps between multiple frames of the target video; Based on the multiple temporal image feature maps, determining a plurality of reference position probability maps and a plurality of reference position offset vector maps for the multiple frames through a three-dimensional reconstruction process, each reference position probability map indicating a confidence that each pixel in a corresponding frame belongs to a reference position of an object, and each reference position offset vector map indicating an offset of a pixel in the corresponding frame relative to a reference position of the object; determining, based on the confidence levels in the plurality of reference position probability maps, a candidate reference position of at least one object in each of the plurality of frames; determining a three-dimensional reference position of at least one object in each of the plurality of frames based on candidate reference positions of the at least one object in each of the plurality of frames and offsets corresponding to the candidate reference positions, the at least one object including a target object; Determine, based on the multiple temporal image feature maps and the multiple optical flow feature maps, at least a plurality of motion bias maps for the multiple frames, each motion bias map indicating a motion bias of the at least one object in a corresponding frame relative to a previous frame in a three-dimensional space; as well as The motion trajectory of the target object in the three-dimensional space is determined based on at least the multiple motion offset maps by referring to the three-dimensional reference position of the target object in each of the multiple frames.

2. The method of claim 1 , wherein determining the plurality of reference position probability maps and the plurality of reference position offset vector maps for the plurality of frames comprises: Determining, based on the multiple time-domain image feature maps, multiple first reference position probability maps and multiple first reference position offset vector maps for the multiple frames at a first viewing angle; Determining, based on the multiple time-domain image feature maps, a plurality of second position probability maps and a plurality of second reference position offset vector maps for the multiple frames at a second viewing angle; For each frame of the plurality of frames, determining the reference position probability map of the frame by combining the first reference position probability map and the second reference position probability map of the frame; as well as For each frame of the plurality of frames, the reference position offset vector map of the frame is determined by combining the first reference position offset vector map and the second reference position offset vector map of the frame. The method of claim 2 , wherein the first perspective comprises a main perspective, and the second perspective comprises a bird's-eye perspective.

4. The method according to claim 1, wherein determining at least a plurality of motion bias maps for the plurality of frames based on the plurality of temporal image feature maps and the plurality of optical flow feature maps comprises: Determining a plurality of first motion bias maps for the plurality of frames based on the plurality of temporal image feature maps and the plurality of optical flow feature maps, each first motion bias map indicating a motion bias of the at least one object in a corresponding frame relative to a previous frame in a camera coordinate system; and The motion trajectory of the target object includes the motion trajectory of the target object in the camera coordinate system.

5. The method according to claim 1, wherein determining at least a plurality of motion bias maps for the plurality of frames based on the plurality of temporal image feature maps and the plurality of optical flow feature maps comprises: Determining a plurality of second motion offset maps for the plurality of frames based on the plurality of temporal image feature maps and the plurality of optical flow feature maps, each second motion offset map indicating a three-dimensional motion offset of the at least one object in a corresponding frame relative to a previous frame in a world coordinate system; as well as A plurality of three-dimensional orientation maps for the plurality of frames are determined based on the plurality of temporal image feature maps and the plurality of optical flow feature maps, each three-dimensional orientation map indicating a three-dimensional orientation of the at least one object in a corresponding frame in a world coordinate system.

6. The method according to claim 5, wherein determining the motion trajectory of the target object in three-dimensional space comprises: The motion trajectory of the target object in the world coordinate system is determined based on the multiple second motion offset maps and the multiple three-dimensional orientation maps by referring to the three-dimensional reference position of the target object.

7. The method according to claim 1, wherein determining the motion trajectory of the target object in three-dimensional space comprises: identifying the target object from the at least one object based on the respective tracking identifiers of the at least one object; extracting a three-dimensional motion offset of the target object from at least the plurality of motion offset maps respectively according to the three-dimensional reference position of the target object in each of the plurality of frames; as well as A motion trajectory of the target object in three-dimensional space is determined based at least on the extracted three-dimensional motion offset of the target object.

8. The method of claim 7, wherein identifying the target object from the at least one object comprises: determining an initial reference position of the target object in a reference frame of the plurality of frames based on a user selection of the target object in the reference frame, the target object being assigned a target tracking identifier; determining a three-dimensional motion offset of the at least one object in each frame based on the multiple temporal image feature maps and the multiple optical flow feature maps, the three-dimensional motion offset indicating a three-dimensional position change of the at least one object across adjacent frames; Determining tracking identifiers of respective objects other than the target object in the at least one object based on matching results of the 3D reference position, confidence, and 3D motion offset of the at least one object in each frame with the initial position of the target object; as well as The target object is determined from the at least one object based on the tracking identification of each of the at least one object.

9. The method according to claim 1, further comprising: Determining a plurality of object morphology feature maps for the plurality of frames based on the plurality of temporal image feature maps and the plurality of optical flow feature maps, each object morphology feature map indicating a shape and a posture of the at least one object in a corresponding frame; as well as Based on the three-dimensional reference position of the target object in the multiple frames, the shape and posture of the target object in the multiple frames are determined from the multiple object morphological feature maps.

10. The method according to claim 7, further comprising: Based on the motion trajectory of the target object in the three-dimensional space and the shape and posture of the target object in the multiple frames, the target object in the multiple frames is three-dimensionally reconstructed.

11. A device for target tracking, comprising: An extraction module is configured to extract a plurality of temporal image feature maps and a plurality of optical flow feature maps between a plurality of frames of a target video; A three-dimensional reference position determination module is configured to Based on the multiple temporal image feature maps, determining a plurality of reference position probability maps and a plurality of reference position offset vector maps for the multiple frames through a three-dimensional reconstruction process, each reference position probability map indicating a confidence that each pixel in a corresponding frame belongs to a reference position of an object, and each reference position offset vector map indicating an offset of a pixel in the corresponding frame relative to a reference position of the object; determining, based on the confidence levels in the plurality of reference position probability maps, a candidate reference position of at least one object in each of the plurality of frames; as well as determining the three-dimensional reference position of the at least one object in each of the plurality of frames based on candidate reference positions of the at least one object in each of the plurality of frames and the offsets corresponding to the candidate reference positions, the at least one object including a target object; a motion bias map determining module configured to determine, based on the multiple temporal image feature maps and the multiple optical flow feature maps, at least a plurality of motion bias maps for the multiple frames, each motion bias map indicating a motion bias of the at least one object in a corresponding frame relative to a previous frame in a three-dimensional space; as well as The motion trajectory determination module is configured to determine the motion trajectory of the target object in the three-dimensional space based on at least the multiple motion offset maps by referring to the three-dimensional reference position of the target object in each of the multiple frames.

12. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Video image smoke detection method based on dense optical flow

    CN107301375A

  • Crowd abnormal behavior detection method based on comprehensive optical flow feature descriptor and track

    CN110287870A