Method, computer readable memory and apparatus for updating trajectory of target
Patent Information
- Application Number
- CN202610342654.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2026-03-19
- Publication Date
- 2026-09-25
AI Technical Summary
然而,目标检测算法可能不总是提供正确的位置信息,并且当多个目标接近时,它们可能被混淆
Smart Images

Figure CN122820764A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target tracking. Specifically, it relates to a method and apparatus for updating the trajectory of a target in relation to a current image frame. Background Technology
[0002] To track targets in a sequence of video images from a camera, tracking filters such as Kalman filters are typically used. The purpose of these filters is to filter out noisy detections of targets from image frames in the video to output a smooth target trajectory. These filters include motion models that model how the target's state, such as its position and velocity, evolves from one point in time to another. When using filters, the motion model is used to predict the target's state in the current image frame based on the target's state in previous image frames. Considering the target detections from a computer-implemented target detector using an object detection algorithm, the predicted states in some or all of the video image sequences can be updated. However, the target detection algorithm may not always provide accurate positional information, and they can become confused when multiple targets are close together. Therefore, updating the predicted positions with positional information from a computer-implemented target detector does not constitute an enhancement. Thus, there is room for improvement. Summary of the Invention
[0003] The object of the present invention is to alleviate the above-mentioned problems and to provide a new method, non-transitory computer-readable storage, and apparatus for updating the trajectory of a target in relation to a current image.
[0004] According to a first aspect, a method for updating the trajectory of a target associated with a current image frame is provided. A first tracker implemented by a computer is used to determine a first trajectory of the target in a sequence of image frames preceding the current image frame. Based on target detection from a target detector implemented by a computer using a target detection algorithm, the first tracker implemented by the computer is updated at least in a first subset of image frames in the image frame sequence. The first trajectory includes a corresponding estimated position of the target in each image frame of the first subset of image frames in the image frame sequence. A motion model of the first tracker implemented by the computer predicts the position of the target in the current image frame. Furthermore, determined corresponding positions of a plurality of detected motion regions in the current image frame are received from a motion detector implemented by a computer using a motion detection algorithm. Using the first tracker implemented by the computer, and based on the predicted position of the target in the current image frame and the determined corresponding positions of the plurality of detected motion regions, a detected motion region among the plurality of detected motion regions is determined to be associated with the target in the current image frame. Using the first tracker implemented by the computer, and based on the determined positions of the detected motion regions determined to be associated with the target in the current image frame, the predicted position of the target in the current image frame is updated. Thus, the estimated position of the target in the current image frame is determined.
[0005] This invention is based on the understanding that there are scenarios where target detection algorithms cannot provide accurate location information, and scenarios where targets may be confused due to proximity, and motion detection may identify targets that target detection has not identified. In such scenarios, the determined location of the motion region detected by the motion detection algorithm can be used to update the predicted location of the target, thereby enhancing target tracking.
[0006] By receiving the determined positions of detected motion regions in the current image frame from a computer-implemented motion detector, and determining that one of the detected motion regions is associated with a target in the current image frame, a computer-implemented first tracker can be used to update the predicted position of the target in the current image frame based on the determined positions of the detected motion regions. This improves the accuracy of the estimated position of the target in the current image determined by the update.
[0007] According to a second aspect, a method for updating the trajectory of a target associated with a current image frame is provided. A first trajectory of the target in a sequence of image frames preceding the current image frame is determined using a computer-implemented first tracker. Based on target detection from a computer-implemented target detector using a target detection algorithm, the computer-implemented first tracker is updated at least in a first subset of image frames in the image frame sequence. The first trajectory includes a corresponding estimated position of the target in each of the first subset of image frames in the image frame sequence. A motion model of the computer-implemented first tracker predicts the position of the target in the first trajectory of the current image frame. Furthermore, a second trajectory is determined to be associated with the first trajectory. A second trajectory is determined in a sequence of image frames preceding the current image frame using a computer-implemented second tracker. Based on motion detection from a computer-implemented motion detector using a motion detection algorithm, the computer-implemented second tracker is updated at least in a second subset of image frames in the image frame sequence. An estimated position in the second trajectory of the target in the current image frame is received from the computer-implemented second tracker. The predicted position of the target in the current image frame is updated using the computer-implemented first tracker and based on the received estimated position in the second trajectory of the target in the current image frame. Thus, the estimated position in the first trajectory of the target in the current image frame is determined.
[0008] By determining that the second trajectory is associated with the first trajectory, and receiving the estimated position of the target in the second trajectory of the current image frame from a computer-implemented second tracker, the predicted position of the target in the current image frame can be updated using the computer-implemented first tracker based on the determined position of the detected motion region. This improves the accuracy of the estimated position of the target in the current image by updating the determined position.
[0009] According to the third aspect, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium including computer program code, which, when executed by a device having processing capability, causes the device to perform one of the methods of the first aspect or the methods of the second aspect.
[0010] According to the fourth aspect, an apparatus is provided for updating the trajectory of a target in relation to a current image frame, the apparatus including circuitry configured to perform one of the methods of the first aspect or the second aspect.
[0011] It should be understood that, because such apparatus and methods can vary, the present invention is not limited to the specific components of the described apparatus or the operation of the described method. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It must be noted that, as used in the specification and appended claims, unless the context clearly specifies otherwise, the words “a,” “the,” and “the” are intended to indicate the presence of one or more elements. Thus, for example, a reference to “a unit” or “the unit” can include several devices, etc. Furthermore, the words “comprising,” “including,” “containing,” and similar wording do not exclude other elements or steps. Attached Figure Description
[0012] The above and additional objects, features and advantages of the invention will be better understood from the following illustrative and non-limiting detailed description of embodiments of the invention with reference to the accompanying drawings, in which the same reference numerals will be used for similar elements, in which:
[0013] Figure 1 The schematic map illustrates an apparatus for updating the trajectory of a target in relation to the current image frame, according to an embodiment.
[0014] Figure 2 This is a flowchart of a method for updating the trajectory of a target in relation to the current image frame, according to an embodiment.
[0015] Figure 3 The schematic map illustrates another apparatus for updating the trajectory of a target in relation to the current image frame, according to an embodiment.
[0016] Figure 4 This is a flowchart of another method for updating the trajectory of a target in relation to the current image frame, according to an embodiment. Detailed Implementation
[0017] In the following description, the invention will now be described more fully with reference to the accompanying drawings, in which embodiments of the invention are illustrated. The systems and apparatus disclosed herein will be described during operation.
[0018] Figure 1 The figure illustrates an apparatus 100 for updating the trajectory of a target in relation to the current image frame. The apparatus includes circuitry 102 configured to perform a method for updating the trajectory of a target in relation to the current image frame. Circuitry 102 is configured to perform various functions of apparatus 100. These functions correspond to a computer-implemented target detector 104, a computer-implemented first tracker 106, and a computer-implemented motion detector 108. Hereinafter, these are referred to as target detector 104, first tracker 106, and motion detector 108, respectively.
[0019] In a hardware implementation, each of functions 104, 106, and 108 may correspond to a dedicated circuit specifically designed to perform that function. The circuit may take the form of one or more application-specific integrated circuits (ASICs) or one or more field-programmable gate arrays (FPGAs). For example, the first tracker 106 may therefore include circuitry for determining the trajectory of a target in a sequence of image frames during use.
[0020] In a software implementation, instead, the circuitry may take the form of a processor, such as a microprocessor, associated with computer code instructions stored on a (non-transitory) computer-readable medium, such as non-volatile memory, to cause device 100 to perform any of the methods disclosed herein. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, and optical discs. In the software case, functions 104, 106, and 108 may therefore each correspond to a portion of the computer code instructions stored on the computer-readable medium that, when executed by the processor, causes device 100 to perform that function.
[0021] To further understand, some of functions 104, 106, and 108 can be implemented entirely in hardware, while others can be implemented in software stored on a computer-readable medium and executed by a processor (not shown) of device 100.
[0022] When in use, image frame sequence 120 is input to device 100. Image frame sequence 120 is input to target detector 104 configured to detect targets in image frames. Target detector 104 can detect targets in every image frame, but can also operate at a lower frame rate to detect targets in a first subset of image frames (e.g., in every nth image frame), where n>1. Target detector 104 can take a single image frame as input and can provide target detection 130 of one or more targets in the image frame as output. Target detection can take the form of regions in the image frame where targets are detected (referred to herein as detection regions) and can be given in the form of bounding boxes. In addition to detection regions, target detector 104 can provide additional information about target detection, such as target category and confidence score of target classification. Target detector 104 can be configured to detect one or more targets of a specific type or target category, such as people, vehicles, etc. To this end, target detector 104 can detect targets by extracting features from the image frames. That is, target detector 104 can detect targets based on their appearance in the image frames. Accordingly, the detection performed by object detector 104 can be described as feature-based object detection or appearance-based object detection. For example, object detector 104 can implement a deep learning model that has been trained to identify features in image frames that correspond to one or more objects of a specific object category of interest. Many such models are known in the art, such as the YOLO object detector (https: / / arxiv.org / abs / 1506.02640) which implements a convolutional neural network for this task.
[0023] Each target detection 130 from target detector 104 is input to first tracker 106, which outputs a trajectory 140 of the target in the image frame sequence 120. The target trajectory 140 includes an estimated location (referred to herein as the estimated location) that the first tracker 106 believes the target is located in the image frames of sequence 120. In one example, the first tracker 106 determines the target location for each image frame in the image frame sequence 120. In another example, the first tracker 106 determines the target location for those image frames in which target detector 104 has already detected the target. Furthermore, when the first tracker 106 receives motion detection 150 from motion detector 108, the first tracker 106 determines the target location for those image frames.
[0024] Generally, the first tracker 106 can implement a tracking filter that predicts the state of the target based on the target detection 130 provided by the target detector 104. For example, the tracking filter can be a Kalman filter or a particle filter. Specifically, the tracking filter can predict the statistical distribution of the target's state, such as, in the case of a Gaussian distribution, a statistical distribution expressed by its mean vector and covariance matrix. The tracking filter models the dynamics of the target using a motion model, such as a linear motion model. The state of the target can be defined by its target region (bounding box) (e.g., its position, width, and height), velocity vector, and rate of change of size. It is understood that other definitions of the state, such as the positions of the two diagonals of the target region and the velocity vector, are also possible.
[0025] A motion model is a model that describes the motion of a target according to a function (e.g., a linear function in the case of a linear motion model). Specifically, it can refer to a model that, for example, models the temporal evolution of the state of a target in a sequence of image frames from a point in time corresponding to one image frame in the sequence to another point in time corresponding to a subsequent image frame in the sequence, using a linear function. For example, a motion model can model the temporal evolution from one image frame to the next in a sequence. In particular, a motion model can be a constant velocity model, i.e., a model that assumes the target moves at a constant velocity. Therefore, a motion model can be used to predict the state of the target in subsequent image frames given the state of the target in the current image frame. Specifically, it can be used to predict the target region in which the target is located in subsequent image frames. The state of the target can be described, for example, by a state vector including the target's position, velocity, size, and rate of change of size in the image frame. The target's position and size together define the target region in which the target is located in the image frame. Sometimes, a motion model may be referred to as a linear kinematic model or a linear dynamic model.
[0026] The first tracker 106 uses a motion model to predict the state of the target at time t corresponding to the current image frame based on the state of the target at the previous time point t-1 corresponding to the previous image frame. Furthermore, if the target detector 104 has already detected the target in the current image frame (i.e., it has observed the state of the target), the first tracker 106 updates the predicted state in consideration of the target detection, thereby determining the estimated position of the target in the image frame.
[0027] Image frame sequence 120 is further input to motion detector 108. Motion detector 108 is configured to detect motion in the image frames and output motion detection 150, specifically in the form of regions in the image frames where motion exists (referred to herein as motion regions). The motion detector can detect motion in every image frame, but typically operates at a lower frame rate to detect motion in each m-th image frame, where m > 1. Object detector 104 and motion detector 108 may operate at different frame rates. For example, object detector 104 may operate at a higher frame rate than motion detector 108, and vice versa. Furthermore, this may cause the motion detector and object detector to be inactive relative to the same frame, or even relative to any identical frame. Motion exists in a region when the pixel value in that region changes over time (e.g., between consecutive image frames). Therefore, motion detector 108 can be configured to find motion regions by detecting changes in the image frames. For example, motion detector 108 can locate motion regions by detecting changes between the current image frame and previous image frames in sequence 120, by detecting changes between the current image frame and the background model (also known as background subtraction), or a combination of these methods.
[0028] Although both motion detector 108 and object detector 104 output regions in an image frame where objects may be located, they do so using different principles. Motion detector 108 looks for moving regions, i.e., pixel regions with changing pixel values, while object detector 104 looks for regions with features corresponding to a specific category of object. Each of these principles has its advantages and disadvantages. For example, motion detector 108 is only sensitive to motion, meaning it will only detect moving objects. This contrasts with object detector 104, which can detect both moving and stationary objects. Furthermore, motion detector 108 detects any moving object regardless of its appearance or object category, while object detector 104 detects objects with a specific appearance or object category.
[0029] Motion detection 150 from motion detector 108 is input to first tracker 106. Motion detection 150 is used to update the predicted position of the target to determine the estimated position included in the output trajectory 140.
[0030] Now refer to Figure 2 Flowchart and further reference Figure 1 This describes the operation of device 100 when method 200, used to update the trajectory of a target in relation to the current image frame, is executed. If several targets are being tracked, it can be understood that method 200 can be applied to each tracked target.
[0031] Method 200 includes a first tracker implemented using a computer (such as...) Figure 1 The first tracker 106 determines the first trajectory of the target in the sequence of image frames preceding the current image frame in S210. This is based on a computer-implemented target detector (e.g., using a target detection algorithm). Figure 1 For target detection by the target detector 104, the first tracker 106 is updated at least in a first subset of image frames in the image frame sequence. The first subset of image frames may correspond, for example, to an update rate of 15, 10, or 3 times per second. For a video with 30 frames per second, this would correspond to each second, third, or tenth frame of the image frame sequence, respectively. Other update rates are, of course, possible depending on what is appropriate in the application. The first track includes the corresponding estimated position of the target in each image frame of the first subset of image frames in the image frame sequence. The estimated position is typically included in the state of the target in each image frame. In addition to position, the state may be defined by width and height, velocity vector, and the rate of change of the target's dimensions.
[0032] Then, the motion model of the first tracker 106 is used to predict the position of the target in the current image frame of S220. The prediction using the first tracker 106 can be performed in every frame, but it can also be performed only in the frame where either motion detection or target detection is performed.
[0033] The corresponding location of the detected motion region in the current image frame is determined from a computer-implemented motion detector (such as...) using motion detection algorithms. Figure 1 The motion detector 108 receives S230. Motion detection using the motion detector 108 can be performed asynchronously with updates using the object detector 104. Thus, motion detection and object detection are typically not performed in the same frame. The corresponding position of the motion region detected in the current image frame can be performed using the motion detector 108 and can be provided to the first tracker 106.
[0034] Using the first tracker, and based on the predicted position of the target in the current image frame and the determined corresponding positions of the detected motion regions, in step S240, the detected motion regions are determined to be associated with the target in the current image frame. Using the first tracker 106, and based on the determined positions of the detected motion regions determined to be associated with the target in the current image frame, in step S250, the predicted position of the target in the current image frame is updated. Thus, the estimated position of the target in the current image frame is determined.
[0035] In summary, the predicted position / state of the current image frame is based on the estimated position / state of the previous image frames. When the predicted position / state of the current image frame is updated, it becomes the estimated position / state of the current frame.
[0036] Only under the condition that criterion C242 is met can the first tracker 106 update the predicted position of the target in the current image in S250 based on the determined position of the detected motion region determined to be associated with the target in the current image frame. C242 is that the detected motion region determined to be associated with the target does not overlap with a trajectory different from the first trajectory. Here, a different trajectory means a trajectory associated with a target different from the target associated with the first trajectory.
[0037] This reduces the risk of updating the predicted location of a target based on detected motion regions associated with different targets.
[0038] Additionally, the first tracker 106 may be used to update the predicted position of the target in the current image based on the determined position of the detected motion region that is determined to be associated with the target in the current image frame, only if one or more of the other criteria C244 are met. The other criteria C244 are:
[0039] The estimated position of the target in each of at least a first predetermined number of image frames in the sequence of image frames preceding the current image frame is not based on target detection from a computer-implemented target detector;
[0040] The spatial confidence level of determining the predicted location of the target in the current image frame is lower than a first spatial confidence threshold; and
[0041] The target determination speed in the current image frame exceeds the speed threshold.
[0042] This further reduces the risk of updating the predicted location of a target based on detected motion regions associated with different targets.
[0043] Method 200 may further include an optional action of receiving, from the first tracker 106, the predicted spatial extension of the target in the current image frame in S232. The predicted spatial extension of the target in the current image frame may be defined, for example, by width and height, polygons, masks, or any other way of defining spatial extension. The predicted spatial extension, together with the predicted position of the target in the current image frame, determines the predicted target region of the target in the current image frame. The predicted spatial extension of the target in the current image frame from the first tracker 106 is typically a predicted spatial extension determined based on image frames preceding the current image frame. The predicted spatial extension in the current image may be included in the state of the first tracker 106 related to the current image. The method then further includes using the first tracker 106 to determine the spatial overlap between the detected motion region determined in S236 to be associated with the target in the current image frame and the predicted target region of the target in the current image frame. In the action of updating S250, the higher the degree of determined spatial overlap, the greater the degree to which the position of the detected motion region determined to be associated with the target in the current image frame will be considered when updating the predicted position of the target in the current image frame.
[0044] The estimated position of the target in each image frame of a first subset of image frames in the action of determining the first trajectory in S210 can be determined by performing the following actions on each image frame of the first subset of image frames: First, the motion model of the first tracker 106 is used to predict the position of the target in the image frame. Then, the target detector 104 is used to determine the target in the image frame. Then, the target detector 104 is used to determine the position of the detected target in the image frame. Then, the first tracker 106 is used to update the predicted position of the target in the image frame based on the determined position of the detected target in the image frame. Thus, the estimated position of the target in the image frame is determined.
[0045] Method 200 may further include preventing the use of the first tracker 106 to update the predicted position of the target in the current image frame based on the determined position of the detected target in the current image frame, provided that the following criteria are met:
[0046] There are two or more target detections from target detector 104 in the current image that may be associated with a target in the current image frame;
[0047] There exists a second tracker implemented using a computer (not in) Figure 1 (as shown in the figure) a second trajectory in the sequence of image frames preceding the current image frame, wherein, based on motion detection from motion detector 108, a computer-implemented second tracker is updated at least in a second subset of image frames in the image frame sequence;
[0048] The second trajectory is associated with the first trajectory; and
[0049] Within a second predetermined number of image frames in the second subset of image frames in the sequence of image frames preceding the current image frame, the second trajectory is not the result of splitting the trajectory into two trajectories or merging two trajectories into one trajectory.
[0050] Splitting a trajectory into two trajectories means that a single trajectory that was previously detected as a single target is now detected as two separate trajectories for two separate targets. Merging two trajectories into a single trajectory means that two separate trajectories that were previously detected as two separate targets are now detected as a single trajectory for a single target.
[0051] Preventing updates to the predicted position of a target in the current image frame is only meaningful when the target detector 104 performs target detection in the current image frame. If the target detector 104 performs target detection on subsequent image frames, updates to the predicted position of the target can be performed for subsequent image frames based on this criterion. This prevention related to subsequent image frames can be limited to image frames preceding the next image frame for which the motion detector 108 performs motion detection. Thereafter, prevention is performed based on this criterion on the next image frame.
[0052] Prevent updates related to target detection based on target detector 104. According to this standard, in a scenario where a target is tracked via a continuous trajectory from a computer-implemented second tracker based on motion detection from motion detector 108, updates to the predicted position of the target detection based on target detector 104 are prevented.
[0053] Figure 3 The diagram illustrates an apparatus 300 for updating the trajectory of a target in relation to the current image frame. The apparatus includes circuitry 302 configured to perform a method for updating the trajectory of a target in relation to the current image frame. Circuitry 302 is configured to perform various functions of apparatus 300. These functions correspond to a computer-implemented target detector 304, a computer-implemented first tracker 306, a computer-implemented motion detector 308, and a computer-implemented second tracker 310. Hereinafter, these are referred to as target detector 304, first tracker 306, motion detector 308, and second tracker 310, respectively.
[0054] In a hardware implementation, each of functions 304, 306, 308, 310, and 312 may correspond to a dedicated circuit specifically designed to perform that function. The circuit may take the form of one or more application-specific integrated circuits (ASICs) or one or more field-programmable gate arrays (FPGAs). For example, the first tracker 306 may therefore include circuitry for determining the trajectory of a target in a sequence of image frames during use.
[0055] In a software implementation, instead, the circuitry may take the form of a processor, such as a microprocessor, associated with computer code instructions stored on a (non-transitory) computer-readable medium, such as non-volatile memory, to cause device 300 to perform any of the methods disclosed herein. Examples of non-volatile memory include read-only memory, flash memory, ferroelectric RAM, magnetic computer storage devices, and optical discs. In the software case, functions 304, 306, 308, 310, and 312 can therefore each correspond to a portion of the computer code instructions stored on a computer-readable medium that, when executed by the processor, causes device 300 to perform that function.
[0056] To further understand, some of functions 304, 306, 308, 310, and 312 can be implemented entirely in hardware, while others can be implemented in software stored on a computer-readable medium and executed by a processor (not shown) of device 300.
[0057] When in use, an image frame sequence 320 is input to the device 300. The image frame sequence 320 is input to a target detector 304 configured to detect targets in the image frames. The target detector 304 can detect targets in every image frame, but can also operate at a lower frame rate to detect targets in a first subset of image frames (e.g., in every nth image frame), where n > 1. The target detector 304 can take a single image frame as input and can provide target detection 330 of one or more targets in the image frame as output. Target detection can take the form of regions in the image frame where targets are detected (referred to herein as detection regions) and can be given in the form of bounding boxes. In addition to detection regions, the target detector 304 can provide additional information about the target detection, such as the target category and a confidence score for the target classification. The target detector 304 can be configured to detect one or more targets of a specific type or target category, such as people, vehicles, etc. To this end, the target detector 304 can detect targets by extracting features from the image frames. That is, the target detector 304 can detect targets based on their appearance in the image frames. Accordingly, the detection performed by object detector 304 can be described as feature-based object detection or appearance-based object detection. For example, object detector 304 can implement a deep learning model that has been trained to recognize features in image frames that correspond to targets of one or more specific target categories of interest. Many such models are known in the art, such as the YOLO object detector (https: / / arxiv.org / abs / 1506.02640), which implements a convolutional neural network for this task.
[0058] Each target detection 330 from target detector 304 is input to first tracker 306, which outputs a trajectory 340 of the target in the image frame sequence 320. Target tracker 340 includes an estimated location (referred to herein as the estimated location) that the first tracker 306 believes the target is located in the image frames of sequence 320. In one example, first tracker 306 determines the target location for each image frame in image frame sequence 320. In another example, first tracker 306 determines the target location for those image frames in which target detector 304 has already detected the target. Furthermore, first tracker 306 determines the target location for these image frames when it receives input 360 from second tracker 310.
[0059] Generally, the first tracker 306 can implement a tracking filter that predicts the state of the target based on the target detection 330 provided by the target detector 304. For example, the tracking filter can be a Kalman filter or a particle filter. Specifically, the tracking filter can predict the statistical distribution of the target's state, such as, in the case of a Gaussian distribution, a statistical distribution expressed according to its mean vector and covariance matrix. The tracking filter models the dynamics of the target using a motion model, such as a linear motion model. The state of the target can be defined by its target region (bounding box) (e.g., its position, width, and height), velocity vector, and rate of change of size. It is understood that other definitions of the state, such as the positions of the two diagonals of the target region and the velocity vector, are also possible.
[0060] The first tracker 306 uses a motion model to predict the state of the target at time t corresponding to the current image frame, based on the state of the target at the previous time point t-1 corresponding to the previous image frame. Furthermore, if the target detector 304 has already detected the target in the current image frame (i.e., it has observed the state of the target), the first tracker 306 updates the predicted state in consideration of the target detection, thereby determining the estimated position of the target in the image frame.
[0061] Image frame sequence 320 is further input to motion detector 308. Motion detector 308 is configured to detect motion in the image frames and output motion detection 350, specifically in the form of motion detection 150 in the form of regions in the image frames where motion exists (referred to herein as motion regions). The motion detector can detect motion in every image frame, but typically operates at a lower frame rate to detect motion in each m-th image frame, where m > 1. Object detector 304 and motion detector 308 may operate at different frame rates. For example, object detector 304 may operate at a higher frame rate than motion detector 308, and vice versa. Motion exists in a region when the pixel value in the region changes over time (e.g., between consecutive image frames). Therefore, motion detector 308 can be configured to find motion regions by detecting changes in the image frames. For example, motion detector 308 can find motion regions by detecting changes between the current image frame and previous image frames in sequence 320, by detecting changes between the current image frame and the background model (also referred to as background subtraction), or a combination of these methods.
[0062] Although both motion detector 308 and object detector 304 output regions in an image frame where objects may be located, they do so using different principles. Motion detector 308 looks for moving regions, i.e., pixel regions with changing pixel values, while object detector 304 looks for regions with features corresponding to a specific category of object. Each of these principles has its advantages and disadvantages. For example, motion detector 308 is only sensitive to motion, meaning it will only detect moving objects. This contrasts with object detector 304, which can detect both moving and stationary objects. Furthermore, motion detector 308 detects any moving object regardless of its appearance or object category, while object detector 304 detects objects with a specific appearance or object category.
[0063] Motion detection 350 from motion detector 308 is input to second tracker 310. Second tracker 310 forms one or more trajectories based on motion detection 350 and provides these motion trajectories 360 as output. Second tracker 310 can typically track moving regions in a sequence of image frames. For example, second tracker 310 can correlate moving regions in different image frames with each other, as they may correspond to the same target motion. This can be simply based on the spatial proximity of moving regions in subsequent image frames, but using a tracking filter such as a Kalman filter is also possible. In the latter case, the tracking filter is preferably configured to operate under process noise conditions greater than that of the first tracker 306 and / or use an acceleration term in the state vector. This is possible because tracking moving regions is generally a easier problem than target detection, which includes static targets in addition to moving targets, increasing the risk of identity switching. Therefore, second tracker 310 will be better at handling nonlinear motion. Thus, each motion trajectory 360 includes a moving region in which motion has been detected (e.g., where changes relative to previous image frames or background models have been detected). In one example, motion trajectory 360 includes the motion region of each image frame in sequence 320. In another example, motion trajectory 306 includes the motion region of an image frame in which motion detector 308 has detected motion.
[0064] A motion trajectory 360, including the motion area from the second tracker 310, is input to the first tracker 306. The motion trajectory 360 is used to update the predicted position of the target to determine the estimated position included in the output trajectory 340.
[0065] Now refer to Figure 4 Flowchart and further reference Figure 3 This describes the operation of the apparatus 300 when the method 400 for updating the trajectory of a target in relation to the current image frame is executed. If several targets are being tracked, it can be understood that the method 400 can be applied to each tracked target.
[0066] Method 400 includes a first tracker implemented using a computer (such as...) Figure 3 The first tracker 306 determines the first trajectory of the target in the sequence of image frames preceding the current image frame in S410. This is based on a computer-implemented target detector (e.g., using a target detection algorithm). Figure 3The target detector 304 performs target detection, and the first tracker 306 is updated at least in a first subset of image frames in the image frame sequence. The first subset of image frames may correspond, for example, to an update rate of 15, 10, or 3 times per second. For a video with 30 frames per second, this would correspond to each second, third, or tenth frame of the image frame sequence, respectively. Other update rates are, of course, possible depending on what is appropriate in the application. The first trajectory includes the corresponding estimated position of the target in each image frame of the first subset of image frames in the image frame sequence. In addition to position, the state may be defined by width and height, velocity vector, and the rate of change of the target's dimensions.
[0067] Then, the motion model of the first tracker 306 is used to predict the position of the target in the first trajectory in the current image frame of S420. The prediction using the first tracker 306 can be performed in every frame, but it can also be performed only in the frame where either motion detection or target detection is performed.
[0068] Then, the association between the S430 second trajectory and the first trajectory is determined. A second tracker implemented using a computer (such as...) Figure 3 The second tracker 310 in the image determines a second trajectory in the sequence of image frames preceding the current frame. This is based on a computer-implemented motion detector (such as...) that uses motion detection algorithms. Figure 3 The motion detection using motion detector 308 is performed by the second tracker 310, which updates at least in a second subset of image frames in the image frame sequence. Motion detection using motion detector 308 can be performed asynchronously with updates using object detector 104. Thus, motion detection and object detection are typically not performed in the same frame.
[0069] Then, the estimated position of the target in the second trajectory in the current image frame of S440 is received from the second tracker.
[0070] The first tracker 306 is used to update the predicted position of the target in the current image frame in S450 based on the received estimated position in the second trajectory of the target in the current image frame. Thus, the estimated position of the target in the first trajectory of the current image frame is determined.
[0071] Only under the condition that standard C442 is met can the first tracker 306 be used to update the predicted position of the target in the current image in S450 based on the received estimated position in the second trajectory of the target in the current image frame. Standard C242 is that the estimated position in the second trajectory of the target in the current image frame does not overlap with a trajectory that is different from the first trajectory in the current image frame. Here, a different trajectory means a trajectory related to a target that is different from the target related to the first trajectory.
[0072] This reduces the risk of updating the predicted location of a target based on detected motion regions associated with different targets.
[0073] Additionally, the predicted position of the target in the current image can be updated using the first tracker 306 based on the received estimated position in the second trajectory of the target in the current image frame, provided that one or more of the other criteria C444 are met. The further criteria C444 are:
[0074] The estimated position of the target in each of at least a first predetermined number of image frames in the sequence of image frames preceding the current image frame is not based on target detection from a computer-implemented target detector;
[0075] The spatial confidence level of determining the predicted location of the target in the current image frame is lower than a first spatial confidence threshold; and
[0076] The target determination speed in the current image frame exceeds the speed threshold.
[0077] This further reduces the risk of updating the predicted location of a target based on detected motion regions associated with different targets.
[0078] Method 400 may further include the optional action of receiving, from the first tracker 106, a predicted spatial extension of the first trajectory of the target in the current image frame in S432. The predicted spatial extension of the target in the current image frame may be defined, for example, by predicted width and height, polygons, masks, or any other way of defining spatial extension. The predicted spatial extension, together with the predicted position of the target in the current image frame, determines the predicted target region of the target in the current image frame. The predicted spatial extension of the target from the first tracker 106 in the current image frame is typically a predicted spatial extension determined based on image frames preceding the current image frame. The predicted spatial extension in the current image is typically included in the state of the first tracker 106 in relation to the current image. Then, method 400 further includes: receiving, from the second tracker 310, an estimated spatial extension in the second trajectory of the target in the current image frame in S434, wherein the estimated spatial extension in the second trajectory of the target in the current image frame, together with the estimated position in the second trajectory of the target in the current image frame, determines the estimated target region in the second trajectory of the target in the current image frame. The estimated spatial range in the second trajectory of the target in the current image frame can be determined based on the corresponding motion region in the current image frame. Then, method 400 further includes using the first tracker 306 to determine the spatial overlap between the estimated spatial extension of the target in the second trajectory of the target in the current image frame and the predicted target region in the first trajectory of the target in the current image frame. In the update S450 action, the higher the degree of spatial overlap determined, the greater the degree to which the received estimated position of the target in the second trajectory of the current image is taken into account when updating the predicted position of the target in the current image frame.
[0079] The estimated position of the target in each image frame of a first subset of image frames in the action of determining the first trajectory of S410 can be determined by performing the following actions on each image frame of the first subset of image frames: First, the motion model of the first tracker 306 is used to predict the position of the target in the image frame. Then, the target detector 304 is used to determine the target in the image frame. Then, the target detector 304 is used to determine the position of the detected target in the image frame. Then, the first tracker 306 is used to update the predicted position of the target in the image frame based on the determined position of the detected target in the image frame. Thus, the estimated position of the target in the image frame is determined.
[0080] The estimated position of the target in each image frame of the second subset of image frames in the image frame sequence can be determined by performing the following actions on each image frame of the second subset of image frames: First, the position of the target in the image frame is predicted using the motion model of the second tracker 310. Then, a motion region is detected in the image frame using a motion detector 308. The position of the detected motion region in the image frame is then determined using the motion detector 308. The predicted position of the target in the image frame is updated using the second tracker 310 based on the determined position of the detected motion region in the image frame. Thus, the estimated position of the target in the second trajectory of the image frame is determined.
[0081] Method 200 may further include preventing the use of the first tracker 106 to update the predicted position of the target in the current image frame based on the determined position of the detected target in the current image frame, provided that the following criteria are met:
[0082] There are two or more target detections that indicate a spatial overlap with the predicted location of the target in the first trajectory of the target in the current image frame;
[0083] The second trajectory is associated with the first trajectory; and
[0084] Within a third predetermined number of image frames in the second subset of image frames in the sequence of image frames preceding the current image frame, the second trajectory is not the result of splitting the trajectory into two trajectories or merging two trajectories into one trajectory.
[0085] Splitting a trajectory into two trajectories means detecting a single trajectory that has already been detected as a single target as two separate trajectories for two separate targets. Merging two trajectories into a single trajectory means detecting two separate trajectories that have already been detected as two separate targets as a single trajectory for a single target.
[0086] Preventing updates to the predicted position of a target in the current image frame is only meaningful when the target detector 104 performs target detection in the current image frame. If the target detector 104 performs target detection on subsequent image frames, updates to the predicted position of the target can be performed for subsequent image frames based on this criterion. This prevention related to subsequent image frames can be limited to image frames preceding the next image frame for which the motion detector 108 performs motion detection. Thereafter, prevention is performed based on this criterion on the next image frame.
[0087] Prevent updates related to target detection based on target detector 104. According to this standard, in a scenario where a target is tracked via a continuous trajectory from a computer-implemented second tracker based on motion detection from motion detector 108, updates to the predicted position based on target detection from target detector 104 are prevented.
[0088] It should be understood that those skilled in the art can modify the embodiments described above in various ways while still utilizing the advantages of the invention as shown in the above embodiments. Therefore, the invention should not be limited to the embodiments shown, but should be defined only by the appended claims. Furthermore, as those skilled in the art will understand, the embodiments shown can be combined.
Claims
1. A method for updating the trajectory of a target in relation to a current image frame, comprising: A computer-implemented first tracker determines a first trajectory of a target in a sequence of image frames preceding the current image frame, wherein the computer-implemented first tracker is updated at least in a first subset of image frames in the image frame sequence based on target detection from a computer-implemented target detector using a target detection algorithm, and wherein the first trajectory includes a corresponding estimated position of the target in each of the first subset of image frames in the image frame sequence. The motion model of the first tracker implemented by the computer is used to predict the position of the target in the current image frame; Receive the determined positions of multiple motion regions detected in the current image frame from a computer-implemented motion detector using a motion detection algorithm; Using the first tracker implemented by the computer, based on the predicted position of the target in the current image frame and the determined corresponding positions of the detected plurality of motion regions, it determines that a detected motion region among the detected plurality of motion regions is associated with the target in the current image frame; and Only if the detected motion region determined to be associated with the target does not overlap with a trajectory different from the first trajectory: The computer-implemented first tracker updates the predicted position of the target in the current image frame based on the determined position of the detected motion region that is determined to be associated with the target in the current image frame, thereby determining the estimated position of the target in the current image frame.
2. The method according to claim 1, wherein, Under the condition that one or more of the following additional criteria are met, the first tracker implemented using the computer updates the predicted position of the target in the current image frame based on the determined position of the detected motion region determined to be associated with the target in the current image frame: The estimated position of the target in each of at least a first predetermined number of image frames in the sequence of image frames preceding the current image frame is not based on target detection from the computer-implemented target detector; The spatial confidence of determining the predicted location of the target in the current image frame is lower than a first spatial confidence threshold; and The speed at which the target is determined in the current image frame exceeds a speed threshold.
3. The method according to claim 1, further comprising: The computer-implemented first tracker receives a predicted spatial extension of the target in the current image frame, wherein the predicted spatial extension, together with the predicted position of the target in the current image frame, determines an estimated target region of the target in the current image frame. The first tracker, implemented using the computer, determines the spatial overlap between the detected motion region identified as associated with the target in the current image frame and the estimated target region of the target in the current image frame. In the update operation, the higher the degree of spatial overlap, the more the detected motion region associated with the target in the current image frame will be considered when updating the predicted position of the target in the current image frame.
4. The method according to claim 1, wherein, For each image frame in the first subset of the image frame sequence, the estimated position of the target in each image frame in the first subset of the image frame sequence is determined by the following steps: The motion model of the first tracker implemented by the computer predicts the position of the target in the image frame; The target detector implemented by the computer detects the target in the image frame; The computer-implemented target detector is used to determine the location of the detected target in the image frame; as well as Using a first tracker implemented in the computer, the predicted position of the target in the image frame is updated based on the determined position of the detected target in the image frame, thereby determining the estimated position of the target in the image frame.
5. The method of claim 1, further comprising: The motion detector implemented using the computer detects multiple motion regions in the current image frame; The computer-implemented motion detector is used to determine the corresponding positions of the detected multiple motion regions in the current image frame; as well as The determined corresponding positions of the detected multiple motion regions in the current image frame are provided to the computer-implemented first tracker.
6. The method of claim 1, further comprising: Under the following conditions, the predicted position of the target in the current image frame is prevented from being updated using the computer-implemented first tracker, based on the estimated position of the detected target in the current image frame: There are two or more target detections from the computer-implemented target detector in the current image that can be associated with the target in the current image frame; There exists a second trajectory in the sequence of image frames preceding the current image frame, determined by a second tracker implemented using a computer, wherein the second tracker implemented using the computer is updated in at least a second subset of image frames in the image frame sequence based on motion detection from a motion detector implemented using the computer. The second trajectory is associated with the first trajectory; and Within a second predetermined number of image frames in the second subset of the image frames preceding the current image frame, the second trajectory is not the result of splitting the trajectory into two trajectories or merging two trajectories into one trajectory.
7. A method for updating the trajectory of a target in relation to the current image frame, comprising: A computer-implemented first tracker determines a first trajectory of a target in a sequence of image frames preceding the current image frame, wherein the computer-implemented first tracker is updated at least in a first subset of image frames in the image frame sequence based on target detection from a computer-implemented target detector using a target detection algorithm, and wherein the first trajectory includes a corresponding estimated position of the target in each of the first subset of image frames in the image frame sequence. The motion model of the first tracker implemented by the computer is used to predict the position of the target in the first trajectory in the current image frame; The second trajectory is determined to be associated with the first trajectory, wherein a computer-implemented second tracker is used to determine the second trajectory in the sequence of image frames preceding the current image frame, wherein the computer-implemented second tracker is updated in at least a second subset of image frames in the sequence of image frames based on motion detection from a computer-implemented motion detector using a motion detection algorithm; Receive the estimated position of the target in the second trajectory in the current image frame from the computer-implemented second tracker; and Only if the estimated position of the target in the second trajectory in the current image frame does not overlap with a trajectory different from the first trajectory in the current image frame: The computer-implemented first tracker updates the predicted position of the target in the current image frame based on the received estimated position in the second trajectory of the target in the current image frame, thereby determining the estimated position of the target in the first trajectory of the target in the current image frame.
8. The method according to claim 7, wherein, Under the condition that one or more of the following additional criteria are met, the first tracker implemented using the computer updates the predicted position of the target in the current image frame based on the received estimated position in the second trajectory of the target in the current image frame: The estimated position of the target in each of at least a first predetermined number of image frames in the sequence of image frames preceding the current image frame is not based on target detection from the computer-implemented target detector; The spatial confidence of determining the predicted location of the target in the current image frame is lower than a first spatial confidence threshold; and The speed at which the target is determined in the current image frame exceeds a speed threshold.
9. The method of claim 7, further comprising: The computer-implemented first tracker receives a predicted spatial extension in the first trajectory of the target in the current image frame, wherein the predicted spatial extension in the first trajectory of the target in the current image frame, together with the predicted position in the first trajectory of the target in the current image frame, determines a predicted target region in the first trajectory of the target in the current image frame. The estimated spatial extension of the second trajectory of the target in the current image frame is received from the second tracker implemented in the computer, wherein the estimated spatial extension of the second trajectory of the target in the current image frame, together with the estimated position in the second trajectory of the target in the current image frame, determines the estimated target region in the second trajectory of the target in the current image frame; The computer-implemented first tracker determines the spatial overlap between the estimated target region in the second trajectory of the target in the current image frame and the predicted target region in the first trajectory of the target in the current image frame. In the update operation, the higher the degree of spatial overlap, the more the estimated position of the second trajectory of the target in the current image frame is taken into account when updating the predicted position of the first trajectory of the target in the current image frame.
10. The method according to claim 7, wherein, For the computer-implemented first tracker, for each image frame in the first subset of image frames in the image frame sequence, the estimated position of the target in each image frame in the first subset of image frames in the image frame sequence is determined by the following steps: The motion model of the first tracker implemented by the computer predicts the position of the target in the image frame; The target detector implemented by the computer detects the target in the image frame; The computer-implemented target detector is used to determine the location of the detected target in the image frame; as well as Using the first tracker implemented by the computer, the predicted position of the target in the image frame is updated based on the determined position of the detected target in the image frame, thereby determining the estimated position of the first trajectory of the target in the image frame.
11. The method according to claim 7, wherein, For the computer-implemented second tracker, for each image frame in the second subset of image frames in the image frame sequence, the estimated position of the target in each image frame in the second subset of image frames in the image frame sequence is determined by the following steps: The motion model of the second tracker implemented by the computer predicts the position of the target in the image frame; The motion detector implemented using the computer detects moving regions in the image frame; The location of the detected motion region in the image frame is determined using a motion detector implemented in the computer. as well as Using the second tracker implemented by the computer, the predicted position of the target in the image frame is updated based on the determined position of the detected motion region in the image frame, thereby determining the estimated position of the second trajectory of the target in the image frame.
12. The method of claim 7, further comprising: Under the following conditions, based on the determined position of the detected target in the current image frame, the first tracker implemented by the computer is prevented from updating the predicted position of the target in the current image frame: There are two or more target detections that indicate an association and have spatial overlap with the predicted location in the first trajectory of the target in the current image frame; The second trajectory is associated with the first trajectory; as well as The second trajectory is not the result of segmentation or merging within a third predetermined number of image frames in the second subset of the image frame sequence preceding the current image frame.
13. A non-transitory computer-readable medium comprising computer program code, which, when executed by a means having processing capability, causes the means to perform the method according to claim 1 or claim 7.
14. An apparatus for updating the trajectory of a target in relation to a current image frame, comprising circuitry configured to perform the functions of the method according to claim 1 or claim 7.