Method, system, and computer-readable medium for prioritizing objects

By prioritizing the object based on motion data and overlap in object tracking applications, the problem of insufficient occlusion and computing resources is solved, and the accuracy and efficiency of object re-identification are improved.

CN118968009BActive Publication Date: 2025-07-11AXIS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410588504.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-05-15
Filing Date
2024-05-13
Publication Date
2025-07-11
Estimated Expiration
2044-05-13

AI Technical Summary

Technical Problem

In object tracking applications, the prior art is difficult to effectively deal with the problem of object occlusion and insufficient computing resources, resulting in inaccuracy and inefficiency of object re-identification.

Method used

By receiving image frames and motion data, determining the region of interest (ROI), and prioritizing objects based on the motion region and overlap, selecting appropriate objects for feature extraction, avoiding future motion prediction and relying on existing trackers, and reducing the risk of feature vector leakage.

Benefits of technology

It improves the accuracy and computing efficiency of object tracking, reduces the possibility of error association, adapts to scene motion, effectively identify potential occlusion objects, and optimizes the use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968009B_ABST
    Figure CN118968009B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, system, and computer-readable medium for prioritizing objects. Regions of interest (ROIs) 110, 116 for object feature extraction are determined based on motion regions in image frames 100a to 100e. Each object 102, 104, 112 detected in the image frame and at least partially overlapping with the ROI is associated with the ROI. For each ROI associated with two or more objects, a list of candidate objects for feature extraction is determined by adding each object among the two or more objects that does not overlap with any of the other objects among the two or more objects by more than a threshold amount. At least one object is selected from the list of candidate objects, and image data of the image frame depicting the selected object is used to determine a feature vector of the selected object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object re-identification, and in particular, to a method, system, and non-transitory computer-readable medium for prioritizing objects for feature extraction for the purpose of object re-identification in an object tracking application. Background Art

[0002] In an object tracking application, re-identification refers to the process of identifying an object in a video sequence as the same object that was previously detected in one or more previous frames of that video sequence or in frames of another video sequence. This is an important task in object tracking because it allows the system to maintain continuous tracking of the movement of the object across the image frames of the video sequence.

[0003] The process of object re-identification can be challenging because objects in a video sequence may undergo various transformations over time (e.g., changes in appearance, orientation, scale, and occlusion). These changes make it difficult to continuously track the object from one image frame to the next.

[0004] To address these challenges, object tracking systems typically use a combination of object detection, feature extraction, and matching techniques to identify and track objects in a video sequence.

[0005] Feature extraction involves extracting a set of descriptive features (feature vectors) from the objects detected in each frame. These features can include color, texture, shape, or other characteristics that can be used to distinguish one object from another. Feature extraction can be sensitive to object occlusion. For example, if two objects partially overlap, the feature extraction process may extract a feature vector representing the combination of the two objects rather than two feature vectors for each object separately. This can lead to confusion during the matching process because the extracted features may not match well with the corresponding features in the previous frame.

[0006] In addition, it may not be possible or suitable to determine feature vectors for every object detected in every image frame of a video sequence because this can increase the computational resources required for both the feature extraction process and the matching process.

[0007] Therefore, improvements are needed in this context. Summary of the Invention

[0008] In view of the above, it would be beneficial to solve or at least reduce one or more of the disadvantages discussed above, as set forth in the appended independent patent claims.

[0009] According to a first aspect of the present invention, there is provided a method for prioritizing objects for feature extraction for the purpose of object re-identification in an object tracking application, comprising the steps of: receiving an image frame depicting a scene including a plurality of objects; receiving object detection data including positioning data indicating the position and spatial extent of each of the plurality of objects in the image frame; receiving motion data indicating one or more motion regions in the image frame, each motion region corresponding to a region in the scene where motion has been detected; for each motion region, determining a region of interest ROI for object feature extraction in the image frame, the ROI for object feature extraction overlapping the motion region; for each ROI for object feature extraction, using the object detection data to determine a list of objects that at least partially overlap the ROI for object feature extraction, and associating the ROI for object feature extraction with the list of objects.

[0010] The method further comprises the steps of: for each ROI for object feature extraction associated with two or more objects, determining a list of prioritized candidate objects for feature extraction by: for each of the two or more objects, adding the object to the list of prioritized candidate objects for feature extraction when it is determined that the object does not overlap any of the other objects in the plurality of objects by more than a threshold amount.

[0011] The method further comprises the steps of: selecting at least one object from the list of prioritized candidate objects, and for each selected object, determining a feature vector of the object based on the image data of the image frame according to the positioning data of the selected object.

[0012] Advantageously, the current method can provide low-complexity and effective selection criteria against which objects can be analyzed to determine feature vectors for object re-identification. For example, the method does not rely on future object motion prediction for object tracking where there may be a risk of lost or misassociated objects. Instead, motion data indicating one or more motion regions in the (currently analyzed) image frame is used. The motion data can be obtained by motion detection techniques such as frame difference that do not require prediction. By avoiding reliance on future object motion prediction, the current method reduces the risk of associating incorrect object tracking with detected objects.

[0013] Furthermore, the method of the present invention does not rely on existing object tracking or the objects associated therewith, thereby increasing the flexibility of the method. The selection of objects for feature vector determination is decoupled from existing object tracking and its associations, which can provide greater freedom to select the most appropriate objects for feature vector determination.

[0014] In addition, the current method can reduce the risk of feature vector leakage between objects, thereby reducing the likelihood of associating incorrect object tracking with the objects detected in the image frame. By reducing the risk of feature vector leakage, the current method improves the accuracy of object tracking and reduces the incidence of incorrect associations between object tracking and the objects in the image frame.

[0015] In addition, by giving a priority order to the motion regions where two or more objects in the image are located, the objects in subsequent image frames that may be occluded or cause occlusion are prioritized for feature extraction. This method prioritizes the objects that the object tracker in subsequent image frames may lose track of.

[0016] In some embodiments, the step of determining the ROI for object feature extraction includes expanding the motion region by a predetermined range in each direction. Advantageously, a low-complexity manner of considering the motion in the scene captured by the image frame can be implemented. The ROI can be determined such that the ROI includes the motion region, which is expanded with an additional margin (predetermined range) around it. The motion region can be expanded to identify objects that may potentially occlude each other in the near future or have been occluded not long ago, and these motion regions are selected for feature vector determination to facilitate object re-identification. By using this method, an effective and efficient means for analyzing scene motion to identify relevant objects for feature vector determination and subsequent object re-identification can be achieved.

[0017] In some embodiments, the motion data further includes an indication of the speed of the motion detected in the corresponding region of the scene for each motion region in the image frame, and the step of determining the ROI for object feature extraction includes expanding the motion region based on the speed. Thus, the ROI for object feature extraction is determined by combining the motion region with an additional margin around it, and the additional margin is based on the characteristics of the detected motion. This method can achieve a flexible and precise analysis of scene motion to identify suitable objects for feature vector determination and subsequent object re-identification. By using this method, the current embodiment can achieve a targeted and adaptive means of selecting the ROI for feature extraction, thereby improving object tracking and re-identification performance.

[0018] In some embodiments, the ROI for object feature extraction is determined by expanding the motion region to a greater extent in the direction corresponding to the direction of the speed than in the direction not corresponding to the direction of the speed. Advantageously, when determining the ROI for object feature extraction, the future possible motion in the scene captured by the first image can be considered, which can improve the identification of suitable objects for feature vector determination and subsequent object re-identification.

[0019] In an example, the shape of the ROI for object feature extraction is one of a pixel mask, a circle, an ellipse, and a rectangle. Advantageously, a flexible method for determining the ROI for object feature extraction can be implemented. For example, a motion area can be used to determine a pixel mask-shaped ROI, and the motion area can be subjected to morphological operations (e.g., dilation) to expand it. Alternatively, a circular-shaped ROI can be obtained by determining the smallest circle enclosing the motion area, and the smallest circle can optionally be further dilated. An ellipse can be determined similarly to the circle, and the ellipse can optionally be expanded according to the direction of the speed. When determining a rectangular-shaped ROI, the same strategy can be utilized.

[0020] In some examples, the threshold amount is 0. This means that no overlap between the objects detected in the ROI for object feature extraction is allowed, which in turn can reduce the risk of feature vector leakage between the objects, thereby reducing the likelihood associated with incorrect object tracking of the objects detected in the image frame. In the context of the present disclosure, feature vector leakage refers to the fact that the pixel data of overlapping objects affects or degrades the calculation of the feature vectors of the objects being overlapped. Pixel leakage between bounding boxes can corrupt the feature vectors and, for example, make the feature vectors more similar than they should be due to the fact that the bounding boxes share pixels. Thus, feature vector leakage can lead to inaccuracies in the feature representation and subsequent inaccuracies in the analysis or tasks (e.g., re-identification) performed on the objects.

[0021] In other examples, a partial overlap is allowed, such that the threshold amount is a predetermined percentage of the spatial extent of one of the overlapping objects. For example, assume that a first object of size X in an image frame overlaps with a second object that is 100 times larger (size 100X) in the same image frame. The second object overlaps with the first object by only 1%, while the first object overlaps with the second object by 100%. In the case where the threshold amount is 1% (including 1%), the second object will be added to the list of candidate objects sorted by priority for feature extraction, while the first object will not be added. Thus, the object size is taken into account, which can act on the impact caused by feature vector leakage between the objects.

[0022] In some embodiments, at most N objects among the list of candidate objects sorted by priority are selected, where N is a predetermined number. Advantageously, N can be selected considering the computational resources available for object re-identification.

[0023] In an example, when the list of prioritized candidate objects consists of more than N objects, the selection step includes: comparing each prioritized candidate object with one or more selection criteria and selecting (at most) N objects that meet at least one of the one or more selection criteria. The selection criteria can be related to object size, object detection probability, object motion, object type, object shape, object orientation, etc. As a result, the current example can provide a flexible method for identifying suitable objects for feature vector determination and object re-identification taking into account any constraints imposed by limited computational resources. This method can allow for efficient use of available resources and improve the performance of object tracking and re-identification tasks.

[0024] In some examples, when the list of prioritized candidate objects consists of multiple N objects, the selection step includes: sorting each prioritized candidate object according to one or more sorting criteria and selecting the N objects with the highest ranking. For example, the N objects can be sorted according to object size, object detection score, object motion, object type, object shape, object orientation, etc. As a result, the current example can provide a flexible method for identifying suitable objects for feature vector determination and object re-identification taking into account any constraints imposed by limited computational resources. This method can allow for efficient use of available resources and improve the performance of object tracking and re-identification tasks.

[0025] In some examples, the method further includes the step of: associating the determined feature vector with at least a portion of the localization data of the selected object for each selected object. When performing object re-identification, the localization data can be used to determine the relevance of the feature vector for a specific tracking.

[0026] According to some embodiments, the method further includes the step of: associating the determined feature vector with a timestamp indicating the capture time of the image frame for each selected object. During the object re-identification process, the timestamp can be used to evaluate the suitability of the feature vector for a specific object tracking. For example, the maximum allowed distance between the position of the selected object and the tracking can be adjusted based on the time difference between the timestamp of the latest object instance in the tracking and the selected object. This method can achieve more accurate and relevant object re-identification taking into account the time-related dynamics of the objects being tracked.

[0027] According to a second aspect of the present invention, the above object is achieved by a system for prioritizing feature extraction for object re-identification in an object tracking application, the system comprising: one or more processors; and one or more non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the system to perform actions including the following: receiving an image frame depicting a scene including a plurality of objects; receiving object detection data including localization data indicating the position and spatial extent of each of the plurality of objects in the image frame; receiving motion data indicating one or more motion regions in the image frame, each motion region corresponding to a region in the scene where motion has been detected; for each motion region, determining a region of interest (ROI) in the image frame for object feature extraction, the ROI for object feature extraction overlapping the motion region; for each ROI for object feature extraction, using the object detection data to determine a list of objects that at least partially overlap the ROI for object feature extraction, and associating the ROI for object feature extraction with the list of objects; for each ROI for object feature extraction associated with two or more objects, determining a list of prioritized candidate objects for feature extraction by: for each of the two or more objects, adding the object to the list of prioritized candidate objects for feature extraction when it is determined that the object does not overlap any of the other objects in the plurality of objects by more than a threshold amount; selecting at least one object from the list of prioritized candidate objects, and for each selected object, determining a feature vector of the object based on image data of the image frame according to the localization data of the selected object.

[0028] According to a third aspect of the present invention, the above object is achieved by one or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations including the following: receiving an image frame depicting a scene including a plurality of objects; receiving object detection data including positioning data indicating the position and spatial extent of each of the plurality of objects in the image frame; receiving motion data indicating one or more motion regions in the image frame, each motion region corresponding to a region in the scene where motion has been detected; for each motion region, determining a region of interest (ROI) in the image frame for object feature extraction, the ROI for object feature extraction overlapping with the motion region; for each ROI for object feature extraction, using the object detection data to determine a list of objects that at least partially overlap with the ROI for object feature extraction, and associating the ROI for object feature extraction with the list of objects; for each ROI for object feature extraction associated with two or more objects, determining a list of candidate objects in a prioritized order for feature extraction by: for each of the two or more objects, adding the object to the list of candidate objects in a prioritized order for feature extraction when it is determined that the object does not overlap with any of the other objects in the plurality of objects by more than a threshold amount; selecting at least one object from the list of candidate objects in a prioritized order, and for each selected object, determining a feature vector of the object based on the image data of the image frame according to the positioning data of the selected object.

[0029] The second and third aspects generally may have the same features and advantages as the features and advantages of the first aspect. Further, it should be noted that the present disclosure relates to all possible combinations of the various features, unless otherwise explicitly stated. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The above and additional objects, features, and advantages of the present invention will be better understood from the following illustrative and non-limiting detailed description of embodiments of the present disclosure with reference to the accompanying drawings, in which like reference numerals will be used for like elements, and in the drawings:

[0031] Figure 1 Illustrates a scenario for prioritizing feature extraction for object re-identification according to an embodiment;

[0032] Figures 2 to 4 Illustrates a schematic example of determining a region of interest for object feature extraction according to an embodiment;

[0033] Figure 5 Illustrates overlapping bounding boxes of objects according to an embodiment;

[0034] Figure 6 A flowchart of a method for prioritizing feature extraction for object re - identification in an object tracking application according to an embodiment is shown;

[0035] Figure 7 A flowchart of a method for determining a prioritized list of candidate objects for feature extraction from one or more regions of interest for object feature extraction according to an embodiment is shown;

[0036] Figure 8 A system implementing a method according to an embodiment is shown Figures 6 to 7 is shown. DETAILED DESCRIPTION

[0037] Object tracking involves the process of detecting and monitoring the movement of objects captured in a series of image frames in a scene. In the case where a tracking algorithm fails or loses track of an object for any reason, re - identification helps to regain the lost track. Due to several factors, re - identification can be computationally complex. For example, to perform re - identification, features that can effectively capture its appearance, pose, and other unique attributes may need to be extracted from the object. When matching objects, the resulting high - dimensional feature space may lead to increased computational requirements. In addition, the need for real - time processing and the use of computationally intensive deep - learning models for feature extraction from objects may result in a situation of insufficient available computational resources, especially when a large number of objects are detected in the scene.

[0038] One reason for the need for re - identification in object tracking applications is occlusion. An object may be partially or completely occluded by other objects, making it difficult for the tracking system to maintain continuous tracking. Re - identification (ReID) helps to identify an object once it reappears from behind the object imposing the occlusion. However, occlusion can complicate the extraction of the feature vectors (appearance vectors, ReID vectors, etc.) of the object imposing the occlusion or the occluded object. For example, when people in a crowd partially occlude each other, the pixel data related to the first person who partially occludes the second person (i.e., the pixel data from the bounding box of the first person) may affect or degrade the calculation of the feature vector of the second person. Such a situation can complicate the process of accurately distinguishing and identifying each individual.

[0039] The present disclosure aims to solve the above two problems by providing methods, systems, and non - transitory computer - readable media for prioritizing feature extraction for object re - identification in object tracking applications.

[0040] Figure 1shows a scene captured by five image frames 100a to 100e. In this scene, three objects 102, 104, 112 appear. In Figure 1 the scene, all the objects are people, but it should be noted that the techniques described herein can be applied to any suitable object type (e.g., cars, bicycles, animals, etc.). The scene can be any type of scene (e.g., an indoor scene (room, corridor, library, convenience store, shopping mall, entrance of a building, etc.) or an outdoor scene (road, park, suburb, beach, etc.)). The techniques discussed here can provide greater advantages in situations with a larger number of objects. However, for simplicity of explanation, Figure 1 demonstrates the case with only three objects 102, 104, 112. In Figure 1 it, the positions and spatial extents of the objects 102, 104, 112 are shown as dashed (long dash) rectangles 106, 108, and 114. Thus, in this example, the positions and spatial extents are shown as bounding boxes. However, depending on the object detection algorithm applied to it, the forms of the positions and spatial extents of the detected objects 102, 104, 112 can be different and can include segmentation masks, bounding polygons, etc.

[0041] In addition, in Figure 1 it, the regions of interest ROI 110, 116 for object feature extraction in the image frames are shown as dashed (short dash) rectangles. Each ROI 110, 116 is determined based on the motion regions (regions in the scene where motion has been detected), which will be further discussed in conjunction with Figures 2 to 4 below. It is important to note that Figure 1 the ROIs 110 and 116 depicted in it are included for illustrative purposes and may not accurately represent them because these ROIs 110 and 116 will be identified in the actual occurring situations.

[0042] Each ROI 110, 116 overlaps with the corresponding moving objects 104, 112 in the scene. The method discussed herein for prioritizing feature extraction in object re-identification focuses on restricting feature extraction to ROIs that are at least partially overlapped by at least two objects, which can limit feature extraction to objects at risk of being occluded or objects that may have recently been occluded. In Figure 1 the background, it means that at least two bounding boxes 106, 108, 114 intersect with the ROIs 110, 116.

[0043] In the first image frame 100a, two persons 102, 104 are shown. The first person 102 is stationary, while the second person 104 is moving. In this first frame, the two persons are not very close to each other, such that feature extraction for re-identification purposes as described herein can be operative. However, it should be noted that according to some embodiments, in the case where sufficient computing power is available for feature extraction and re-identification, feature extraction can be performed on the two persons 102 and 104 even if they are not determined to be objects prioritized for feature extraction as described herein.

[0044] In the second image frame 100b, a third moving person 112 enters the scene. Additionally, the first person 102 and the second person 104 are closer to each other. More specifically, the two persons 102, 104 at least partially overlap with the ROI 110 (i.e., the bounding boxes 106, 108 of the persons 102, 104 at least partially overlap with the ROI 110). Thus, re-identification may be required in the future to maintain continuous tracking of the persons 102, 104. However, to avoid the risk that the feature vector computed for the first person 102 based on its visual appearance in the image frame 100b may "leak" into the feature vector computed for the second person 104 and vice versa, the overlap between the persons 102, 104 (and also the person 112) is checked. In the image frame 100b, the persons do not overlap, i.e., the bounding boxes 106, 108 of the persons 102, 104 do not overlap. Thus, the risk of "leakage" is determined to be low, and thus, the persons 102, 104 are determined to be prioritized candidates for feature extraction (shown by the Figure 1 bold rectangles 108, 106). The third person 112 is determined not to be a prioritized candidate for feature extraction because, as discussed herein, the third person is far enough away from any of the other objects 102, 104.

[0045] In the third image frame 100c, both the first person 102 and the second person 104 still at least partially overlap with the ROI 110. However, unlike the second image frame 100b, the persons 102, 104 now overlap with each other (i.e., the bounding boxes 106, 108 of the persons 102 and 104 overlap). As will be discussed below in connection with Figure 5 objects that overlap beyond a threshold range are not considered to be prioritized candidates for feature extraction because the risk of leakage of the visual appearance of the overlapping objects is considered to be too high when determining the feature vectors of the overlapping objects. In Figure 1In the example of , in the third image frame 100c, the first person 102 and the second person 104 are considered to have an overly high overlap range, and thus, the persons 102, 104 are not determined to be candidates after prioritization for feature extraction. Since the third person 112 is far enough away from any of the other objects 102, 104, the third person 112 is determined not to be a candidate after prioritization for feature extraction, as discussed herein.

[0046] In the fourth image frame 100d, both the first ROI 110 and the second ROI 116 are associated with two objects in the scene, which means that these two objects at least partially overlap with each of the ROIs 110, 116. Both the first person 102 and the second person 104 still at least partially overlap with the first ROI 110. In addition, both the first person 102 and the third person 112 at least partially overlap with the second ROI 116. For the first ROI 110, it is determined that the associated objects 102, 104 do not overlap with each other (or do not overlap with the object 112), and thus, the objects 102, 104 are added to the list of candidates after prioritization. For the second ROI 116, it is determined that the associated objects 102, 112 do not overlap with each other (or do not overlap with the object 104), and thus, the objects 102, 112 are added to the list of candidates after prioritization. It should be noted that in some embodiments, an object cannot be added again if it has already been added to the list of candidates after prioritization. In other embodiments, when at least one object in the list of candidates after prioritization is selected for determining the feature vector, any duplicates will be removed, as further described below.

[0047] In the fifth image frame 100e, all the objects are far enough away such that it is determined that there are no candidates after prioritization for feature extraction. In other words, none of the ROIs 110, 116 are associated with at least two objects, as described herein.

[0048] Figures 2 to 4 An example is shown of how to determine the ROI 110 based on the motion region 202 in the image frame. In Figures 2 to 4 , a cropped portion of the image frame is shown. This cropped portion includes the object 104, along with the motion region 202 from the image frame. In Figures 2 to 4 's example, the motion region 202 is shown as a pixel mask 202. However, other ways of representing the motion region in the image frame can be used (e.g., polygons, rectangles, etc.). The motion region 202 can be determined by identifying changes over time in the scene (e.g., using frame differences, background subtraction, temporal filters, etc.), typically by comparing consecutive or adjacent frames.

[0049] In Figure 2 are shown four examples of how to determine the ROI 110 based on the motion region 202. The upper left example represents the case where the ROI 110 has the shape of the pixel mask 110. In Figure 2 's upper left example, the ROI 110 has the same form and extent as the motion region 202. Figure 2 The lower left example in Figure 2 represents the case where the ROI 110 has a circular form and completely overlaps with the motion region 202. Figure 2 The upper right example in represents the case where the ROI 110 has an elliptical form and completely overlaps with the motion region 202.

[0050] The lower right example in Figure 3 Figure 3 Figure 3 represents the case where the ROI 110 has a rectangular form and completely overlaps with the motion region 202. Figure 2 are shown four other examples of how to determine the ROI 110 based on the motion region 202. In some embodiments, determining the ROI for object feature extraction includes expanding the motion region by a predetermined extent in each direction. In Figure 2 's example, the motion region 202 expands uniformly in each direction. However, in other examples, other strategies can be used to expand the motion region 202, for example, expanding to a greater extent in the horizontal direction compared to the vertical direction. Generally, the motion in the scene has a greater speed in the horizontal direction compared to the speed in the vertical direction, which can be advantageously captured by such embodiments. In Figure 3 each of the four examples corresponds to the example in

[0051] Figure 4 However, compared to the ROI 110 in

[0052] · Performing motion detection via optical flow: This method tracks the way different parts of the image move between frames and calculates the speed of the motion.

[0053] · Spatiotemporal motion detection: This method separates foreground elements from background elements, looks for gradients that change over time in the "image volume", and uses the motion information and direction to identify blobs to calculate the speed.

[0054] · Solutions via neural networks: The method includes training a neural network to identify and track motion in an image and then using the network to detect and calculate the speed of the motion region.

[0055] · Output from a tracker from a previous frame: The method includes associating the tracked object from a previous frame with the motion region in the current frame to assume that the speed of the object is equal to the speed of the motion region.

[0056] In Figure 4 , the speed is represented by arrow 402. Thus, Figure 4 In the example of , the speed is horizontal and has a direction to the right. Thus, the ROI 110 can be expanded based on the speed. For example, it can be expanded to a greater extent in the direction corresponding to the direction of speed 402 than in the direction not corresponding to the direction of speed 402. In another example, the ROI can be expanded in the direction of the speed and in the direction opposite to the direction of the speed so as to prioritize the object 104 both before and after occlusion. In yet another embodiment, the ROI can be expanded based on both the direction and magnitude of the speed such that a larger magnitude causes a greater expansion than a smaller magnitude.

[0057] In Figure 4 , each of the four examples corresponds to the example of Figure 3 . However, compared to the ROI 110 in Figure 3 being uniformly expanded in each direction, the ROI 110 in Figure 4 is all expanded according to the speed. Thus, the size of the ROI expansion can be a function of the speed of the object, i.e., how far the object is expected to move within a particular time frame. Additionally, the analysis speed of the system can be considered, i.e., how quickly the system can perform the necessary analysis. To ensure that the system has enough time to analyze the object before the object is occluded (or, leaves the scene captured by the video sequence including the image frames), a margin can be added to the size of the ROI expansion. This margin allows the system more time to perform the necessary analysis before the object is lost or occluded.

[0058] Figure 5Three overlapping objects are schematically represented by bounding boxes 106, 108, 114 by way of example. As discussed herein, visual appearance leakage may occur when overlapping objects are not properly separated during the feature extraction process. For example, if one object partially covers another object, the feature extraction process may inadvertently include information from both objects in the overlapping area. As a result, the feature vector of the partially occluded object may contain the visual appearance characteristics of the occluding object. In some cases, an overlap below a specific threshold may still produce a sufficiently accurate feature vector. This is because despite the overlap, most of the image data used to determine the feature vector of the object can still be derived from the correct object. In Figure 5 , the bounding box 106 of the first object overlaps the bounding box 108 of the second object by a specific extent 506, which means that the area 504 does not overlap with any other object. The bounding box 108 of the second object overlaps the bounding boxes 106, 114 of the two objects, and the overlapping extents are represented by Figure 5 the areas 506 and 510 in, which means that the area 502 does not overlap with any other object. Thus, the bounding box 114 of the third object overlaps the bounding box 108 of the second object by a specific extent 510, which means that the area 508 does not overlap with any other object.

[0059] According to an embodiment, when it is determined that an object does not overlap with any other object by more than a threshold amount, the object can be considered a candidate object in a prioritized order for feature extraction. In some embodiments, the threshold amount is 0. In some embodiments, the threshold amount (i.e., the allowed overlap extent) is a predetermined percentage of the spatial extent of one of the overlapping objects. Examples of the predetermined percentage can be 1%, 5%, 7%, 10%, 25%, etc. For example, for the case of the overlap between the first object and the second object, the overlapping extent 506 can be considered too high compared to the total area of the bounding box 106 of the first object, which means that the first object is not included among the candidate objects in the prioritized order described herein. For the case of the overlap between the second object and the third object, the overlapping extent 510 (i.e., the overlapping area 510 compared to the total area of the bounding box 114 of the third object) can be considered not to exceed the threshold, which means that the third object is included among the candidate objects in the prioritized order described herein. For the second object that overlaps with two objects, in some embodiments, the total overlapping area (506 + 510) is compared with the entire area of the bounding box 108 to determine whether the second object should be included among the candidate objects in the prioritized order. In other embodiments, the second object is not included among the candidate objects in the prioritized order when one of the areas 506, 510 is considered higher than the threshold compared to the total area of the bounding box 108 of the second object.

[0060] Figure 6 A flowchart of a method 600 for prioritizing feature extraction for object re-identification in an object tracking application is shown by way of example. The method 600 will now be described in conjunction with a system 800 that implements the method 600 by way of example. Figure 8 The method 600 includes a step of receiving, by a receiving component 808, an image frame 100a that depicts a scene including a plurality of objects, as depicted in S602. The image frame can be, for example, a frame from a video sequence 802 that depicts a scene captured by a camera.

[0061] The method 600 further includes a step of receiving, by the receiving component 808, object detection data 805 (the object detection data 805 includes, for each object of the plurality of objects, positioning data for indicating the position and spatial extent of the object in the image frame), as depicted in S604. The object detection data can be received from an object detector component 804. The object detector component can implement any suitable conventional algorithm for object detection (e.g., Haar cascade classifier, histogram of oriented gradients (HOG), region-based convolutional neural network (R-CNN), etc.). The object detector component can include an artificial intelligence (AI) or machine learning (ML) algorithm that is trained to detect objects of interest in an image. AI / ML is a suitable technique for detecting objects in an image and can be relatively easily trained using a large labeled / annotated image dataset that includes the types of objects of interest. It should be noted that there are several object detection algorithms that do not rely on neural networks. These methods can be based on classical computer vision such as scale-invariant feature transform (SIFT), template matching, etc.

[0062] The method 600 further includes a step of receiving, by the receiving component 808, motion data 807 that indicates one or more motion regions in the image frame, each motion region corresponding to a region in the scene where motion has been detected, as depicted in S606. In some embodiments, the motion data further includes an indication of the speed of the motion detected in the corresponding region of the scene for each motion region in the image frame. The motion data 807 can be received from a motion detector component 806. The motion detector can generally implement any suitable algorithm (as described herein) for detecting motion in a scene (captured by a video sequence) by comparing consecutive or neighboring frames.

[0063] The method 600 further includes, for each motion region, determining, by a region of interest (ROI) determination component 810, an ROI in the image frame for object feature extraction, the ROI for object feature extraction overlapping the motion region, as depicted in S608. The ROI can be determined by the ROI determination component 810 (as described above in connection with Figures 2 to 4S610 (further described below) is further extended (e.g., by extending the ROI obtained from the motion area).

[0064] Method 600 further includes, for each ROI used for object feature extraction, using the object detection data to determine a list of objects that at least partially overlap with the ROI used for object feature extraction, and associating, by ROI association component 812, the ROI used for object feature extraction with the list of objects S612.

[0065] Method 600 further includes determining S614, by candidate object determination component 814, a prioritized list of candidate objects for feature extraction, as described further below in conjunction with Figure 7 to be further described.

[0066] Method 600 further includes selecting S616, by candidate object selection component 816, at least one object from the prioritized list of candidate objects. The selection S616 may include removing duplicates from the prioritized list of candidate objects. The selection S616 may limit the number of objects used to determine the feature vector to a specific number (N). Such a limitation may further include comparing each prioritized candidate object with one or more selection criteria and selecting (at most) N objects that meet at least one of the one or more selection criteria. The selection criteria may be configured based on the application or requirements of the application. For example, the selection criteria may include selecting objects of a specific object type and / or of a specific size and / or objects associated with the speed of motion, etc. In an embodiment, an object may be associated with the speed of motion if the object overlaps with the motion area associated with the speed of motion by a high range (complete overlap, 95% overlap, etc.). In some embodiments, such a limitation may include sorting each prioritized candidate object according to one or more sorting criteria and selecting the N objects with the highest sorting. The sorting criteria may include, for example, sorting based on object size, sorting based on the confidence score from the object detector, sorting based on the speed of the object, etc.

[0067] Method 600 further includes, for each selected object, determining S618, by feature vector determination component 818, the feature vector of the object based on the image data of the image frame according to the positioning data of the selected object. Any suitable algorithm for determining the feature vector from the image data may be used. Example algorithms include:

[0068] · Scale-invariant feature transform (SIFT): SIFT is a feature extraction algorithm that detects and describes local features in an image. It can be considered robust to changes in scale and rotation.

[0069] · Speeded Up Robust Features (SURF): SURF is a feature extraction algorithm similar to SIFT, but SURF can be considered faster and more robust to changes in illumination and viewpoint.

[0070] · Histogram of Oriented Gradients (HOG): HOG is a feature extraction algorithm that counts the occurrences of gradient orientations in local parts of an image.

[0071] · Convolutional Neural Network (CNN): CNN is a deep learning algorithm that can be used to extract features from images. CNN learns features from images by using multiple layers of convolution and pooling.

[0072] · Local Binary Pattern (LBP): LBP is a texture descriptor that extracts local features from an image by comparing each pixel with its surrounding neighbors.

[0073] · Feature extraction using object or sub - object classification capabilities: This technique includes detecting, for example, a person's clothing, accessories, etc. (color of a sweater, wearing a handbag, wearing glasses, wearing a hat, etc.), the type of vehicle (red pickup truck, black SUV, etc.), or similar features of an object.

[0074] Subsequently, the feature vector can be used by the re - identification component 820, for example, to identify a selected object in an image frame as the same object previously detected in one or more previous frames of the same video sequence 802 or detected in another video sequence. In another embodiment, the feature vector can be transmitted to a tracker, which can use the feature vector to create a more robust tracking or recreate a tracking that has been lost due to occlusion. An example of using the feature vector for such purposes is described in "Simple Online and Realtime Tracking with a Deep Association Metric" (published by Wojke et al. at https: / / arxiv.org / abs / 1703.07402 at the time of filing this application).

[0075] In some embodiments, the determined feature vectors may be associated with at least a portion of the localization data of the selected corresponding objects. For example, the position of the selected object may be associated with the feature vectors. Such information may be used by the re-identification component 820 to determine the relevance of the feature vectors of the tracks that have been lost in the video sequence. For example, certain object movements may be considered unreasonable. Very large movements (such as crossing the image in only one frame) may instead appear to be two different objects that are similar. Trackers including the re-identification component 820 typically have a limited area based on the most recent detection of the object (where the detection expects the object to reappear), based on the speed of the object, and a certain error margin. If the object is lost (e.g., when a person walks behind a bush), the radius is typically increased to achieve re-association. Association with a given track based on detections outside of this radius is typically "expensive" (requiring a very secure match) or prohibited.

[0076] In some embodiments, the determined feature vectors may be associated with a timestamp indicating the capture time of the image frame. Such information may be used by the re-identification component 820 to determine the relevance of the feature vectors of the tracks that have been lost in the video sequence. For example, the timestamp may be used to determine an acceptable distance between the position of the selected object and the last instance of the lost track. Additionally, when performing long-term tracking, the timestamp may be used to determine the relevance of the feature vectors.

[0077] Figure 7 is shown by way of example Figure 6 more details of the step of determining S614 the list of prioritized candidate objects for feature extraction in. For all ROIs determined in step S608 (and potentially extended in S610) as in Figure 6 the method of, step S614 is performed. Thus, when all ROIs have been examined in S702, the step of determining S614 the list of prioritized candidate objects for feature extraction ends in S712. For each ROI, it is checked in S704 whether the ROI is associated with two or more objects (i.e., the number of objects that at least partially overlap the ROI as determined in the association step S612 above). If so, then each of the two or more objects associated with the ROI is analyzed in S708 to determine whether the object should be added in S710 to the list of prioritized candidate objects. Thus, the method includes an analysis step S708 of checking whether the object overlaps any other object in the plurality of objects by more than a threshold amount. The plurality of objects detected in the image frame (i.e., as in Figure 6When there are no other objects among the objects identified from the object detection data received in step S604 that overlap with the currently analyzed object by more than a threshold amount, the currently analyzed object is added in S710 to the list of candidate objects sorted by priority. The analysis step S708 (and the addition step S710, if applicable) is performed until all the objects (such as the objects determined in step S706) have been analyzed. When it is determined in S706 that all the objects associated with the ROI have been analyzed, the next ROI is checked and so on.

[0078] Figure 6 and Figure 7 The methods shown in and any other methods or functions described herein can be stored as instructions on a non - transitory computer - readable storage medium such that when these instructions are executed on a device or system having processing capabilities, these methods are implemented. Such a device or system (e.g., as shown in Figure 8 can include one or more processors. In an example, system 800 can be implemented in a single device such as a camera. In other examples, some or all of the different components (modules, units, etc.) 804 to 820 can be implemented in a server or in the cloud. In some examples, the camera implements Figure 8 all components of except for the re - identification component 820 that can be implemented by, for example, a server. In some examples, the camera implements Figure 8 all components of except for the feature - vector determination component 818 and the re - identification component 820 that can be implemented by, for example, a server. In other examples, Figure 8A subset of the components (e.g., the image capture and motion detector 806 that captures the video sequence 802) is implemented on the camera, while other components are implemented in the server. Generally, one or more devices (such as cameras, servers, etc.) that implement components 804 to 820 may each include circuitry configured to implement components 804 to 820 and more specifically to implement their functions. Thus, the features and methods described herein may be advantageously implemented in one or more computer programs that are executable on a programmable system, which may include at least one programmable processor coupled to receive data and instructions from and to transmit data and instructions to a data storage system, at least one input device such as a camera for capturing image frames / video sequences, and at least one output device such as a display for displaying images, potentially implemented in an object tracking application including the tracked objects described herein. By way of example, suitable processors for executing the instructions include both general purpose and special purpose microprocessors, as well as the sole processor or one of multiple processors or cores of any type of computer. The processor may be supplemented by, or incorporated in, an ASIC (Application Specific Integrated Circuit).

[0079] The embodiments above should be understood as illustrative examples of the present invention. Other embodiments of the present invention are contemplated. For example, motion data may be received from the object tracker component and, based on the movement of the object, tracking may be performed between image frames. It should be understood that any feature described with respect to any one embodiment may be used alone or may be combined with other features described, and may also be combined with one or more features of any other embodiment or any combination of any other embodiment. Additionally, equivalents and modifications not described above may also be employed without departing from the scope of the present invention as defined in the appended claims.

Claims

1. A computer-implemented method for prioritizing objects for feature extraction for object re-identification purposes in an object tracking application, the method comprising the steps of: Receiving an image frame depicting a scene including a plurality of objects; Receiving object detection data, the object detection data including positioning data for each of the plurality of objects indicating the position and spatial extent of the object in the image frame; Receiving motion data indicating one or more motion regions in the image frame, each motion region corresponding to a region in the scene where motion has been detected, wherein the motion data further includes an indication of the speed of the motion detected in the corresponding region of the scene for each motion region in the image frame; For each motion region, determining a region of interest (ROI) for object feature extraction in the image frame and expanding the ROI based on the speed, the ROI for object feature extraction overlapping the motion region; For each ROI for object feature extraction, using the object detection data to determine a list of objects that at least partially overlap the ROI for object feature extraction and associating the ROI for object feature extraction with the list of objects; For each ROI for object feature extraction associated with two or more objects, determining a list of prioritized candidate objects for feature extraction by: for each of the two or more objects, adding the object to the list of prioritized candidate objects for feature extraction when it is determined that the object does not overlap any of the other objects in the plurality of objects by more than a threshold amount; Selecting at least one object from the list of prioritized candidate objects and, for each selected object, determining a feature vector of the object based on image data of the image frame according to the positioning data of the selected object; 2. The method according to claim 1, wherein, The step of determining the ROI for object feature extraction includes expanding the motion region by a predetermined extent in each direction.

3. The method according to claim 1, wherein The ROI for object feature extraction is determined by expanding the motion region to a greater extent in the direction corresponding to the direction of the speed than in the direction not corresponding to the direction of the speed.

4. The method according to claim 1, wherein, The shape of the ROI for object feature extraction is one of a pixel mask, a circle, an ellipse, and a rectangle.

5. The method according to claim 1, wherein, The threshold amount is 0.

6. The method according to claim 1, wherein, The threshold amount is a predetermined percentage of the spatial extent of one of the overlapping objects.

7. The method according to claim 1, wherein At most N objects from the list of prioritized candidate objects are selected, where N is a predetermined number.

8. The method according to claim 7, wherein, When the list of prioritized candidate objects consists of more than N objects, the selecting step includes comparing each prioritized candidate object with one or more selection criteria and selecting the N objects that satisfy at least one of the one or more selection criteria.

9. The method according to claim 7, wherein When the list of prioritized candidate objects consists of more than N objects, the selection step includes sorting each of the prioritized candidate objects according to one or more sorting criteria and selecting the N objects with the highest sorting.

10. The method according to claim 1, further comprising the following steps: For each selected object, associating at least a portion of the determined feature vector with the localization data of the selected object.

11. The method according to claim 1, further comprising the following steps: For each selected object, associating the determined feature vector with a timestamp indicating the capture time of the image frame.

12. A system for prioritizing feature extraction for object re-identification in an object tracking application, comprising: One or more processors; And One or more non-transitory computer-readable media storing computer-executable instructions, which when executed by the one or more processors cause the system to perform operations including the following items: Receiving an image frame depicting a scene including a plurality of objects; Receiving object detection data, the object detection data including localization data indicating the position and spatial extent of each of the plurality of objects in the image frame; Receiving motion data indicating one or more motion regions in the image frame, each motion region corresponding to a region in the scene where motion has been detected, wherein the motion data further includes an indication of the speed of the motion detected in the corresponding region of the scene for each motion region in the image frame; For each motion region, determining a region of interest ROI in the image frame for object feature extraction and expanding the ROI based on the speed, the ROI for object feature extraction overlapping with the motion region; For each ROI for object feature extraction, using the object detection data to determine a list of objects that at least partially overlap with the ROI for object feature extraction, and associating the ROI for object feature extraction with the list of objects; For each ROI for object feature extraction associated with two or more objects, determining a list of prioritized candidate objects for feature extraction by: for each of the two or more objects, adding the object to the list of prioritized candidate objects for feature extraction when it is determined that the object does not overlap with any of the other objects in the plurality of objects by more than a threshold amount; Selecting at least one object from the list of prioritized candidate objects, and for each selected object, determining a feature vector of the object based on the image data of the image frame according to the localization data of the selected object.

13. One or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein, The instructions, when executed, cause the one or more processors to perform operations including the following items: Receiving an image frame depicting a scene including a plurality of objects; Receive object detection data, the object detection data including localization data for each of the plurality of objects indicating the position and spatial extent of the object in the image frame; Receive motion data indicating one or more motion regions in the image frame, each motion region corresponding to a region in the scene in which motion has been detected, wherein the motion data further includes an indication of the speed of the detected motion in the corresponding region of the scene for each motion region in the image frame; For each motion region, determine a region of interest ROI for object feature extraction in the image frame and expand the ROI based on the speed, the ROI for object feature extraction overlapping with the motion region; For each ROI for object feature extraction, use the object detection data to determine a list of objects that at least partially overlap with the ROI for object feature extraction, and associate the ROI for object feature extraction with the list of objects; For each ROI for object feature extraction associated with two or more objects, determine a list of candidate objects in priority order for feature extraction by: for each of the two or more objects, adding the object to the list of candidate objects in priority order for feature extraction when it is determined that the object does not overlap with any of the other objects in the plurality of objects by more than a threshold amount; Select at least one object from the list of candidate objects in priority order, and for each selected object, determine a feature vector of the object based on the image data of the image frame according to the localization data of the selected object.

Citation Information

Patent Citations

  • System and method for boosting object detection performance in videos

    CN103914702A

  • Object detection and tracking method and device

    CN110458861A