Frame-based video segmentation

By extracting the mask features of the object in each frame of the video data and storing it in the feature library, the problem of inaccurate object tracking between video frames in the prior art is solved, and fast and accurate video instance segmentation and tracking are realized, which is suitable for online application scenarios.

CN120051798APending Publication Date: 2025-05-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073817.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-25
Filing Date
2023-09-12
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing video instance segmentation technology is difficult to accurately track objects between video frames, especially in the presence of occlusion, lighting changes, rescaling and object deformation, and the calculation time is long, making it difficult to be applicable to online application scenarios.

Method used

By receiving each frame of video data, the initial mask of the extracted object is combined with the extracted feature, a mask feature representation is generated, and stored in the feature library. The object is then tracked in a continuous frame, and by comparing the mask feature representation with the stored representation in the feature library, the feature library is updated and the frame is adjusted to achieve continuous tracking of the object.

Benefits of technology

Accurate object tracking between video frames is realized, the accuracy of personalized objects in the video is improved, the calculation time is reduced, and the technology can process video streams in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051798A_ABST
    Figure CN120051798A_ABST
Patent Text Reader

Abstract

Methods and systems for frame-based image segmentation are provided. For example, a method for feature object tracking between frames of video data is provided. The method comprises: receiving a first frame of video data; extracting a mask feature of each of the one or more objects of the first frame; the first frame is adjusted by applying each initial mask and a corresponding identification to a respective object of the first frame, and the adjusted first frame is output. The method also includes tracking the one or more objects in one or more consecutive frames. The tracking includes: extracting a mask feature of each of one or more objects in the consecutive frames; the consecutive frame is adjusted by applying each initial mask and a corresponding identification of the consecutive frame to a respective object of the one or more objects of the consecutive frame, and the adjusted consecutive frame is output.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims the benefit of U.S. Patent Application No. 18 / 049,473, filed on October 25, 2022, entitled "FRAME - BASED VISIO SEGMENTATION", which is hereby incorporated by reference in its entirety. Technical Field

[0002] This disclosure relates to frame - based video instance segmentation. In particular, this disclosure relates to methods and systems for frame - based video instance segmentation, including the segmentation of objects (or object instances, as used interchangeably herein) in a video and the tracking of objects between frames. Background Art

[0003] In recent years, object segmentation in images and videos has received increasing attention in both academia and industry. Given the popularity of video applications and the need for precise and personalized processing of objects within videos, video instance segmentation (VIS) is gradually becoming a key research topic.

[0004] Different from image segmentation (which only segments), the goal of video instance segmentation is to simultaneously segment and track object instances in a video. Specifically, given an input video, a VIS model needs to segment objects from each individual video frame (pixel - wise classification for multiple things, such as people, pets, etc.) and associate the segmented objects between frames (instance - level tracking function). Therefore, VIS is more complex and challenging than most other computer vision tasks. Summary of the Invention

[0005] Aspects for improving the prior art and systems for video instance segmentation are presented below. The techniques of this disclosure at least improve the segmentation of objects in a video and the tracking of objects between frames, thereby allowing masks (generated from object segmentation) to be consistently applied to objects frame - by - frame. This provides improved accuracy for the personalized processing of objects within a video.

[0006] According to a first aspect, an embodiment provides a method for object tracking between frames of video data. The method includes receiving a first frame of video data. The first frame includes one or more objects. For each of the one or more objects in the first frame, a mask feature of the object is extracted by combining an initial mask of the object with one or more extracted features of the object. A representation of each mask feature is generated, the representation indicating the one or more extracted features of the mask feature. Each representation and an associated identifier are stored in a feature library. The first frame is adjusted by applying each initial mask and the corresponding identifier to the corresponding object among the one or more objects in the first frame, and the adjusted first frame is output. The method further includes tracking the one or more objects in one or more consecutive frames. For each consecutive frame, tracking includes: extracting a mask feature of each of the one or more objects in the consecutive frame by combining an initial mask of the object with one or more extracted features of the object. Tracking further includes generating a representation of each mask feature of the consecutive frame, the representation indicating the one or more extracted features of the mask feature. Additionally, it is determined whether the representation of each mask feature of the consecutive frame corresponds to a representation stored in the feature library. In response to determining that the representation of the mask feature of the consecutive frame corresponds to a representation stored in the feature library, the corresponding initial mask of the consecutive frame is associated with the identifier of the corresponding stored representation; and the corresponding stored representation is updated in the feature library with the corresponding representation of the mask feature of the consecutive frame. Tracking further includes adjusting the consecutive frame by applying each initial mask and the corresponding identifier of the consecutive frame to the corresponding object among the one or more objects in the consecutive frame, and outputting the adjusted consecutive frame.

[0007] The one or more objects in a frame may include multiple objects.

[0008] The method may further include: in response to determining that the representation of the mask feature of the consecutive frame does not correspond to a representation stored in the feature library, storing the representation of the mask feature of the consecutive frame as a new entry in the feature library.

[0009] Determining whether the representation of the mask feature of the consecutive frame corresponds to a representation stored in the feature library may include: comparing the representation of the mask feature of the consecutive frame with each stored representation; determining a similarity metric for each stored representation, the similarity metric indicating the similarity between the representation of the mask feature of the consecutive frame and the stored representation; and based on the similarity metric, determining whether the representation of the mask feature of the consecutive frame corresponds to a representation stored in the feature library.

[0010] Extracting a mask feature of each of the one or more objects in the consecutive frame may include extracting a mask feature of each of the multiple objects in the consecutive frame.

[0011] The representation of the mask feature may be a vector.

[0012] Determining whether the representation of the mask features of consecutive frames corresponds to the representation stored in the feature library may include calculating the cosine similarity between the representation of the mask features of the consecutive frames and the stored representation.

[0013] The correspondence between the representation of the mask features of consecutive frames and the representation stored in the feature library may indicate the similarity between the features of the mask features of consecutive frames and the features of the mask features stored in the feature library.

[0014] The features of the mask features may include one or more of the color of the object, the edge of the object, or the corner of the object.

[0015] Extracting the mask features of each of one or more objects in a first frame may include: inputting the first frame into a convolutional neural network to extract semantic data indicating one or more features of each of the one or more objects; segmenting the first frame to generate an initial mask for each of the one or more objects; and combining the initial mask with the semantic data to generate the mask features of each of the one or more objects. One or more features of the mask features may include semantic data indicating one or more features of the corresponding object. One or more features of the corresponding object may include one or more of the color of the object, the edge of the object, or the corner of the object. Each mask feature may include one or more extracted features. These extracted features may be considered low-level semantic features herein. Intermediate semantic features may also be used. For example, if the use case involves a person as an example object, the intermediate semantic features may include hands, arms, and heads. Both low-level features and intermediate features can be used to distinguish objects within the same category (e.g., distinguish two different people or two different vehicles, etc.). In certain use cases, high-level semantic features such as categories like person, vehicle, pet, building may also be used - such features may be suitable for use cases that require discrimination between categories (e.g., discrimination between a person and a car) rather than discrimination within a category (e.g., discrimination between different people). Extracting semantic data indicating one or more features of each of the one or more objects may include: applying one or more functions to the data of the first frame, the one or more functions including a convolutional function, an activation function, a batch normalization function, and a dropout function. Segmenting the first frame to generate an initial mask for each of the one or more objects may include: performing a pixel-by-pixel comparison of the pixels of the first frame, the pixel-by-pixel comparison identifying groups of similar pixels, each group representing one of the one or more objects in the frame. According to certain embodiments, segmenting includes reducing the resolution of the first frame to produce a reduced frame (a frame with reduced size) and performing a pixel-by-pixel comparison of the pixels of the reduced frame. This can be achieved by arranging the resolution-reduced frame along a single dimension and then performing a pixel-by-pixel comparison.

[0016] According to this aspect, embodiments also provide a system for object tracking between frames of video data, the system including one or more processors configured to: receive a first frame of video data, the first frame including one or more objects; extract mask features for each of the one or more objects in the first frame by combining an initial mask of each object with one or more extracted features of the object; generate a representation of each mask feature, the representation indicating the one or more extracted features of the mask feature and store each representation in a feature library; adjust the first frame by applying each initial mask and corresponding identification to a respective one of the one or more objects in the first frame and output the adjusted first frame; and track the one or more objects in one or more consecutive frames, for each consecutive frame, the tracking including: extracting mask features for each of the one or more objects in the consecutive frame by combining an initial mask of each object with one or more extracted features of the object; generating a representation of each mask feature of the consecutive frame, the representation indicating the one or more extracted features of the mask feature; determining whether the representation of each mask feature of the consecutive frame corresponds to a representation stored in the feature library; in response to determining that the representation of the mask feature of the consecutive frame corresponds to a representation stored in the feature library, associating the corresponding initial mask of the consecutive frame with the identification of the corresponding stored representation; and updating the corresponding stored representation in the feature library with the corresponding representation of the mask feature of the consecutive frame; adjusting the consecutive frame by applying each initial mask and the corresponding identification of the consecutive frame to a respective one of the one or more objects in the consecutive frame and output the adjusted consecutive frame.

[0017] The one or more objects of the frame may include multiple objects.

[0018] The one or more processors of the system may also be configured to, in response to determining that the representation of the mask feature of the consecutive frame does not correspond to a representation stored in the feature library, store the representation of the mask feature of the consecutive frame as a new entry in the feature library.

[0019] Determining whether the representation of the mask feature of the consecutive frame corresponds to a representation stored in the feature library may include: comparing the representation of the mask feature of the consecutive frame with each stored representation; determining a similarity metric for each stored representation, the similarity metric indicating the similarity between the representation of the mask feature of the consecutive frame and the stored representation; and determining based on the similarity metric whether the representation of the mask feature of the consecutive frame corresponds to a representation stored in the feature library.

[0020] Extracting mask features for each of the one or more objects in the consecutive frame may include extracting mask features for each of the multiple objects in the consecutive frame.

[0021] The representation of the mask feature may be a vector.

[0022] Determining whether the representation of the mask features of consecutive frames corresponds to the representation stored in the feature library may include calculating the cosine similarity between the representation of the mask features of consecutive frames and the stored representation.

[0023] The correspondence between the representation of the mask features of consecutive frames and the representation stored in the feature library may indicate the similarity between the features of the mask features of consecutive frames and the features of the mask features of the representation stored in the feature library.

[0024] The features of the mask features may include one or more of the color of the object, the edges of the object, or the corners of the object.

[0025] Extracting the mask features of each object in one or more objects of the first frame may include: inputting the first frame into a convolutional neural network to extract semantic data indicating one or more features of each object in the one or more objects; segmenting the first frame to generate an initial mask for each object in the one or more objects; and combining the initial mask with the semantic data to generate the mask features of each object in the one or more objects. One or more features of the mask features may include semantic data indicating one or more features of the corresponding object. One or more features of the corresponding object may include one or more of the color of the object, the edges of the object, or the corners of the object. Extracting semantic data indicating one or more features of each object in the one or more objects may include: applying one or more functions to the data of the first frame, the one or more functions including a convolutional function, an activation function, a batch normalization function, and a dropout function. Segmenting the first frame to generate an initial mask for each object in the one or more objects may include: performing a pixel-by-pixel comparison of the pixels of the first frame, the pixel-by-pixel comparison identifying groups of similar pixels, each group representing one object in one or more objects of the frame. According to certain embodiments, segmenting includes reducing the resolution of the first frame to produce a reduced frame (a frame with reduced size) and performing a pixel-by-pixel comparison of the pixels of the reduced frame. This may be achieved by arranging the resolution-reduced frame along a single dimension and then performing a pixel-by-pixel comparison.

[0026] According to a further aspect, an embodiment provides a computer-readable medium having instructions thereon that are configured to cause one or more processors to perform the method according to any one of the first aspect, the second aspect, or the third aspect.

[0027] In cases where functional modules or units are mentioned in the device embodiments for performing various functions or steps of the described methods, it should be understood that these modules or units can be implemented in hardware, software, or a combination of both. When implemented in hardware, these modules can be implemented as one or more hardware modules, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). When implemented in software, these modules can be implemented as one or more computer programs executed on one or more processors. Description of the Drawings

[0028] The drawings are presented to assist in describing its various aspects.

[0029] Figure 1 is a schematic diagram of an exemplary existing video instance segmentation system;

[0030] Figure 2 is an example of a video instance segmentation output frame;

[0031] Figure 3 shows multiple frames that illustrate examples of object changes in a video stream;

[0032] Figure 4 is an exemplary failure case of existing video instance segmentation;

[0033] Figure 5 is a schematic diagram of an exemplary video instance segmentation system according to aspects of the present disclosure;

[0034] Figure 6 is a schematic diagram of an exemplary backbone network module for implementing aspects of the present disclosure;

[0035] Figure 7 is a schematic diagram of an exemplary instance segmentation module for implementing aspects of the present disclosure;

[0036] Figure 8 shows a comparison of output frames generated without video instance segmentation using the Mask2Former-VIS process and using an exemplary process of the present disclosure;

[0037] Figure 9 is a flowchart illustrating a video instance segmentation process according to aspects of the present disclosure; and

[0038] Figure 10 is a flowchart illustrating the steps of a video instance segmentation process according to aspects of the present disclosure. Detailed Description of the Invention

[0039] This disclosure relates to improvements in techniques such as video instance segmentation (VIS). VIS is the process of segmenting objects (also referred to herein as object instances, or simply instances) in a video stream and then tracking those objects within the video stream. This requires segmenting the objects from each video frame. This can be done by classifying the pixels of the objects. For example, a set of pixels can be classified as an object such as a person, a pet, a vehicle, etc., which can be done by various techniques further discussed herein. When objects are segmented from each frame, associations need to be determined between the frames in order to track the same object between the frames. This allows, for example, a mask (such as a color mask) to be consistently applied to the same object between the frames (e.g., the same object can be highlighted in blue in all frames in which the object appears).

[0040] Accordingly, VIS is more complex and challenging than many other computer vision tasks. Various solutions have been proposed for VIS, such as DETR, VisTR, IFC, SeqFormer, Mask2Former-VIS, etc. These solutions treat video clips as 3D spatio-temporal volumes and directly predict 3D masks for each instance. Mask2Former-VIS 100 is an example of the prior art and is shown in Figure 1 where the input video passes through a backbone network 101, a pixel decoder 102, and a transformer decoder 103 to perform instance segmentation to generate a mask for the video and output an adjusted video stream, where the mask covers the instances in the video.

[0041] Various technical challenges have been recognized.

[0042] Challenge 1: Accurate segmentation of each individual frame. This is a problem of per-pixel classification. The process requires classifying each pixel on the frame into one of the potential classes (as well as the background). Referring to Figure 2 , this particular example is for a person (in a real-world application scenario, more classes may be included). Thus, the segmentation process requires classifying each pixel as either a "person" or a "background" class. In Figure 2 , a color mask is used to highlight all person pixels, while the brightness of the background pixels is dimmed. For the sake of illustration, in Figure 2Different colors among people are shown using different hash / line patterns, with each different pattern representing a different color. For example, the people on the right - hand side of the basketball court are highlighted with a straight - line pattern, which can exemplify, for instance, a red highlight. A diagonal - line pattern can represent a green highlight, a horizontal straight - line pattern can represent a blue highlight, and a dotted pattern can represent a yellow highlight. It should be understood that the mapping between colors and patterns is for illustrative purposes only, and different colors can be used. Alternatively, hash / line / dot patterns can actually be used instead of colors to visually highlight these people, in which case the output will look visually equivalent to Figure 2 corresponding. In the following text, if "color" is mentioned as being used to highlight an object in an embodiment (which is exemplified by line patterns, dot patterns, etc. in the figure), it should be understood that the line / dot patterns are used in the figure to represent specific colors.

[0043] Challenge 2: Precise instance separation. The process also needs to distinguish different individuals belonging to the same class. As Figure 2 shown, there are four people playing basketball in this image. Among all the pixels predicted to be "people", the process also needs to know which pixel belongs to which person. In Figure 2 , different colors are used to indicate different people.

[0044] Challenge 3: The segmentation and instance separation results of consecutive video frames should be as consistent as possible. Since people in the video are usually moving, this results in significant changes in their postures, facing directions, lighting, sizes, etc. (as Figure 2 shown). The model needs to generate consistent segmentation and instance separation results (robust to all these changes).

[0045] Challenge 4: The model needs to associate the segmented instances between frames. In most scenarios, especially in sports videos, people move very fast, and there are many occlusions, color changes, rescaling, etc. Different examples of this situation are exemplified in Figure 3 , including examples of occlusion, color change, and rescaling (e.g., due to changes in lighting) and object deformation (e.g., due to changes in people's postures between frames) from left to right. Therefore, matching people between frames is challenging.

[0046] Existing solutions do not address all these challenges. Some of the disadvantages of existing solutions are as follows.

[0047] The ability to process only short video clips. For example, Mask2Former-VIS can only process 70 frames (2.8 seconds long, 1080×1920 resolution) at a time on an NVIDIA RTX A6000 GPU (more frames will result in an out-of-memory error). It is worth emphasizing that the NVIDIA RTX A6000 GPU is a powerful GPU with 48GB of memory, and Mask2Former-VIS can only process video clips with a maximum of 8 frames (0.32 seconds long, 1080×1920 resolution) running on an NVIDIA GeForce RTX 2080 GPU (11GB of memory). For more details, please refer to the performance section below.

[0048] Unable to handle changes, such as occlusion, lighting changes, etc. In other words, when changes occur, segmentation and tracking result systems such as Mask2Former-VIS may be suboptimal (as Figure 4 shown). The performance degradation may be caused by process defects. For example, the model in use may treat all input video frames as a volume and extract general global features from all frames simultaneously (the features of one frame are related to the features of other frames). The final prediction output is based on the extracted global features. When changes occur (the features change significantly), the model cannot adapt to these special features, resulting in worse performance.

[0049] Object instances may be lost. Some existing solutions may only generate segmentation masks for a small number of objects (e.g., ten objects). When the video clip (instead of each frame) includes more objects than the limit of ten objects, the system may lose instances in its output. Additionally, when the global features extracted from the input video clip have degraded, such as when changes occur, existing solutions also lose instances. As Figure 4 shown, the Mask2Former-VIS process ignores the second person on the right.

[0050] Long computation time. For example, Mask2Former-VIS extracts 3D (spatial + temporal) features from video clips. Extracting such features and making predictions takes a long time.

[0051] Some solutions may only be applicable to offline application scenarios. For example, Mask2Former-VIS is only applicable to offline application scenarios. It cannot start working until the video clip is recorded. Therefore, Mask2Former-VIS cannot be used in online application scenarios.

[0052] To address the above problems, a frame-based video instance segmentation process and system according to the present disclosure are provided.

[0053] System Overview

[0054] In Figure 5 it, embodiments according to the present disclosure are shown. Specifically, the VIS system 500 is shown to include the following modules: (1) a backbone network module 501, (2) an instance segmentation module 502, (3) a feature adaptation module 503, (4) an instance feature generation module 504, (5) an instance library 505, and (6) a similarity calculation module 506. Each module may be implemented by suitable hardware and / or software (e.g., one or more processors configured to perform the functions of each module). Alternatives are possible, for example, each module may be implemented by dedicated hardware, or the modules may be grouped together with two or more modules implemented in a single hardware module. In the following sections, we will describe these modules in detail. Note that for each frame of the video stream, the functions of each module described herein are repeated.

[0055] The backbone network module 501 and the feature adaptation module 503 can receive the first frame of video data (e.g., frame (t - 1)), which includes one or more objects. Here, the object can refer to a data object, that is, a set of pixels corresponding to / representing a real-world object. The instance segmentation module 502 and the feature adaptation module 503 can extract the mask feature (e.g., mask feature 509) of each object in one or more objects of the first frame by combining the initial mask 507 of each object with one or more extracted features 508 of the frame. The extracted features 508 of the frame can be referred to as the feature map of the frame under discussion in this article. The instance feature generation module 504 can generate a representation (e.g., a feature vector) of each mask feature, which indicates one or more features of the mask feature. The instance library 505 (also referred to as the feature library in this article) stores each representation and the associated identifier. One or more processors of the system are also configured to adjust the first frame by applying each initial mask and the corresponding identifier to the corresponding object in one or more objects of the first frame, and output the adjusted first frame. The system tracks one or more objects in one or more consecutive frames by utilizing the modules of the system. More specifically, for each consecutive frame, the instance segmentation module 502 and the feature adaptation module 503 extract the mask feature of each object in one or more objects of the consecutive frame by combining the initial mask of each object with one or more extracted features of the object. The instance feature generation module 504 generates a representation of each mask feature of the consecutive frame, which indicates one or more extracted features of the mask feature. The similarity calculation module 506 determines whether the representation of each mask feature of the consecutive frame corresponds to the representation stored in the instance library 505. In response to determining that the representation of the mask feature of the consecutive frame corresponds to the representation stored in the instance library 505, one or more processors associate the initial mask of the consecutive frame with the identifier of the corresponding stored representation, and update the corresponding stored representation in the feature library with the corresponding representation of the mask feature of the consecutive frame. The consecutive frame is adjusted by applying each initial mask and the corresponding identifier of the consecutive frame to the corresponding object in one or more objects of the consecutive frame, and the adjusted consecutive frame is output by the system.

[0056] Backbone Network Module

[0057] The backbone network module 501 extracts spatial features (2D) from each input video frame. The features are the features of the objects in the frame (e.g., object color, corners or edges of the object, appearance of the object, such as the arm or leg of a person, etc.). Mask2Former-VIS uses ResNet50 as its default backbone network. In Figure 5 the implementation, MobileNetV3 is used. Due to its lightweight arrangement (faster speed and less resource usage) and good performance, it has been pre-trained on the ImageNet dataset. This is in Figure 6is shown, where the Mask2Former-VIS backbone network 601 is shown on the left, and the MobileNetV3 backbone network 602 used in this embodiment is shown on the right. However, it should be understood that any suitable convolutional neural network (CNN) backbone network can be used.

[0058] The MobileNetV3 backbone network includes a plurality of functional layers, which have been grouped into four functional blocks 602, and each functional block includes one or more functional layers. Note that the grouping of the layers into Figure 6 4 blocks is for illustration and to facilitate comparison with the ResNet50 arrangement, which includes more functional blocks than MobileNetV3. Each block includes a plurality of functions (layers), and each block may include different combinations of functions that result in different outputs, and thus can be selected for use based on the desired output type. Each layer can be assigned a numerical I.D. to distinguish the layers and their respective functions. For example, block 603 may include a convolutional function (which has I.D. layer 1) as the first layer, an activation function layer 2, a batch normalization function layer 3, and a dropout function layer 4. Each layer operates on the image data and produces a specific output. Block 604 may include a convolutional function layer 5, an activation function layer 6, and a batch normalization function layer 7. Blocks 605 and 606 may have different combinations of functions represented by additional layers (e.g., layer 8 forward). Generally, each block represents a combination of functions that can be appropriately grouped together to produce a specific output. For example, block 603 combines the functions of layers 1 to 4 to produce an output that, for example, extracts low-level semantic meaning features, such as the edges or corners of an object in an image frame. Occupying block 603 requires the operation of block 603 on the data (so layers 1 to 4). Then, the functional block of block 603 receives the data produced by block 603 and operates on this data using layers 5 to 7 to produce intermediate-level semantic meaning features, such as data about the arms or legs of the frame object. The lower layers, i.e., layer 8 forward, provide higher-level semantic meaning, such as a person, background, sky, etc. However, as in this embodiment, it is necessary to distinguish a person, and higher-level semantic meaning is not required because it does not distinguish individuals, only the category of a person from the background or sky, etc. Therefore, the outputs from layers 4 and 7 are selected because their output data facilitates tracking different people between frames - extracting certain attributes (corners, edges, arms, legs, etc.) of each person, which helps to identify individuals between frames.

[0059] In this embodiment, layers 4 and 7 blocks are used, as Figure 5As shown. Note that the "usage" of layer 4 implies the occupancy of the entire block 603, because layer 4 operates on the data output from layer 3, while layer 7 operates on the output from layer 6. In other words, the usage of layer 4 means operating on the data of block 603, and the usage of layer 7 means operating on the data of block 603 (and then block 604). It should be understood that other blocks may be used depending on the usage. It should also be understood that the blocks may have a different set of functions and / or may have different combinations of functions. For example, a block may have one, two, or more functions. The functions of the layers may be implemented by a convolutional neural network or by suitable software implemented by one or more processors. If the last two layers (average pooling layer and classification layer) of MobileNetV3 are used for an image classification task, these last two layers may be removed. It is also possible to use separate layers. For example, only layer 1 may be used, or layer 3 may be used (which requires the use of layer 1 and layer 2), etc.

[0060] As Figure 6 shown, the batch size of the image is N (set here to N = 1), the number of channels is 3, and the pixel height and width are 960×720. This may vary depending on the usage. The data provided for each functional block may be compressed or reduced relative to the entire image size. For example, for the first block of MobileNetV3, the number of channels is increased, but the height and width are reduced. It is also noted that compared to Mask2Former, the first block 603 of the MobileNetV3 backbone network has a reduced number of channels and dataset height and width. This is advantageous because it reduces the computational complexity and computational time, and thus improves the execution speed.

[0061] Object Instance Segmentation Module

[0062] The instance segmentation module 502 generates a semantic mask for each detected instance (as Figure 5 shown). Figure 7 Shows a preferred instance segmentation module 502 for implementing aspects of the present disclosure. It should be understood that any suitable image instance segmentation is feasible.

[0063] The outputs from blocks 603, 604, and 606 of the backbone network 600 are fed into the segmentation module 502.

[0064] The convolution and group normalization function 701 is applied to the outputs from block 604 and block 606 to produce block 702 and block 703. Block 702 has 128 channels, a height of 60, and a width of 46. Note that block 604 and block 606 can be downsampled versions of the input frame (1 / 16 and 1 / 32 of the input resolution respectively). Block 703 has 128 channels, a height of 30, and a width of 23. Other channel numbers, heights, and widths for block 702 or block 703 can be used. Then each of block 702, 703 is flattened along a single dimension. This produces block 704 with a width of 2760 and block 705 with a width of 690. This allows these blocks to then be combined into block 706 with a height of 3450, which can be simply achieved by concatenating each of block 704, 705 along a single dimension.

[0065] Then block 706 is fed into a transformer 707 which performs a pixel-by-pixel similarity calculation to determine which pixels pertain to the same object (e.g., the same person) in the image frame. For example, this can be done through appropriate matrix calculations. The similarity calculation can compare, for example, pixel colors to identify pixels of the same color as identifying the same object. Then these pixels are extracted to extract the object of interest from the image data. Note that this pixel-by-pixel similarity calculation is performed on the downsampled frame, which is concatenated along a single dimension compared to the original image frame. In Figure 7 an embodiment, the similarity calculation is performed on a block having 3450 pixels along a single dimension compared to the original image which is a 2D block of 960×720. Once the transformer 707 has identified which pixels belong to the object of the frame, the reverse operation is performed to generate block 708 and block 709 which have the same dimensions as block 702 and block 703.

[0066] Then block 708 is concatenated with the output from block 604 to which convolution has been applied to provide a block having the same resolution as block 708. Then a further convolution operation is applied to the concatenated block to produce a block having the same resolution as block 603. Then this block is concatenated with the block which is the convolution of block 603 and which has the same properties as block 603 but the number of channels has been increased to match the number of channels of the mask module. This produces the final mask feature 710 which includes the mask of the object identified by the transformer but is identified at the resolution of block 603.

[0067] Then blocks 708 and 709 are passed to the transformer decoder 711, where appropriate initialization values (here vectors such as 128-dimensional vectors) are associated with the mask to identify the mask from other masks generated for the frame. The different values associated with a given mask are depicted by the four boxes to the left of the transformer decoder 711. These boxes illustrate that each mask can be associated with a different value. Each box can correspond to an instance of the final output. In Figure 7 , four boxes (before the transformer decoder) are shown, indicating that there are four possible output instances (masks). If there are fewer than four instances in the input frame, the model will still output four instance masks (some of them can be empty, meaning no instance is found for the corresponding initialization vector). The initialization vectors / values serve as implicit cues for the following instance segmentation. These values / vectors can be pre-trained. Then the output of the transformer decoder 711 is combined with the full-resolution mask features through a matrix multiplication operation to produce mask 507, which is output from the instance segmentation module 502, as Figure 5 shown. Mask 507 can be referred to as the initial mask in this article. This initial mask has not yet been combined with the semantic meaning features extracted by the backbone network module 501 and the RGB data from the input frame. The result of this combination is the mask feature 509 discussed further below. Note that in other embodiments, it may not be the color applied by the transformer decoder, but another property that supplements the mask, such as object weighting, to highlight one object more than other objects. The category of the mask feature is also provided, for example, whether the mask feature is a person, an object, the sky, etc.

[0068] Note that this process generally follows the structure used in the Mask2Former segmentation process. Figure 7 The difference between the instance segmentation of the embodiment of Figure 6 and the instance segmentation of Mask2Former is that, compared to Mask2Former, the block resolution is reduced (as discussed regarding Figure 7 ), and block 605 is ignored in the embodiment of

[0069] Feature Adaptation Module

[0070] because it has the same resolution as block 606 and is thus largely redundant. These differences are beneficial because they reduce the computational complexity and computational time. However, it should be understood that alternatives are possible depending on the usage and implementation details.The Feature Adaptation Module 503 normalizes the features before feeding them to subsequent modules. Specifically, as discussed above, features from the 4th and 7th layers of the MobileNetV3 backbone network 501 are extracted, and then these features along with the normalized input RGB frame (e.g., normalized to [-1, 1]) are fed into the Feature Adaptation Module 503. After obtaining the features, the Feature Adaptation Module first resizes the dimensions of all the features to the target resolution (e.g., 1 / 8 of the input frame resolution). This is necessary because the input RGB frame, layer 4 input, and layer 7 input may all have different resolutions. After that, two activation functions (Tanh and ReLU) are used to process the features, such as removing negative values from the input dataset. The output is the concatenation of all the features (rgb_tanh+rgb_relu+4th_tanh+4th_relu+7th_tanh+7th_relu). In Figure 5 different lines are used to highlight different features (original RGB, low-level features, mid-level features, etc.). The reason for choosing to use the RGB frame, the 4th layer features, and the 7th layer features is that: (1) the RGB frame provides color information, (2) the 4th layer features consist of low-level semantic information (e.g., edges, corners, etc.), and (3) the 7th layer features include slightly higher semantic information than the 4th layer (e.g., textures, etc.). As previously discussed, features from higher layers (above the 7th layer) are not helpful for this instance segmentation task because they are high-level semantic information and thus indistinguishable within a category.

[0071] The output of this module is the concatenation of data from layer 7, layer 4, and the RGB frame, as shown by the different layers of the feature map 508 - the front layer corresponding to layer 7 data, the middle layer corresponding to layer 4 data, and the thinnest back layer corresponding to the RGB data of the frame. As mentioned above, layers other than layer 4 and layer 7 can be appropriately used.

[0072] Note that any appropriate feature adaptation can be used depending on the usage. For example, the feature adaptation may only involve rescaling different input features to the same target resolution, or different layers other than layer 4 and layer 7 can be used as previously discussed.

[0073] Instance Feature Generation Module

[0074] The instance feature generation module 504 generates instance-level distinguishable mask features, also referred to herein as mask features 509. The mask features 509 are referred to as mask features herein because they are generated by performing an element-wise multiplication of an "instance mask" 507 (each mask corresponding to an instance, which can be regarded as a binary mask with foreground pixels being 1 and background pixels being 0) and a feature map 508 including "adapted features" of the frame, where the "adapted features" include low-level semantic meaning features and mid-level semantic meaning features (and in some cases high-level semantic meaning features) of the entire frame. The 1 of each mask (representing the part of the frame corresponding to the object under discussion) is multiplied by the corresponding feature of the feature map, and the feature map extracts the features of the corresponding object. Thus, 509 is a "mask feature" composed of the foreground semantic meaning features of the corresponding object, and all background features and the features of other objects are ignored. Therefore, it should be understood that there is a "mask feature" 509 for each object in the frame, where the specific mask feature includes various semantic meaning features (edges, corners, colors, arms, legs, etc.) that have been extracted for that object. In other words, the instance feature generation module 504 generates features that can (they include the necessary semantic meanings) identify individual objects of the same category in the frame. For example, these features allow two individuals to be distinguished from each other. As Figure 5 shown, after receiving the outputs from the instance segmentation module 502 and the feature adaptation module 503 and the original RGB data of the frame, an element-wise multiplication (matrix multiplication) is performed, where the instance segmentation module provides a basic mask for each object identified for the frame, and the feature adaptation module provides the semantic meaning (details of the edges, or object attributes such as arms, legs, etc.) for each object corresponding to the extracted mask. This allows the foreground features to be retained while ignoring the background (such as the sky). This means that each mask of 507 is multiplied by the feature block 508 to generate one mask feature of 509. The output of this combination is the mask feature 509, which represents the mask 507 of the identified object of the frame combined with the data (i.e., features that allow tracking of objects between frames). For example, the color, shape, size, etc. of the object (as provided by the output of the feature adaptation module 503) allow two objects of the same category to be distinguished from each other and then tracked frame by frame. In this case, the instance feature generation module 504 only performs global average pooling within the instance (foreground) to obtain the object instance-level feature vector 510. The global average pooling operation takes a feature of shape (B, C, H, W) as its input and outputs a (B, C, 1, 1) feature vector, where B is the batch size (in this case B = 1), C is the number of channels, and H and W are the spatial dimensions. In this way, it is possible to convert the features into vectors that can be used for similarity calculations as discussed below. In other words, these vectors are representations of the mask features 509.

[0075] These feature vectors 510 are stored in the instance library 505.

[0076] To summarize the process so far, frame (t-1) is received by the backbone network module 501 and the feature adaptation module 503. The instance segmentation module 502, the feature adaptation module 503, and the instance feature generation module 504 extract the mask features of each object in the frame. The mask features 509 are then pooled into feature vectors 510.

[0077] Instance Library

[0078] The instance library 505 (also referred to as the feature library in this document) is maintained by a suitable memory unit to store the features of all detected object instances. The instance library 505 is updated after each frame. As Figure 5 shown, initially (e.g., for the first frame of a video stream), the instance library 505 is empty. The process described in this document segments three instances from the first frame (video frame (t-1)). They are all added to the instance library 505.

[0079] On the second frame (video frame (t)), the process described herein is repeated. Briefly, the frame is passed to the backbone network 501, which outputs to the instance segmentation module 502, which generates an initial mask 507 for frame (t). The feature adaptation module 503 receives the input RGB frame and the outputs of layers 4 and 7 from the backbone network 501 to generate a hierarchical block 508 representing the extracted features of frame (t). The hierarchical block 508 of frame (t) and the initial mask 507 of frame (t) are combined through a matrix multiplication operation to generate mask features 509 for frame (t), which are then pooled to obtain four object instance vectors 510. Three of the four instances have vectors already stored in the instance library 505, so the vectors in the instance library 505 are updated with the new vectors. Here, the instance vectors from frame (t - 1) are simply replaced by the corresponding vectors of frame (t), where the vectors of frame (t) are assigned the same identification I.D. in the feature library as their corresponding vectors from frame (t - 1). For example, the vector of the first segmented object in frame (t - 1) is stored in the instance library with the I.D. "instance 1". The same object is then identified and segmented in frame (t), matches the vector instance 1 stored in the instance library (i.e., it is determined that the two vectors refer to the same object), and is then saved in the instance library as instance 1, replacing the existing instance 1 vector currently stored in the library. In some cases, the identification can be the association of the vector with a color, which is then the color applied to the corresponding mask on the output frame. Other suitable ways of assigning an I.D. to each vector can be used. Although updating the feature library 505 has been described as involving replacing the existing vector with the corresponding vector of the current frame, other update processes are possible. For example, the existing vector itself can be updated based on the vector of the current frame, rather than being completely replaced. For example, a weighted sum such as weight * previous_vector+(1 - weight)*current_vector can be used. Alternatively, a CNN module can be trained to update the vector.

[0080] The fourth instance is new and does not correspond to any vector in the instance library 505, so its corresponding vector is added to the instance library 505.

[0081] The determination of whether the feature is a feature that already has a corresponding vector in the feature library 505 or whether it is a new feature (i.e., a feature corresponding to an object not in the previous frame) is performed by the similarity calculation module 506. In other words, the tracking of objects between frames is achieved by comparing the similarity of the vectors extracted for the current frame with the vectors stored in the feature library 505.

[0082] Similarity Calculation Module

[0083] The similarity calculation module 506 calculates the similarity between the feature vectors of the instances of the current frame and the feature vectors stored in the instance library 505. Determining whether the representations of the mask features of consecutive frames correspond to the representations stored in the instance library 505 includes: comparing the representations of the mask features of consecutive frames with each of the stored representations; determining a similarity metric for each of the stored representations, the similarity metric indicating the similarity between the representation of the mask features of consecutive frames and the stored representation; and determining based on the similarity metric whether the representation of the mask features of consecutive frames corresponds to the representation stored in the feature library. Here, cosine similarity is used to measure the similarity between two feature vectors, but it should be understood that other suitable similarity metrics may also be used. The equation is

[0084]

[0085] where is and is the dot product of, and calculates the magnitude. As an example, vector A may represent a feature vector stored in the feature library, such as the instance 1 vector determined for frame (t - 1). Vector B may represent the corresponding vector determined for frame (t).

[0086] The vector indicates one or more features of the mask feature. Thus, the correspondence between the representation of the mask features of consecutive frames and the representation stored in the feature library indicates the similarity between the features of the mask features of consecutive frames and the features of the mask features of the representation stored in the feature library. The features of the mask feature include the features of the object associated with the mask feature. More specifically, the one or more features are features that allow one mask feature to be distinguished from each of the other mask features. In Figure 5 the case of, the one or more features may be the object color (e.g., the color of a particular person's clothing), semantic edge data indicating the edges of a particular object, and semantic arm and leg data of the object. Other features may be used, such as object height, width, depth, etc., and they may depend on the usage scenario.

[0087] The similarity calculation module 506 compares each vector extracted for a frame with each vector stored in the feature library 505. In Figure 5In an embodiment, for each vector extracted for a frame, the cosine similarity is calculated for each vector stored in the feature library 505 (four vectors and three vectors respectively). If the cosine similarity between the vectors of the first frame (e.g., frame (t - 1)) and the vectors of the second frame (e.g., frame (t)) meets the threshold, it is determined that the vectors in these two frames relate to the same object. If the comparison of the vectors of the current frame with the vectors stored in the feature library results in more than one cosine similarity meeting the threshold, the vector in the feature library that has the maximum cosine similarity with the vectors of the current frame is identified as relating to the same object as the current vectors. Once the feature vectors of the current frame have been identified with the vectors in the feature library, the I.D. of the vectors in the feature library can be assigned to the vectors of the current frame. For example, if the identified vector in the feature library has the color green, green is selected for the vectors of the current frame. This operation is performed for each vector of the current frame such that the color identifying each initial mask of the current frame matches the color identifying the corresponding initial mask of the previous frame (and, on output, the mask and the associated color are applied to the corresponding object of the frame, as discussed in more detail below). Note that any suitable identifier other than color can be used - for example, numbers, or other types of visual identifiers such as visual patterns (hash, dotted lines, etc.). As explained above, for the sake of illustration, dots and line patterns are used in the figures, and each pattern can represent a color highlight, or it can itself be a visual highlight for tracking an object. In other words, after the similarity calculation, it is known which instance generated for the current frame corresponds to which instance in the instance library (or the previous frame). Then, the masks 507 can be reordered such that their order meets the instance order stored in the instance library 505. Functionally, the final output of the model can be a mask, just like 507, but with the correct order, as shown by the mask 512 in Figure 5 shown in

[0088] If no vector comparison reaches the threshold, the vectors of the current frame are determined to relate to an object for which there is no corresponding vector in the instance library, i.e., it relates to an object that is identified and segmented for the first time in the current frame.

[0089] The similarity calculation module 506 performs this vector comparison for each frame to ensure that any identified and segmented objects that appear in each frame are assigned the same feature vector I.D. frame by frame. This I.D. can be derived from the order in the instance library (and the order of the reordered masks for the current frame), or the I.D. can simply be the order. In some cases, the I.D. can include an associated visualization, such as a color or an appropriate hash or other means of visually highlighting the relevant object to the user. This allows the masks to be consistently applied to objects from one frame to the next, even if the objects change position or orientation between frames. For example, instance 1 can be associated with green (e.g., shown as a diagonal pattern in the figure). Correctly mapping instance 1 to the same object (e.g., the same person) that appears frame by frame ensures that the green mask is consistently applied to the same object frame by frame. Here, the I.D. is green associated with the first entry in the instance library 505. This allows for effective tracking of objects of the same category between frames.

[0090] Output Frame

[0091] Once the vectors 510 and the associated I.D. have been generated for a frame, the initial mask 507 is applied to the original input frame together with the I.D. to produce the output frame 511. The masks are applied to each of their corresponding objects within the frame. In Figure 5 the example, the objects are people, and in the output frame, each person is covered with the corresponding mask and the associated color (the identity of the mask).

[0092] The instance segmentation results can be visualized on the frame by coloring each instance using the corresponding color for each instance (each ID has a pre-designed color) and then overlaying it on the RGB frame. It is worth emphasizing that this visualization is for the user to check the quality of the output masks. Any other processes can be based on the model output (ordered instance masks), such as background blurring, removing a specific instance from the video, automatically zooming in or out for a specific person, etc.

[0093] These frames are output, where each frame has the updated mask and I.D. applied to it. As Figure 8 shown, a sequence of output frames is shown. The left - hand column represents the original RGB frames. The middle column represents the output frames produced due to the Mask2Former - VIS process. The right - hand column represents the results of the technology of the present disclosure. As can be seen in the right - hand column, in frames a) to d), the same frame objects (here people) have the same mask (e.g., on the object, highlighted by a red mask in each frame) covering them. This is independent of the changes in the position and orientation of the objects between frames.

[0094] For Figure 8Each frame of the sports video is 30 seconds long (750 frames). The test video consists of different variations, such as multiple occlusions, large deformations, fast movements, changes in lighting and scale, etc. To generate results using Mask2Former-VIS, the test video needs to be divided into multiple short segments. In particular, the video is divided into 11 segments, and each segment has up to 70 frames.

[0095] In contrast, according to the solution presented in this paper, the video is processed frame by frame without any special preprocessing as in Mask2Former-VIS. Figure 8 A visual comparison between the results according to the present disclosure and Mask2Former-VIS on the test video is shown. Different colors are used to indicate different instances throughout the video, which means the same color is used to highlight the same instance on different frames. The results according to the present disclosure are an improvement over Mask2Former-VIS. Specifically, (1) Mask2Former-VIS cannot correctly track instances in long videos, (2) Figure 8 (c) Mask2Former-VIS loses instances, and (3) Figure 8 (d) Mask2Former-VIS generates suboptimal segmentation results for the previous person. However, according to the technology of the present disclosure, almost perfect video instance segmentation results are achieved for this challenging test video.

[0096] It is also noted that the computational time of the technology of the present disclosure is reduced compared to Mask2Former-VIS. All results are collected by running the model on an NVIDIA RTX A6000 GPU. The computational time is measured for processing video segments (70 frames) and is the average of 10 runs. The solution presented in this paper runs 10 times faster than Mask2Former-VIS.

[0097] Processing

[0098] Each module of the Figure 5 VIS system has been discussed in turn. Now, the end-to-end process according to aspects of the present disclosure will be discussed. According to aspects of the present disclosure, the method shown in the Figure 9 and Figure 10 flowcharts is performed. This can be implemented by the Figure 5 VIS system 500 shown or another suitable system.

[0099] At step 91, the first frame of video data is received. At Figure 5In an embodiment, the first frame (t-1) is received by the backbone network 501 and the feature adaptation module 503. The first frame includes one or more objects. At step 92, a mask feature of each object in the one or more objects of the first frame is extracted by combining an initial mask of each object with one or more extracted features of the object. This can be the mask feature 509 generated by combining the output of the instance segmentation module 502 with the feature adaptation module 503 through a matrix multiplication operation. At step 93, a representation of each mask feature is generated, which indicates one or more extracted features of the mask feature. This representation can be the feature vector 510 obtained through a pooling operation of the mask feature 509. The feature vector indicates the features of the mask feature that allow differentiation between the mask features. For example, one or more features are attributes that allow one mask feature to be differentiated from each of the other mask features. In Figure 5 the case of, one or more features can be the object color (e.g., the color of a particular person's clothing), semantic edge data indicating the edges of a particular object, and semantic arm and leg data of the object. Other features can be used, such as object height, width, depth, etc., and they can depend on the usage. Each representation and the associated identifier are stored in a feature library (e.g., the instance library 505). At step 94, the first frame is adjusted by applying each initial mask and the corresponding identifier to the corresponding object in the one or more objects of the first frame, and the adjusted first frame is output to produce, for example, an output frame, such as Figure 5 the frame 511. At step 95, the method further includes tracking the one or more objects in one or more consecutive frames. Generally speaking, this step includes repeating the steps performed for the first frame for each consecutive frame, and for each consecutive frame, updating the feature library with the representation of the mask feature calculated for the consecutive frame, so that the initial mask tracks the object frame by frame.

[0100] The tracking process is shown in Figure 10 and includes: for each consecutive frame, a mask feature of each object in the one or more objects in the consecutive frame is extracted by combining an initial mask of each object with one or more extracted features of the object (step 1001). This can be performed in the same manner as for the first frame. Tracking also includes step 1002: generating a representation of each mask feature of the consecutive frame, which indicates one or more extracted features of the mask feature. Again, this is performed in the same manner as Figure 9 step 92 discussed above. In addition, at step 1003, it is determined whether the representation of each mask feature of the consecutive frame corresponds to the representation stored in the feature library. This can be done by Figure 5The similarity calculation module 506 is performed in the manner discussed above. In response to determining that the representations of the mask features of consecutive frames correspond to the representations stored in the feature library, the corresponding initial masks of the consecutive frames are associated with the identifiers of the corresponding stored representations. This is shown at step 1004 in Figure 10 Tracking also includes updating the corresponding stored representations in the feature library with the corresponding representations of the mask features of consecutive frames.

[0101] Returning to Figure 9 , the method further includes: at step 96, adjusting the consecutive frames by applying each initial mask and the corresponding identifier of the consecutive frames to the corresponding objects in one or more objects of the consecutive frames, and outputting the adjusted consecutive frames.

[0102] Advantages

[0103] Fast frame-based video instance segmentation. In particular, the techniques of the present disclosure provide frame-by-frame processing, rather than treating video clips as 3D spatio-temporal volumes as in legacy techniques. Providing frame-based VIS (i.e., VIS capable of tracking objects frame by frame) allows for faster computations when solving the video instance segmentation task.

[0104] Innovative instance re-identification algorithm. Aspects relate to a training-free and dataset-free instance re-identification method. The method extracts low-level features from a backbone network module and employs, for example, two activation functions to normalize these features. The re-identification function is achieved by calculating the cosine similarity between the instance features of two frames.

[0105] Online solution capability. The techniques of the present disclosure can process video streams in real time and can process recorded videos without any modification. Whether it is a video stream or a pre-recorded video, the process is frame by frame, and the instance segmentation results of each frame are output. However, Mask2Former-VIS can only process pre-recorded videos.

[0106] Hardware-friendly. Since this is a frame-based solution, it will not have an unacceptably large impact on resource usage regardless of the length of the input video (it only requires a small amount of storage to be added to the instance library). However, Mask2Former-VIS requires more GPU usage to process longer videos.

[0107] Dataset-friendly. As a frame-based solution, we only need to train the CNN backbone network on a dataset with frame-level annotations (instead of video-level annotations). However, Mask2Former-VIS can only be trained on a dataset with video annotations (obtaining such datasets is labor-intensive and time-consuming). Accessing a dataset with frame-level annotations is much more efficient than accessing a dataset with video-level annotations.

[0108] Flexibility. There is flexibility during modification. In theory, any frame instance segmentation model (as well as backbone network) can be used. Only the backbone network and the instance segmentation module need to be trained, while the feature adaptation module, the instance feature generation module, the instance library, and the similarity calculation module are training-free. Therefore, the instance matching function can be easily embedded into other backbone networks, instance segmentation modules, or other tasks (e.g., tracking).

Claims

1. A method for tracking an object between frames of video data, the method comprising: include: receiving a first frame of video data, the first frame comprising one or more objects; extracting mask features for each of the one or more objects of the first frame by combining the initial mask of each object with the one or more extracted features of the object; generating a representation of each mask feature, the representation indicating the one or more extracted features of the mask feature, and storing each representation and associated identification in a feature repository; adjusting the first frame by applying each initial mask and corresponding identification to a corresponding object of the one or more objects of the first frame, and outputting the adjusted first frame; as well as Tracking the one or more objects in one or more consecutive frames, for each consecutive frame, the tracking comprising: extracting mask features for each of the one or more objects in the consecutive frames by combining an initial mask for each object with one or more extracted features of the object; generating a representation of each mask feature of the consecutive frames, the representation indicating the one or more extracted features of the mask feature; determining whether the representation of each mask feature of the consecutive frames corresponds to a representation stored in the feature library; In response to determining that the representations of mask features of the consecutive frames correspond to representations stored in the feature library, associating respective initial masks of the consecutive frames with identities of corresponding stored representations; and updating in the feature library corresponding stored representations of the mask features of the consecutive frames with corresponding representations of the mask features of the consecutive frames; The continuous frames are adjusted by applying each initial mask and the corresponding identification of the continuous frames to corresponding objects of the one or more objects of the continuous frames, and the adjusted continuous frames are output. The method of claim 1 , wherein the one or more objects of the frame comprises a plurality of objects.

3. The method according to claim 2, further comprising: include: In response to determining that the representation of the mask features of the consecutive frames does not correspond to a representation stored in the feature library, the representation of the mask features of the consecutive frames is stored as a new entry in the feature library.

4. The method according to claim 1 , wherein determining whether the representation of the mask features of the consecutive frames corresponds to the representation stored in the feature library include: comparing the representation of the mask features of the consecutive frames to each stored representation; determining a similarity measure for each stored representation, the similarity measure indicating a similarity between the representation of the mask feature of the consecutive frames and the stored representation; as well as A determination is made based on the similarity measure whether the representations of the mask features of the consecutive frames correspond to representations stored in the feature library. 5 . The method according to claim 1 , wherein extracting a mask feature of each of one or more objects in the consecutive frames comprises extracting a mask feature of each of a plurality of objects in the consecutive frames. The method of claim 1 , wherein the representation of mask features is a vector.

7. The method of claim 6, wherein determining whether the representation of the mask features of the consecutive frames corresponds to a representation stored in the feature library comprises calculating a cosine similarity between the representation of the mask features of the consecutive frames and the stored representation.

8. The method of claim 1, wherein the correspondence between the representation of the mask features of the consecutive frames and the representation stored in the feature library indicates similarity between the features of the mask features of the consecutive frames and the features of the mask features of the representation stored in the feature library.

9. The method of claim 8, wherein the characteristics of a mask feature include one or more of a color of the object, an edge of the object, or a corner of the object.

10. The method according to claim 1, wherein a mask feature of each of the one or more objects of the first frame is extracted include: inputting the first frame into a convolutional neural network to extract semantic data indicative of one or more features of each of the one or more objects; segmenting the first frame to generate the initial mask of each of the one or more objects; The initial mask is combined with the semantic data to generate the mask features for each of the one or more objects. The method of claim 10 , wherein the one or more features of the mask features include the semantic data indicative of one or more features of a corresponding object. 12 . The method of claim 11 , wherein the one or more features of the corresponding object include one or more of a color of the object, an edge or corner of the object, a feature of the object, and an appearance of the object.

13. The method of claim 10, wherein extracting semantic data indicating one or more features of each of the one or more objects include: One or more functions are applied to the video data of the first frame, the one or more functions comprising a convolution function, an activation function, a batch normalization function, and a dropout function.

14. The method of claim 10, wherein the first frame is segmented to generate an initial mask for each of the one or more objects. include: A pixel-by-pixel comparison is performed on pixels of the first frame, the pixel-by-pixel comparison identifying groups of similar pixels, each group representing one of the one or more objects of the frame.

15. A system for object tracking between frames of video data, the system comprising one or more processors, the one or more processors being configured to: receiving a first frame of video data, the first frame comprising one or more objects; extracting mask features for each of the one or more objects of the first frame by combining the initial mask of each object with the one or more extracted features of the object; generating a representation of each mask feature, the representation indicating the one or more extracted features of the mask feature, and storing each representation and associated identification in a feature repository; adjusting the first frame by applying each initial mask and corresponding identification to a corresponding object of the one or more objects of the first frame, and outputting the adjusted first frame; as well as Tracking the one or more objects in one or more consecutive frames, for each consecutive frame, the tracking comprising: extracting mask features for each of the one or more objects in the consecutive frames by combining an initial mask for each object with one or more extracted features of the object; generating a representation of each mask feature of the consecutive frames, the representation indicating the one or more extracted features of the mask feature; determining whether the representation of each mask feature of the consecutive frames corresponds to a representation stored in the feature library; In response to determining that the representations of mask features of the consecutive frames correspond to representations stored in the feature library, associating respective initial masks of the consecutive frames with identities of corresponding stored representations; and updating in the feature library corresponding stored representations of the mask features of the consecutive frames with corresponding representations of the mask features of the consecutive frames; The continuous frames are adjusted by applying each initial mask and the corresponding identification of the continuous frames to corresponding objects of the one or more objects of the continuous frames, and the adjusted continuous frames are output.

16. The system of claim 15, wherein the one or more objects of the frame comprises a plurality of objects.

17. The system according to claim 16, further comprising: include: In response to determining that the representation of the mask features of the consecutive frames does not correspond to a representation stored in the feature library, the representation of the mask features of the consecutive frames is stored as a new entry in the feature library.

18. The system of any one of claims 15, wherein determining whether the representation of mask features of the consecutive frames corresponds to a representation stored in the feature library include: comparing the representation of the mask features of the consecutive frames to each stored representation; determining a similarity measure for each stored representation, the similarity measure indicating a similarity between the representation of the mask feature of the consecutive frames and the stored representation; as well as A determination is made based on the similarity measure whether the representations of the mask features of the consecutive frames correspond to representations stored in the feature library.

19. The system of any one of claims 15, wherein extracting a mask feature of each of one or more objects in the consecutive frames comprises extracting a mask feature of each of a plurality of objects in the consecutive frames.

20. The system of claim 15, wherein the representation of mask features is a vector.

21. The system of claim 20, wherein determining whether the representation of the mask features of the consecutive frames corresponds to a representation stored in the feature library comprises computing a cosine similarity between the representation of the mask features of the consecutive frames and the stored representation.

22. The system of claim 15, wherein the correspondence between the representation of the mask features of the consecutive frames and the representation stored in the feature library indicates similarity between the features of the mask features of the consecutive frames and the features of the mask features of the representation stored in the feature library.

23. The system of claim 22, wherein the characteristics of a mask feature include one or more of a color of the object, an edge of the object, or a corner of the object.

24. The system of claim 15, wherein a mask feature of each of the one or more objects of the first frame is extracted include: inputting the first frame into a convolutional neural network to extract semantic data indicative of one or more features of each of the one or more objects; segmenting the first frame to generate the initial mask of each of the one or more objects; The initial mask is combined with the semantic data to generate the mask features for each of the one or more objects.

25. The system of claim 24, wherein the one or more features of the mask features include the semantic data indicative of one or more features of a corresponding object.

26. The system of claim 25, wherein the one or more features of the corresponding object include one or more of a color of the object, an edge or corner of the object, a feature of the object, and an appearance of the object.

27. The system of claim 24, wherein semantic data indicative of one or more features of each of the one or more objects is extracted include: One or more functions are applied to the video data of the first frame, the one or more functions comprising a convolution function, an activation function, a batch normalization function, and a dropout function.

28. The system of claim 24, wherein the first frame is segmented to generate an initial mask for each of the one or more objects. include: A pixel-by-pixel comparison is performed on pixels of the first frame, the pixel-by-pixel comparison identifying groups of similar pixels, each group representing one of the one or more objects of the frame.

29. The system of claim 15, further comprising a memory connected to the one or more processors.

30. A computer-readable medium comprising instructions which, when executed, cause one or more circuits of an apparatus for processing data to perform the method of claim 1.