Artificial-intelligence powered ground truth generation for object detection and tracking on image sequences

IN598656BActive Publication Date: 2026-08-11ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
IN202014033989
Authority / Receiving Office
IN · IN
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-08
Filing Date
2020-08-07
Publication Date
2026-08-11
Estimated Expiration
2040-08-07

AI Technical Summary

Technical Problem

Current methods for generating ground truth data for safety-critical systems, such as object detection and tracking, are inefficient due to the high complexity and cost of human annotation, which limits scalability and data availability for machine learning algorithms.

Method used

A human-machine collaboration system that decomposes annotation tasks into micro-tasks, assigning them to either human annotators or machine learning models based on confidence levels and skill types, using AI-interactive and machine-learning pre-labels to improve efficiency and accuracy.

Benefits of technology

This approach reduces the cognitive load on human annotators, increases the speed and accuracy of annotation, and enables the generation of high-precision ground truth data at scale, enhancing the training of machine learning models for object detection and tracking tasks.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A storage (110) maintains raw image data (132) including video having a sequence of frames, and annotations of the frames that indicate aspects of objects identified in the respective frames. A processor (106) determines, for annotation of key frames of the raw image data (132), a task type for key frames and an agent type for key frames, receives annotations of objects identified in the key frames of the raw image data (132) according to the key frame task type and key frame agent type, selects to review the key frames based on a confidence level of the annotations of the key frames, determines, for annotation of intermediate frames of the raw image data (132), a task type for intermediate frames and an agent type for intermediate frames.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present disclosure relates to aspects of an artificial-intelligence (AI)powered ground truth generation for object detection and tracking on image sequences.BACKGROUND ART5

[0002] Ground truths used in safety-critical systems often require high precision ingeometric shape annotation and complex attributes compared to stereotypes ofannotation used in public datasets. For example, a stereotype of pedestrian annotationused in public datasets requires loose bounding box to cover a visible part of apedestrian. However, annotation for the safety-critical systems often requires an10 estimated bounding box with a pixel level accuracy together with centerline of the bodyand additional attributes such as body pose and head angle. Due to this complexity andlengthy requirements along with various scenes, it takes too long for human annotatorsto be aware of all requirements needed for annotation tasks. This prevents scaling outthe number of annotators due to the high learning curve to understand the requirements.15 Furthermore, the high cost of human-only annotation is an obstacle to producing largeamounts of annotation data, which is a pre-requisite to data driven machine learningalgorithms such as deep learning.SUMMARY

[0003] In one or more illustrative examples, a system for human-machine collaborated20 high-precision ground truth data generation for objects identification, localization, andtracking in a sequence of images includes a user interface; a storage configured tomaintain raw image data including video having a sequence of frames, and annotationsof the frames that indicate aspects of objects identified in the respective frames; and aprocessor, in communication with the storage and the user interface. The processor is25 programmed to determine, for annotation of key frames of the raw image data, a tasktype for key frames and an agent type for key frames; receive annotations of objectsidentified in the key frames of the raw image data according to the key frame task type 3and key frame agent type; select to review the key frames based on a confidence levelof the annotations of the key frames; determine, for annotation of intermediate framesof the raw image data, a task type for intermediate frames and an agent type forintermediate frames; and receive annotations of objects identified in the intermediate5 frames of the raw image data according to the intermediate frame task type andintermediate frame agent type.

[0004] In one or more illustrative examples, a method for human-machine collaboratedhigh-precision ground truth data generation for objects identification, localization, andtracking in a sequence of images, includes maintaining raw image data including video10 having a sequence of frames, and annotations of the frames that indicate aspects ofobjects identified in the respective frames, the objects including one or more ofpedestrians, cyclists, animals, vehicles, animals, and moving objects in an indoorenvironment, the annotations include one or more of geometric shapes around theobjects, centerlines of the objects, or directions of travel of the objects; determining,15 for annotation of key frames of the raw image data, a task type for key frames and anagent type for key frames, the task type including one of a human-only annotation tasktype, an AI-interactive task type, or a human task with machine-learning pre-labels tasktype, the agent type including one of a worker with average annotation skill, a workerwith expert skill, or a machine using a machine-learning model; receiving annotations20 of objects identified in the key frames of the raw image data according to the key frametask type and key frame agent type; selecting to review the key frames based on aconfidence level of the annotations of the key frames; determining, for annotation ofintermediate frames of the raw image data, a task type for intermediate frames and anagent type for intermediate frames; and receiving annotations of objects identified in25 the intermediate frames of the raw image data according to the intermediate frame tasktype and intermediate frame agent type.

[0005] In one or more illustrative examples, a computer-readable medium includesinstructions that, when executed by a processor, cause the processor to maintain raw 4image data including video having a sequence of frames, and annotations of the framesthat indicate aspects of objects identified in the respective frames, the objects includingone or more of pedestrians, cyclists, animals, vehicles, animals, and moving objects inan indoor environment, the annotations include one or more of geometric shapes5 around the objects, centerlines of the objects, or directions of travel of the objects;determine, for annotation of key frames of the raw image data, a task type for keyframes and an agent type for key frames, the task type including one of a human-onlyannotation task type, an AI-interactive task type, or a human task with machinelearning pre-labels task type, the agent type including one of a worker with average10 annotation skill, a worker with expert skill, or a machine using a machine-learningmodel; receive annotations of objects identified in the key frames of the raw image dataaccording to the key frame task type and key frame agent type; select to review the keyframes based on a confidence level of the annotations of the key frames, the confidencelevel being based on one or more of (i) performance of a worker performing the15 annotation task, (ii) overall performance of the worker across a plurality of annotationtasks, (iii) a prediction score determined based on a machine-identification of theannotations, or (iv) an analysis of the image quality of the raw image data; determine,for annotation of intermediate frames of the raw image data, a task type for intermediateframes and an agent type for intermediate frames; and receive annotations of objects20 identified in the intermediate frames of the raw image data according to theintermediate frame task type and intermediate frame agent type.BRIEF DESCRIPTION OF DRAWINGS

[0006] FIG. 1 illustrates an example annotation system for the capture and annotationof image data;25

[0007] FIG. 2 illustrates an example of a data diagram for the annotation of image data;

[0008] FIG. 3 illustrates an example workflow of the annotation tasks;

[0009] FIG. 4 illustrates an example of qualification task;5

[0010] FIG. 5 illustrates an example of a user interface for performing manualannotations;

[0011] FIG. 6 illustrates an example of further aspects of a user interface forperforming manual annotations;5

[0012] FIG. 7 illustrates an example of the annotation of pedestrian direction;

[0013] FIG. 8 illustrates an example of pedestrian ID matching;

[0014] FIG. 9 illustrates an example of frames of raw image data with respect topedestrian ID matching;

[0015] FIG. 10 illustrates an example of a user interface for performing AI-assisted10 manual annotations;

[0016] FIG. 11 illustrates an architecture of an AI-based annotator;

[0017] FIG. 12 illustrates an example of a procedure for extracting image patches fortraining;

[0018] FIG. 13 illustrates an example of a final review task for a set of frames;15

[0019] FIG. 14 illustrates an example workflow of a two-level review;

[0020] FIG. 15 illustrates an example of a review question used to perform finalreview; and

[0021] FIG. 16 illustrates an example of a process for the decomposition andperformance of annotation tasks as multiple tasks.20 DETAILED DESCRIPTION

[0022] Implementations of the present disclosure are described herein. It is to beunderstood, however, that the disclosed implementations are merely examples andother implementations can take various and alternative forms. The figures are notnecessarily to scale; some features could be exaggerated or minimized to show details25 of particular components. Therefore, specific structural and functional details disclosed 6herein are not to be interpreted as limiting, but merely as a representative basis forteaching one skilled in the art to variously employ the implementations. As those ofordinary skill in the art will understand, various features illustrated and described withreference to any one of the figures can be combined with features illustrated in one or5 more other figures to produce implementations that are not explicitly illustrated ordescribed. The combinations of features illustrated provide representativeimplementations for typical applications. Various combinations and modifications ofthe features consistent with the teachings of this disclosure, however, could be desiredfor particular applications.10

[0023] This disclosure relates to systems and methods for human and machinecollaborated high-precision ground truth data generation for objectdetection / localization / tracking tasks on sequence of images at scale. Target objects forannotation include pedestrians, cyclists, animals, various types of vehicles in theoutdoor environment and people, animals and any moving objects in the indoor15 environment. The disclosed methods enable decomposition of annotation tasks intomultiple tasks automated in a workflow and dynamically assign each to humanannotator(s) or machine in order to deliver ground truths efficiently. The ground truthsgenerated from the annotation process are used to re-train machine learning modelsused in the annotation process to improve machine prediction over time. The annotation20 process is divided into two major steps: key frame annotations and intermediate frameannotations, which have different tasks in the workflow but share similar tasks.

[0024] Machine learning models, optionally with humans in the loop that are used forground truths generation can be used for other purposes such as prediction of objectrecognition services.25

[0025] To address the scalability and efficiency for annotation tasks, the system in thisdisclosure decomposes a complex annotation task into multiple micro- / machine tasks.Each micro-task is designed to be performed by any workers who have passed basictraining and qualification, without remembering full requirements. Machine tasks are 7designed to apply cutting-edge machine learning models to make the annotationprocess efficient. To improve efficiency more over time, the machine learning modelsfrom previous annotations are used as-is or after re-trained properly for exampleapplying transfer learning. Depending on the characteristics of images collected from5 various cameras, mounted locations and deployed environment, the machine learningmodels from previous annotations are used as-is or re-trained properly for exampleapplying transfer learning.

[0026] As explained in detail herein, the disclosure provides for a human-machinecollaboration for data annotation at scale. Efficiency in large scale data annotation may10 be provided by integrating machine learning models and humans in the loop for theannotation process, and by improving machine prediction by retraining with the datafrom previous batches over time. A manual / human annotation may be time consumingand expensive. This disclosure, accordingly, provides for systems and methods toreduce human annotation efforts by increasing the number of accurate machine15 annotation that does not require adjustment of geometric shape annotation.

[0027] Additionally, the described systems and methods provide for a reduction ofcognitive load in complex annotation tasks. Indeed, it may take significant time for anovice human annotator to learn complex annotation requirements to be able toannotate an entire video without errors. The proposed systems and methods reduce20 learning time for human annotators by dividing a full annotation job into multiplemachine tasks that can be done by many people. Thus, the systems and methods arescalable to quickly recruit and train human annotators.

[0028] With respect to a machine and interactive annotation user interface, machinelearning models are used for some tasks in the annotation work flow to generate25 annotation automatically or collaboratively with human annotators through aninteractive UI. With respect to quality control in crowdsourcing tasks, quality controlmechanisms are embedded in the design of tasks and workflow.8

[0029] As described herein, efficient and scalable ground truths generation system andmethods produce high precision (pixel level accuracy) annotations that are used todevelop object detection / localization, object tracking. This disclosure provides forsystems and methods for human and machine collaborated high precision ground truth5 data generation for object detection / localization / tracking tasks on sequence of imagesat scale. As some examples, objects for annotation include pedestrians, cyclists,animals, various types of vehicles in the outdoor environment and people, animals andany moving objects in the indoor environment.

[0030] FIG. 1 illustrates an example annotation system 100 for the capture and10 annotation of image data 132. The annotation system 100 includes a server 102 thathosts an annotation web application 124 that is accessible to client devices 104 over anetwork 122. The server 102 includes a processor 106 that is operatively connected toa storage 110 and to a network device 118. The server 102 further includes an imagedata input source 130 for the receipt of image data 132. The client device 104 includes15 a processor 108 that is operatively connected to a storage 112, a display device 114,human-machine interface (HMI) controls 116, and a network device 120. It should benoted that the example annotation system 100 is one example, and other systems 100may be used. For instance, while only one client device 104 is shown, systems 100including multiple client devices 104 are contemplated. As another possibility, while20 the example implementation is shown as a web-based application, alternate systemsmay be implemented as standalone systems or as client-server systems with thick clientsoftware.

[0031] Each of the processor 106 of the server 102 and the processor 108 of the clientdevice 104 may include one or more integrated circuits that implement the functionality25 of a central processing unit (CPU) and / or graphics processing unit (GPU). In someexamples, the processors 106, 108 are a system on a chip (SoC) that integrates thefunctionality of the CPU and GPU. The SoC may optionally include other componentssuch as, for example, the storage 110 and the network device 118 or 120 into a single 9integrated device. In other examples, the CPU and GPU are connected to each othervia a peripheral connection device such as PCI express or another suitable peripheraldata connection. In one example, the CPU is a commercially available centralprocessing device that implements an instruction set such as one of the x86, ARM,5 Power, or MIPS instruction set families.

[0032] Regardless of the specifics, during operation, the processors 106, 108 executestored program instructions that are retrieved from the storages 110, 112, respectively.The stored program instructions accordingly include software that controls theoperation of the processors 106, 108 to perform the operations described herein. The10 storages 110, 112 may include both non-volatile memory and volatile memory devices.The non-volatile memory includes solid-state memories, such as NAND flash memory,magnetic and optical storage media, or any other suitable data storage device thatretains data when the annotation system 100 is deactivated or loses electrical power.The volatile memory includes static and dynamic random-access memory (RAM) that15 stores program instructions and data during operation of the annotation system 100.

[0033] The GPU of the client device 104 may include hardware and software fordisplay of at least two-dimensional (2D) and optionally three-dimensional (3D)graphics to a display device 114 of the client. The display device 114 may include anelectronic display screen, projector, printer, or any other suitable device that reproduces20 a graphical display. In some examples, the processor 108 of the client 104 executessoftware programs using the hardware functionality in the GPU to accelerate theperformance of machine learning or other computing operations described herein.

[0034] The HMI controls 116 of the client 104 may include any of various devices thatenable the client device 104 of the annotation system 100 to receive control input from25 workers or other users. Examples of suitable input devices that receive human interfaceinputs may include keyboards, mice, trackballs, touchscreens, voice input devices,graphics tablets, and the like. 10

[0035] The network devices 118, 120 may each include any of various devices thatenable the server 102 and client device 104, respectively, to send and / or receive datafrom external devices over the network 122. Examples of suitable network devices 118,120 include a network adapter or peripheral interconnection device that receives data5 from another computer or external data storage device, which can be useful forreceiving large sets of data in an efficient manner.

[0036] The annotation web application 124 be an example of a software applicationexecuted by the server 102. When executed, the annotation web application 124 mayuse various algorithms to perform aspects of the operations described herein. In an10 example, the annotation web application 124 may include instructions executable bythe processor 106 of the server 102 as discussed above. Computer-executableinstructions may be compiled or interpreted from computer programs created using avariety of programming languages and / or technologies, including, without limitation,and either alone or in combination, JAVA, C, C++, C#, VISUAL BASIC,15 JAVASCRIPT, PYTHON, PERL, PL / SQL, etc. In general, the processor 106 receivesthe instructions, e.g., from the storage 110, a computer-readable medium, etc., andexecutes these instructions, thereby performing one or more processes, including oneor more of the processes described herein. Such instructions and other data may bestored and transmitted using a variety of computer-readable media.20

[0037] The web client 126 may be a web browser, or other web-based client, executedby the client device 104. When executed, the web client 126 may allow the client device104 to access the annotation web application 124 to display user interfaces of theannotation web application 124. The web client 126 may further provide input receivedvia the HMI controls 116 to the annotation web application 124 of the server 102 over25 the network 122.

[0038] In artificial intelligence (AI) or machine learning systems, model-basedreasoning refers to an inference method that operates based on a machine learningmodel 128 of a worldview to be analyzed. Generally, the machine learning model 128 11is trained to learn a function that provides a precise correlation between input valuesand output values. At runtime, a machine learning engine uses the knowledge encodedin the machine learning model 128 against observed data to derive conclusions such asa diagnosis or a prediction. One example machine learning system may include the5 TensorFlow AI engine made available by Alphabet Inc. of Mountain View, CA,although other machine learning systems may additionally or alternately be used. Asdiscussed in detail herein, the annotation web application 124 and machine learningmodel 128 may be configured to recognize and annotate features of the image data 132for use in the efficient and scalable ground truths generation system and methods to10 produce high precision (pixel level accuracy) annotations that are used to developobject detection / localization, object tracking.

[0039] The image data source 130 may be a camera, e.g., mounted on a moving objectsuch as car, wall, pole, or installed in a mobile device, configured to capture image data132. In another example, the image data input 132 may be an interface, such as the15 network device 118 or an interface to the storage 110, for the retrieval of previouslycaptured image data 132. The image data 132 may be video, e.g., a sequence of images.Each image in the image data 132 may be referred to herein as a frame. For privacyconcerns, faces and license plates may be blurred from the image data 132 for certainannotation tasks.20

[0040] FIG. 2 illustrates an example 200 of a data diagram for the annotation of imagedata 132. As shown, the raw image data 132, such as videos, is stored in a data lake(e.g., the storage 110, drive, or other storage device). In an annotations task workflow202, the raw image data 132 is provided to a human annotation task 204 to createannotations 206. The annotations 206 may include, for example, weak annotations 208,25 machine-learned annotations 210, and final annotations 212. Additional metadata 214may also be stored with respect to the raw image data 132. For instance, this additionalmetadata 214 may include weather conditions during which the raw image data 132was captured, geographic locations of where the raw image data 132 was captured, 12times during which the raw image data 132 was captured, etc. As discussed in furtherdetail below, a training data selector 216 may be used to select raw image data 132 andannotations 206 from the memory 110 as shown at 218. A machine-learning algorithmat 220 receives the selected raw image data 132 and annotations and creates a revision5 of the trained model at 222. This trained model is then used by the annotation taskworkflow 202 to provide the machine-learned annotations 210. The machine-learnedannotations 210 may also be overseen by the human annotation task 204.

[0041] FIG. 3 illustrates an example 300 workflow of the annotation tasks. In general,the workflow of the annotation is divided into two phases: 1) key frame annotation and10 2) intermediate frame annotation. If a key frame interval is one, then all frames areannotated like a key frame without intermediate frame annotation. With respect to keyframe annotation, key frame interval selection is performed, then annotation type andtasks are performed, then a review of the key frame annotations may be performed.With respect to the intermediate frame annotation, the intermediate frame annotations15 are generated by machine first then human annotators validate correctness of machineannotations and provide feedback to machine by correcting annotations.

[0042] Regarding key frame annotation, a key frame interval selection is performed.In an example, the key frame interval may be selected or configured in the annotationsystem 100 as a static value (e.g., key frame interval is 5). In another example, the key20 frame interval may be defined pursuant to a formula or function with importantparameters in the domain. For instance, a key frame interval selection function inautonomous driving domain may be represented with parameters of car speed, steeringwheel, road type (city, highway), etc. and / or scene similarity / difference across nearbyframes. The key frame interval may be dynamically selected by annotators to maximize25 their efficiency and annotation accuracy when they interact with underlying machinethat auto-generates annotations for subsequent frames upon annotators input.

[0043] With respect to annotation types and tasks, the annotation system 100 may bedynamically configured for different annotation tasks depending on efficiency or other 13reasons. The annotation system 100 may also be configured for the maximum andminimum number of tasks assigned to one annotator.

[0044] Regarding manual or human-only annotation, the annotations may be donemainly by human annotator(s). Depending on task complexity, human annotators may5 be qualified for a given task type by performing an online training class and by passinga qualification test.

[0045] FIG. 4 illustrates an example 400 of qualification task. As shown, the example400 illustrates an annotation task interface, in which an aspect of bounding box trainingis shown where it is the human operator’s turn to draw a bounding box around the same10 pedestrian shown in a previous example. User interface buttons are provided in thedisplayed user interface to allow the human to draw the box. Once the box has beendrawn, the check my work control may be selected to allow the user to have the workchecked.

[0046] The annotation task interface may include an instruction pane in which the15 instructions may be provided, and an annotation pane, where annotators can draw ageometric shape over a target object on an image / frame. The geometric shapes mayinclude bounding boxes, centerlines, cuboids, L-shapes, single or multi-points, lines orfree drawing forms. It should be noted that these are common examples of shapes thatmay be used in annotation tasks, but different shapes may be used. The annotation task20 may include providing various attributes such as body pose of a person, head angle ofperson, etc. Note that these are specific requirements that may be used only in certainexamples, and some implementations do not have these attributes for pedestrianannotation. Moreover, different object types may have different attributes. Forinstance, a pedestrian may have a walking direction or body pose attributes, but a25 vehicle may not.

[0047] The annotation task may include an object matching task which is asked toidentify same object on different frames. If one annotator is asked to draw a geometric 14shape for the same object across frames, an object ID may be assigned to each of theannotations across frames. If the frames are divided into multiple annotation tasks, IDmatching tasks may be created to find a same object across frames. Optionally IDmatching may be preliminarily done by machine before creating an ID matching task5 for human annotator. Examples of manual annotations are illustrated with respect toFIGS. 5, 6, and 7. Examples of ID matching are illustrated with respect to FIGS. 8 and9.

[0048] FIG. 5 illustrates an example 500 of a user interface for performing manualannotations. In an example, the user interface may be provided to the display 114. The10 title of the user interface indicates that the user interface is for the drawing of abounding box around one person for five frames of raw image data 132. As shown, themanual annotation user interface includes an instruction pane instructing the user tolook at the frame to the right and pick one pedestrian larger than the rules shown belowand without a yellow box around it. If such a pedestrian is found, the user may select a15 control to indicate that a pedestrian was found to annotate. If not, the user may select acontrol to indicate that all pedestrians already have a box. The user interface may alsoinclude an annotation page displaying a frame of the raw image data 132 from whichthe user may attempt to identify pedestrians.

[0049] FIG. 6 illustrates an example 600 of further aspects of a user interface for20 performing manual annotations. As indicated, a goal of the user interface is theidentification of accurate bounding boxes and centerlines for all pedestrians. Theoperations that may be used to identify a new pedestrian include to click on a “newpedestrian” control of the user interface, and then select the outermost points of thefour sides of the pedestrian’s body. If a pedestrian is partially occluded, then the points25 should be entered to estimate the covered outermost points. Next, the user may adjustthe centerline to cross the center of the hip of the pedestrian. This process may berepeated until all the pedestrians have accurate boxes and centerlines. Extra attentionmay be paid to dark areas as well as to small pedestrians in the image. In one example, 15the human operator may be compensated per box entered or per pedestrian identified.If so, an indication of the human operator’s earnings may be included in the userinterface.

[0050] In an example, if the machine learning aspect of the annotation system 1005 determines that there are more than a predefined number of pedestrians in the image(e.g., 20 pedestrians), the user interface may provide the users with a choice to stop andsubmit after a portion of the pedestrians are located (e.g., once 20 are located).

[0051] In an example, if an image includes a large number of pedestrians, then theimage may be divided into different patches to be manually annotated separately. The10 user may be able to view one of the patches initially, and may be able to click on theadditional patches to show that portion of the image. In some implementations, once anew patch is shown, the user may not be able to return to the previous patch.

[0052] FIG. 7 illustrates an example 700 of the annotation of pedestrian direction. Inaddition to bounding box and centerline, additional attributes of the pedestrians such15 as the direction the pedestrian is walking may be annotated in the images. As shown,the title of the user interface indicates that the user interface is for identification ofpedestrian direction. Additionally, the annotation user interface includes an instructionpane instructing the user to identify which direction best describes a direction that apedestrian highlighted in the annotation pane is walking. The choices may include that20 the pedestrian is walking to the left, is walking to the right, is walking towards thedriver’s direction (roughly), or is walking away from the driver’s direction (roughly).The user interface may further ask which angle best reflects the angle that thehighlighted pedestrian is walking, and may provide some example angles.Additionally, the instruction pane indicates that if a highlighted pedestrian is not25 present in the current frame, that any answer may be selected, and that the reviewer’swork will be reviewed by another worker. The user interface may also include anannotation page displaying a frame of the raw image data 132 from which the user mayattempt to identify the direction of walk of the pedestrian.16

[0053] FIG. 8 illustrates an example 800 of pedestrian ID matching. As shown, theuser interface provides for selection of one or more of a set of pedestrians that areidentified in a base frame (illustrated on the left). The user interface also provides amatch frame (illustrated on the right) from which the user may map the same5 pedestrians. In this way, the same ID may be used for the same pedestrian acrossframes.

[0054] FIG. 9 illustrates an example 900 of frames of raw image data with respect topedestrian ID matching. As shown, five key frames are requested for a worker A toannotate of one pedestrian. (If there are n pedestrians, then a total of n different tasks10 with workers are to be done for the same key frames). Also as shown, a next set of fivekey frames are requested for a worker B to annotate of the same pedestrian. Betweenthe two sets of frames, there should be a matching of the same pedestrian to generate acoherent pedestrian ID.

[0055] Turning to AI-assisted annotation, this annotation type may be designed to be15 interactively done with annotators and machine. A human annotator may provide aweak label on a target object (e.g., a single point click on a center of the target object,a rough bounding box covering a target object, etc). A machine task may provide arefined / accurate geometric shape annotation (e.g., precise bounding box on a targetpedestrian). If a machine-generated geometric shape annotation is inaccurate and / or20 not within a tolerance range, then the annotator may provide feedback simply bycorrecting any incorrect parts through a user interface. The weak label, originalmachine predicted annotation, and human corrected annotation may be saved to theannotation system 100 to retrain the machine online or offline.

[0056] Machine-generated annotations may be achieved by various approaches. In one25 example, an approach may be utilized that takes an image patch as input and estimatesa tight bounding box around the main object in the image patch as output. The inputimage patch may be a cropped region from the original image based on estimated objectlocation by an object detector or tracker. A deep convolutional network architecture 17may be utilized that efficiently reduces errors of the computer vision pre-labels andthat is easily adopted to datasets that are different from the training data. The latterproperty may be useful in annotation systems 100 since they usually encounter datawith different metadata characteristics (e.g., different camera parameters, different5 conditions in road and weather, etc.).

[0057] Different geometric shape annotations may also leverage other machinelearning algorithms. In an example, to have an accurate bounding box, a machinegenerated semantic segmentation annotation may be used by selecting the outmostpoints of (x, y) coordinates of a target segment. Similarly, a center of body annotation10 may be generated by leveraging key points of a body prediction.

[0058] FIG. 10 illustrates an example 1000 of a user interface for performing AIassisted manual annotations. As shown, the user interface may indicate that the goal isthe annotation of accurate bounding boxes and centerlines for all pedestrians with AIassist. The steps that may be performed to do so may include to first click two points15 top left and bottom right to cover a pedestrian. The AI may then generate a boundingbox. The AI may then be taught by the user clocking the correct outmost point(s) ofthe pedestrian, if necessary. If a pedestrian is partially occluded, then the points shouldbe entered to estimate the covered outermost points. The AI may then generate thecenterline. The AI may be taught by the user clicking to correct the centerline, if20 necessary. This process may be repeated until all the pedestrians have accurate boxesand center lines. Extra attention may be paid to dark areas as well as to smallpedestrians in the image. In one example, the human operator may be compensated perbox entered or per pedestrian identified. If so, an indication of the human operator’searnings may be included in the user interface. Additionally, the user interface may25 provide an indication of how much the AI learned from the user.

[0059] A third type of annotation, beyond manual annotation and AI-assistedannotation, is machine-initiated annotation with pre-labels. This annotation type mayuse an object detector to detect an object with object class / category. These detected 18objects may be input to the AI as image patches cropped from the video frames basedon pre-labeled bounding boxes from computer vision algorithms, either object trackingor object detection. Before cropping the images, the four edges of the pre-labeledbounding boxes may be expanded to ensure that the visible part of the object is included5 in the image patch. The AI may then predict precise bounding boxes for the objects. Avideo annotation approach is one example, but the AI approach may be applied to anyannotation systems 100 that utilize computer vision pre-labels, or be used to makerough bounding boxes drawn by annotators more precise. In such an annotationpipeline, an input video sequence of raw image data 132 is first divided into key frames10 (sampled every K frames, where K can be determined by the speed of car movementsand environmental changes) and intermediate frames. Pre-labels in key frames may beinitialized by object detectors, and then refined by the AI. Such key frame pre-labelsmay be reviewed and then corrected or re-drawn by a human worker to ensure keyframe labels are precise. Annotated key frames may be used to populate pre-labels for15 intermediate frames using object trackers. Pre-labels for intermediate frames may thenbe refined by the AI, and may go through human annotators to correct them. The resultmay be that the detected objects input to the AI are refined into tight bounding boxesfor the detected objects. Regardless, pre-labels from step 1 may be validated and / orcorrected by the human annotator(s).20

[0060] A review of key frame annotations may be performed by different humanannotator(s) having high quality profiles before generating intermediate frameannotation to have high quality annotation. The review process may be oneconsolidated task, or it may be divided into multiple steps that can be performed bymore than one human annotator. The review may be performed for all annotations, or25 may instead be done for targets-only annotations with a low confidence level. FIG. 6illustrates an example.

[0061] Moving from key frame annotations to intermediate frame annotations, theintermediate frame annotations may be generated by a machine first. Then, human 19annotators may validate correctness of the machine annotations and provide feedbackto machine by correcting the annotations.

[0062] Geometric shape and object ID generation may be performed by the machine.A pre-label location / position of a target object may be identified based on the position5 of a start key frame and an end key frame of selected frames. The position may becalculated with interpolation and computer vision technique such as kernelizedcorrelation filters. The generated annotation may have the same ID.

[0063] A patch of a target object may be created (pre-labels from machine learningalgorithms with extra regions) to get estimation of fine-grained (high precision)10 geometric shape annotation. A detailed description of an example estimator isdiscussed in detail herein.

[0064] FIG. 11 illustrates an architecture 1100 of an AI-based annotator. As shown,the first group of layers are the feature extractors, followed by three fully connectedlayers. In this example, the output dimension is four, which corresponds to the (x, y)15 coordinates of bottom left corner and upper right corner of the bounding box. Here,feature extractors refer to the convolutional layers of well-known deep neural networkarchitectures for computer vision, such as VGG16 (described in detail in KarenSimonyan and Andrew Zisserman. Very deep convolutional networks for large-scaleimage recognition. arXiv preprint arXiv:1409.1556, 2014), ResNet50 (described in20 Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning forimage recognition,” Proceedings of the IEEE conference on computer vision andpattern recognition, pages 770–778, 2016), and MobileNet (Andrew G Howard,Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand,Marco Andreetto, and Hartwig Adam, “Mobilenets: Efficient convolutional neural25 networks for mobile vision applications.” arXiv preprint arXiv:1704.04861, 2017),each of which is incorporated herein by reference in its entirety.20

[0065] Regarding the overall architecture, the AI-based annotator first learns andextracts features from images, and then uses the extracted features to estimate the fourextreme coordinates of the main object in the image patch. Hence, the architecture isdivided into a feature extractor and a coordinate estimator. For the feature extractor, a5 design goal is to represent the main object in the image patch. This may make use, insome examples, of feature extractors shown effective in various computer vision tasks,such as the first thirteen layers of VGG16, the first forty-nine layers of ResNet50, andthe first eighteen layers of MobileNet. Another design metric of the feature extractoris transferability, as data annotation systems usually encounter images with different10 characteristics; therefore, the AI should be easily transferable to different datasets.

[0066] Regarding direct transferability, at the beginning of annotating a new set ofimages, the AI model used in the annotation pipeline would have been trained with adifferent dataset. The AI may estimate the four extreme coordinates of the objects asprecisely as possible, even though more error is expected since the AI model was15 trained with a different dataset. This is to say, the AI model should have as low varianceas possible, while maintaining estimation power. Given the same size of training dataand training strategies, reducing variance calls for reducing the number of parameters.

[0067] Regarding minimal data required for adaptation, as much as the AI model isintended to be directly transferable, the model will never perform as well as if it was20 trained on the same dataset. A common practice in data annotation is to re-train orfinetune computer vision algorithms used in the system after a portion of the data isannotated. One can see that if the algorithms can be fine-tuned in the earlier stageduring annotation, the more the algorithms can assist human annotators, and hence thelower cost of annotating the whole dataset. Therefore, the AI model should require25 minimal size of data to be fine-tuned to new datasets.

[0068] Moving to the coordinate estimator, top layers of the AI model are configuredto learn mappings between extracted image features and the four extreme coordinatesof the main object in the image patch. The estimator may be a deep neural network 21with an architecture including feature extractors and coordinate estimators. The featureextractors may be consistent with the architecture proven to be useful in computervision literature. Examples of such feature extractors are VGG16, ResNet50, andMobileNet. The coordinate estimators may be configured to learn mappings between5 extracted image features and the four extreme coordinates of the main object in theimage patch. To learn such a nonlinear mapping, more than one layer of regressionis needed because the coordinate estimator is inherently more difficult than an objectdetector (which usually has only one fully connected layer after pooling). A lossfunction may be defined as well. For the loss function, the purpose of the AI is to make10 bounding box boundaries as precise as possible, meaning that there are as few aspossible pixels between the object extreme points and the bounding box. For instance,L1 distance may be a natural choice for measuring the performance of the estimation.To make the optimization, Huber loss may be adopted. To mimic error in pre-labels fortraining such an estimator, error statistics may be gathered that were produced in pre15 labels and inject such error when generating image patch for training. For a centerlineof a body annotation, to use a target object type of a pedestrian, the centerline of thebody may be predicted by leveraging machine models for key-points of body and / orinterpolation of center line based on two key frames. Attribute annotation generationmay also be performed by the machine. The attributes may be generated based on20 interpolation of two key frames.

[0069] The first step of training the AI model is to estimate distributions of the errorthat the AI model is going to correct. Such error distribution helps the coordinateestimator to localize the main object in the image patch and the training procedure tomimic errors that the AI model needs to correct in the annotation pipeline. In addition,25 training with error distribution instead of real error from the computer visionalgorithms better isolates the AI model from how exactly computer vision algorithmsperform in the annotation pipeline, and improve transferability of the AI model to anew dataset.22

[0070] Using the annotation pipeline in FIG. 11 as an example, the error that the AImodel corrects may include object detector error and object tracker error. Ideally,statistics may be collected of both algorithms, but to reduce training effort the AI modelmay first train with the worse error. In this case, since the object tracker is initialized5 every K frames and the objects of interest (vehicles and pedestrians) usually do nothave sudden change of motion, as long as K is not too large, bounding box boundaryerror of the object tracker should be smaller than that of the object detector. Hence theAI may operate the object detector over the whole training dataset and match boundingboxes with the ground truth to collect bounding box boundary error statistics10 introduced by the object detector.

[0071] FIG. 12 illustrates an example 1200 of a procedure for extracting image patchesfor training. After obtaining computer vision error statistics, the next step is to extractimage patches containing each object and ground truth coordinates of the object withineach patch – image patches will be the input of the AI model, and ground truth15 coordinates will be used to compute the loss value. As shown in the example 1200,given an image and ground truth bounding box of one fully-visible object, the fouredges of the ground truth bounding box are first expanded by a fixed ratio to ensurethat the object is fully included in the image patch; then the four edges are shifted(depending on the number drawn from the distribution, each edge can be moved inward20 or outward) randomly based on error statistics collected from the object detector; thenimage patch is cropped, the patch is normalized to a fixed size as input to train the AImodel. The normalization procedure maintains aspect ratio of the original patch andthe empty pixels are filled with 0 for all channels. Note that the training patches aregenerated on the fly during each epoch of training, so size of the cropped patch could25 be different when an object is used multiple times during training.

[0072] There may be certain validation tasks for human annotators. The validationtasks may be divided into three different steps: i) deleting any incorrect annotation thatcovers untargeted object(s), ii) adding a new annotation that does not cover a targeted 23object, and iii) adjusting geometric shape annotations if the machine generatedannotation do not satisfy a precision requirement. If machine confidence level exists,the validation tasks may be targeted for annotations with a low confidence level.

[0073] Then, final review tasks may be performed for all frames. The review process5 may be one consolidated task, or it may be divided into multiple steps that can beperformed by more than one human annotator. The selection of review may be donefor annotations with low confidence level. The review may be done interactively withmachine. After review, the geometric shape of annotation for all target objects and IDtracking may be done by machine again. Further annotation (e.g., grouping of objects)10 may be added by machine by calculating overlapping of bounding boxes and attributesof two or more objects.

[0074] FIG. 13 illustrates an example 1300 of a final review task for a set of frames.As shown, three frames are in progress, while three other frames are waiting. Of thethree frames in progress, the first frame is indicated as having completed the annotation15 of bounding boxes and the generation of IDs, while the second and third frames areonly indicated as having completed the annotation of bounding boxes.

[0075] Continuous training of the machine learning models with outcomes from theannotation system may be performed. The annotation system 100 stores all (includingintermediate and final) annotation results to continuously train machine learning20 algorithms used in the annotation process. Referring back to FIG. 2, the continuoustraining pipeline architecture has access to the data lake (a storage repository)containing raw image data 132, all annotation data and meta data. The training dataselector determines which data shall be used for the next training cycle. The trainingdata selector has functions and logic programmed to statistically analyze distribution25 of the meta data, compute differences between machine annotations and finalannotations, and select the target training data based on the analysis results in order tomaximize learning in machine learning algorithms. For example, if final annotationsfor an object with height > 500 pixel in night scenes have < 70% IoU (Intersection of 24Union) with machine annotation, frames with those annotation may be selected astarget training data.

[0076] Referring to quality control of crowd workers, the annotation system 100 maycategorize the human annotators into two different roles depending on their quality5 profile. One of these roles is that of the average worker, a worker who passed trainingand qualification, if necessary, to perform annotation task. Another of these roles is thetrust nodes / workers, who have done great work in the past and trusted workers are incharge of reviewing the other worker’s tasks.

[0077] The annotation system 100 may have three different review processes.10 Depending on task type, one or more of these different review processes may beapplied. A first of these processes is a two-level review between workers themselvesfor the same task. A second of these processes is an independent review / validation task.A third of these review tasks is the final review tasks (for key frames and final results)performed by experts.15

[0078] FIG. 14 illustrates an example 1400 workflow of a two-level review. In the twolevel review, a worker (human annotator) review happens after each worker (e.g.,Worker A) submitted his / her task. The annotation system 100 may assign a review andan annotation task to another worker (e.g., Worker B) to review the accuracy of WorkerA and to provide feedback before Worker B works on his / her own annotation task.20

[0079] Referring more specifically to the example 1400, if the Worker B is not a trustnode / worker, then the annotation system 100 may create a review task for trust workerto review Worker B’s task. If the review of Worker B’s task is negative, the annotationsystem 100 may send the task to the original worker (Worker B) and ask him / her torevise. If Worker B does not provide a revision until the deadline, the task may be25 rejected and another review and annotation task may be created. Otherwise the WorkerB’s task may be approved. If the review of Worker B’s task is positive, then the resultby Worker B is valid.25

[0080] If the review of Worker A’s task is negative, then the annotation system 100may send the task to the original worker (Worker A) for revision. If Worker A does notprovide a revision until the deadline, the task may be rejected and another review andannotation task may be created. Otherwise, the Worker A’s task may be approved.5 Upon task approval or rejection, the worker’s quality profile is updated.

[0081] For the independent review / validation task, for annotations that are mostly doneby machine, instead of a two-level review process, an independent task may be utilizedwhere review is only done by n number of workers. For the final review tasks (e.g., forkey frames and final results of intermediate frames) by experts: before10 publishing / finalizing ground truths, experts (whose quality profile is higher or equal tominimum quality condition for trust nodes) may be engaged to correct any incorrectannotations.

[0082] FIG. 15 illustrates an example 1500 of a review question used to perform finalreview. As shown, the user interface is requesting the worker look at the frame on the15 right (the annotation pane) and identify pedestrians without a bounding box aroundthem. The instruction may continue to remind the worker to pay particular attention tothe dark areas and small pedestrians. The review question may be to ask how manypedestrians lack bounding boxes around them. In answering this question, theannotation system 100 may receive additional input on the quality of the annotation.20

[0083] FIG. 16 illustrates an example of a process 1600 for the decomposition andperformance of annotation tasks as multiple tasks. In an example the process 1600 maybe performed by the annotation web application 124 in the context of the annotationsystem 100. The process 1600 may include a flow for the annotation of key frames, aswell as a flow for the annotation of intermediate frames.25

[0084] With respect to key frames, the process may begin at operation 1602 with theannotation web application 124 selecting key frames for annotation. In an example, the 26annotation web application 124 may identify the key frames in an input video sequenceof raw image data 132.

[0085] At operation 1604, the annotation web application 124 may identify a task typefor the annotation to be performed, and also an agent type for the annotation. The task5 type may include, for example, a human-only annotation task type, an AI-interactivetask type, or a human task with machine-learning pre-labels task type. The agent typemay include, for example, a worker with average annotation skill, a worker with expertskill, or a machine. In some examples, these identifications may be performedaccording to user input. In other examples, these identifications may be performed by10 the annotation web application 124 based on the raw image data 132 available forannotation.

[0086] A test instance may be launched by the annotation web application 124 for theselected frames to perform the annotation at 1606. Example user interfaces for theannotation and / or review of annotations are discussed in detail above with respect to15 FIGS. 4, 5, 6, 7, 8, 10, 13, and 15. The review may then be performed by the indictedagent type for the indicated task type. Once the review is completed or aborted, atoperation 1608 control may return to operation 1604 to select another annotation task.

[0087] Additionally or alternately, the process 1600 may continue from operation 1608to operation 1610, wherein frames of the raw image data 132, as annotated, are selected20 for expert review. This selection may be performed based on a confidence level of theannotation. For instance, the confidence level may be based on one or more of theperformance of the worker performing the annotation task (or the worker’s overallperformance in all annotation tasks), a prediction score of the machine determinedbased on the machine identification of the annotations, an analysis of the image quality25 of the raw image data 132, and / or based on other difficulties in performing theannotation (e.g., human operator or machine lack of ability to identify objects in theraw image data 132).27

[0088] At operation 1612, the annotation web application 124 launches a review taskinstance for selected frames of the annotated raw image data 132. In an example, thereview may be an AI-assisted review, or a human-only review. Example user interfacesfor the annotation and / or review of annotations are discussed in detail above with5 respect to FIGS. 4, 5, 6, 7, 8, 10, 13, and 15.

[0089] At operation 1614, the annotation web application 124 completes the reviewfor the selected frames and target objects. For instance, the annotation web application124 may confirm that review of the task type has been completed for all objects, for allobjects with a lower than a threshold value confidence, etc.10

[0090] Next, at 1616, the annotation web application 124 determines whetheradditional annotation and tasks remain to be performed for key frames. If so, controlpasses to operation 1604. If not, the process 1600 ends.

[0091] With respect to annotation of intermediate frames, at operation 1618 theannotation web application 124 may perform an automatic generation of annotations15 for the intermediate frames. At operation 1620, similar to operation 1604 but forintermediate frames, the annotation web application 124 may select a task type, anagent type for the intermediate frames. This task may be, for example, a human-onlytask, or an AI-interactive task.

[0092] At operation 1622, similar to operation 1606 but for intermediate frames, the20 annotation web application 124 launches a task instance for the intermediate frames.The review of the intermediate frames may accordingly be performed according to thetask type and agent type. After operation 1622, at operation 1624 the annotation webapplication 124 determines whether there is additional review to be performed ofintermediate frames. For instance, there may be some intermediate frames that are to25 be reviewed using a different task type. If so, control passes to operation 1620. If not,once annotation and review of the key frames and also the intermediate frames is 28completed, control passes to operation 1626 to indicate the completion of theannotation.

[0093] In general, the processes, methods, or algorithms disclosed herein can bedeliverable to / implemented by a processing device, controller, or computer, which can5 include any existing programmable electronic control unit or dedicated electroniccontrol unit. Similarly, the processes, methods, or algorithms can be stored as data andinstructions executable by a controller or computer in many forms including, but notlimited to, information permanently stored on non-writable storage media such asROM devices and information alterably stored on writeable storage media such as10 floppy disks, magnetic tapes, CDs, RAM devices, and other magnetic and opticalmedia. The processes, methods, or algorithms can also be implemented in a softwareexecutable object. Alternatively, the processes, methods, or algorithms can beembodied in whole or in part using suitable hardware components, such as ApplicationSpecific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), state15 machines, controllers or other hardware components or devices, or a combination ofhardware, software and firmware components.

[0094] While exemplary implementations are described above, it is not intended thatthese implementations describe all possible forms encompassed by the claims. Thewords used in the specification are words of description rather than limitation, and it is20 understood that various changes can be made without departing from the spirit andscope of the disclosure. As previously described, the features of variousimplementations can be combined to form further implementations of the presentsubject matter that may not be explicitly described or illustrated. While variousimplementations could have been described as providing advantages or being preferred25 over other implementations or prior art implementations with respect to one or moredesired characteristics, those of ordinary skill in the art recognize that one or morefeatures or characteristics can be compromised to achieve desired overall systemattributes, which depend on the specific application and implementation. These 29attributes can include, but are not limited to cost, strength, durability, life cycle cost,marketability, appearance, packaging, size, serviceability, weight, manufacturability,ease of assembly, etc. As such, to the extent any implementations are described as lessdesirable than other implementations or prior art implementations with respect to one5 or more characteristics, these implementations are not outside the scope of thedisclosure and can be desirable for particular applications.

Claims

I / We Claim:

1. A system for human-machine collaborated high-precision ground truth datageneration for object identification, localization, and tracking in a sequence of images,comprising:5 a user interface;a storage (110) configured to maintain raw image data (132) including videohaving a sequence of frames, and annotations of the frames that indicate aspects ofobjects identified in the respective frames; anda processor (106), in communication with the storage (110) and the user10 interface, programmed todetermine, for annotation of key frames of the raw image data (132), a task typefor key frames and an agent type for key frames,receive annotations of objects identified in the key frames of the raw image data(132) according to the key frame task type and key frame agent type,15 select to review the key frames based on a confidence level of the annotationsof the key frames,determine, for annotation of intermediate frames of the raw image data (132),a task type for intermediate frames and an agent type for intermediate frames, andreceive annotations of objects identified in the intermediate frames of the raw20 image data (132) according to the intermediate frame task type and intermediate frameagent type.

2. The system as claimed in claim 1, wherein the task type includes one of a humanonly annotation task type, an AI-interactive task type, or a human task with machinelearning pre-labels task type.

313. The system as claimed in claim 1, wherein the agent type includes one of a workerwith average annotation skill, a worker with expert skill, or a machine using a machinelearning model.

4. The system as claimed in claim 1, wherein the processor (106) is further programmed5 to, when operating using an agent type of a machine using a machine-learning model,detect objects with precise tight bounding geometric shapes using the machine-learningmodel, the machine-learning model having a deep convolutional network architectureincluding a feature extractor configured to identify features of the objects, followed bya coordinate estimator configured to identify coordinates of the objects using the10 identified features.

5. The system as claimed in claim 1, wherein the objects for annotation include one ormore of pedestrians, cyclists, animals, vehicles, animals, or moving objects in an indoorenvironment.

6. The system as claimed in claim 1, wherein the annotations include one or more of15 geometric shapes around the objects, bounding boxes around the objects, centerlinesof the objects, object-specific attributes, or directions of travel of the objects.

7. The system as claimed in claim 1, wherein the confidence level is based on one ormore of (i) performance of a worker performing the annotation task, (ii) overallperformance of the worker across a plurality of annotation tasks, (iii) a prediction score20 determined based on a machine-identification of the annotations, or (iv) an analysis ofthe image quality of the raw image data (132).

8. The system as claimed in claim 1, wherein the processor (106) is further programmedto:select frames from the raw image data (132) and corresponding manual25 annotations of the frames;revise training of a machine-learning model configured to identify objects inthe frames using the manual annotations; andprovide machine-learned annotations of the frames for receiving manualcorrections via the user interface.5 9. The system as claimed in claim 8, wherein the raw image data (132) is associatedwith additional metadata including one or more elements of context information, thecontext information specifying one or more of weather conditions during which theraw image data (132) was captured, geographic locations of where the raw image data(132) was captured, or times during which the raw image data (132) was captured, and10 the metadata is used as an input to aid in the revised training of the machine-learning10. The system as claimed in claim 8, wherein the manual corrections received via theuser interface are used as at least a portion of the manual annotations to revise thetraining of the machine-learning model.15 11. The system as claimed in claim 8, wherein the manual annotations of the framesinclude clicks identifying estimated centers of objects regardless of whether the objectis occluded, and the machine-learned annotations include bounding geometric shapesaround the objects as identified by the centers.

12. The system as claimed in claim 8, wherein the manual annotations of the frames20 include identifying estimated outmost points of the objects regardless of whether theobject is occluded, and the machine-learned annotations include centerlines of theobjects as identified by the outmost points.

13. The system as claimed in claim 1, wherein the processor (106) is furtherprogrammed to, as review of the annotations of the key frames, receive validation input25 from the user interface, the validation input including one or more of (i) manualdeletion of incorrect annotations that cover untargeted objects, (ii) manual addition ofnew annotations that do not cover a targeted object, and (iii) adjustment of geometric 33shape annotations for machine-generated annotations that fail to satisfy a precisionrequirement.

14. A method for human-machine collaborated high-precision ground truth datageneration for objects identification, localization, and tracking in a sequence of images,5 comprising:maintaining raw image data (132) including video having a sequence of frames,and annotations of the frames that indicate aspects of objects identified in the respectiveframes, the objects including one or more of pedestrians, cyclists, animals, vehicles,animals, and moving objects in an indoor environment, the annotations include one or10 more of geometric shapes around the objects, centerlines of the objects, or directionsof travel of the objects;determining, for annotation of key frames of the raw image data (132), a tasktype for key frames and an agent type for key frames, the task type including one of ahuman-only annotation task type, an AI-interactive task type, or a human task with15 machine-learning pre-labels task type, the agent type including one of a worker withaverage annotation skill, a worker with expert skill, or a machine using a machinelearning model;receiving annotations of objects identified in the key frames of the raw imagedata (132) according to the key frame task type and key frame agent type;20 selecting to review the key frames based on a confidence level of theannotations of the key frames;determining, for annotation of intermediate frames of the raw image data (132),a task type for intermediate frames and an agent type for intermediate frames; andreceiving annotations of objects identified in the intermediate frames of the raw image25 data (132) according to the intermediate frame task type and intermediate frame agent3415. The method as claimed in claim 14, wherein the confidence level is based on oneor more of (i) performance of a worker performing the annotation task, (ii) overall5 the image quality of the raw image data (132).

16. The method as claimed in claim 14, wherein the processor (106) is furtherprogrammed to:annotations of the frames,10 revise training of a machine-learning model configured to identify objects inthe frames using the manual annotations, and17. The method as claimed in claim 16, wherein the raw image data (132) is associated15 with additional metadata including one or more elements of context information, the20 model.

18. The method as claimed in claim 16, wherein the manual corrections received viathe user interface are used as at least a portion of the manual annotations to revise the19. The method as claimed in claim 16, wherein the manual annotations of the frames25 include clicks identifying estimated centers of objects regardless of whether the object 35is occluded, and the machine-learned annotations include geometric shapes around the20. The method as claimed in claim 16, wherein the manual annotations of the frames5 object is occluded, and the machine-learned annotations include centerlines of the21. The method as claimed in claim 14, wherein the processor (106) is further10 deletion of incorrect annotations that cover untargeted objects, (ii) manual addition of22. A computer-readable medium comprising instructions that, when executed by a15 processor (106), cause the processor (106) to:maintain raw image data (132) including video having a sequence of frames,20 more of geometric shapes around the objects, centerlines of the objects, or directions25 machine-learning pre-labels task type, the agent type including one of a worker with36of the key frames, the confidence level being based on one or more of (i) performance5 of a worker performing the annotation task, (ii) overall performance of the workeracross a plurality of annotation tasks, (iii) a prediction score determined based on amachine-identification of the annotations, or (iv) an analysis of the image quality ofthe raw image data (132);10 a task type for intermediate frames and an agent type for intermediate frames; and23. The medium as claimed in claim 22, further comprising instructions that, when15 executed by the processor (106), cause the processor (106) to:corrections via the user interface; and20 utilize the manual corrections received via the user interface as at least a portionof the manual annotations to revise the training of the machine-learning model toidentify the objects,wherein the raw image data (132) is associated with additional metadataincluding one or more elements of context information, the context information25 specifying one or more of weather conditions during which the raw image data (132)was captured, geographic locations of where the raw image data (132) was captured, 37or times during which the raw image data (132) was captured, and the metadata is usedas an input to aid in the revised training of the machine-learning model.

24. The medium as claimed in claim 22, further comprising instructions that, whenexecuted by the processor (106), cause the processor (106) to, as review of the5 annotations of the key frames, receive validation input from the user interface, thevalidation input including one or more of (i) manual deletion of incorrect annotationsthat cover untargeted objects, (ii) manual addition of new annotations that do not covera targeted object, and (iii) adjustment of geometric shape annotations for machinegenerated annotations that fail to satisfy a precision requirement.