A computer implementation method for generating at least one dataset associated with at least one image of a real-world environment for annotating at least one image frame to train a machine learning model.
A computer-implemented method using multisensor data for automated annotation addresses the inefficiencies of manual video data annotation, improving the speed and quality of training machine learning models for augmented reality and autonomous driving.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RAMBLR GMBH
- Filing Date
- 2024-04-29
- Publication Date
- 2026-05-26
AI Technical Summary
Existing image-based systems for augmented reality and autonomous driving require large-scale manual annotation of video data, which is labor-intensive, time-consuming, and resource-intensive, with inconsistencies in quality and efficiency due to the complexity of identifying relevant objects of interest.
A computer-implemented method that utilizes multisensor data, including audio and eye-tracking, to automate the annotation process by generating a dataset for training machine learning models, reducing ambiguity and improving annotation quality and speed through audiovisual ground truth and ranked object of interest lists.
The method enhances the efficiency and consistency of annotating image frames, enabling faster and more accurate training of machine learning models for augmented reality and autonomous driving applications by focusing on relevant objects of interest.
Smart Images

Figure 2026516860000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer implementation method for generating at least one dataset, particularly egocentric video data, associated with one or more objects in at least one image of a real-world scene. The at least one dataset thus generated is configured for input to a computer implementation process for annotating at least one image frame of a real-world scene, particularly for input to and training of at least one image-based machine learning model. The present invention also relates to a computer implementation method for annotating at least one image frame of a real-world scene, particularly egocentric video data, configured for input to and training of at least one image-based machine learning model, and a corresponding computer program. [Background technology]
[0002] In image-based hardware systems, such as intelligent wearables that operate based on real-time object recognition, as applied to augmented reality applications, it is crucial that the system can identify and categorize a large number of highly diverse objects in the real environment surrounding the user. Such a system could be relied upon, for example, to identify where keys were left, how to change a car battery, or how to avoid dangers while cycling.
[0003] However, in order to ensure that such image-based systems in augmented reality applications and the like function correctly, the system needs to be trained on real images or dynamic image sequences on which images are annotated to define objects of material relevant to a particular system, such as various real objects, especially interacted objects, visually prominent objects, or universally significant objects.
[0004] In this context, traditionally, datasets have been manually annotated, meaning that human operators manually label or assign clusters of pixels to specific target categories or instances. When datasets are manually tagged, a massive workforce is required to generate the enormous amount of labeled data needed to drive computer models used, for example, in augmented reality applications. Annotating such large datasets often involves a trade-off between consistency and labeling quality. Users of such technologies often avoid large-scale annotation tasks due to a lack of staff, time, and resources.
[0005] Video data annotation presents its own challenges; it is extremely difficult, time-consuming, and expensive. Video annotation is typically known as the task of labeling and tagging video footage to train computer vision models. Annotating video (or image sequences) is more complex and labor-intensive than annotating still images because the target subject is moving. Video annotation is processed as frame-by-frame image data. Therefore, a 10-second video consists of hundreds of frames, meaning that completing a single video annotation task takes a considerable amount of time.
[0006] In addition, not every frame in a video sequence contains useful or relevant object information for a specific machine learning task. Time and resources are better spent annotating relevant objects, or more appropriately, so-called objects of interest ("OoI" can be both singular and plural). For example, when training a computer model to recognize the tools needed to make coffee and the range of human movement, recognizing and labeling objects such as kitchen tiles, light bulbs, ovens, dishwashers, and refrigerator magnets is unnecessary for this task. In the example above, the objects of interest (OoI) would be the coffee maker, coffee pot, water, coffee filter, coffee beans, and coffee grounds.
[0007] There are conventional techniques available in the field of OoI detection based on so-called bottom-up models, such as visual splendor or visual indicators. For example, see Ko, BC, & Nam, JY (2006). Object-of-interest image segmentation based on human attention and semantic region clustering. JOSA A, 23(10), 2462-2470.
[0008] Systems have also been created that utilize bottom-up saliency and pointing gestures to identify pointed-at objects for machine learning recognition and learning tasks. For example, see Schauerte, B., Richarz, J., & Fink, GA (2010, October). Saliency-based identification and recognition of pointed-at objects. In 2010 IEEE / RSJ International Conference on Intelligent Robots and Systems (pp. 4638-4643). IEEE.
[0009] There is also research on modeling top-down attention, such as predicting where a person will look in an image by reflecting factors like prior knowledge, current goals, and commands; see, for example, Burra, N., Mares, I., & Senju, A. (2019). The influence of top-down modulation on the processing of direct gaze. Wiley Interdisciplinary Reviews: Cognitive Science, 10(5), e1500. Top-down attention can represent the spontaneous assignment of attention to a specific object, feature, or region in space. For example, a subject might choose to focus on all the yellow items or an empty area of space.
[0010] Several studies have considered the correlation between object recognition and the development of attention. For example, see 't Hart, BM, Schmidt, HC, Roth, C., & Einhauser, W. (2013). Fixations on objects in natural scenes: dissociating importance from salience. Frontiers in psychology, 4, 455.
[0011] In particular, there are systems that fuse gaze mapping and attention mapping to improve object detection. For example, see U.S. Patent No. 9,740,949, which presents a system and method for detecting objects of interest in dynamic images by fusing cognitive algorithms with human analysts to improve detection accuracy.
[0012] Human eye-tracking data has been used in conventional techniques to generate gaze maps from input video data, thereby graphing gaze points to indicate potential OoI locations, as people tend to gaze at specific objects of interest. The relationship between gaze and object is a recognized tool for understanding natural scene processing and has the potential to improve the video annotation process. Specifically for video annotation, work has been done on object of interest segmentation in video sequences based on gaze-based interactions. See, for example, Krieger, L., Heidemann, G., & Schoning, J. (2018, December). Object of Interest Segmentation in Video Sequences with Gaze Data. In 2018 IEEE International Conference on Image Processing, Applications and Systems (IPAS) (pp. 104-109). IEEE.
[0013] With regard to target mask prediction and human annotator modification, European Patent Application Publication No. 3 540 691 provides a method for image segmentation and annotation that performs pixel-level clustering of an image. These clusters are then selected based on their correspondence to regions of interest ("ROIs") defined by a given image mask, and the proposed ROIs are compared to the actual ROIs.
[0014] An annotation system is available, for example, TORAS from the University of Toronto (Kar, A., Kim, SW, Boben, M., Gao, J., Li, T., Ling, H., …Fidler, S. (2021). Toronto Annotation Suite. Retrieved from https: / / aidemos.cs.toronto.edu / toras), which enables state-of-the-art, rapid manual annotation of images. While instructions on how to annotate a particular image can be manually assigned by the task manager, most context and task-specific information is not available to the annotator. This leads to consistency and quality issues across different annotation projects. [Prior art documents] [Patent Documents]
[0015] [Patent Document 1] U.S. Patent No. 9,740,949 [Patent Document 2] European Patent Application Publication No. 3 540 691 [Non-patent literature]
[0016] [Non-Patent Document 1] Ko, BC, & Nam, JY(2006).Object-of-interest image segmentation based on human attention and semantic region clustering.JOSA A,23(10),2462-2470 [Non-Patent Document 2] Schauerte, B., Richarz, J., & Fink, GA (2010, October).Saliency-based identification and recognition of pointed-at objects.In 2010 IEEE / RSJ International Conference on Intelligent Robots and Systems(pp.4638-4643).IEEE [Non-Patent Document 3] Burra, N., Mares, I., & Senju, A. (2019). The influence of top-down modulation on the processing of direct gaze. Wiley Interdisciplinary Reviews: Cognitive Science, 10(5), e1500 [Non-Patent Document 4] 't Hart,BM,Schmidt,HC,Roth,C.,&Einhauser,W.(2013).Fixations on objects in natural scenes:dissociating importance from salience.Frontiers in psychology,4,455 [Non-Patent Document 5] Krieger, L., Heidemann, G., & Schoning, J. (2018, December). Object of Interest Segmentation in Video Sequences with Gaze Data. In 2018 IEEE International Conference on Image Processing, Applications and Systems (IPAS) (pp. 104-109). IEEE [Non-Patent Document 6] Kar, A., Kim, S. W., Boben, M., Gao, J., Li, T., Ling, H., … Fidler, S. (2021). Toronto Annotation Suite. Retrieved from https: / / aidemos.cs.toronto.edu / toras
Summary of the Invention
Means for Solving the Problems
[0017] Therefore, a customized computer-implemented method is needed that is configured for input to a computer-implemented process for annotating at least one image frame of a scene in a real-world environment, enables more efficient classification of digital images or videos and training of image-based machine learning models for use in augmented reality applications or autonomous driving, and generates at least one dataset associated with one or more objects in at least one image of a scene in a real-world environment.
[0018] The present disclosure relates to a computer-implemented method and a computer program according to the appended claims. Embodiments are disclosed in the dependent claims.
[0019] According to one aspect, a computer-implemented method for generating at least one dataset associated with one or more objects in at least one image of a scene in a real-world environment, configured for input to a computer-implemented process for annotating at least one image frame of a scene in a real-world environment for input and training of at least one machine learning model, comprising: receiving, by at least one computing device, at least one image of a scene in a real-world environment captured by an image sensor operated by a user; At least one computing device receives audio data from the user detected by at least one audio sensor, and eye-tracking data of the user detected by at least one human eye-tracking sensor, while the user is capturing at least one image. The system includes, by at least one computing device, performing a speech analysis process which involves analyzing at least a portion of speech data received from a user into text format and generating a set of terms determined by the analysis process that indicate the user's articulated tasks or activities while operating an image sensor, The process involves pre-assigning objects of interest from a set of terms generated in the analysis process, associating a probability with each of the pre-assigned objects of interest, and generating a scene objects of interest list associated with at least one image, each containing its respective object category and associated probability. To generate an eye-tracking heatmap from received eye-tracking data, The process involves processing at least a portion of at least one image using a panoptic segmentation model to generate a set of segmentation masks, wherein each segmentation mask in the set of segmentation masks has its respective assigned target category and panoptic segmentation confidence level. The method involves generating at least one dataset comprising a set of ranked interests, each of which is ranked according to at least one aggregate score determined based on a generated list of scene interests, a generated heatmap, and a set of generated segmentation masks, wherein the at least one dataset is configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model. A method comprising the above is disclosed.
[0020] In another embodiment, a computer implementation method for generating at least one dataset associated with one or more objects of at least one image of a real-world scene, configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model, The system includes receiving at least one image of a real-world scene captured by a user-operated image sensor via a computing device, and At least one computing device receives audio data from the user detected by at least one audio sensor, and eye-tracking data of the user detected by at least one human eye-tracking sensor, while the user is capturing at least one image. The process involves using at least one computing device to analyze at least a portion of the audio data received from the user and to perform an audio analysis process that generates analyzed audio data indicating the user's articulated tasks or activities or areas of interest in the scene while the image sensor is operating, To generate an eye-tracking heatmap from received eye-tracking data, The panoptic segmentation model processes at least a portion of at least one image and generates a set of segmentation masks, wherein each segmentation mask in the set of segmentation masks has its own assigned target category. To generate audio-visual ground truth from analyzed audio data and gaze heatmaps, which includes location information for at least one object in at least one image and the respective object category, The method involves matching the generated audiovisual ground truth with the spatially corresponding segmentation masks and their assigned target categories in a set of generated segmentation masks, thereby generating at least one dataset of audiovisual ground truth having at least one data label indicating each target category and the matched segmentation mask, wherein the generated at least one dataset is configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model. A method comprising the above is disclosed.
[0021] Therefore, according to the present invention, the decision-making process for training image-based machine learning models can be automated in determining what is eligible as an object of interest and what is not. Specifically, the proposed method reduces the selection range of a human annotator and thus reduces ambiguity and deliberation time when determining which objects in a given keyframe are objects of interest. Only objects relevant to the hand-held annotation task are annotated. By reducing the mental burden on the annotator and their arbitrary selection through guidance signals, the quality and speed of task-specific annotations can be improved, and therefore, the training of image-based machine learning models, such as those used in augmented reality (AR) applications or autonomous driving, can be improved.
[0022] According to one embodiment, the ranked set of objects of interest is configured to be used in a computer implementation process for annotating at least one image frame of a real-world scene for input and training of at least one machine learning model for an augmented reality application.
[0023] According to one embodiment, the method further comprises providing a list of unprocessed universal objects of interest from at least one image, and extending a list of scene objects of interest with the list of universal objects of interest.
[0024] According to one embodiment, the method further comprises generating an audiovisual ground truth from the user's received audio data and eye-tracking data, comprising an object category and location information for at least one object in at least one image, and matching the generated audiovisual ground truth with the spatially corresponding segmentation masks and their assigned object categories in a set of generated segmentation masks.
[0025] According to one embodiment, at least one aggregate score is determined based on (i) the probability of each object of interest included in the scene object of interest list, (ii) the gaze heat for each corresponding segmentation mask having the object category corresponding to each object of interest, and (iii) the weighted sum of the panoptic segmentation confidence scores for each corresponding segmentation mask having the object category corresponding to each object of interest.
[0026] According to one embodiment, determining at least one aggregate score includes matching the target categories included in the scene interest list with each target category in the set of generated segmentation masks, and matching the segmentation masks and gaze heatmaps of the set of generated segmentation masks, and optionally determining at least one aggregate score includes assigning the score from the scene interest list to each target category that could not be matched with the target categories in the generated segmentation masks.
[0027] According to one embodiment, the method further comprises adding to a scene interest list at least one region of interest resulting from gaze heat exceeding a predetermined threshold that was not matched with the generated segmentation mask, the at least one region of interest being assigned an undefined interest category.
[0028] According to one embodiment, an image sensor and optionally a human eye-tracking sensor are part of a wearable computing device worn by a user while viewing a scene in a real environment. For example, such a wearable computing device may be or comprise wearable glasses.
[0029] In another embodiment, a computer implementation method for annotating at least one image frame of a real-world scene, configured for input and training of at least one machine learning model, The system includes receiving at least one image frame of a real-world scene to be annotated by at least one computing device, and providing at least one image frame to a display device for presentation to the user, Receiving at least a portion of a set of ranked interests provided in accordance with a method for generating at least one dataset comprising the set of ranked interests described herein, by at least one computing device, Presenting at least a portion of the ranked set of interests received on the display device while displaying at least one image frame to be annotated on the display device, The system receives instructions from a human annotator to instruct at least one computing device to add or improve one or more target categories or segmentation masks to at least one image frame, and extends at least one image frame with corresponding annotation information to at least one annotated dataset configured for input and training to at least one machine learning model. A method is provided that includes the following:
[0030] According to one embodiment, at least a portion of the ranked set of objects of interest is presented on a display device after receiving an instruction from an annotator to instruct at least one computing device to verify or correct any stored audiovisual ground truth. Optionally, at least a portion of the ranked set of interests presented on the display device comprises only those interests not included in the verified or modified audiovisual ground truth.
[0031] In a further embodiment, a computer implementation method for annotating at least one image frame of a real-world scene, configured for input and training of at least one machine learning model, The system includes receiving at least one image frame of a real-world scene to be annotated by at least one computing device, and providing at least one image frame to a display device for presentation to the user, Receiving, by at least one computing device, at least one dataset of audiovisual ground truth generated according to the method for generating at least one dataset of audiovisual ground truth described herein, The system receives instructions from a human annotator to instruct at least one computing device to add or improve one or more target categories and segmentation masks to at least one image frame, and to augment at least one image frame with corresponding annotation information to at least one annotated dataset configured for input and training to at least one machine learning model. A method is provided that includes the following:
[0032] According to one embodiment, the method further comprises presenting one or more segmentation masks of a set of segmentation masks on a display device so as to overlap with at least one image frame, and receiving an instruction from an annotator to instruct at least one computing device to verify or modify at least one dataset of audiovisual ground truth by verifying or improving each of the subject categories and one or more of the presented segmentation masks.
[0033] According to one embodiment, at least one image frame corresponds to at least one image captured by the image sensor, or, if at least one image is part of a video stream, corresponds to an image captured by the image sensor before or after the at least one image captured by the image sensor.
[0034] In a further embodiment, a computer program is provided which, when the program is executed by at least one computing device, comprises instructions causing at least one computing device to perform the method described in any one of the preceding claims. The computer program may be contained in a non-temporary computer-readable medium such as memory or a data storage device, and may be stored, for example.
[0035] According to aspects of the present invention, beyond image-based systems, annotation of images and video images can be supported by speech-to-text solutions based on existing technologies that convert voice tags created by human collectors during data acquisition into machine-readable forms. For example, a human collector might say "tomato" when they see, point to, or touch a tomato in a scene. These voice tags can be further reduced to category labels.
[0036] According to one embodiment, at least one image-based machine learning model is configured for use by one or more processes in an augmented reality application. Sensor data may include egocentric sensor data collected from the viewpoint of a user of a wearable computing device. According to one embodiment, sensor data is multi-sensor data provided by one or more sensors. In one embodiment, a video annotation solution derived from natural language processing (NLP) may exist.
[0037] For example, data can be collected by a wearable hardware device (such as "Tobii Pro Smart Glasses") without an AR application. However, according to another embodiment, an AR application running on a wearable hardware device can simultaneously provide an AR experience while collecting egocentric sensor data.
[0038] This invention may be applicable to other fields of use beyond augmented reality, such as autonomous driving.
[0039] According to aspects of the present invention, the sensor data includes: one or more images, such as one or more video streams; audio data; and collector eye line-of-sight data (in particular corneal reflection, stereo geometry, eye orientation, and / or motion).
[0040] Herein, aspects of the present invention will be further described in reference to the following figures illustrating exemplary embodiments. [Brief explanation of the drawing]
[0041] [Figure 1] This figure shows a computer implementation method according to an embodiment of the present disclosure, which in particular discloses one embodiment of a multimodal data processing and annotation workflow. [Figure 2]This figure shows an exemplary scenario in an exemplary use case of OoI enhancement and guidance signals regarding the potential data collection process for the "cooking pasta" scene. [Figure 3] This figure shows an exemplary preprocessing scheme for obtaining a ranked OoI list based on analyzing gaze, panoptic segmentation, and audio into a scene OoI list. [Figure 4] This figure illustrates one embodiment of Stage 3-1 of the workflow shown in Figure 1, which verifies AV ground truth. [Figure 5] Figure 1 illustrates one embodiment of Stage 3-2 of the workflow shown, which demonstrates a ranked OoI review. [Modes for carrying out the invention]
[0042] Herein, aspects of the present invention will be described in more detail with reference to the drawings. It should be noted that the present invention is not limited to the disclosed embodiments provided primarily to illustrate specific aspects of the present invention in an exemplary manner. A detailed description of one or more embodiments of the present invention is provided below, along with accompanying drawings illustrating the present invention. The present invention will be described in relation to such embodiments, but the present invention is not limited to any embodiment. For clarity, technical materials or terms known to those skilled in the art will not be described in detail so as not to unnecessarily obscure the present invention.
[0043] As used herein, machine learning (ML) is understood as a term commonly known in the art. Techniques related to ML may also be designated as deep learning (DL), artificial intelligence (AI), etc. ML is used as a term to describe all different forms of algorithms based on machine learning or deep learning. This could be image classification, object detection, or other methods of interpreting sensor, task, and / or process data. For example, deep learning analysis is performed using machine learning models such as artificial neural networks.
[0044] A machine learning model is understood to be a model ready for use by AI or ML algorithms, such as an artificial neural network.
[0045] (Image) A frame is a common term known in the art, used for a single still image. A keyframe is a frame having one or more properties, such as contrast exceeding a certain threshold, and is suitable for being selected from a group of frames for further steps according to aspects of the invention described herein. For example, in the case of video, keyframes of video are typically intended to be a set of frames or the smallest set of frames that provide an accurate or most accurate representation of the video.
[0046] Annotated or annotated data refers to data that has a structure that can be used as input to and trained on at least one machine learning model for generating ML-based processes or components of such processes. In particular, annotated or annotated data includes, for example, machine-readable information about the data that is used as input to and trained on machine learning models. Typically, annotated or annotated data includes a description (annotation) of the ingested data in a format that can be further processed by a machine.
[0047] In possible embodiments, at least one computing device can be used, which may comprise one or more components, such as one or more microprocessors in a local, mobile, and / or distributed configuration. For example, the at least one computing device may comprise a wearable device, such as wearable glasses 11, having one or more microprocessors for processing data, in combination with a local and / or remote computer 12, such as a personal computer or server computer, which communicates with the wearable device 11 by wired or wireless means (as will be described in more detail below with reference to Figure 2). The wearable glasses 11 are worn by the user while viewing a scene in a real environment, as will be described further below herein. At least one computing device 11, 12 and / or a portion of the data sources used (such as one or more sensors) can be implemented in software and / or hardware, discretely or distributedly, in any suitable processing device, such as one or more microprocessors in a local device, such as a mobile or wearable computer system, and / or in one or more server computers accessible through a local network and / or the Internet.
[0048] Unless otherwise specified, components such as processors described as being configured to perform a task may be implemented as general components temporarily configured to perform a task at a given time, or as specific components manufactured to perform a task. As used herein, the terms processor, processing device, or computing device refer, respectively, to one or more devices, circuits, microprocessors, servers, and / or processing cores configured to process data such as computer program instructions.
[0049] A method according to an embodiment of the present invention, carried out by at least one computing device or computer system or computer apparatus (as schematically shown in 11, 12), involves at least one computing device 11, 12 receiving sensor data (also called multisensor data) from one or more sensors, according to an embodiment from multiple sensors, which includes receiving at least one image of a scene of a real environment captured by an image sensor such as a digital camera (not explicitly shown by itself, for example, implemented in wearable glasses 11 illustrated in Figure 2, as will be described in more detail below). In one embodiment, the at least one image may comprise, for example, multiple images in discrete-time instances, or multiple images in the form of a stream of images, for example, a video stream.
[0050] The method further involves at least one computing device receiving audio data from the user detected by at least one audio sensor, and user eye-tracking data detected by at least one human eye-tracking sensor while the user is capturing at least one image. In one embodiment, at least one audio sensor (such as one or more microphones) is implemented in the wearable glasses 11 or is a separate device. At least one human eye-tracking sensor may be implemented as a sensor such as a camera for detecting the eye gaze of a collector (e.g., corneal reflection, stereo geometry, eye orientation, and / or movement), or may comprise such a sensor. In one embodiment, at least one human eye-tracking sensor is also implemented in the wearable glasses 11 or is a separate device.
[0051] The embodiments described herein are, for example, multisensor data acquisition and processing to facilitate simpler, faster, and more consistent annotation by human annotators. The inventors have found that average keyframe annotation time can be reduced and consistency improved by providing efficient instructions to the collector and guidance signals to the human annotator. Specifically, apart from providing object masks derived from initial deep learning models that are validated and refined by human annotators, the invention described herein leverages audiovisual (AV) ground truth and ranked OoI list guidance signals to significantly enhance the process.
[0052] Module M1 - Data Acquisition: Referring more closely to Figure 1, the first module (designated as module M1) comprises data acquisition. In particular, multi-sensor data is collected by a collector 10 (see Figure 2) of a person wearing the hardware device, and in this embodiment, comprises multiple sensors including a world-facing video camera, a collector-facing video camera for capturing line of sight (potentially, attention through pupil dilation, for example), and an audio recorder for recording voice. Such sensor stacks are very popular, for example, in the field of AR. Thus, in one embodiment, the acquisition process can be adapted to most currently available acquisition devices.
[0053] In one embodiment, during data acquisition, at least one computing device provides instructions to a human collector, for example, via a display signal or an audible signal, and performs two levels of articulation and voice caption using, for example, a codeword (preferably a simple codeword; underlined in the following example).
[0054] Level 1 (L1): Contextual description (specific phrase): For example, at the start of the activity, the collector says, This sessionRegarding cooking pasta It becomes something It articulates as follows: In another example, the collector says, Right now I I'm cutting tomatoes. By the way "; Right now I Pouring milk By the way Articulate during the activity to introduce a task or sub-action such as "[...]". These sequences of words are captured by a voice sensor to retrieve and organize contextual information.
[0055] Level 2 (L2): Object identification (specific phrases combined with gaze): For example, before or during an interaction with an object, the collector articulates, "This is a tomato." Here again, this sequence of words is captured by a speech sensor. At this level, the collector is commanded by at least one computing device, for example, via a display signal or an audible signal, and sees the OoI as the collector speaks what it is.
[0056] Figure 2 illustrates an exemplary scenario in an exemplary use case of OoI enhancement and guidance signals regarding a potential data collection process for a “cooking pasta” scene. Such objects of interest can be used in AR applications, for example, in an AR application where a user wearing translucent glasses is instructed by a virtual object blended into a view of the real environment on what steps should be taken to cook pasta. Such an AR application may use an image-based machine learning model trained accordingly on an annotated dataset, as described herein, to ensure that the virtual object within the AR application is blended into the correct context.
[0057] Figure 2 shows an example of an image frame 30 captured by the image sensor of the wearable glasses 11 with potential objects of interest while the user (here, the collector) is viewing a scene in the real environment. In this example, the potential objects of interest are a tomato 31, a knife 32, and a cutting board 33. As described above, at the start of the activity, the collector articulates "This session is going to be about cooking pasta" and "I'm making pasta sauce now," which are analyzed in stage 2-1 of speech analysis 201 (described in more detail below). While articulating "This is a tomato," the collector gazes at the tomato 31, which is detected by the eye-tracking sensor of the wearable glasses 11 and illustrated in an image frame (illustrated by "x" in Figure 2) which is processed in stage 2-1 of eye-tracking mapping 204.
[0058] Module M2 - Preprocessing: Referring again to Figure 1, the second module of the method (designated as module M2) comprises preprocessing. After the raw multisensor data is collected, it is preprocessed before being presented during the annotation stage. In one embodiment, the second module M2 comprises stage 2-1 and / or stage 2-2.
[0059] Stage 2-1: Prepare audio, gaze, and egocentric video. A: Speech analysis: In processing stage 2-1, the method according to an aspect of the present invention is performed by at least one of the above-described computing devices (e.g., wearable glasses 11 and / or computer 12), or different computing devices, on a speech analysis process 201 which analyzes at least a portion of the speech data received from the user into text format and generates a set of terms determined by the analysis process that indicate the user's articulated task or activity while operating the image sensor.
[0060] In one embodiment, audio or speech data can be parsed into text format through a well-established and readily available speech-to-text model as a ready-to-use solution. The parsing yields (i) a word list that can be utilized for scene OoI proposals through a predefined table or classification query (from level 1 of the audio captions), and (ii) object identification for establishing audiovisual (AV) ground truth (from level 2 of the audio captions).
[0061] In the next step, a method according to an aspect of the present invention performs pre-allocation of objects of interest from the set of generated terms determined in the analysis process, using at least one computing device (e.g., wearable glasses 11 and / or computer 12) or different computing devices, associates a probability with each of the pre-allocated objects of interest, and generates a list of scene objects of interest (OoI) associated with at least one image, each having a corresponding object category and associated probability for each of the pre-allocated objects of interest. In particular, in one embodiment, the following is performed:
[0062] (i) Scene OoI List Predefined Tables: In its simplest embodiment, keywords can be used to query a predefined table representing the likelihood that a particular object is present in a given video scene (search engine function). Thus, it provides a context-based method for suggesting objects of interest, as determined by belonging to a task-specific group of objects. The principle behind this engine is a probabilistic mapping between activities and sets of objects, so for any given activity, there exists a set of objects that are likely to be seen near the person performing this activity and are therefore suggested to the annotator.
[0063] For example, the analyzed vocabulary might represent a specific task / activity such as "brew coffee," and this task / activity might be pre-assigned a defined set of objects of interest, such as a coffee maker; coffee pot; water; coffee filter; coffee beans; and ground coffee.
[0064] Expected subject of interest: Further embodiments of the present invention include, for example, learning probabilities between activities and objects based on a deep learning model. In such embodiments, the model can predict what the current object is and what the probability is based on the identified activity. This can also incorporate a time element as the frames progress, resulting in the probability of finding a particular object increasing / decreasing over time.
[0065] In any case, the result of this stage is the provision of a so-called scene OoI list 203. In one embodiment, the resulting scene OoI list can be supplemented by a universal OoI list 202 (see below in more detail).
[0066] (ii) Target identification Using Level 2 audio caption analysis, it's also possible to generate a list of "ground truth" objects present in the scene, such as "This is a tomato / knife / cutting board / grater." Thus, objects can be directly identified (see Stage 2-2 for more details).
[0067] B: Eye-tracking mapping In Stage 2-1, the method according to an aspect of the present invention further comprises a gaze mapping 204 that generates a gaze heatmap from received gaze tracking data. In particular, human gaze is strongly related to the underlying cognitive and attentional processes that drive it. These processes aim to maximize the amount of information extracted from a scene, partly by fixing the gaze (where the eyes stop looking) and partly by regulating attention, which is reflected in the amount of pupil dilation.
[0068] For example, by tracking eye movements and pupil dilation, a gaze map or "heatmap" can be generated by accumulating the gaze positions on the viewed image or video frame (effectively as a two-dimensional histogram) and weighting each point by the degree of pupil dilation. This map then directly highlights potential objects of interest from a cognitive and attentional perspective.
[0069] C: Egocentric Images / Videos In Stage 2-1, a method according to an aspect of the present invention further comprises a panoptic segmentation 205, which in particular processes at least a portion of at least one image with a panoptic segmentation model to generate a set of segmentation masks, wherein each segmentation mask in the set of segmentation masks has its respective assigned object category and panoptic segmentation confidence. In particular, for example, after the standardization of the received data items, the image or video data is processed by the panoptic segmentation model. In one embodiment, the model performs two tasks, namely, creating masks so that foreground objects in a scene can be separated from each other (instance segmentation), and preferably assigning an object category to each of these masks (semantic segmentation). Preferably, such a model is created by training with a large amount of annotated video data to make the model robust to multiple visual scenarios. The masked items represent potential objects of interest and their respective categories.
[0070] D: Universal OoI Scene OoI list 203 includes OoIs directly related to the task at hand. However, in some cases, there may also be items that are generally and almost universally relevant in modern human life. Such items can be classified as “universal OoIs,” and such a list may include, for example, “wallet,” “cell phone,” “glasses,” “keys,” and “headphones.” This universal OoI list 202 may optionally be used to enhance or extend the Scene OoI list according to the respective definitions of OoIs. The universal OoI list 202 is a list of universal objects of interest that have not been processed from at least one image.
[0071] Stage 2-2: Audiovisual Ground Truth In a further aspect of the present invention, an “audiovisual (“AV”) ground truth” 210 is established for each object by combining gaze / gaze on the object with an audio caption (e.g., a Level 2 audio caption as described in Module 1). For example, a trained human collector gazes at a tomato and says, “This is a tomato.”
[0072] In particular, the generated AV ground truth includes location information and respective object categories for at least one object in each captured image from the analyzed audio data and gaze heatmap. These AV ground truths are matched with spatially corresponding masks from panoptic segmentation and their associated categories. Specifically, this involves matching the generated AV ground truth with spatially corresponding segmentation masks from the generated set of segmentation masks and their assigned object categories. As a result of this process, at least one dataset of AV ground truths is generated, each having one or more data labels indicating the respective object category and the matched segmentation mask. The dataset thus generated is configured and appropriate to be used as input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one image-based machine learning model.
[0073] Stage 2-3: Ranked OoI List In a further aspect of the present invention, in addition to or instead of Stage 2-2 and AV ground truth generation, at least one dataset comprising a ranked set of objects of interest is derived. Such a set or list of ranked objects of interest ("ranked OoI list") 220 comprises objects of interest ranked according to at least one aggregate score determined based on a set of generated scene OoI lists, generated heatmaps, and generated segmentation masks. The dataset thus generated is configured and appropriate to be used as input to a computer implementation process for annotating at least one image frame of a real-world scene for input and training of at least one image-based machine learning model.
[0074] In particular, determining the aggregate score is based on the following information sources: - Probability of Scene OoI (from Scene OoI list) - Gaze "Heat" - Panoptic segmentation confidence ("PS confidence") Such determined scores combine multiple information sources available for annotation and propose the most likely OoIs present in the scene, even if they are not included in the panoptic segmentation predictions and / or AV ground truth. In one embodiment, the results are presented as a list of objects paired with aggregated scores, in descending order, for example, where a higher score indicates a higher probability of an object being an OoI in the scene.
[0075] In effect, there are three types of sources.
[0076] 1) Location only: Information associated with a location in the image but not with a category, i.e., gaze.
[0077] 2) Category only: Information that is associated with a specific target category but not with any given location within the image, i.e., a scene OoI list.
[0078] 3) Combination: Information associated with both category and location in space, i.e., a panoptic segmentation mask.
[0079] Intuitively, since objects appearing in all sources are more likely to be relevant, in one embodiment, the first step in ranking is to match categories present on the scene OoI list with categories from panoptic segmentation. The same can be done between the segmentation masks from panoptic segmentation and their respective gaze heatmaps. The aggregated score for these objects ("OoI score") is, in one embodiment, a weighted sum of (i) the object probability for that particular category, (ii) the gaze heat for the corresponding mask, and (iii) the PS confidence score.
[0080] In the second step, combinations of other modalities are evaluated. Specifically, target categories that are not found in the segmentation output but have high scene OoI probabilities receive a score based on this probability multiplied by the weights used in the above scoring. For example, areas with high gaze heat are assigned generic labels ("gaze ROI-A", "-B", ...) and added to a list associated with their weighted scores. The same can be done if there are no segmentation results in the scene OoI list and / or their respective heatmaps.
[0081] In the final step, the proposed objects are then ranked by their OoI scores, which, as mentioned above, can be seen as the probability that a particular object is an OoI in the scene.
[0082] OoI score - Example for the AR application "Cooking Pasta": Table 1 below shows exemplary OoI scores for a simple example based on a cooking scene as illustrated in Figure 2.
[0083] The OoI score was calculated as shown in Equation 1 below: s ooi =w op p op + w g g+w c c formula 1 Here, s ooi is the OoI score, p op is the scene OoI probability, g is the gaze heat, c is the panoptic segmentation confidence, and w op w g w c is the weight coefficient for each respective metric normalized to 1.
Table 1
[0084] Table 1: Exemplary calculation of the ranked OoI list.
[0085] In this example, the OoI score is calculated as a weighted average with a uniform weight of 1 / 3 for each of the three score components (scene OoI probability, gaze heat, PS confidence). In one embodiment, as further described below, "tomato" and "garlic" are corrected at annotation stage 3-1 (AV ground truth checked against the panoptic segmentation mask) and are thus excluded from the table.
[0086] In a further aspect, the present invention also provides a computer-implemented method for annotating at least one image frame of a scene in a real environment, configured for input and training of at least one machine learning model, one embodiment of which is also shown in FIG. 1 according to module M3.
[0087] Figure 3 shows an exemplary preprocessing scheme for obtaining a ranked OoI list 220 based on parsing gaze, panoptic segmentation, and speech into a scene OoI list, following the example shown in relation to Figure 2 and using exemplary probability values as described above in relation to Table 1. In speech analysis 201, a natural language processing (NLP) model can generate terms, e.g., “pasta,” “cutting board,” “garlic,” “tomato,” and “knife” from the collector’s articulated phrases in the speech captions described above. The generated terms are determined by the NLP model (which may, for example, be pre-trained accordingly) to be related to “cooking pasta” and “pasta sauce.” From the NLP model, an exemplary scene OoI list 203 is generated.
[0088] Module M3 - Notes: An annotation process according to one aspect of the present invention comprises one or both of stages 3-1 and 3-2 for each image frame to be annotated. In one embodiment in which the annotation process comprises both stages 3-1 and 3-2 for each image frame, the annotator first encounters stage 3-1 and proceeds to stage 3-2 only after stage 3-1 is completed. After stage 3-2 is completed, the image frame can be considered fully annotated, and the annotator can move on to the next frame.
[0089] Stage 3-1 - AV Ground Truth Verification In the first stage of annotation, the human annotator reviews or corrects the segmentation output associated with each available AV ground truth (as described in Stage 2-2). Specifically, the annotator is presented with two information sources.
[0090] 1) Annotations extracted from AV audio components that correspond to the regions in the image that were focused on during the recording of the annotations.
[0091] 2) Categories and masks generated by a panoptic segmentation model that spatially overlap with AV annotations.
[0092] Based on these two sources, the annotator checks whether the segmentation mask and category are correct for each AV annotation. If there is an error in the mask or if a mask is unavailable, the annotator corrects the mask by reassigning the incorrect pixels within the corresponding mask, creating a new mask, or clearing the false mask. If the category is incorrect, the annotator corrects the label. Once all AV ground truth annotations have been corrected, the annotator moves to stage 3-2, if present, or to the next frame accordingly.
[0093] Following the "cooking pasta" scenario, the collector recorded an AV ground truth annotation about "tomatoes," which is then presented to the annotator along with the gaze locations and category labels extracted from the audio components during the annotation. Simultaneously, any overlapping panoptic segmentation categories and masks with the annotation are displayed, and the annotator reviews or modifies the AV category labels and panoptic segmentation masks accordingly.
[0094] Stage 3-2 - Ranked OoI Reviews In the second stage of annotation, based on the ranked OoI list 220, the human annotator is presented with the ranked OoI list 220, preferably with only those objects that are not present in the AV ground truth, if Stage 3-1 is included in a particular pipeline embodiment. Features with high scene OoI probability, gaze heat, and PS confidence scores receive higher scores. The scores of less likely features gradually decrease (as described in Stage 2-3). According to the present invention, in contrast to other known methods, objects that were not detected (not seen and not commented on) by the human collector are still listed for annotator review if they are relevant to the task or universally relevant. The annotator goes through the presented ranked OoI list, which is supplemented with highlighted masks where available, and adds or modifies (improves) object categories and / or object masks as per the specific case (see differences in presentation in Stage 2-3).
[0095] By the end of Stage 3-2, the annotator has traversed all possible OoIs in the current image frame, which is determined based on the context and experience of the data collector, the operator performing the task, rather than requiring out-of-context decisions.
[0096] Video annotations Methods according to aspects of the present invention are suitable for application to image streams such as video, even while annotating still images. This is achieved by aggregating temporal information from multiple frames into annotated keyframes. Because this method relies on the fact that scenes containing more OoIs remain in the field of view for a longer period, the calculation system can aggregate the AV ground truth into a single frame (in the case of Stage 3-1) and also aggregate information about the ranked OoI list (Stage 3-2), thus minimizing the total number of frames that need to be annotated while still covering all OoIs in the scene. The interval and frequency of keyframes can be adjusted to suit the quality of the video and propagation method. Thus, if at least one image acquired in data acquisition is part of a video stream, the at least one image frame to be annotated (e.g., a keyframe) can correspond to an image acquired by the image sensor in module M1, which is before or after at least one image acquired in module M1 and preprocessed in module M2.
[0097] Note - Example: "Cook pasta" Figure 4 illustrates one embodiment of Stage 3-1 for verifying AV ground truth. According to Figure 4, at least one computing device, such as a personal computer with a display device, receives at least one image frame 301 of a scene in a real environment to be annotated, and provides at least one image frame to the display device for presentation to the user (here, the annotator). The computer further receives at least one dataset of audiovisual ground truth generated according to Stage 2-2 as described above. In principle, ground truth generated by a scheme different from Stage 2-2 may also be appropriate. In this example, the annotator is presented with a segmentation mask ("target mask") labeled "garlic" determined by a panoptic segmentation model in panoptic segmentation 205, which overlaps with the AV annotation "tomato" from speech analysis 201 so as to overlap with image frame 301.
[0098] The computer then receives instructions from the annotator to add or refine one or more of the presented target categories and segmentation masks to image frame 301. In this example, the annotator identifies the target category "tomato" and refines / modifies the segmentation mask, which was pre-labeled "garlic," to cover "tomato" (at least more strictly), as shown in the annotated image frame 302. The image frame is then augmented with corresponding annotation information for at least one annotated dataset configured for input and training to at least one machine learning model. In another example (not shown) where the segmentation masks and / or target categories are correctly determined by panoptic segmentation 205, the annotator verifies the audiovisual ground truth by identifying one or more of the respective target categories and presented segmentation masks.
[0099] Figure 5 shows one embodiment of Stage 3-2, illustrating a ranked OoI review in relation to an example AR application called "Cooking Pasta."
[0100] As shown in Figure 5, at least one computing device, such as a personal computer having a display device, receives at least one image frame 501 of a real-world scene to be annotated and provides at least one image frame to the display device for presentation to the user (here, the annotator). The computer also receives at least a portion of the set of ranked objects of interest (here, the ranked OoI list 220) provided in stages 2-3. At least a portion of the ranked OoI list 220 is presented on the display device while displaying at least one image frame 501 to be annotated on the display device.
[0101] The first annotation task presented to an annotator with a ranked OoI list of 220 is the object with the highest score, i.e., "tomato". If information from Stage 3-1 is included and "tomato" is confirmed as AV ground truth, the first annotation task presented to an annotator with a ranked OoI list of 220 is the object with the highest score that does not exist in AV ground truth, i.e., "knife". Each object is associated with a segmentation mask and object category proposal from a panoptic segmentation model, and the annotator is asked to refine the proposed annotation.
[0102] For the next element in the ranked OoI list 220, the Region of Interest (ROI) arising from the heatmap ("Eyes-View ROI-A"), the annotator is presented with the highlighted ROI in each image frame. For the task, the annotator is instructed by the computing device to create a segmentation mask, in this case "Cutting Board," and find a matching target category label from a quick selection in the ranked OoI list 220 (associated scores indicate a high probability). As a result, the "Cutting Board" entry is blacked out and considered complete (see annotated image frame 502).
[0103] The last two elements are rejected by human annotators because they are either not included in image frame 501 ("pasta") or they represent an incorrect mask proposal by the panoptic segmentation model ("sword").
[0104] The annotators then (via computer) add or refine their respective target categories and / or segmentation masks to at least one image frame, and augment at least one image frame with corresponding annotation information to at least one annotated dataset configured for input and training to at least one machine learning model.
[0105] Accordingly, according to aspects of the present invention, the provision of eligible objects of interest that can be presented to a human annotator for the annotation and training of image-based machine learning models is customized, in particular, to solve the difficulties of large-scale image or video annotation for the training of task-specific image-based machine learning models.
[0106] The definitive method for OoI according to the present invention encompasses subjects in one or more of the following categories: 1) Objects related to the current task (contextual relevance); 2) Prominent objects (high contrast); and / or 3) An object that is on the list of universal OoI (ubiquitous associations).
[0107] The method according to the present invention is designed to automate the decision-making process regarding what is eligible as an OoI and what is not. Specifically, the method reduces the selection range of a human annotator and thus reduces ambiguity and deliberation time when determining which objects within a given image frame or keyframe of a video stream are OoIs. Only objects relevant to the hand-held annotation task are annotated. By reducing the mental burden on annotators and their arbitrary selection through guidance signals, the method can improve the quality and speed of task-specific annotations and thus effectively train image-based machine learning models and improve their potential implementation in applications such as augmented reality or autonomous driving.
[0108] A particular aspect of the method according to the present invention relates to the collection of (multisensor) egocentric data. A solution is provided for collector-assisted OoI image or video annotation that leverages the collector's top-down or goal-oriented cognitive processing. Specifically, the collector can be instructed on how to provide custom audio captions and object selection through fixation. The collector is the one that performs the task and interacts with the actual scene, thus defining what the OoI is, and this information is then processed through module M2 and converted into guidance signals for a human annotator. Thus, the annotation process according to the present invention is unique in that it establishes a bridge between the collector and the annotator and provides context for both the task and the scene by leveraging existing sensors for higher quality annotation.
Claims
1. A computer implementation method for generating at least one dataset associated with one or more objects of at least one image of a real-world scene, configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model, The system includes receiving at least one image of a real-world scene captured by a user-operated image sensor via at least one computing device, The at least one computing device receives, while the user is capturing the at least one image, voice data from the user detected by at least one voice sensor, and eye-tracking data of the user detected by at least one human eye-tracking sensor. Performing a speech analysis process by at least one computing device, which involves analyzing at least a portion of the received speech data from the user into text format and generating a set of terms determined by the analysis process that indicate the user's articulated tasks or activities while the image sensor is operating; The process involves pre-assigning objects of interest from the set of terms generated in the analysis process, associating a probability with each of the pre-assigned objects of interest, and generating a list of scene objects of interest associated with at least one image, each of which has a target category and associated probability for each of the pre-assigned objects of interest. The process involves generating a gaze heatmap from the received gaze tracking data, The process involves processing at least a portion of the at least one image using a panoptic segmentation model to generate a set of segmentation masks, wherein each segmentation mask in the set of segmentation masks has its own assigned target category and panoptic segmentation confidence level. The method involves generating at least one dataset comprising a set of ranked interests, each of which interests are ranked according to at least one aggregate score determined based on the generated scene interest list, the generated heatmap, and the generated segmentation mask set, wherein the at least one dataset is configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model. A method that includes [a certain feature].
2. The method according to claim 1, wherein the ranked set of interests is configured to be used in a computer implementation process for annotating at least one image frame of a real-world scene for input and training of at least one machine learning model for an augmented reality application.
3. To provide a list of unprocessed universal objects of interest from the at least one image, and to extend the scene object list with the list of universal objects of interest. The method according to any one of claims 1 or 2, further comprising:
4. To generate an audiovisual ground truth from the user's received audio data and eye-tracking data, comprising the object category and location information for at least one object in the at least one image, The generated audiovisual ground truth is compared with the spatially corresponding segmentation masks of the generated set of segmentation masks and their assigned target categories. The method according to any one of claims 1 to 3, further comprising:
5. The method according to any one of claims 1 to 4, wherein the at least one aggregate score is determined based on (i) the probability of each of the objects of interest included in the scene object of interest list, (ii) the gaze heat for each of the corresponding segmentation masks having object categories corresponding to each of the objects of interest, and (iii) the weighted sum of the panoptic segmentation confidences for each of the corresponding segmentation masks having object categories corresponding to each of the objects of interest.
6. Determining the aforementioned at least one aggregate score is Matching the target categories included in the aforementioned scene interest list with each target category in the set of generated segmentation masks, Matching the segmentation masks of the generated set of segmentation masks with the gaze heatmap, optionally determining the at least one aggregate score, includes assigning the score from the scene interest list to each interest category that could not be matched with the interest categories of the generated segmentation masks. The method according to any one of claims 1 to 5, comprising:
7. The method according to claim 6, further comprising adding to the scene interest list at least one region of interest arising from gaze heat exceeding a predetermined threshold that was not matched with a generated segmentation mask, wherein the at least one region of interest is assigned an undefined interest category.
8. A computer implementation method for generating at least one dataset associated with one or more objects of at least one image of a real-world scene, configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model, The system includes receiving at least one image of a real-world scene captured by a user-operated image sensor via at least one computing device, The at least one computing device receives, while the user is capturing the at least one image, voice data from the user detected by at least one voice sensor, and eye-tracking data of the user detected by at least one human eye-tracking sensor. The process involves using at least one computing device to analyze at least a portion of the received audio data from the user and to perform an audio analysis process that generates analyzed audio data indicating the user's articulated tasks or activities or areas of interest in the scene while the image sensor is operating, The process involves generating a gaze heatmap from the received gaze tracking data, The process involves processing at least a portion of the at least one image using a panoptic segmentation model to generate a set of segmentation masks, wherein each segmentation mask in the set of segmentation masks has its own assigned target category. From the analyzed audio data and gaze heatmap, generate an audiovisual ground truth comprising location information for at least one object in the at least one image and the respective object category. The method involves matching the generated audiovisual ground truth with the spatially corresponding segmentation masks and their assigned target categories in the set of generated segmentation masks, and generating at least one dataset of the audiovisual ground truth comprising at least one data label indicating each of the target categories and the matched segmentation mask, wherein the generated at least one dataset is configured for input to a computer implementation process for annotating at least one image frame of a real-world scene for input to and training of at least one machine learning model. A method that includes [a certain feature].
9. The method according to any one of claims 1 to 8, wherein the image sensor and optionally the human gaze tracking sensor are part of a wearable computing device, in particular wearable glasses, worn by the user while viewing the scene in the real environment.
10. A computer implementation method for annotating at least one image frame of a real-world scene, configured for input and training of at least one machine learning model, The system includes receiving at least one image frame of a real-world scene to be annotated by at least one computing device, and providing the at least one image frame to a display device for presentation to the user, The at least one computing device receives at least a portion of the ranked set of interests provided according to the method of any one of claims 1 to 7 or 9, While displaying the at least one image frame to be annotated on the display device, presenting at least a portion of the received ranked set of interests on the display device, Receiving instructions from a human annotator to the at least one computing device to add or improve one or more of the target categories or segmentation masks to the at least one image frame, and extending the at least one image frame with corresponding annotation information to at least one annotated dataset configured for input and training to at least one machine learning model, A method that includes [a certain feature].
11. At least a portion of the ranked set of interests is presented on the display device after receiving an instruction from the annotator to instruct the at least one computing device to verify or correct any stored audiovisual ground truth. Optionally, at least a portion of the ranked set of interests presented on the display device comprises only those interests not included in the verified or modified audiovisual ground truth. The method according to claim 10.
12. A computer implementation method for annotating at least one image frame of a real-world scene, configured for input and training of at least one machine learning model, The system includes receiving at least one image frame of a real-world scene to be annotated by at least one computing device, and providing the at least one image frame to a display device for presentation to the user, The at least one computing device receives at least one dataset of audiovisual ground truth generated according to the method of claim 8 or 9, Receiving instructions from a human annotator to the at least one computing device to instruct it to add or improve one or more of the respective target categories and segmentation masks to the at least one image frame, and extending the at least one image frame with corresponding annotation information to at least one annotated dataset configured for input and training to at least one machine learning model, A method that includes [a certain feature].
13. Presenting one or more segmentation masks from the set of segmentation masks on the display device so as to overlap with at least one image frame, Receiving instructions from the annotator to instruct the at least one computing device to verify or modify the at least one dataset of audiovisual ground truth by verifying or improving one or more of the aforementioned target categories and presented segmentation masks. The method according to claim 12, further comprising:
14. The method according to any one of claims 10 to 13, wherein the at least one image frame corresponds to the at least one image captured by the image sensor, or, if the at least one image is part of a video stream, corresponds to an image captured by the image sensor before or after the at least one image captured by the image sensor.
15. A computer program, wherein when the program is executed by at least one computing device, it includes instructions causing the at least one computing device to perform the method according to any one of claims 1 to 14.