Online incremental real-time learning for tagging and labeling data streams for deep neural networks and neural network applications.

JP7914264B2Active Publication Date: 2026-09-01NEURALA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025027950
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-03-17
Filing Date
2025-02-25
Publication Date
2026-09-01
Estimated Expiration
2038-03-19

Smart Images

  • Figure 0007914264000001
    Figure 0007914264000001
  • Figure 0007914264000002
    Figure 0007914264000002
  • Figure 0007914264000003
    Figure 0007914264000003
Patent Text Reader

Abstract

To provide a method and system for tagging sequences of images.SOLUTION: A smart tagging utility automatically learns tags and tag images using a feature extraction module and a fast learning classifier module. The feature extraction module and fast learning classifier module can be implemented as an artificial neural network that associates labels with features extracted from images and tags similar features or other images from the image by the same label. The smart tagging utility can further learn from user adjustments to suggested tagging. This reduces tagging time and errors.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-Reference to Related Patent Applications This application claims the priority benefit under 35 U.S.C. § 119(e) of U.S. Patent Application No. 62 / 472,925 filed on March 17, 2017, the content of which is incorporated herein by reference in its entirety.

Background Art

[0002] Conventional deep neural networks (DNNs), including convolutional neural networks (CNNs) that include multiple layers of neurons disposed between an input layer and an output layer, require thousands or even millions of repeated cycles for training on a specific data set. Before this training can occur, all images in the data set must be tagged by a human user. The tagging process may require labeling the entire image for classification or labeling of individual regions of each image, as specific objects for classification and detection / segmentation of individual objects.

[0003] Conventional image tagging is a slow and time-consuming process. A human views images on a computer, tablet, or smartphone, identifies one or more objects in the image, and tags those objects with descriptive tags (e.g., “tree”, “house”, or “car”). The main difficulties in manually tagging objects of interest include slowness and susceptibility to human error caused by distraction and fatigue. These problems give rise to two types of issues: first, data preparation for training can take unacceptably long anywhere outside of academic settings; and second, improperly tagged data means the quality of tagging directly affects the quality of subsequent learning, to the extent that a DNN cannot reach acceptable performance standards. [Overview of the Initiative]

[0004] An embodiment of the technology of the present invention includes a method and system for tagging a sequence of images. An exemplary method includes a user tagging a first instance of an object representation in a first image of the sequence of images via a user interface. At least one processor learns the object representation tagged by the user in the first image and tags a second instance of the object representation in the sequence of images. The user performs tagging and / or position adjustments on the second instance of the object representation created by the processor. The processor then tags a third instance of the object representation in the sequence of images based on the adjustments.

[0005] A second instance of an object's representation may be in the first image of a sequence of images or in another image of the sequence of images.

[0006] In some cases, the user may perform tagging and / or position adjustments on a third instance of the object representation created by the processor, and the processor may tag a fourth instance of the object representation in a sequence of images based on the tagging and / or position adjustments on the third instance of the object representation.

[0007] An example of this method may include classification via a fast learning classifier running on a processor, where the object representation is the representation of the object tagged by the user in the first image. In this case, tagging a third instance of the object representation may include extracting a convolutional output representing the features of the third instance of the object representation using a neural network operably coupled to the fast learning classifier. The fast learning classifier classifies the third instance of the object representation based on the convolutional output.

[0008] Examples of this method may include tagging a second instance of an object representation by extracting a convolutional output that represents the features of the second instance of the object representation using a neural network running on a processor, and classifying the second instance of the object representation based on the convolutional output using a classifier operablely coupled to the neural network.

[0009] A system for tagging a sequence of images may include a user interface and at least one processor operablely coupled to the user interface. During operation, the user interface allows the user to tag a first instance of an object representation in a first image of the sequence of images. The processor also learns the user-tagged object representation in the first image and tags a second instance of the object representation in the sequence of images. The user interface allows the user to perform tagging and / or position adjustments on the second instance of the object representation created by at least one processor, and the processor tags a third instance of the object representation in the sequence of images based on the adjustments.

[0010] Other embodiments of this technology include methods and systems for tagging objects in a data stream. An exemplary system includes at least one processor configured to implement a neural network and a fast learning module, and a user interface operablely coupled to the processor. In operation, the neural network extracts a first convolutional output from a data stream containing at least two representations of a first category of objects. This first convolutional output represents the features of the first representation of the first category of objects. The fast learning module classifies the first representation into a first category based on the first convolutional output and learns tags and / or the positions of the first representation of objects based on user adjustments. The user interface also displays the tags and / or positions for the first representations and allows the user to perform adjustments to the tags and / or positions of the first representations.

[0011] In some cases, the tag is the first tag and the position is the first position. In these cases, the neural network may extract a second convolutional output from the data stream. This second convolutional output represents the features of the second representation of the first category of the object. Then, in these cases, the classifier classifies the second representation into the first category based on the second convolutional output and the adjustment of the tag and / or position for the first representation. The user interface may display the second tag and / or second position based on the first category.

[0012] If necessary, the classifier can determine a confidence level that the tags and / or positions of the first representation are correct. The user interface may display this confidence level to the user.

[0013] If the object is the first object and the tag is the first tag, the classifier can learn a second tag for the second category of the object represented in the data stream. In these cases, the neural network can extract a subsequent convolution output from a subsequent data stream containing at least one other representation of the second category of the object. This subsequent convolution output represents the features of the other representation of the second category of the object. Based on the subsequent convolution output and the second tag, the classifier classifies the other representation of the second category of the object into the second category. The user interface also displays the second tag. In these cases, the neural network can extract the first convolution output by generating multiple segmented sub-areas of the first image in the data stream and coding each of the segmented sub-areas.

[0014] A further embodiment of this technology includes a method for tagging multiple instances of an object. An example of this method includes using a feature extraction module to extract a first feature vector representing a first instance of an object of multiple instances. The user tags the first instance of the object with a first label via a user interface. The classifier module associates the first feature vector with the first label. The feature extraction module extracts a second feature vector representing a second instance of the object of multiple instances. The classifier module calculates the distance between the first and second feature vectors and performs a comparison of the difference with a predetermined threshold to classify the second instance of the object based on the comparison. If necessary, the second instance of the object may be tagged with the first label based on the comparison. The classifier module may then determine the confidence of the classification based on the comparison.

[0015] Naturally, all combinations of the aforementioned concepts and any additional concepts discussed in more detail below (assuming such concepts are not contradictory) are considered to be part of the subject matter of the invention disclosed herein. In particular, all combinations of the subject matter described in the last claims of this disclosure are considered to be part of the subject matter of the invention disclosed herein. Naturally, any terms explicitly used in any disclosure incorporated by reference must be given meanings that are most consistent with the specific concepts disclosed herein. [Brief explanation of the drawing]

[0016] Those skilled in the art will understand that the drawings are for illustrative purposes only and are not intended to limit the scope of the subject matter of the invention described herein. The drawings are not necessarily to a fixed proportion, and in some examples, various aspects of the subject matter of the invention disclosed herein may be exaggerated or enlarged in the drawings to facilitate the understanding of different features. In the drawings, similar reference letters generally mean similar features (e.g., functionally similar and / or structurally similar elements).

[0017] [Figure 1A] Figure 1A illustrates the online incremental real-time learning process for tagging and labeling data streams using a processor-implemented convolutional neural network and a fast learning classifier.

[0018] [Figure 1B] Figure 1B shows an operational workflow for tagging objects in one or more images (e.g., video or image frames in a database) using the smart tagging system of the present invention.

[0019] [Figure 2] Figures 2A to 2D show the operational workflow of the invention's smart tagging system from the user's perspective.

[0020] [Figure 3] Figure 3 shows a screenshot of an exemplary implementation after a user tag has tagged a first instance of a first object.

[0021] [Figure 4] Figure 4 shows a screenshot of an exemplary implementation after the system has learned from the previous step and provided a suggestion to a user for tagging a second instance of the first object.

[0022] [Figure 5] Figure 5 shows a screenshot of an exemplary implementation after the user has corrected the proposal generated by the system.

[0023] [Figure 6] Figure 6 is a schematic outline of an implementation of a smart tagging utility that can automatically tag objects similar to a user-selected object within the same frame or different frames of video or other image data. DETAILED DESCRIPTION OF EMBODIMENTS OF THE INVENTION

[0024] The capability of backpropagation-based neural networks, including deep neural networks and convolutional neural networks, depends on the availability of large amounts of training and testing data to enhance and verify the performance of these architectures. However, generating a large amount of labeled or tagged data is a manual, cumbersome and high-cost process.

[0025] This application relates to automatically tagging, annotating, or labeling objects of interest that are identified and to be located in a data stream (e.g., red / green / blue (RGB) images, point cloud data, IR images, hyperspectral images, or a combination of these or other data). One use of these tagged data streams is to create training and ground truth data that is used during the training and testing of controlled neural networks, including backpropagation-based deep neural networks that use thousands of images for proper training. The terms “annotating,” “labeling,” and “tagging” are used interchangeably in this document. The term “fast learning” is used in this application to describe a method that, unlike backpropagation, can be updated incrementally without requiring the entire system to be retrained on all of the previously presented data, for example, from a single example. Fast learning is in contrast to “batch” training, which requires the iterative presentation of a large data corpus even to learn a single new instance of an object.

[0026] The techniques described herein accelerate and improve the accuracy of manually labeling data by introducing automated, real-time, high-speed learning steps. This high-speed learning proposes candidate tags for subsequent occurrences of a tagged item in the same image as the first instance of the tagged item, subsequent images, or both. Conversely, current techniques for tagging data rely on a human labeling each of the objects of interest in each frame (e.g., frames in a video stream).

[0027] The present invention introduces interactive assistance to the user. This interactive assistance takes the form of a neural network-based automated assistant, also called a smart tagging system or utility, which has the ability to rapidly learn new tags during the human process of labeling data. The automated assistant labels or suggests labels for new data, receives corrections from the user to the suggested labels, and iteratively refines the quality of the automated labeling as the user continues to correct any errors made by the automated assistant. This has the advantage that the system does more work on itself and moves what it has learned away from the user, allowing the user to focus on new objects of interest and verify the automated tags. As a result, the tagging process becomes faster as more images are processed. Our research shows up to a 40% improvement in tagging speed compared to a simple human tagger; in other words, assistance in the invention's smart tagging utility is 40% faster than manual tagging by someone who has never tagged images before.

[0028] Figure 1A shows an example of this tagging process 10. 1. The user tags the first instance of an object in Frame 1 (for example, in Figure 1A, 12). For example, the user can draw a boundary polygon and tag an object on Frame 1, such as a tree. 2. The classifier, running on one or more processors of the smart tagging utility, learns the first instance of an object (for example, in Figure 1A, 14) by associating features representing the tagged object with user-defined tags. In the example of Step 1 (tree tagging), the classifier learns the tree immediately after tagging. 3. The processor can tag subsequent instances of an object (for example, in Figure 1A, 16). For example, a convolutional neural network run by the processor can extract features from a frame (for example, in Figure 1A, 18). The classifier (for example, in Figure 1A, 20) classifies the extracted features based on their similarity to user-defined tags associated with the extracted features. For example, the processor can tag other trees in frame 1. If different, the neural network extracts features from other trees in frame 1, and the classifier classifies them appropriately. Each tree has a confidence value associated with its boundary polygon and can be distinguished from other, manually labeled objects through some visual representation (e.g., boundary polygon color, dotted or dashed boundary polygon contour, etc.). The confidence value can be assigned by several methods for classifying or tracking objects of interest, where the confidence value can be a scalar between 0 and 1, for example, where 1 indicates absolute confidence that the object of interest belongs to a particular class. A particular object may have a probability distribution associated with a hierarchy of classes. For example, an object can be classified as both a tree and a plant at the same time. 4. If necessary, in 22, the user adjusts the labels or edits the position or shape of the machine-generated boundary polygons in Frame 1, and the classifier uses rapid learning to update its knowledge and update the suggestions for Frame 1 objects that have not yet been validated by the user. This allows the tagged objects in Frame 1 to be automatically updated as needed until the objects are sufficiently tagged in Frame 1. 5. The user loads frame 2. 6. The smart tagging utility automatically tags objects that appear in Frame 2 if they were learned in Frame 1. In other words, the smart tagging utility automatically tags objects in Frame 2 if any of the tags in Frame 1 tag an object in Frame 2, taking user adjustments into account. 7. If necessary, the user may add new objects that are not present in frame 1 or proceed to the next frame. In either case, the process described in steps 1-7 is repeated until the desired objects in the provided image are tagged or the user terminates the process.

[0029] This method can be applied to any area of ​​interest, such as rectangular, polygonal, or pixel-based tagging. In polygonal or pixel-based tagging, object shading is depicted in more detail than the image area, and background areas can be included in the tagged object, increasing the "target pixels" count for rectangular or polygonal tags.

[0030] By employing various technologies, a high-speed learning architecture can be introduced, including, for example, • A combination of DNNs with a fast classifier, such as another neural network, support vector machine, or decision tree, which provides a set of functions that the DNN acts as input to the fast classifier. • A feature detection and tracking process that can be initialized quickly on a target subset of an image (e.g., a keypoint tracker), This may include, but is not limited to, any combination of the technologies described above.

[0031] The technology of the present invention enables the efficient and cost-effective preparation of datasets for training backpropagation-based neural networks, particularly DNNs, and more generally, streamlines the learning of parallel, distributed equation systems for purposes such as controlling autonomous vehicles, drones, or other robots in real time.

[0032] More specifically, embodiments of the technology of the present invention improve or replace the process of manually tagging each occurrence of a particular object in a data stream (e.g., frames) or a sequence of frames, and the process of optimally selecting objects by reducing the manual work and costs associated with dataset preparation.

[0033] [Incremental real-time learning process for tagging and labeling data streams] Figure 1B shows a flowchart of the operation process of the smart tagging system of the invention. In this example, the system operates on images, but it can operate equally well on any data that may have human-readable 2D representations of different objects or regions of interest to be tagged. The user starts the system (100) and loads a set of images or a video file into the system. In the case of a video file, an additional processing step is to decompose the video into a sequence of keyframes to reduce the number of redundant images. This decomposition can be done automatically using keyframe information encoded in the video file, manually by the user, or both. Once the sequence of images is ready, the system checks if there are any untagded frames (105), and if there are none remaining, it terminates (110). If there are frames to be tagged, the system loads the first frame (115) and performs feature extraction on this frame (120). In other words, the system loads the first frame and a neural network on one or more processors in the system extracts the convolutional output to create a feature vector representing the features on the first frame. Details of the feature extraction process are outlined in the corresponding sections below.

[0034] Simultaneously, the system checks whether it already possesses knowledge (125). This knowledge may include a previously learned set of associations between extracted feature vectors and their corresponding labels. If the previously learned knowledge includes associations between extracted feature vectors and their corresponding labels, a classifier running on one or more processors classifies the extracted feature vectors by their respective labels. To classify the extracted feature vectors, the system performs feature matching. For example, the system compares the extracted feature vectors to features (and feature vectors) known to the system (e.g., previously learned knowledge). The comparison is performed based on a distance metric (e.g., the Euclidean norm of the related feature space) that measures the distance in the feature space between the extracted feature vectors and features known to the system. The system then classifies objects based on the difference. If the difference between the extracted feature and the feature for the first object in the system's existing knowledge is less than a threshold, the system classifies the feature as a potential first object. The actual distance or difference between the distance and the threshold may be, or can be used to derive, a confidence value indicating the quality of the match.

[0035] The system can store such knowledge after a tagging session, and this stored knowledge can be loaded by the user at the start of a new session. This can be particularly useful if the set of images currently being tagged comes from the same domain that the user and the system have previously tagged. If no knowledge is preloaded, the system displays a frame to the user (130) and waits for user input (135). For the first frame, which the system has no prior knowledge of, the user manually tags one or more instances of objects in the first image via the user interface (140). Once the user tags the first instance of the first object in the image, the system learns the features of the tagged object and the associated labels (145). Details of the fast learning classifier involved in this stage are described in the corresponding sections below.

[0036] After the system learns the features of tagged objects in a frame, it processes the frame to check if it can find other instances of the same object in the frame (150). Note that if the system has preloaded knowledge from a previous session, it may attempt to find known objects (150) before the first frame is displayed to the user through the same process. For instances of objects the system has found in the image, the system creates a bounding polygon with attached labels (155), superimposes the bounding polygon onto the image (160), and displays the image with the superimposed bounding polygon and tags to the user (130). In some examples, if the user is not satisfied with the tags the system creates, the user can adjust the tags via the user interface. The classifier learns the adjusted tags and updates its knowledge. The inner loop (170) then continues to allow the user to add new objects and revise the system's predictions until the user is satisfied with the tagging for this frame. Once the user is satisfied, the system checks if there are still frames to tag (105), and if so, loads the next frame (115), performs feature extraction (120), and re-enters the inner loop (170). Note that in this case, the system has prior knowledge from at least the previous frame, so the inner loop (170) enters through the lower branch of the workflow, and the system makes predictions (150, 155, 160) before displaying the frame to the user (130).

[0037] The entire process continues until the image is tagged or the user finishes the workflow. Before finishing, the system can save the knowledge gained from the user during this session so that it can be reused in the next session.

[0038] [Operation procedure from the user's perspective] Figures 2A-2D illustrate the operational workflow of Figure 1B from the user's perspective. The system can provide the user with several input methods (e.g., mouse (200) or touchscreen (210)) for controlling system operation and for image tagging, as shown in Figure 2A. The sequence of frames to be tagged is loaded by the user either as a directory on the local computer where the system is installed containing images, a remote directory containing images, or a plain text or markup document containing local or remote filenames for the images, or as one or more video files. In the latter case, the video is divided into a sequence of keyframes by the system. In any case, the system is ready to operate once the sequence of frames to be processed (220) is defined and the first frame has been loaded into the system, resulting in the user viewing the image on the screen. If the user so desires, they can load a trained version of the system described herein to further speed up the tagging process and even reduce the manual portion of the task.

[0039] In Figure 2B, the user selects a tree in frame 1, for example, one with a rectangular bounding box (230). Alternatively, other methods of selecting a candidate region are possible; for example, the user can draw polygons as shown in Figures 3-5. The system learns combinations of features within the bounding polygon and associates those combinations of features with labels provided by the user (e.g., "tree"). The learning process is a fast learning procedure that completes in less than 100 milliseconds in the exemplary implementation. Because the process is so fast, the system can provide the user with suggestions about other untagged instances of the same object in the frame. In this exemplary implementation, these suggestions also take less than 100 milliseconds to compute, so very seamlessly for the user, the system suggests tagging other instances (240) of the same object in the image immediately after the user has finished tagging one object.

[0040] Firstly, especially if the system is not pre-loaded with previously trained data, the suggestions (240) generated by the system may be quite far removed from the user's perspective. Secondly, the user can reject completely incorrect predictions, adjust incorrect labels for the correct boundary polygons, adjust the boundary polygons suggested by the classifier for the correct labels as shown in Figure 2C, and / or accept the correct suggestions. To simplify the process, the boundary polygons suggested by the classifier (240, 260) can be shown as dotted lines, and the user's original tagging (230, 270) and accepted corrections (250) can be shown as dashed or solid lines as shown in Figures 2C and 2D. The accepted corrections constitute another input that the system can use to further refine and retrain a specific class of objects to improve further suggestions for tagging.

[0041] Next, the process continues for subsequent frames, as shown in Figure 2D, where new objects (270) can be tagged by the user, and previously tagged objects (260) can be visualized by their associated class and confidence value, indicating how confident the system is in its suggestion of the objects it marks in the image. As the system is increasingly trained, the user interaction process shifts from primarily correcting suggestions to accepting suggestions made by the system, which is significantly less labor-intensive and much faster than manually tagging images. For simple taggers, the overall tagging speed on the test dataset increased by up to 40%, and for expert taggers, a less abrupt but significant increase was observed.

[0042] Figures 3-5 provide screenshots from the graphical user interface of an illustrative smart tagging system. The user launches the manual tagging tool (300), selects a label from a list of existing labels, or creates a new label, and draws a polygon around the first cyclist (310). The color of the polygon represents the label selected in this example. Since this polygon was created manually, it automatically obtains the approved status (320). The results of these actions are shown in Figure 3. The smart tagging system's fast classifier then learns the features within the polygon (310) and associates them with the user's label for the first cyclist. Next, the system looks at the features of the entire frame and attempts to find another cyclist in the image and create a polygon around it (400), showing the user the polygon superimposed on the other cyclist as shown in Figure 4. Note that since the polygon is created by the system, it is marked as "proposed" (410) along with a confidence value for the proposal. Finally, Figure 5 shows that the user has restarted the manual correction tool (300) to update the polygon for the second cyclist (500), and that it now has the approved label (510). The user can now tag other objects in the frame as needed, use the arrow buttons (530) to go to the next frame, or complete the session by pressing the "Tagging Complete" button (540).

[0043] [Smart Tagging System] Figure 6 shows a smart tagging system, which is a hardware configuration for implementing the smart tagging utility described herein. Sensory information (e.g., RGB, infrared (IR), and / or LiDAR images) originates from robots, drones, autonomous vehicles, toy robots, industrial robots, and / or other devices (610). Alternative embodiments may derive data from isolated or networked sensors (e.g., cameras, LiDAR 620). Furthermore, the data may be organized within an existing database (630). The data is transmitted to a computing device equipped with a user interface (650) that a user (640) can use to visualize the data (images), tag the images, visualize the tagged images, and adjust the tags as needed. The computing device can be a mobile device (e.g., a smartphone, tablet, laptop), desktop, or server, comprising one or more processors or processing units (670), such as a digital signal processing unit, a field-programmable gate array (FPGA), a central processing unit (CPU), a graphics processing unit (GPU), and / or a combination of these processors sufficient to perform the workflow shown in Figure 1B. These processing units (670) may be used to implement feature extractors and classifiers that learn and automatically apply and adjust tags. An architecture including several software modules may be stored in memory (680) and loaded into the processing units (670) for execution. Tagged images may be visualized in a UI (650) for display to a user (640) and stored in a separate database (660).

[0044] [Feature Extraction Module] The feature analysis module (Figure 6-120, Figure 1-672) is one of the two core components of the automated smart tagging system described herein. It receives input to the system in a displayed format. For the convenience of visualization and tagging in the system's graphical user interface, a 2D representation of the raw external input is preferred, but it does not limit the input format to visual images only. Sound can be represented in 2D after the Fast Fourier Transform, and other input formats can be represented in 2D by other means. Text input can be displayed as text and supplied to the system. Advances in 3D displays, such as virtual and augmented reality systems, will enable the system to accept and tag 3D data.

[0045] The output of the feature extraction module is a set of feature vectors. Depending on the nature of the tagging, the set of feature vectors can be one feature vector per image (for example, in the simple case where the entire image is tagged at once for scene recognition), or multiple feature vectors representing the relevant regions of the image where these features are found. Depending on the ultimate goal of the tagging process, these regions can be as simple as a rectangular bounding box, a more complex polygon, or even a pixel-like mask.

[0046] The exemplary implementations described herein use a deep convolutional neural network for feature extraction. The convolutional neural network (CNN) uses convolutional units, where the receptive field of the unit filter (weight vector) is stepped across the entire height and width dimensions of the input. Because each filter is small, the number of parameters is significantly reduced compared to a fully connected layer. If a set of features can be extracted for an object when it is at one spatial location, the same set of features can be extracted for the same object when it appears at any other spatial location, since the features containing the object are independent of the object's spatial location. These invariances provide a feature space in which the encoding of the input has improved stability with respect to changes in vision, meaning that as the input changes (e.g., an object is slightly translated and rotated in the image frame), the output values ​​change only very little more than the input values.

[0047] Convolutional neural networks are also adept at generalization. Generalization means that within the form in which it was trained, the network can produce similar outputs for test data that is not identical to the training data. Learning the main regularities that define a set of class-specific features requires a large amount of data. If the network is trained on many classes, the lower layers, where filters are shared across classes, provide a good set of regularities for inputs of the same form. Therefore, a CNN trained on one task can deliver excellent results when used as an initialization for other tasks, or when the lower layers are used as processors for new, higher-level representations. For example, natural images share a common set of statistical properties. Since the learned features in the lower layers are fairly class-independent, they can be reused even if the class the user is trying to tag is not one of the classes the CNN was involved with. It is sufficient to take these feature vectors and feed them as input to the fast learning classifier part of the system (150 in Figure 1B).

[0048] Depending on the target final outcome of the tagging process, different CNNs can function as feature extractors. For whole-scene recognition, versions of Alexnet, GoogLeNet, or ResNet can be used depending on the computing power of the available hardware. An average pooling layer must be added after the last feature layer of these networks to pool across the entire image location to create the entire scene feature vector. If the system is used to create a dataset of trained detection networks, the same networks can be used without average pooling, or a region-proposed network such as fRCNN can be used instead for better spatial accuracy. If image segmentation is the target, segmentation networks such as Mask RCNN, FCN, or U-Net can be used for feature extraction and mask generation. These masks can be converted to polygons for display and modification. A custom-made version of FCN was used for the exemplary implementation shown in Figures 3-5.

[0049] Alternative implementations of the feature extraction module include, but are not limited to, scale-invariant feature transformation (SIFT), accelerated robust feature (SURF), Haar-like feature detectors, dimensionality reduction, component analysis, and others. Any suitable technique for feature extraction can be used as long as it produces feature vectors that are sufficiently clear and fast enough to handle different objects that require system tagging, so that the user does not have to wait an unnoticeable amount of time while the system computes the feature set.

[0050] [High-speed learning classifier module] The fast learning classifier module (150 in Figure 1B, 674 in Figure 6) takes in feature vectors created by the feature extraction module (120 in Figure 1B), performs classification or feature matching, and outputs a class label for each given set of features. One exemplary implementation can be simple template matching, where the system stores template feature vectors obtained based on user input for each known object. When an input set of feature vectors is presented, each of these vectors is compared to the template vector based on some distance metric (e.g., the Euclidean norm of the relevant feature space) to measure the difference between the current input and the template. The classifier then marks feature vectors with a difference metric smaller than a set threshold as possible objects. The reciprocal value of the distance can serve as a confidence measure for this classification scheme.

[0051] Other techniques that follow fast learning can be replaced by classifiers for this template matching technique, and these techniques include regression analysis methods (e.g., linear regression, logistic regression, minimax analysis), kernel methods (e.g., support vector machines), Bayesian models, ensemble methods (e.g., expert ensemble), decision trees (e.g., incremental decision trees, ultrafast decision trees and their derivatives), adaptive resonance theory-based models (e.g., Fuzzy ARTMAP), and linearly discriminative online algorithms (e.g., online passive-active algorithms). For example, the implementation shown in Figures 3-5 uses a modified ARTMAP for fast learning. The modification reduces the order dependency of ARTMAP, improves statistical consistency, and reduces the growth of the system memory footprint as it learns new objects.

[0052] [Conclusion] While various inventive embodiments have been described and illustrated herein, those skilled in the art will readily be able to develop various other means and / or structures to implement and / or achieve the functions and / or benefits described herein, and each such variation and / or modification will be considered within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are illustrative, and that actual parameters, dimensions, materials, and / or configurations will depend on the particular application or the application in which the teachings of the present invention are used. Those skilled in the art will be able to recognize or confirm numerous equivalents to the particular inventive embodiments described herein using ordinary experimentation. Therefore, it should be understood that the embodiments described herein are presented only as examples, and embodiments of the present invention may be implemented within the scope of the appended claims and their equivalents, beyond those specifically described and described in the claims. The inventive embodiments of this disclosure cover the individual features, systems, articles, materials, kits, and / or methods described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present invention as long as such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0053] The embodiments described above can be implemented in any of a number of ways. For example, embodiments of the technology disclosed herein may be implemented using hardware, software, or a combination thereof. If implemented in software, the software code can run on any suitable processor or set of processors, whether provided on a single computer or distributed across multiple computers.

[0054] Furthermore, it should be understood that computers can be embodied in any of many forms, such as rack-mounted computers, desktop computers, laptop computers, or tablet computers. In addition, computers may be embedded in devices that are generally considered computers but have sufficient processing power, including personal digital assistants (PDAs), smartphones, or any other suitable portable or fixed electronic devices.

[0055] Furthermore, a computer may have one or more input and output devices. These devices can, among other things, be used to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for a visual representation of the output and speakers, or for an audible representation of the output of other speech-generating devices. Examples of input devices that can be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitizer tablet. As another example, a computer may receive input information in speech recognition or other audible format.

[0056] These computers may be interconnected by one or more networks of any suitable form, such as a local area network, a wide area network such as an enterprise network, an intelligent network (IN), or the Internet. These networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0057] The various methods or processes outlined herein may be coded as software executable on one or more processors using any one of a variety of operating systems or platforms. Furthermore, such software may be written using any of a number of suitable programming languages ​​and / or programming or scripting tools, and may be compiled as executable machine code or intermediate code that runs on a framework or virtual machine.

[0058] In this regard, various concepts of the invention may be embodied as computer-readable storage media (or multiple computer-readable storage media) (e.g., computer memory, one or more floppy disks, compact disks, optical disks, magnetic tapes, flash memory, field-programmable gate arrays or other semiconductor device circuit configurations or other non-temporary or tangible computer storage media) coded by one or more programs that perform methods for carrying out the various embodiments of the invention described above when executed on one or more computers or other processors. The computer-readable media or media may be movable, as described above, so that the programs stored thereon or the programs can be loaded onto one or more different computers or other processors to carry out the various aspects of the invention.

[0059] The terms “program” or “software” are used herein to mean, as described above, any type of computer code or set of computer executable instructions that can be used to program a computer or other processor. Furthermore, naturally, according to one aspect, one or more computer programs that perform a method of the present invention do not have to belong to a single computer or processor, but may be distributed in a modular form among a number of different computers or processors for carrying out various aspects of the present invention.

[0060] Computer executable instructions can take many forms, such as program modules, which are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. Typically, the functions of program modules may be combined or distributed as desired in various embodiments.

[0061] Furthermore, the data structure may be stored in a computer-readable medium in any suitable form. For the sake of simplicity of illustration, the data structure may be shown to have fields relating to the location of the data structure. Such relationships can similarly be achieved by allocating storage for fields having locations in a computer-readable medium that convey the relationships between fields. However, any suitable mechanism may be used to establish relationships between the information in the fields of the data structure, including the use of pointers, tags, or other mechanisms to establish relationships between data elements.

[0062] Furthermore, various inventive concepts may be embodied in one or more methods, and examples of such methods are provided. The actions performed as part of the method can be ordered in any suitable manner. Thus, embodiments can be constructed in which the actions are performed in a different order than those exemplified, and this may include performing several actions simultaneously, even if they are shown as sequential actions in the exemplary embodiments.

[0063] All definitions defined and used herein should be understood to govern dictionary definitions, definitions incorporated by reference in documents, and / or the ordinary meanings of the defined terms.

[0064] As used herein and in claims, the indefinite articles "a" and "an" should be understood to mean "at least one" unless explicitly stated otherwise.

[0065] As used herein and in the claims, the phrase “and / or” should be understood to mean “either or both” of the combined elements, that is, elements that exist jointly in some cases and separately in others. The multiple elements listed in “and / or” should be interpreted in the same form, that is, “one or more” of the combined elements. Other elements may exist as they see fit, whether related to or unrelated to the elements specifically identified, in addition to the elements identified by the “and / or” clause. Thus, as a non-restrictive example, a reference to “A and / or B” when used in conjunction with unrestrictive language such as “including” may, in one embodiment, refer to A only (optionally including elements other than B), in another embodiment, refer to B only (optionally including elements other than A), and in yet another embodiment, refer to both A and B (optionally including other elements), and so on.

[0066] Where used herein and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” is interpreted as inclusive, that is, including at least one of a number of elements or a list of elements, and optionally additional items not on the list, but also including two or more. Conversely, only terms that explicitly indicate this, such as “one of” or “exactly one of” or, when used in a claim, the term “consisting of,” refers to the inclusion of exactly one element from a number of elements or a list of elements. In general, where used herein, the term “or” is interpreted only as indicating exclusive substitutes (i.e., “either one or the other” when preceded by terms of exclusivity, such as “either,” “one of,” “one of” or “exactly one of”). “Consisting of” where used in a claim shall have the usual meaning as used in the field of patent law.

[0067] As used herein, the terms “about” and “approximately” generally mean plus or minus 10% of the stated value.

[0068] As used herein and in claims, the phrase “at least one” referring to a list of one or more elements should be understood to mean at least one element selected from any one or more elements of the list of elements, but not necessarily including at least one of each of the elements specifically enumerated within the list of elements, nor excluding any combination of elements from the list of elements. This definition also allows for the existence of elements other than those specifically identified within the list of elements to which the phrase “at least one” refers, whether related to or unrelated to the specifically identified elements. Therefore, as a non-restrictive example, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") may, in one embodiment, refer to A (optionally including elements other than B) in which B does not exist and at least one, optionally including two or more elements other than B; in another embodiment, refer to B (optionally including elements other than A) in which A does not exist and at least one, optionally including two or more elements other than A; and in yet another embodiment, refer to A (optionally including two or more elements) and at least one, optionally including two or more elements other than B (and optionally including other elements).

[0069] In the claims and the above specification, all transitional phrases, such as “comprising,” “including,” “carry,” “have,” “containing,” “involving,” “hold,” “constitute,” and similar terms, are understood to be unrestrictive, meaning they include but are not limited to them. Only the transitional phrases “consist of” and “substantially constitute” are closed or semi-closed transitional phrases, respectively, as described in Section 2111.03 of the U.S. Patent and Trademark Examination Guidelines.

Claims

1. A method for tagging instances of objects that appear in a sequence of image frames, The steps include: drawing a first boundary polygon around a first instance of an object appearing in the sequence of image frames, via a user interface; The steps include: the user tagging a first instance of the object with a first label through the user interface; The steps include: extracting a first feature vector representing the contents of the first boundary polygon using a feature extraction module; The classification module includes the step of associating the first label of the object with the first feature vector based on the tagging by the user, The classification module includes the step of generating a second boundary polygon in the sequence of image frames, The steps include: extracting a second feature vector representing the contents of the second boundary polygon using the feature extraction module; The steps include: calculating the similarity between the first feature vector and the second feature vector using the classification module; The steps include comparing the similarity between the first feature vector and the second feature vector with a predetermined threshold using the classification module, Using the classification module, the step of classifying the contents of the second boundary polygon as a second instance of the object based on the comparison, The steps of displaying to the user, through the user interface, the second boundary polygon and a proposed label for the content of the second boundary polygon that is the same as the first label, The steps include receiving a modification from the user to at least one of the second boundary polygon and the proposed label through the user interface, The steps of updating the classification module based on the aforementioned modifications: A method that includes this.

2. The method according to claim 1, further comprising the step of determining a confidence value for correctly classifying the contents of the second boundary polygon as the second instance, based on a comparison of the distance between the first feature vector and the second feature vector and the predetermined threshold.

3. The method according to claim 2, further comprising the step of displaying the confidence value to the user through the user interface.

4. The aforementioned object is the first object, The steps include: the user tagging a first instance of a second object with a second label through the user interface; The steps include: using the feature extraction module to extract a third feature vector representing the contents of a third boundary polygon including a first instance of the second object; The classification module includes the step of associating the second label with the third feature vector. The method according to claim 1, further comprising:

5. The steps include: extracting a fourth feature vector from the content of the fourth boundary polygon in the sequence of image frames using the feature extraction module; The steps include: calculating the distance between the third feature vector and the fourth feature vector using the classification module; The steps include using the classification module to compare the distance between the third feature vector and the fourth feature vector with the predetermined threshold, Using the classification module, the step of classifying the contents of the fourth boundary polygon as a second instance of the second object based on the comparison, When a second instance of the second object is classified, the user interface is used to display the second instance of the second object along with the second label. The method according to claim 4, further comprising:

6. The method according to claim 1, wherein the extraction of the first feature vector includes the extraction of a first convolution output of a first image in the sequence of image frames, the first convolution output includes the first feature vector representing a first instance of the object.

7. The extraction of the first convolution output is as follows: In the first image described above, multiple divided sub-areas are generated, The feature extraction module encodes each of the plurality of divided sub-areas. The method according to claim 6, including the method described in claim 6.

8. The method according to claim 1, wherein the feature extraction module extracts the second feature vector from the first image in the sequence of image frames.

9. The method according to claim 1, wherein the feature extraction module extracts the second feature vector from the second image in the sequence of image frames.

10. The steps include: extracting a third feature vector from the content of a third boundary polygon in the sequence of image frames using the feature extraction module; The steps include: calculating the distance between the first feature vector and the third feature vector using the updated classification module; The steps include comparing the distance between the first feature vector and the third feature vector with the predetermined threshold using the updated classification module, Using the updated classification module, the contents of the third boundary polygon are classified as a third instance of the object based on a comparison of the distance between the first feature vector and the third feature vector and a predetermined threshold. The method according to claim 1, further comprising:

11. The method according to claim 1, wherein the step of receiving the modification includes receiving a change from the user through the user interface to at least one of the shape and position of the second boundary polygon.

12. The method according to claim 1, wherein the step of receiving the modification includes receiving a change from the user through the user interface to the first label relating to the content of the second boundary polygon.

13. The method according to claim 1, wherein the sequence of image frames includes keyframes extracted from a video file.

14. The method according to claim 1, wherein the first boundary polygon or the second boundary polygon includes a rectangular bounding box or a pixel-based mask.

15. A system for tagging instances of objects that appear in a sequence of image frames, A user interface that enables a user to tag a first instance of the representation of the object appearing in the sequence of image frames, At least one processor operably connected to the user interface and It has, The aforementioned at least one processor is The steps include: drawing a first boundary polygon around a first instance of an object appearing in a sequence of image frames, via the user interface; The steps include: the user tagging a first instance of the object with a first label through the user interface; The steps include: extracting a first feature vector representing the contents of the first boundary polygon using a feature extraction module; The classification module includes the step of associating the first label of the object with the first feature vector based on the tagging by the user, The classification module includes the step of generating a second boundary polygon in the sequence of image frames, The steps include: extracting a second feature vector representing the contents of the second boundary polygon using the feature extraction module; The steps include: calculating the similarity between the first feature vector and the second feature vector using the classification module; The steps include comparing the similarity between the first feature vector and the second feature vector with a predetermined threshold using the classification module, Using the classification module, the step of classifying the contents of the second boundary polygon as a second instance of the object based on the comparison, The steps of displaying to the user, through the user interface, the second boundary polygon and a proposed label for the content of the second boundary polygon that is the same as the first label, The steps include receiving a modification from the user to at least one of the second boundary polygon and the proposed label through the user interface, The steps of updating the classification module based on the aforementioned modifications: To do system.

16. The system according to claim 15, wherein the extraction of the first feature vector includes the extraction of a first convolution output of a first image in the sequence of image frames, the first convolution output includes the first feature vector representing a first instance of the object.

17. The extraction of the first convolution output is as follows: In the first image described above, multiple divided sub-areas are generated, The feature extraction module encodes each of the plurality of divided sub-areas. The system according to claim 16, including the system described in claim 16.

18. The system according to claim 15, wherein the feature extraction module extracts the second feature vector from the first image in the sequence of image frames.

19. The system according to claim 15, wherein the feature extraction module extracts the second feature vector from the second image in the sequence of image frames.

Citation Information

Patent Citations

  • Activity process reflection support system

    JP2009229605A

  • Learning structured prediction models for interactive image labeling

    US20120269436A1

  • Method for annotating images

    WO2013182298A1