Negative Sampling Algorithm for Enhancing Image Classification
Through the role recognition engine in the media indexer, the problem of character recognition relies on manual tags in animated videos is solved, and efficient and accurate role classification is achieved.
Patent Information
- Application Number
- CN202080058773.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-26
- Filing Date
- 2020-06-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-06-17
AI Technical Summary
In the prior art, character identification in animated videos relies on manual marking, resulting in a cumbersome process and unscalable, making it difficult to effectively search and retrieve.
The role identification engine in the media indexer automatically detects role instances in multi-frame animation media files and groups them. The image classification model is used to automatically classify, reducing the need for manual annotation.
It realizes the scalability and accuracy of animation character detection, reduces the computational complexity, reduces the amount of manual markers, and improves the accuracy of character classification.
Smart Images

Figure CN114287005B_ABST
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure relate to the fields of machine learning and artificial intelligence, and more particularly to the automatic identification and grouping of characters in multi-frame media files (e.g., animated videos) for semi-supervised training of machine learning image classification models. Background Art
[0002] Animation is an extremely large business globally and is a major product of many of the largest media companies. However, animated videos typically contain very limited metadata, so effective search and retrieval of specific content is not always possible. For example, a key component in animated media is the animated characters themselves. In fact, the characters in an animated video must first be indexed, e.g., detected, classified, and annotated, in order to be able to effectively search for and retrieve those characters in the animated video.
[0003] Various services can utilize artificial intelligence or machine learning to understand images. However, these services typically rely on large amounts of manual labeling. For example, character identification in animated videos currently involves manually drawing bounding boxes around each character and annotating (or labeling) the character contained within the bounding box with, e.g., the name of the character. This manual annotation process is repeated for each character in each frame of a multi-frame animated video. Unfortunately, this manual annotation process is cumbersome and severely limits the scalability of these services.
[0004] In general, examples of some existing or related systems and their associated limitations herein are intended to be illustrative and not exclusive. By reading the following, other limitations of existing or current systems will become apparent to those skilled in the art. Summary of the Invention
[0005] Among other benefits, one or more embodiments described herein solve one or more of the foregoing or other problems in the art by providing systems, methods, and non-transitory computer-readable media that are capable of automatically detecting instances (or occurrences) of characters in multi-frame animated media files and grouping them such that each group contains images associated with a single character. The character groups themselves can then be labeled and used to train an image classification model for automatically classifying animated characters in subsequent multi-frame animated media files.
[0006] While multiple embodiments are disclosed, other embodiments of the invention will become apparent to those skilled in the art from the following detailed description, which illustrates and describes illustrative embodiments of the invention. As will be recognized, the invention is capable of modification in various aspects, all of which do not depart from the scope of the invention. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive.
[0007] This summary is provided to introduce some concepts in a simplified form that will be further described in the following technical disclosure. It is understood that this summary is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of such exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] To describe the manner in which the above and other advantages and features can be obtained, a more specific description will be set forth and will be represented by reference to specific examples illustrated in the accompanying drawings. It should be understood that these drawings only depict typical examples and should not be considered as limiting the scope thereof, and the embodiments will be described and explained with additional features and details by using the drawings.
[0009] Figure 1A A block diagram depicting an exemplary animation character recognition and indexing framework according to some embodiments, the animation character recognition and indexing framework for training an artificial intelligence-based (AI-based) image classification model to automatically classify characters in a multi-frame animation media file for indexing.
[0010] Figure 1B A block diagram depicting an exemplary animation character recognition and indexing framework according to some embodiments, the animation character recognition and indexing framework applying (and retraining as needed) the AI-based image classification model trained in the example of Figure 1A .
[0011] Figure 2 A data flow diagram graphically depicting the operations and data flow between modules of a media indexer according to some embodiments.
[0012] Figure 3 A flowchart depicting an exemplary process for indexing a multi-frame animation media file using the automatic character detection and grouping techniques discussed herein.
[0013] Figure 4 A flowchart depicting an exemplary process for training or refining an AI-based image classification model using grouped character training data.
[0014] Figure 5 A flowchart depicting an exemplary process for grouping (or clustering) characters automatically detected in a multi-frame animation media file according to some embodiments.
[0015] Figure 6Depicts a graphical user interface including various menus for selecting options to upload video files according to some embodiments.
[0016] Figure 7 Depicts a graphical user interface showing an exemplary video that has been indexed using a media indexer according to some embodiments.
[0017] Figure 8 Depicts a graphical user interface showing an exemplary video that has been indexed using a media indexer according to some embodiments.
[0018] Figure 9 Depicts a flowchart showing an exemplary process for indexing a multi-frame animated media file using the automatic character detection and grouping techniques discussed herein according to some embodiments.
[0019] Figure 10 Depicts a flowchart showing another exemplary process for grouping (or clustering) characters automatically detected in a multi-frame animated media file according to some embodiments.
[0020] Figure 11 Depicts a flowchart showing an exemplary process for identifying and classifying negative examples of target content according to some embodiments.
[0021] Figure 12 Depicts an exemplary scenario for identifying negative samples of target content according to some embodiments.
[0022] Figure 13 Depicts a flowchart showing another exemplary process for identifying and classifying negative samples of target content according to some embodiments.
[0023] Figure 14 Depicts an exemplary scenario for identifying negative samples of target content according to some embodiments.
[0024] Figure 15 Depicts an exemplary scenario for identifying negative samples of target content according to some embodiments.
[0025] Figure 16 Depicts a block diagram of an exemplary computing system suitable for implementing the techniques disclosed herein, the exemplary computing system including any applications, architectures, elements, processes, and operational scenarios and sequences shown in the figures and discussed in the following technical disclosure.
[0026] The accompanying drawings are not necessarily to scale. Similarly, for purposes of discussing some embodiments of the present technology, some components and / or operations may be divided into different blocks or combined into a single block. Additionally, while the technology may have various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail below. However, the intention is not to limit the technology to the particular embodiments described. Rather, the technology is intended to cover all modifications, equivalents, and alternatives falling within the scope of the technology as defined by the appended claims. Detailed Description
[0027] Examples are discussed in detail below. While specific implementations are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the spirit and scope of the subject matter of this disclosure. These implementations may include machine-implemented methods, computing devices, or computer-readable media.
[0028] Identifying animated characters in video can be challenging for a variety of reasons such as the non-regular nature of the animated characters themselves. In fact, animated characters can have many different forms, shapes, sizes, etc. In many cases, content producers (e.g., companies that generate or manipulate animated media content) want to index the characters included in their animated media content. However, as described above, this is currently a very difficult and non-scalable process that requires manually annotating each frame of a multi-frame animated media file for each character.
[0029] The technology described herein relates to a media indexer that includes a character identification engine that can automatically detect instances (or occurrences) of characters in a multi-frame animated media file and group them such that each group contains images associated with a single character. The character groups themselves are then labeled, and the labeled groups are used to train an image classification model for automatically classifying animated characters in subsequent multi-frame animated media files.
[0030] Various technical effects can be achieved by the technology discussed herein. Among other benefits, the technology discussed herein provides a scalable solution for training an image classification model that has a minimal impact on character detection or character classification accuracy. Additionally, the use of key frames reduces the amount of data that needs to be processed while maintaining a high data variance. Furthermore, automatic character identification eliminates the need for manually annotating bounding boxes, and the automatic grouping of characters produces accurate annotations while significantly reducing the amount of manual effort, such as semi-supervised training via group labeling rather than per-character annotation.
[0031] As used herein, the term "animated character" refers to an object that exhibits humanoid characteristics included or detected in an animated multi-frame media file. For example, an "animated character" can be a living or inanimate anthropomorphic object that exhibits any human form or attribute (including but not limited to human characteristics, emotions, intentions, etc.).
[0032] Describe a general overview and exemplary architecture of an animated character recognition and indexing framework for training an AI-based image classification model related to Figure 1A Then,[[]] Figure 1B Depict an example of applying (and retraining or refining as needed) the trained AI-based image classification model with the animated character recognition and indexing framework. Thereafter, a more detailed description of the components and processes of the animated character recognition and indexing framework is provided with respect to the subsequent figures.[[]]
[0033] Figure 1A Depict a block diagram showing an exemplary animated character recognition and indexing framework 100 according to some embodiments. The animated character recognition and indexing framework 100 is used to train an AI-based image classification model to automatically classify characters in a multi-frame animated media file for indexing. In fact, the exemplary animated character recognition and indexing framework 100 includes a media indexer service 120 that can automatically detect instances (or occurrences) of characters in a media file and group them such that each group contains images associated with a single character. Then, the character groups are correspondingly identified (recognized) and labeled. As Figure 1A shown in the example of,[[]] the labeled character groups (or grouped character training data) can then be used to train an AI-based image classification model to automatically classify animated characters in subsequent multi-frame animated media files.[[]]
[0034] As Figure 1A shown in the example of,[[]] the animated character recognition and indexing framework 100 includes an image classifier 110, a media indexer 120, and a user 135 of an operating computing system 131 that can provide user input to manually label (or recognize) character groups. Additional or fewer systems or components are possible.[[]]
[0035] The image classifier 110 can be any image classifier of an image classification service. In some embodiments, the image classifier 110 can be implemented by the Azure Custom Vision Service provided by Microsoft. The Custom Vision Service uses machine learning algorithms to apply tags to images. Developers typically submit groups of tagged images, which may or may not have the features under discussion. The machine learning algorithm is trained using the submitted data and calculates its own accuracy by testing itself on these same images. Once the machine learning algorithm (or model) is trained, the image classifier 110 can test, retrain, and use the model to classify new images.
[0036] As Figure 1A and Figure 1B shown in the example of, the media indexer 120 includes a character recognition engine 122, a media indexer database 128, and an indexing engine 129.
[0037] The character recognition engine 122 includes a key frame selection module 123, a character detection module 124, a character grouping module 125, and a group tagging module 126. The functions represented by the components, modules, managers, and / or engines of the character recognition engine 122 can be implemented, in whole or in part, alone or in any combination thereof, in hardware, software, or a combination of hardware and software. Additionally, although shown as discrete components, the operations and functions of the components, modules, managers, and / or engines of the character recognition engine 122 can be partially or fully integrated into other components of the animated character recognition and indexing framework 100.
[0038] In operation, an unindexed (or unstructured) multi-frame animated media file 105a is fed into the media indexer 120 for character recognition and indexing. The media indexer 120 includes a character recognition engine 122, a media indexer database 128, and an indexing engine 129. Additional or fewer systems or components are possible.
[0039] The key frame selection module 123 is configured to select or otherwise identify a small subset of the total frames of the multi-frame animated media file to reduce the computational complexity of the character recognition process with minimal or limited impact on accuracy. In practice, the key frame selection module 123 is configured to identify and select important or meaningful frames (e.g., frames with the highest likelihood of observing a character) from the multi-frame animated media file. In some embodiments, the key frame is determined at least in part based on the individual importance of the key frame in determining a micro-scene or shot segment. In some embodiments, each frame can be assigned an importance value, and frames with an importance value greater than a threshold are selected as key frames. Alternatively or additionally, a percentage of the total frames, such as the top one percent of frames with the highest evaluated importance values, can be selected as key frames.
[0040] As discussed herein, key frames typically constitute a small fraction (e.g., one percent) of the total frames in a multi-frame animation media file (e.g., an animated video). However, the performance difference between tagging every frame in a multi-frame animation media file and tagging only key frames is nominal for detecting each character in the multi-frame animation media file. Thus, key frames allow the media indexer 130 to maintain character detection accuracy while reducing computational complexity.
[0041] The character detection module 124 is configured to process or analyze key frames to detect (or propose) instances (or occurrences) of characters in the key frames of a multi-frame animation media file. In practice, the character detection module 124 can process key frames and provide character region proposals (also referred to as bounding boxes). For example, the character detection module 124 can capture each character region proposal as an image.
[0042] As discussed herein, detecting animated characters can be difficult because characters can take the form of almost any living (e.g., human, animal, etc.) or inanimate (e.g., robot, car, candle, etc.) object. Thus, in some embodiments, the character detection module 124 includes an object detection model that is trained to detect bounding boxes of animated characters (e.g., cars, humans, robots, etc.) of different styles, themes, etc.
[0043] In some embodiments, the character detection module 124 can be trained to detect objects that exhibit humanoid features. That is, the character detection module 124 is designed to detect any anthropomorphic object within the key frames. As discussed herein, the term "anthropomorphic object" refers to any living or inanimate object that exhibits any human form or attribute (including but not limited to human features, emotions, intentions, etc.).
[0044] The character grouping module 125 is configured to compare and group character region proposals based on image similarity such that each group contains images associated with a single character. In some cases, more than one of the resulting character groups can be associated with the same character. For example, the first group includes images of SpongeBob wearing a hat, and the second group includes images of SpongeBob not wearing a hat.
[0045] In some embodiments, the character grouping module 125 applies a clustering algorithm using the embeddings of the detected character region proposals to determine character groups. In practice, character groups can be determined by embedding the features of the character region proposals (or images) into a feature space to simplify image comparison. Refer to Figure 5 Examples illustrating methods of applying a clustering algorithm are shown and discussed in more detail, which include embedding character region proposals (or images) into a feature space and comparing the embeddings to identify character groups.
[0046] The group tagging module 126 is configured to tag (annotate or classify) the role groups without using a classification model. As discussed herein, tagging the role groups is useful for the initial training of the classification model as well as for refining the trained classification model (as shown and discussed in more detail with reference to Figure 1B .
[0047] In some embodiments, the group tagging module 126 may present each role group to the user 135 as an image cluster. The role groups can then be classified using input from the user 135. For example, the user 135 may provide an annotation or label for the group. Alternatively or additionally, the user 115 may provide a canonical image of the role(s) desired to appear in the multi-frame animated media file. In this case, the canonical role(s) can be compared to the role groups to identify and tag the role groups. In other embodiments, the user 115 may provide the name of a movie or series of the multi-frame animated media file. In this case, the group tagging module 126 may query a data store, such as Satori (Microsoft's Knowledge Graph), to obtain information about the movie and / or series and extract the names of the roles and any available canonical images.
[0048] Figure 1B FIG. depicts a block diagram showing an exemplary animated character recognition and indexing framework 100 according to some embodiments, the animated character recognition and indexing framework 100 applying (and retraining as needed) the Figure 1A AI-based image classification model trained in the example of. In fact, the trained AI-based image classification model is trained to automatically identify and index the animated characters in the multi-frame animated media file 106a. The multi-frame animated media file 106a is related to the multi-frame animated media file 105a (e.g., the same series or having one or more overlapping characters).
[0049] As discussed herein, in some embodiments, the user may specify the trained AI-based image classification model for indexing the multi-frame animated media file. Referring to Figure 6 , an example showing a graphical user interface including various menus for selecting the trained AI-based image classification model is shown and discussed in more detail.
[0050] In operation, the media indexer 120 can classify character groups using a trained AI-based image classification model and refine (or tune) the trained AI-based image classification model using newly grouped character training data (e.g., new characters or existing characters with new or different appearances or features). As discussed herein, the media indexer 120 interfaces with the image classifier 110 to utilize, train, and / or refine the AI-based (one or more) image classification models 116.
[0051] As discussed above, the image classifier 110 can be implemented by Azure Custom Vision Service, which can be applied to each cluster (or character group). In some embodiments, a smoothing operation can be applied to handle cases where a single character is split into two or more different clusters (or character groups), such as a group of images of SpongeBob with a hat and a group of images of SpongeBob without a hat. The smoothing operation can operate to merge two or more different clusters (or character groups) and provide grouped character training data to refine the trained AI-based image classification model such that future classifications are categorized as the same character.
[0052] Figure 2 A data flow diagram depicts, in graphical form, the operations and data flows between modules of a media indexer 200 according to some embodiments. As Figure 2 shown in the example of, the media indexer 200 includes Figure 1A and 1B a key frame selection module 123, a character detection module 124, a character grouping module 125, and a group tagging module 126. Additional or fewer modules, components, or engines are possible.
[0053] Figure 3 A flow chart depicts an exemplary process 300 for indexing a multi-frame animated media file (e.g., an animated video) using the automatic character detection and grouping techniques discussed herein. The exemplary process 300 can be performed in various implementations by a media indexer (e.g., Figure 1A and Figure 1B the media indexer 120) or one or more processors, modules, engines, or components associated therewith.
[0054] First, at 310, the media indexer presents a user interface (UI) or an application programming interface (API). As discussed herein, the user can specify a multi-frame animated media file to be indexed and an AI-based image classification model to use for indexing (if trained) or training (if not trained). Refer to Figure 6, an example showing a graphical user interface is presented in more detail and discussed. The graphical user interface includes various menus for selecting an AI-based image classification model for training.
[0055] At 312, the media indexer receives a multi-frame animated media file (e.g., an animated video) for indexing. At 314, the media indexer extracts or identifies key frames. At 316, the media indexer detects characters in the key frames. At 318, the media indexer groups the characters automatically detected in the multi-frame animated media file. Refer to Figure 5 An example illustrating character grouping is presented in more detail and discussed. At 320, the media indexer determines whether a trained classification model is specified. If so, at 322, the media indexer classifies the character groups using the trained classification model, and at 324, smooths (or merges) the classified character groups.
[0056] Finally, at 326, the multi-frame animated media file (e.g., an animated video) is indexed using the identified (classified) and unidentified (unknown) characters. At Figure 8 An exemplary graphical user interface illustrating a multi-frame animated media file indexed by identified and unidentified characters is shown in the example of . As discussed herein, the user can specify or label groups of unidentified characters to refine the AI-based image classification model.
[0057] Figure 4 A flowchart depicting an exemplary process 400 for training or refining an AI-based image classification model using grouped character training data according to some embodiments is shown. The exemplary process 400 can be executed by a media indexer (e.g., Figure 1A and Figure 1B the media indexer 120 of ) or one or more processors, modules, engines, or components associated therewith in various implementations.
[0058] First, at 412, the media indexer identifies (or otherwise obtains) label or classification information.
[0059] At 414, the media indexer...
[0060] Finally, at 416, the media indexer...
[0061] Figure 5 A flowchart depicting an exemplary process 500 for grouping (or clustering) characters automatically detected in a multi-frame animated media file according to some embodiments is shown. The exemplary process 500 can be executed by a media indexer (e.g., Figure 1A and Figure 1B the media indexer 120 of ) or one or more processors, modules, engines, or components associated therewith in various implementations.
[0062] First, at 412, the media indexer identifies (or otherwise obtains) the label or classification information of an unknown (or unclassified) group of animated characters. As discussed herein, the media indexer can identify label information, such as the name of an individual animated character associated with each group of animated characters, and classify (or annotate) the group of animated characters using the identified label information, thereby generating at least one annotated group of animated characters.
[0063] At 414, the media indexer collects the identified (or annotated) groups of animated characters in the media indexer database. Finally, at 416, the media indexer trains or refines an image classification model by feeding the annotated groups of animated characters to an image classifier to train the image classification model.
[0064] Figure 5 A flowchart depicting an exemplary process 500 for grouping (or clustering) characters automatically detected in a multi-frame animated media file according to some embodiments is shown. Exemplary process 500 can be performed in various implementations by a media indexer (e.g., Figure 1A and Figure 1B media indexer 120) or one or more processors, modules, engines, or components associated therewith.
[0065] First, at 510, the media indexer accesses the next identified character. As discussed herein, each character region proposal includes a bounding box or subset of key frames that contain the proposed animated character. At 512, the media indexer extracts the features of the next identified character contained in the character region proposal, and at 514, embeds the features into a feature space.
[0066] At decision 516, the media indexer determines whether more character region proposals have been identified. If so, it returns to step 510. As discussed herein, multiple key frames from a multi-frame animated media file are identified first. Each key frame can include one or more character region proposals. Once each character region proposal has been traversed, at 518, the media indexer selects the clustered groups of characters in the feature space. For example, the media indexer can determine the similarity between character region proposals by comparing the embedded features within the feature space and apply a clustering algorithm based on the determined similarity to identify groups of animated characters.
[0067] Figures 6 to 8 Various graphical user interfaces that can be presented to a user are depicted. First referring to the example of Figure 6 Figure 6 A graphical user interface according to some embodiments is depicted, which includes various menus for selecting various options for uploading a video file. More specifically, Figure 6Depicts a graphical user interface including various menus for selecting various options for uploading video files and (optionally) selecting a trained AI-based image classification model to index the video files (or alternatively for training).
[0068] Next, referring to Figure 7 as an example, Figure 7 depicts a graphical user interface showing an exemplary video that has been indexed using the media indexer discussed herein. In fact, Figure 7 the example of
[0069] Similarly, Figure 8 depicts a graphical user interface showing an exemplary video that has been indexed using the media indexer discussed herein. More specifically,
[0070] Figure 9 depicts a flowchart showing an exemplary process 900 according to some embodiments for indexing a multi-frame animated media file (e.g., an animated video) using the automatic role detection and grouping techniques discussed herein. The exemplary process 900 can be performed in various implementations by a media indexer (e.g., Figure 1A and Figure 1B media indexer 120) or one or more processors, modules, engines, or components associated therewith.
[0071] The exemplary process 900 is similar to the exemplary process 300, except that the exemplary process 900 includes steps for style adaptation. For example, an AI-based image classification model can be trained using a first type (or style) of animation (e.g., computer-generated graphics (CGI)) and then applied to an input including a second type (or style) of animation (e.g., hand-drawn animation) without retraining the model. In other potential options, key frames can be adjusted or transformed (as shown in the example of Figure 9 or features can be transformed before being embedded into the feature space (as shown in the example of Figure 10 ).
[0072] Referring again to Figure 9, in some embodiments, an additional network for style adaptation can be added to the detector (e.g., the character detection module 124) for online adaptation of unseen (or unknown) animation styles. The additional network can be trained offline in a variety of ways. For example, the training data can be based on a labeled dataset that is used to train the detector for unseen movies (e.g., trailers) and the dataset. The style adaptation network can learn to propagate local feature statistics from the dataset used for training to the unseen data. The training can be based on minimax global optimization, which maximizes the confidence of the character detector in the characters detected in unseen images while minimizing the distance between the deep learning embeddings of the images before and after style conversion (thus preserving similar semantic information). The deep learning embeddings that can be used are the same as those used for feature extraction and grouping.
[0073] Figure 10 FIG. shows a flowchart depicting another exemplary process 1000 for grouping (or clustering) characters automatically detected in a multi-frame animation media file. The exemplary process 1000 can be performed in various implementations by a media indexer (e.g., Figure 1A and Figure 1B the media indexer 120) or one or more processors, modules, engines, or components associated therewith.
[0074] The exemplary process 1000 is similar to Figure 5 the exemplary process 500, except that the exemplary process 1000 includes steps for style adaptation. Specifically, the exemplary process 1000 can adapt or transform features rather than entire key frames (as discussed in the example of Figure 9 ).
[0075] Figure 11 FIG. shows a process 1100 for sampling negative examples of images that are provided as training data to an image classifier. Sampling negative examples for image classification provides classification enhancement. Customizable image classification in any particular domain (e.g., animated characters) requires teaching a machine learning model to distinguish known classes from other things. Background sampling is a good way to generate bounding boxes that do not intersect with the character bounding boxes. The technical problem is computational complexity because the nature of the problem is a non-convex hard problem (NP-complete problem) in a mathematical sense. For example, the number of possible background (BG) boxes grows exponentially with the number of regions of interest (bounding boxes).
[0076] Process 1100 begins by identifying a region of interest around target content in a frame (step 1101). The region of interest can be formed by a rectangular bounding box drawn around the content of interest in an image, a series of images (video frames), etc. Examples of such content include characters in an animated video, components in a layout (e.g., circuits and components on a circuit board or furniture in an office layout).
[0077] Next, process 1100 identifies potential empty regions adjacent to the region of interest (step 1103). The region of interest can be one of the regions of interest identified in the context of step 1101. Each potential adjacent region includes a side adjacent to the central region of interest and has an axis parallel to the axis of the central region of interest. In an example, the central region of interest can be a rectangle, and the potential empty regions are also rectangles that are axially aligned with the central rectangle and have one side adjacent to a side of the central rectangle. Other shapes are also possible, such as triangles, squares, parallelograms, and trapezoids, or even circles that lack the straight sides of the previous examples.
[0078] Process 1100 then continues by identifying at least one empty region from the potential empty regions that meets one or more criteria (step 1105). The empty region that meets the one or more criteria can be classified (or designated) as a negative example of the target content that is the subject of the image classifier (step 1107). The negative example can be grouped with other negative examples in the collection and provided to the classifier as training data along with positive examples of the target content.
[0079] Returning to step 1105, identifying empty regions can be accomplished in a variety of ways. In one example, process 1100 can find (one or more) empty regions by employing a maximum empty rectangle algorithm (step 1105A). The maximum empty rectangle algorithm (or maximum empty rectangle) can quickly find the largest rectangular region in an image without target content. Thus, the empty rectangle will avoid or not overlap with any other rectangle containing target content.
[0080] Alternatively, process 1100 can employ a recursive analysis of regions adjacent to the central region to find empty regions (step 1105). This recursive analysis first identifies regions adjacent to the central region and designates those empty (and optionally of a satisfactory size) regions as negative examples. The analysis then recursively performs the same operation on any non-empty adjacent regions. That is, the analysis identifies other regions adjacent to the adjacent region and examines those empty (and optionally of a satisfactory size) regions (or sub-regions).
[0081] Figure 12Shows the results produced by the implementation of process 1100 in exemplary scenario 1200. In this scenario, process 1100 exemplifies 1201 including two rectangular bounding boxes represented by bounding box 1203 and bounding box 1205. Bounding box 1203 is drawn around one animated character, while bounding box 1205 is drawn around another animated character.
[0082] When applied to image 1203, process 1100 using recursive analysis will identify four empty regions adjacent to bounding box 1203, and these four empty regions are counted as negative examples of the target content represented by region 1211, region 1212, region 1213, and region 1214. Using the maximum rectangle method, process 1100 will only identify a single rectangle, such as region 1214, because region 1214 is the largest among the four rectangles. Whether to use one method instead of another will depend on operational constraints. For example, the maximum rectangle method may be faster than recursive analysis, but the resulting negative samples may inherently have less coding information than the set of negative samples produced by the recursive method. However, from a practical perspective, the speed obtained by the maximum rectangle method may be considered a worthy compromise.
[0083] Figure 13 Shows a recursive process 1300 for sampling negative examples of images provided as training data to an image classifier. A quadtree-like branch and bound recursion is proposed by process 1300, which produces the largest possible bounding boxes under certain complexity constraints. The recursion takes the most centered bounding box and splits the frame four times, namely the upper sub-frame, the lower sub-frame, the right sub-frame, and the left sub-frame. The stopping criterion is that there are no more bounding boxes or when the sub-frames are too small. Even when the image has many bounding boxes, this mechanism allows the indexer and classifier to be integrated to optimize the generation of negative examples, which makes the naive method practically unsolvable.
[0084] More specifically referring to Figure 13 The process starts by identifying the bounding boxes in the frame (step 1301). In some examples, the first box may be the most centered box in the frame. The frame presumably includes one or more bounding boxes around potential animated characters or other objects / regions of interest for which negative examples are needed.
[0085] The process continues by splitting the frame into multiple sub-frames around the bounding box (step 1303). In one example, four sub-frames can be formed from each of the four sides (left, right, top, and bottom) of the bounding box. Each of the four sub-frames will extend from one side of the bounding box to the edge of the frame itself. In other examples, the bounding box under discussion may provide fewer than four sides for which sub-frames are formed.
[0086] In step 1305, the process identifies one of the analyzed sub - frames as a potentially acceptable negative example and then compares the size of the sub - frame with the size of the size constraint (step 1307). If the size of the sub - frame does not meet the minimum size (e.g., is less than the threshold size), the sub - frame is rejected as a potential negative sample. However, if the size of the sub - frame meets the minimum size, the process determines whether the sub - frame includes one or more other bounding boxes or overlaps with one or more other bounding boxes (step 1309).
[0087] If no other bounding boxes are found within the sub - frame, the sub - frame is considered a negative example and can be classified or labeled as a negative example (step 1311). However, if the sub - frame includes one or more other bounding boxes within it, the process returns to step 1303 to split the sub - frame into more sub - frames again.
[0088] Assuming the sub - frame is counted as a negative example, the process continues to determine whether there are any remaining sub - frames for the parent frame to which the sub - frame belongs (step 1313). If so, the process returns to step 1305 to identify and analyze the next sub - frame.
[0089] If there are no remaining other sub - frames, all the identified negative examples can be provided as training data to the classifier (step 1315). This step can be performed individually, in batch mode, or in some other way after each negative example is identified.
[0090] Figure 14 Illustrates Figure 13 An exemplary implementation of the negative example sampling process. In Figure 14 the frame includes two bounding boxes around two characters. These characters are referred to here as "red" and "yellow". The larger box 1403 is drawn around the red character, while the smaller of the two boxes 1405 is drawn around the yellow character.
[0091] When applied to the image in Figure 14 the negative example sampling process of Figure 13 first identifies the most central bounding box, which, for exemplary purposes, is assumed to be the larger box around the red character. The frame around the bounding box is divided into four sub - frames, namely the right - hand side sub - frame 1409 of the bounding box, the left - hand side sub - frame 1411 of the bounding box, the top - hand side sub - frame 1407 of the bounding box, and the bottom - hand side sub - frame 1413 of the bounding box.
[0092] The top - hand side sub - frame is determined to meet the minimum size criteria and has no bounding boxes within it. Therefore, the top - hand side sub - frame is counted as a negative example. The right - hand side sub - frame is also large enough and has no other bounding boxes within it, so it can also be counted as a negative example of the character. However, the bottom - hand side sub - frame is not large enough and is therefore rejected as a negative example candidate.
[0093] On the other hand, the left sub-frame is large enough but includes at least a part of the bounding box, that is, the smaller box surrounding the yellow character. Therefore, the process recursively operates on the part of the bounding box of the yellow character that falls within the left sub-frame of the parent frame.
[0094] Like the parent frame, the left sub-frame is split into multiple sub-frames, but in this case there are only three sub-frames because the right side of the bounding box around the yellow character is excluded from the left sub-frame. The top sub-frame 1415 (of the sub-sub-frame) is counted as a negative example because it is large enough and there are no other bounding boxes in it. For the same reason, the left sub-frame 1417 (of the sub-sub-frame) is also counted as a negative example. However, it is possible that there is no sub-frame on the right side, and the bottom sub-frame 1419 is ignored because it is too small.
[0095] Since there are no other sub-frames at the child or parent level of the image frame, all negative examples have been identified and can be presented to the image classifier to enhance its training. Figure 15 An enlarged view 1500 of the final four negative examples resulting from applying the Figure 13 sampling process to the Figure 14 image is shown.
[0096] Figure 16 A computing system 1601 is shown, which represents any system or collection of systems in which the various processes, programs, services, and scenarios disclosed herein can be implemented. Examples of the computing system 1601 include, but are not limited to, server computers, cloud computing platforms, and data center devices, as well as any other type of physical or virtual server, container, and any variants or combinations thereof. Other examples include desktop computers, laptop computers, desktop computers, Internet of Things (IoT) devices, wearable devices, and any other physical or virtual combinations or variants thereof.
[0097] The computing system 1601 can be implemented as a single device, system, or apparatus, or can be implemented in a distributed manner as multiple devices, systems, or apparatuses. The computing system 1601 includes, but is not limited to, a processing system 1602, a storage system 1603, software 1605, a communication interface system 1607, and a user interface system 1609 (optional). The processing system 1602 is operatively coupled to the storage system 1603, the communication interface system 1607, and the user interface system 1609.
[0098] Processing system 1602 loads and executes software 1605 from storage system 1603. Software 1605 includes and implements process 1606, which represents the processes discussed with respect to the previous figures. When executed by processing system 1602 to provide packet rerouting, software 1605 instructs processing system 1602 to operate as described herein at least for the various processes, operation scenarios, and sequences discussed in the foregoing embodiments. Computing system 1601 may optionally include additional devices, features, or functions not discussed for the sake of brevity.
[0099] Continue Figure 16 As an example, processing system 1602 may include a microprocessor and other circuitry that retrieves and executes software 1605 from storage system 1603. Processing system 1602 may be implemented in a single processing device, but may also be distributed across multiple processing devices or subsystems that cooperate to execute program instructions. Examples of processing system 1602 include general-purpose central processing units, graphics processing units, specialized processors, and logic devices, as well as any other type of processing device, combinations, or variations thereof.
[0100] Storage system 1603 may include any computer-readable storage medium that can be read by processing system 1602 and is capable of storing software 1605. Storage system 1603 may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Examples of storage media include random access memory, read-only memory, magnetic disks, optical disks, flash memory, virtual memory and non-virtual memory, magnetic tape cartridges, tapes, magnetic disk storage devices, or any other suitable storage medium. In any case, a computer-readable storage medium is not a propagated signal.
[0101] In addition to the computer-readable storage medium, in some embodiments, storage system 1603 may also include a computer-readable communication medium through which at least some of software 1605 may communicate internally or externally. Storage system 1603 may be implemented as a single storage device, but may also be implemented across multiple storage devices or subsystems that are co-located or distributed relative to each other. Storage system 1603 may include additional elements, such as controllers, that are capable of communicating with processing system 1602 or possibly other systems.
[0102] Software 1605 (including learning process 1606) may be implemented in program instructions and, among other functions, when executed by processing system 1602, software 1605 may direct processing system 1602 to operate as described in connection with the various operation scenarios, sequences, and processes shown herein. For example, software 1605 may include program instructions for implementing a reinforcement learning process to learn an optimal scheduling strategy as described herein.
[0103] In particular, the program instructions may include various components or modules that cooperate or otherwise interact to perform the various processes and operation scenarios described herein. The various components or modules may be embodied in compiled or interpreted instructions, or in some other variant or combination of instructions. The various components or modules may execute in a synchronous or asynchronous manner, serially or in parallel, in a single-threaded environment or in a multi-threaded environment, or according to any other suitable execution paradigm, variant, or combination thereof. Software 1605 may include additional processes, programs, or components, such as operating system software, virtualization software, or other application software. Software 1605 may also include firmware or some other form of machine-readable processing instructions executable by processing system 1602.
[0104] Generally, when software 1605 is loaded into and executed by processing system 1602, software 1605 may transform a suitable apparatus, system, or device (wherein computing system 1601 is representative) from a general-purpose computing system as a whole into a specialized computing system customized to provide motion learning. In fact, the encoded software 1605 on storage system 1603 may transform the physical structure of storage system 1603. The specific transformation of such physical structure may depend on various factors in different implementations of this specification. Examples of such factors may include, but are not limited to, the technology of the storage medium used to implement storage system 1603 and whether the computer storage medium is characterized as main storage or secondary storage, among other factors.
[0105] For example, if the computer-readable storage medium is implemented as a semiconductor-based memory, software 1605 may transform the physical state of the semiconductor memory when encoding program instructions therein, such as by transforming the state of transistors, capacitors, or other discrete circuit elements that make up the semiconductor memory. Similar transformations may occur for magnetic or optical media. Other transformations of the physical medium are possible without departing from the scope of this specification, and the foregoing examples are provided only to facilitate this discussion.
[0106] The communication interface system 1607 may include communication connections and devices that allow communication with other computing systems (not shown) via a communication network (not shown). Examples of connections and devices that together allow inter-system communication may include network interface cards, antennas, power amplifiers, RF circuits, transceivers, and other communication circuits. The connections and devices may communicate via a communication medium to exchange communications with other computing systems or system networks, such as metal, glass, air, or any other suitable communication medium. The aforementioned communication networks and protocols are well known and need not be discussed in detail herein. However, some communication protocols that may be used include, but are not limited to, Internet Protocol (IP, IPv4, IPv6, etc.), Transmission Control Protocol (TCP), and User Datagram Protocol (UDP), as well as any other suitable communication protocol, variants thereof, or combinations thereof.
[0107] Communication between the computing system 1601 and other computing systems (not shown) may occur via one or more communication networks and in accordance with various communication protocols, combinations of protocols, or variants thereof. Examples include intranets, the Internet, the World Wide Web, local area networks, wide area networks, wireless networks, wired networks, virtual networks, software-defined networks, data center buses and backplanes, or any other type of network, combinations of networks, or variants thereof. The aforementioned communication networks and protocols are well known and need not be discussed in detail herein.
[0108] As will be understood by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects (commonly referred to herein as "circuitry," "module," or "system"). Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0109] Phrases such as "in some embodiments," "according to some embodiments," "in the illustrated embodiments," "in other embodiments," "in some embodiments," "according to some embodiments," "in the illustrated embodiments," "in other embodiments," etc. generally mean that the particular feature, structure, or characteristic following the phrase is included in at least one embodiment or implementation of the present technology and may be included in more than one embodiment or implementation. Moreover, such phrases do not necessarily refer to the same or different embodiments or implementations.
[0110] The functional block diagrams, operational scenarios and sequences, and flowcharts provided in the drawings represent exemplary systems, environments, and methods for performing the novel aspects of the present disclosure. Although, for simplicity of explanation, the methods included herein may be in the form of functional diagrams, operational scenarios or sequences, or flowcharts and may be described as a series of acts, it should be understood and appreciated that these methods are not limited by the order of acts, as some acts may occur in a different order than shown and described herein and / or concurrently with other acts. For example, those skilled in the art will understand and appreciate that, for instance, in a state diagram, a method may alternatively be represented as a series of interrelated states or events. Additionally, not all acts shown in the methods are required for novel implementations.
[0111] The included description and drawings describe specific embodiments to teach those skilled in the art how to make and use the best mode. Some conventional aspects have been simplified or omitted for the purpose of teaching the principles of the invention. Those skilled in the art will recognize variations of these embodiments that fall within the scope of the present disclosure. Those skilled in the art will also recognize that the above features may be combined in various ways to form multiple embodiments. Accordingly, the present invention is not limited to the above specific embodiments, but is defined only by the claims and their equivalents.
Claims
1. A method, comprising: Identifying a bounding box around a character in a frame of a video, wherein the character includes an animated character in the video; Identifying a plurality of sub - frames within the frame and non - overlapping with respect to the bounding box; And For at least one sub - frame of the plurality of sub - frames: Determining that the sub - frame meets a plurality of criteria, including determining that the size of the sub - frame reaches or exceeds a size threshold, and in response to determining that the size of the sub - frame reaches or exceeds the size threshold, determining that the sub - frame is empty; And In response to determining that the sub - frame meets the plurality of criteria, designating the sub - frame as a negative example of one or more animated characters in the video, the one or more animated characters including the animated character in the frame.
2. The method according to claim 1, further comprising: Determining that the sub - frame is empty based at least on the absence of target content in the sub - frame, wherein the target content includes the one or more animated characters.
3. The method according to claim 2, further comprising: In response to determining that the sub - frame is empty but does not reach or exceed the size threshold, discarding the sub - frame without designating the sub - frame as any type of example.
4. The method according to claim 3, further comprising: In response to determining that the sub - frame is not empty; Identifying other sub - frames within the sub - frame and adjacent to a rectangular portion of the sub - frame, the rectangular portion of the sub - frame including at least a part of another bounding box around another animated character; Identifying at least one other sub - frame among the other sub - frames that meets the plurality of criteria; And Classifying the at least one other sub - frame as the negative example of the one or more animated characters in the video.
5. The method according to claim 4, further comprising: Including the negative example in a negative example set of the one or more animated characters; And Training a machine learning model based on training data to identify instances of the one or more animated characters, the training data including the negative example set and positive example set of the one or more characters.
6. A method of identifying samples used to train a machine learning model, comprising: Identifying one or more regions of interest around target content in a frame of a video, wherein the target content includes an animated character in the video; In a portion of the frame outside the region of interest, identifying potential empty regions adjacent to the region of interest, wherein the region of interest includes a bounding box around the animated character; Identifying at least one empty region in the potential empty regions that meets one or more criteria, including: determining that the size of the potential empty region reaches or exceeds a size threshold, and in response to determining that the size of the potential empty region reaches or exceeds the size threshold, determining that the potential empty region is empty; And Classifying the at least one empty region as a negative sample of the target content.
7. The method according to claim 6, further comprising: Including the negative sample of the target content in a negative sample set of the target content; And Training a machine learning model based on training data including the negative sample set to identify instances of the target content.
8. The method according to claim 6, wherein: The potential empty regions adjacent to the region of interest include rectangles, and each potential empty region in the potential empty regions has a side adjacent to the bounding box.
9. The method according to claim 8, wherein: The rectangle includes an empty rectangle that does not overlap with any of the one or more regions of interest surrounding the target content; And Identifying the at least one empty region that meets the one or more criteria includes: identifying the largest empty rectangle among the empty rectangles.
10. The method according to claim 8, wherein for each rectangle in the rectangles, the one or more criteria include: Whether a given rectangle is counted as empty because it does not include any of the one or more regions of interest surrounding the target content; And Whether the size of the given rectangle reaches a size threshold.
11. The method according to claim 10, further comprising: For a rectangle that is counted as empty but does not reach the size threshold, the rectangle is discarded and not classified as any type of sample of the target content.
12. The method according to claim 10, further comprising: For a rectangle that is not counted as empty: Identify potential empty rectangles adjacent to the following rectangular portion of the rectangle, the following rectangular portion including at least a part of another bounding box surrounding another animated character; Identify at least one empty rectangle among the potential empty rectangles that is counted as empty and reaches the size threshold; And Classify the at least one empty rectangle as a negative sample of the target content.
13. The method according to claim 6, wherein: The target content includes an animated character in the video; The one or more regions of interest include bounding boxes drawn around instances of the animated character in the frame; The portion of the frame outside the region of interest includes a boundary region defined by the boundary of the region of interest and the boundary of the frame; and The region of interest includes the most central bounding box among the bounding boxes.
14. A computing device, comprising: One or more computer-readable storage media; One or more processors, the one or more processors operatively coupled to the one or more computer-readable storage media; And Program instructions stored on the one or more computer-readable storage media and used to identify samples utilized to train a machine learning model, which when executed by the one or more processors, instruct the computing device to at least: Identify one or more regions of interest around a target content in a frame of a video, wherein the target content includes an animated character in the video; In a portion of the frame outside the region of interest, identify potential empty regions adjacent to the region of interest; Identify at least one empty region among the potential empty regions that meets one or more criteria, including: determining that the size of the potential empty region reaches or exceeds a size threshold, and in response to determining that the size of the potential empty region reaches or exceeds the size threshold, determining that the potential empty region is empty; And Classify the at least one empty region as a negative sample of the target content.
15. The computing device according to claim 14, wherein the program instructions further instruct the computing device to: include the negative samples of the target content in the negative sample set of the target content; and train a machine learning model based on the training data including the negative sample set to identify instances of the target content.
16. The computing device according to claim 14, wherein: the region of interest includes a bounding box around the animated character; and the potential empty regions adjacent to the region of interest include rectangles, with one side of each rectangle adjacent to the bounding box.
17. The computing device according to claim 16, wherein: the rectangles include empty rectangles that do not overlap with any of the one or more regions of interest around the target content; and to identify at least one empty region that meets the one or more criteria, the program instructions instruct the computing device to identify the largest empty rectangle among the empty rectangles.
18. The computing device according to claim 16, wherein an empty region meets the one or more criteria based on: the empty region lacks any of the one or more animated characters and the size of the empty region meets a size threshold.
19. The computing device according to claim 18, wherein: for rectangles that are not counted as empty, the program instructions further instruct the computing device to: identify potential empty rectangles adjacent to a rectangular portion of the rectangle that includes at least a portion of another bounding box around another animated character; identify at least one empty rectangle among the potential empty rectangles that is counted as empty and meets the size threshold; and classify the at least one empty rectangle as a negative sample of the target content; and for a rectangle that is counted as empty but does not meet the size threshold, discard the rectangle without classifying it as any type of sample of the target content.
20. The computing device according to claim 14, wherein: the one or more regions of interest include bounding boxes drawn around instances of the animated character in the frame; the portion of the frame outside the region of interest includes a boundary region defined by the boundary of the region of interest and the boundary of the frame; and the region of interest includes the most central bounding box among the bounding boxes.
Citation Information
Patent Citations
Improved online Boosting target tracking method
CN108564598A