Embedded adaptive content evaluation based on machine learning models
Patent Information
- Application Number
- CN202280019761.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-16
- Filing Date
- 2022-03-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-03-17
Smart Images

Figure CN117015809B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims the benefit and priority of pending provisional patent application No. 63 / 165,924, filed on March 25, 2021, entitled “VideoEmbedding for Classification”, which is incorporated herein by reference in its entirety. Background Technology
[0003] As visual media has become an almost universally popular content medium, an increasing amount of visual media content is being produced and delivered to consumers. Therefore, the efficiency of visual image analysis, classification, and processing is becoming increasingly important for producers, owners, and publishers of visual media content.
[0004] A major challenge in the effective classification and processing of visual media content is that entertainment and media studios produce many different types of content with distinct characteristics, such as varying visual textures and forms of motion. For example, in the case of audio-video (AV) film and television content, the content produced can include live-action content with realistic computer-generated imagery (CGI) elements, highly complex 3D animation, and even 2D hand-drawn animation. Furthermore, each different type of content produced may require different processing before, after, or both.
[0005] For example, consider the post-production processing of AV or video content. Different types of AV or video content may benefit from different streaming encoding schemes or different localization workflows. In traditional techniques, categorizing content into specific types is often done manually through manual inspection, and in example use cases of video encoding, even after manual inspection, the most suitable workflow may be unidentifiable, requiring trial and error to determine how to categorize content for encoding purposes. This categorization process can be particularly challenging for hybrid content types (such as animation embedded within other live-action content) or for visually complex 3D animation, where visually complex 3D animation may be better suited to post-production using a live-action content workflow than a traditional animation workflow. Attached Figure Description
[0006] Figure 1 A schematic diagram of an example system for performing embedded adaptive content evaluation based on a machine learning (ML) model, according to one embodiment, is shown.
[0007] Figure 2A The description of an embodiment applicable to [the study] is shown. Figure 1A diagram illustrating an exemplary training process for the ML model of the system;
[0008] Figure 2B A description of an application according to another embodiment is shown. Figure 1 A diagram illustrating an exemplary training process for the ML model of the system;
[0009] Figure 3A An exemplary two-dimensional (2D) subspace of a continuous multidimensional vector space according to one embodiment is shown, which includes embedded vector representations of content with respect to a particular similarity measure;
[0010] Figure 3B An example of embedding clustering is shown. Figure 3A The subspace, each cluster relative to Figure 3A The mapping is based on a similarity metric to identify different content categories; and
[0011] Figure 4 A flowchart illustrating an exemplary method for performing embedded adaptive content evaluation based on an ML model, according to one embodiment, is shown. Detailed Implementation
[0012] The following description contains specific information relating to the embodiments in this disclosure. Those skilled in the art will recognize that this disclosure may be implemented in ways other than those specifically discussed herein. The accompanying drawings and their detailed descriptions in this application are for exemplary embodiments only. Unless otherwise stated, the same or corresponding elements in the drawings may be indicated by the same or corresponding reference numerals. Furthermore, the drawings and illustrations in this application are generally not drawn to scale and are not intended to correspond to actual relative dimensions.
[0013] As mentioned above, entertainment and media studios produce many different types of content with varying characteristics, such as different visual textures and forms of motion. For example, in the case of audio-video (AV) or video content, the content produced can include live-action content with realistic computer-generated imagery (CGI) elements, highly complex three-dimensional (3D) animation, or even two-dimensional (2D) hand-drawn animation. Each different type of content produced may require different processing before, after, or both.
[0014] As mentioned above, in the post-production processing of AV or video content, different types of video content can benefit from different streaming encoding schemes or different localization workflows. In traditional techniques, content categorization into specific types is done manually through manual inspection, and in example use cases of video encoding, even after manual inspection, the most suitable workflow may be unidentifiable, requiring trial and error to determine how to categorize the content for encoding. This categorization process can be particularly challenging for hybrid content types (such as animation embedded within other live-action content) or for visually complex 3D animation, where visually complex 3D animation may be better suited to post-production using a live-action content workflow than a traditional animation workflow.
[0015] This application discloses systems and methods for performing embedded adaptive content evaluation based on machine learning (ML) models. It should be noted that the disclosures provided in this application focus on optimization within the encoding pipeline of video streams. Examples of tasks considered include 1) selection of preprocessing and post-processing algorithms or algorithm parameters, 2) automatic encoding parameter selection for each title or segment, and 3) automatic bitrate ladder selection for adaptive streaming for each title or segment. However, the ML model-based adaptive evaluation solution of this invention is task-independent and can be used in environments other than those specifically described herein.
[0016] Therefore, although the adaptive content evaluation solution of the present invention has been described in detail below with reference to exemplary use cases of video coding for clarity of concept, the novel and inventive principles of the present invention can be more generally applied to a variety of other content post-production and pre-production processes, such as colorization, color correction, content restoration, mastering, and audio cleanup or synchronization, to name just a few. Furthermore, the adaptive content evaluation solution disclosed in this application can be advantageously implemented as an automated process.
[0017] As defined herein, the terms “automation,” “automated,” and “automating” refer to systems and processes that do not require human intervention. While in some embodiments human editors may review content evaluations performed by the system and use the methods described herein, human involvement is optional. Therefore, the methods described herein can be executed under the control of the hardware processing components of the disclosed automated system.
[0018] Furthermore, as defined in this application, the term "ML model" can refer to a mathematical model used for making future predictions based on patterns learned from data samples or "training data." Various learning algorithms can be used to map the correlations between input and output data. These correlations form a mathematical model that can be used to make future predictions on new input data. Such a predictive model can include one or more logistic regression models, Bayesian models, or neural networks (NNs).
[0019] In the context of deep learning, "deep neural network" can refer to a neural network that utilizes multiple hidden layers between the input and output layers, which allows learning based on features not explicitly defined in the original data. As used in this application, the features identified as neural networks refer to deep neural networks.
[0020] Figure 1 A system 100 for performing embedded adaptive content evaluation based on a machine learning model is shown according to an exemplary embodiment. Figure 1 As shown, system 100 includes a computing platform 102 having processing hardware 104 and a system memory 106 implemented as a computer-readable non-transitory storage medium. According to this exemplary embodiment, system memory 106 stores software code 110, one or more ML models 120 (hereinafter referred to as "ML models 120"), and a content and classification database 112 storing category assignments determined by ML models 120.
[0021] like Figure 1 As further shown, system 100 is implemented in a usage environment that includes a communication network 114 providing a network communication link 116, a training database 122, a user system 130 including a display 132, and a user 118 of the user system 130. Figure 1 The diagram also illustrates training data 124, input data 128, and content classification 134 of input data 128 determined by system 100. It should be noted that although the exemplary use cases of system 100 described below involve system 110 performing classification of input data 128, in some embodiments, system 100 may perform regression rather than classification on content 100, as these processes are known and distinguishable in the art.
[0022] Although for clarity of concept, this application describes one or more of the software code 110, ML model 120, and content and classification database 112 as stored in system memory 106, more generally, system memory 106 can take the form of any computer-readable non-transitory storage medium. The expression "computer-readable non-transitory storage medium" as defined in this application refers to any medium, excluding carrier waves or other transient signals that provide instructions to the processing hardware 104 of the computing platform 102. Therefore, computer-readable non-transitory storage medium can correspond to various types of media, such as volatile and non-volatile media. Volatile media can include dynamic memory, such as dynamic random access memory (dynamic RAM), while non-volatile memory can include optical, magnetic, or electrostatic storage devices. Common forms of computer-readable non-transitory storage media include, for example, optical discs, RAM, programmable read-only memory (PROM), erasable PROM (EPROM), and flash memory.
[0023] Furthermore, despite Figure 1 The software code 110, ML model 120, and content and classification database 112 are depicted as residing together in system memory 106, but this illustration is provided merely for conceptual clarity. More generally, system 100 may include one or more computing platforms 102, such as computer servers, which may be located in the same location or may form an interconnected but distributed system, such as a cloud-based system. As a result, processing hardware 104 and system memory 106 may correspond to distributed processor and memory resources within system 100. Therefore, in some embodiments, one or more of the software code 110, ML model 120, and content and classification database 112 may be stored remotely to each other on the distributed memory resources of system 100. It is also noted that in some embodiments, ML model 120 may take the form of one or more software modules included in the software code 110.
[0024] Furthermore, despite Figure 1 The training database 122 is shown remotely from system 100, but this illustration is merely exemplary. In some embodiments, the training database 122 may be included as a feature of system 100 and may be stored in system memory 106.
[0025] Processing hardware 104 may include multiple hardware processing units, such as one or more central processing units, one or more graphics processing units, one or more tensor processing units, one or more field-programmable gate arrays (FPGAs), and application programming interface (API) servers. As a limitation, the terms "central processing unit" (CPU), "graphics processing unit" (GPU), and "tensor processing unit" (TPU) as used herein have their conventional meanings in the art. That is, the CPU includes an arithmetic logic unit (ALU) for performing arithmetic and logical operations of the computing platform 102, and a control unit (CU) for retrieving programs (e.g., software code 110) from system memory 106, while the GPU may be implemented to reduce the CPU's processing power by performing computationally intensive graphics processing tasks or other processing tasks. A TPU is an application-specific integrated circuit (ASIC) specifically configured for artificial intelligence (AI) processes such as machine learning.
[0026] In some embodiments, computing platform 102 may correspond to one or more web servers, for example, and may be accessed via a communication network 114 in the form of a packet-switched network such as the Internet. Furthermore, in some embodiments, communication network 114 may be a high-speed network suitable for high-performance computing (HPC), such as a 10 GigE network or an Infiniband network. In some embodiments, computing platform 102 may correspond to one or more computer servers supporting a private wide area network (WAN), a local area network (LAN), or be included in another type of limited distributed or private network. Alternatively, in some embodiments, system 100 may be implemented virtually, for example, virtually in a data center. For example, in some embodiments, system 100 may be implemented in software or as a virtual machine.
[0027] Despite user system 130 in Figure 1 The illustration is shown as a desktop computer, but this is merely an example. More generally, user system 130 can be any suitable mobile or fixed computing device or system that includes a display 132 and implements sufficient data processing capabilities to provide a user interface, support connectivity to communication network 114, and implement the functions attributed herein to user system 130. For example, in other embodiments, user system 130 may take the form of, for example, a laptop computer, tablet computer, or smartphone.
[0028] Regarding the display 132 of the user system 130, the display 132 can be implemented as a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot (QD) display, or any other suitable display screen that performs the physical conversion of signals to light. Furthermore, the display 132 can be physically integrated with the user system 130, or it can be communicatively coupled to the user system 130 but physically separate from it. For example, when the user system 130 is implemented as a smartphone, laptop computer, or tablet computer, the display 132 will typically be integrated with the user system 130. In contrast, when the user system 130 is implemented as a desktop computer, the display 132 can take the form of a display separate from the desktop host form of the user system 130.
[0029] Input data 128 and training data 124 may include segmented content in the form of video clips (e.g., sampled frames), which may include raw frames, encoded frames, or both. Furthermore, in some embodiments, input data 128, training data 124, or both may be augmented with additional data, such as one or more of encoded statistics, distortion maps or metrics, pre-computed features such as per-pixel noise or texture information, or any combination thereof. Therefore, in various embodiments, input data 128 and training content 124 may be 3D (e.g., in the case of video), 2D (e.g., in the case of a single frame), 1D (e.g., in the case of per-frame values), or even a single variable of a clip.
[0030] In the case of AV or video content, input data 128 and training data 124 may include content segmented by shots, scenes, timecode intervals, or as content as individual video frames. Regarding the term "shot," as defined for the purposes of this application, a "shot" refers to a continuous series of video frames captured from a unique camera perspective, without cuts or other cinematic transitions, and a scene typically includes multiple shots. Optionally, input data 128 and training data 124 may include content segmented using the techniques described in U.S. Patent Application Publication No. 2021 / 0076045, published March 11, 2021, entitled "Content Adaptive Boundary Placement for Distributed Encodes," which is incorporated herein by reference in its entirety. It should be noted that, in various embodiments, input data 128 and training data 124 may include video content without audio, audio content without video, AV content, text, or content in any other format.
[0031] ML model 120 includes an ML-based embedder trained using training data 124 selected based on one or more similarity metrics. This similarity metric can be a quantitative similarity metric, i.e., objectively similar, or a perceptual similarity metric under human inspection, i.e., subjectively similar. Examples of perceptual similarity metrics for AV and video content can include texture, motion, and perceptual coding quality, etc. Examples of quantitative similarity metrics for AV and video content can include bitrate distortion curves, pixel density, computed optical flow, and computed coding quality, to name just a few.
[0032] refer to Figure 2A and 2B , Figure 2A Figure 200A illustrates an exemplary contrastive learning process for training an ML model-based embedder 226 included in an ML model 120, according to one embodiment. Figure 2B Figure 200B illustrates an exemplary contrastive learning process for training an embedding 226 based on an ML model, according to another embodiment.
[0033] like Figure 2A As shown, in addition to the ML model-based embedder 226, Figure 200A also includes training content fragments 224a, 224b, and 224d, wherein the similarity metric of training content fragment 224b has the same value as the same similarity metric of training content fragment 224a, and the similarity metric of training content fragment 224d has a different value than the same similarity metric of training content fragment 224a. Furthermore, Figure 2A The embedding vectors 240a, 240b, and 240d of training content fragment 224a, training content fragment 224b, and training content fragment 224d are shown. Figure 2A The diagram also shows the distance function 242 for comparing embedding vectors 240a and 240b, and for comparing embedding vectors 240a and 240d. It should be noted that training content fragments 224a, 224b, and 224d typically correspond to... Figure 1 The training data is 124.
[0034] like Figure 2A As shown, a contrastive learning process is used to train an ML-based embedding 226 to identify training content fragments 224a and 224b as similar, while minimizing the distance function 242. Figure 2A As further shown, the ML-based embedder 226 is also trained using a contrastive learning process to identify training content fragments 224a and 224d as dissimilar, while maximizing the distance function 242.
[0035] like Figure 2BAs shown, in addition to the ML model-based embedder 226, Figure 200B also includes training content fragments 224e, 224f, 224g, and 224h. The similarity metric of training content fragment 224f has the same value as the same similarity metric of training content fragment 224e; the similarity metric of training content fragment 224g has a different value than the same similarity metric of training content fragments 224e and 224f; and the similarity metric of training content fragment 224h has the same value as the same similarity metric of training content fragment 224g, but a different value than the same similarity metric of training content fragments 224e and 224f. Furthermore, Figure 2B The embedding vectors 240e, 240f, 240g, and 240h of training content fragment 224e, training content fragment 224f, training content fragment 224g, and training content fragment 224h are shown. Figure 2B The diagram also shows a classification or regression block 260, configured to perform either classification or regression processing on the embedding vectors 240e, 240f, 240g, and 240h received from the ML model-based embedder 226, respectively. It should be noted that the training content fragments 224e, 224f, 224g, and 224h typically correspond to... Figure 1 The training data is 124.
[0036] The ML-based embedder 226 is responsible for mapping content fragments to embedding vectors. Example implementations of the ML-based embedder 226 include, but are not limited to, one or more 1D, 2D, or 3D convolutional neural networks (CNNs) with early or late fusion, which can be trained from scratch or pre-trained to leverage transfer learning. Depending on the target task, features extracted from different layers of the pre-trained CNN, such as the last layer of a Visual Geometry Group (VGG) CNN or a Residual Network (ResNet), can be used to form the embedding vectors.
[0037] The classification or regression block 260 is responsible for performing classification or regression tasks, such as selecting preprocessing or postprocessing algorithms or parameters, bitrate distortion prediction (where distortion can be measured using various quality metrics), automatic encoding parameter selection for each title or segment, and automatic bitrate ladder selection for each title or segment, thereby enabling the prediction of the highest bitrate in adaptive streaming to achieve a specific perceived quality for a given title or segment that needs to be encoded, to name just a few examples.
[0038] The classification or regression block 260 can be implemented as, for example, a similarity metric with an added threshold to determine which embedding cluster a particular embedding vector belongs to among the different clustering groups available in the continuous vector space to which the embedding vector is mapped. Alternatively, the classification or regression block 260 can be implemented as a neural network (NN) or other ML model architecture included in the ML model 120, wherein the ML model 120 is trained to classify or regress the embedding vectors to a base-true result. Furthermore, in some embodiments, the classification or regression block 260 can be integrated with the ML model-based embedder 226 and can act as, for example, one or more layers of the ML model-based embedder 226.
[0039] Figure 2A and Figure 2B The contrastive learning process described herein can be replicated on corpora of similar and dissimilar training content fragment pairs. For example, in a use case where the content to be classified is AV or video content, similar training content fragments could both be live-action clips or hand-drawn animation, while each of the different training content fragments could be a different type of video content. In a specific use case of video coding, content fragments can be labeled as similar for training purposes based on both training content fragments performing well on a particular coding scheme (according to a performance threshold), regardless of whether one fragment includes live-action content and the other includes animation. In this use case, if training content fragments perform well on different coding schemes, they can be labeled as dissimilar even if they share the same content type, such as live-action or animation.
[0040] The training of the ML model-based embedder 226, the classification or regression block 260, or both the ML model-based embedder 226 and the classification or regression block 260, can be performed by software code 110 executed by the processing hardware 104 of the computing platform 102. In some use cases, the ML model-based embedder 226 and the classification or regression block 260 can be trained independently of each other.
[0041] Optionally, in some embodiments, an ML model-based embedder 226 can be trained first, and the embedding vectors provided by the ML model-based embedder 226 can then be used for different downstream classification or regression tasks. In this case, the ML model-based embedder 226 can be trained by: a) identifying content segments considered similar and feeding them as training data while minimizing the distance function between them; b) identifying two content segments considered dissimilar and feeding them as training data while maximizing the distance function between them; and c) repeating steps a) and b) during training.
[0042] As another alternative, in some embodiments, the ML model-based embedder 226 and the classification or regression block 260 may be trained together. Furthermore, in some embodiments where the ML model-based embedder 226 and the classification or regression block 260 are trained together, for example where the classification or regression block 260 is integrated with the ML model-based embedder 226 as one or more layers of the ML model-based embedder 226, the ML model-based embedder 226 including the classification or regression block 260 may be trained using end-to-end learning.
[0043] After the training of both the ML model-based embedder 226 and the classification or regression block 228, or the ML model-based embedder 226 and the classification or regression block 260, the processing hardware 104 can execute software code 110 to receive input data 128 from the user system 130 and use the ML model-based embedder 226 to transform the input data 128 or fragments thereof into vector representations of the contents mapped to a continuous one-dimensional or multi-dimensional vector space (hereinafter referred to as "embedded vectors"), thereby producing embedded vector representations of the contents in that vector space.
[0044] Figure 3A An exemplary 2D subspace 300 of a continuous multidimensional vector space 350 according to one embodiment is shown. It should be noted that the continuous multidimensional vector space 350 can be a relatively low-dimensional space, such as a 64, 128, or 256-dimensional space. Alternatively, in some embodiments, the continuous multidimensional vector space 350 can be a relatively high-dimensional space with tens of thousands of dimensions, such as 20,000 (20k) dimensions. Figure 3A The diagram also shows the embedding vectors 352a, 352b, 352c, 352d, 352e, 352f, 352g, 352h, 352i, and 352j (hereinafter referred to as "embedded vectors 352a-352j") of the mapping in the continuous multidimensional vector space 350, each embedded vector corresponding to... Figure 1 The input data 128 includes different content samples.
[0045] In addition to using an ML model-based embedder 226 to map the embedding vectors 352a-352j onto a continuous multidimensional vector space 350, the software code 110 can also perform an unsupervised clustering process when executed by the processing hardware 104 to identify each cluster corresponding to a different content category relative to a similarity metric used for comparing the content. Figure 3B A subspace 300 is shown, comprising the continuous multidimensional vector space 350 containing embedded vectors 352a-352j. Figure 3BDifferent clusters, 354a, 354b, 354c and 354d (hereinafter referred to as "clusters 354a-354d"), are also shown, each cluster identifying different content categories relative to a specific similarity metric.
[0046] In the case of AV and video content, for example, the embedding vectors 352a and 352b of live-action content are mapped to regions of the multidimensional vector space 350 identified by cluster 354c, while the embedding vectors 352c, 352d, 352f, and 352h of low-complexity animation content such as hand-drawn and other 2D animations are mapped to different regions of the continuous multidimensional vector space 350 identified by clusters 354a and 354d. The embedding vectors 352e, 352i, and 352j of high-complexity animation content such as 3D animation are shown to be mapped to yet another different region of the continuous multidimensional vector space 350 identified by cluster 354b. It should be noted that the embedding vector 352g of mixed content types, such as animation mixed with live-action, can be mapped to the boundaries of animation clusters, the boundaries of live-action clusters, or the boundaries between these clusters.
[0047] In cases where the content corresponding to embedding vectors 352a-352j is AV or video content and the process of classifying this content is video encoding, for example, each cluster 354a-354d may correspond to a different codec. For instance, cluster 354c may identify content requiring a high bitrate codec, while cluster 354a may identify content for which a low bitrate codec is sufficient. Clusters 354b and 354d may use other specific codecs to identify content. In one such embodiment, where a new codec is introduced or an existing codec is deactivated, system 100 may be configured to automatically re-evaluate embedding vectors 352a-352j relative to the changed set of available codecs. Similar re-evaluation may be performed on any other process to which this concept is applied.
[0048] It should be noted that the continuity of the multidimensional vector space 350 advantageously enables the adjustment of how the embedding vectors corresponding to the content are clustered into categories for each individual use case by utilizing different clustering algorithms and thresholds. Compared to conventional classification methods that rely on prior knowledge of the number of classification labels to be trained, this novel and inventive embedding method is applicable to a variety of classification schemes. Furthermore, due to the unsupervised nature of the clustering performed as part of this adaptive content evaluation solution, the method disclosed in this application can generate unexpected insights into the similarity between seemingly dissimilar content items.
[0049] Reference Figure 4 Further describe the functions of system 100. Figure 4Flowchart 470 is shown, which illustrates an exemplary method according to one embodiment for performing ML model-based embedding vectors for adaptive content evaluation. About Figure 4 The method outlined in the flowchart is shown in the figure. It should be noted that certain details and features have been omitted from the flowchart 470 so as not to hinder the discussion of the features of the invention in this application.
[0050] Now combine Figure 1 And refer to Figure 4 Flowchart 470 includes receiving input data 128 (action 471) comprising multiple content fragments. For example... Figure 1 As shown, system 100 can receive input data 128 from user system 130 via communication network 114 and network communication link 116. As described above, input data 128 may include video content without audio, wherein the video content includes raw frames, encoded frames, or both. Optionally, input data 128 may include audio content without video, AV content, text, or content in any other format. As further discussed above, in some embodiments, input data 128 may also include one or more of encoded statistics, distortion maps or metrics, pre-computed features such as per-pixel noise or texture information, or any combination thereof. Input data 128 may be received by software code 110 in action 471 and executed by processing hardware 104 of computing platform 102.
[0051] Combination Figure 1 and Figure 4 And refer to Figure 2A , Figure 2B and Figure 3A Flowchart 470 also includes using an ML model-based embedder included in ML model 120 to map each of the plurality of content fragments received in action 471 to a corresponding embedding vector in continuous vector space 350, to provide embedding vectors (e.g., embedding vectors 352a-352j) corresponding to the plurality of content fragments respectively (action 472). The mapping of the plurality of content fragments received in action 471 to the embedding vectors in continuous vector space 350 can be executed in action 472 by software code 110, which is executed by processing hardware 104 of computing platform 102, and uses ML model-based embedder 226.
[0052] As described above, the ML-based embedder 226 can be trained using contrastive learning based on one or more similarity metrics. Also as described above, the one or more similarity metrics upon which the contrastive learning of the ML-based embedder 226 is based may include quantitative similarity metrics, perceptual similarity metrics, or both. In some embodiments, as described above, the ML-based embedder 226 may include one or more of, for example, a 1D CNN, a 2D CNN, or a 3D CNN with early or late fusion. Furthermore, regarding the continuous vector space 350, it should be noted that in some embodiments, the continuous vector space 350 may be multidimensional, such as... Figure 3A as well as Figure 3B As shown.
[0053] Flowchart 470 also includes performing one of the classification or regression of content fragments using the mapped embedding vectors (e.g., embedding vectors 352a-352j) (action 473). In some embodiments, as referenced above... Figure 3B The classification performed in action 473 may include grouping each of at least one mapped embedding vector (e.g., embedding vectors 352a-352j) into one or more clusters (e.g., clusters 354a-354d), each cluster corresponding to a different category of the similarity metric on which the contrastive learning of the embeddinger 226 based on the ML model is based. As described above, such clustering can be performed as an unsupervised process.
[0054] As also described above, in some embodiments, the classification or regression performed in action 473 may be performed using classification or regression block 260, which may take the form of a trained neural network (NN) or other ML model architecture included in ML model 120. Furthermore, as further discussed above, in some embodiments, classification or regression block 260 may be integrated with ML model-based embedder 226 and may act as, for example, one or more layers of ML model-based embedder 226. Action 473 may be performed by software code 110, executed by processing hardware 104 of computing platform 102, and in some embodiments, performed using classification or regression block 260.
[0055] Flowchart 470 also includes finding at least one new label (action 474) based on one of the classifications or regressions performed in action 473 to characterize the multiple content fragments received in action 471. For example, as referenced above. Figure 3BAs described, due to the unsupervised nature of the clustering performed as part of this adaptive content evaluation solution, the method outlined by flowchart 470 can generate unexpected insights into the similarity between seemingly dissimilar content items. At least one new label discovered in action 474 can be stored in content and classification database 112. Furthermore, or alternatively, at least one new label discovered in action 474 can be transferred to training database 122 for storage and inclusion in training data 124.
[0056] For example, at least one new tag discovered in action 474 can advantageously lead to the implementation of new, more efficient AV or video encoding parameters. Furthermore, or alternatively, the information discovered as part of action 474 can be used to selectively enable one or more currently unused encoding parameters and selectively disable one or more currently used encoding parameters. As yet another example, the information discovered as part of action 474 can be used to selectively enable or disable certain preprocessing elements in the code conversion pipeline, such as content-feature-based denoising or debanding. Action 474 can be executed by software code 110 executed by the processing hardware 104 of computing platform 102.
[0057] In some embodiments, the method outlined in flowchart 470 may end with action 474 described above. However, in other embodiments, the method outlined in flowchart 470 may further include using contrastive learning and at least one new label discovered in action 474 to further train the ML-based embedding 226 (action 475). That is, action 475 is optional. When included in the method outlined in flowchart 470, action 475 may be executed by software code 110, by the processing hardware 104 of computing platform 102, and may advantageously result in refinement and improvement of the future classification or regression performance of system 100. Regarding the actions included in flowchart 470, note that actions 471, 472, 473, and 474 (hereinafter referred to as "actions 471-474") or actions 471-474 and 475 may be performed as automated processes, where human intervention may be omitted.
[0058] Therefore, this application discloses a system and method for performing ML model-based embedding vectors for adaptive content evaluation. The solutions disclosed in this application advance the prior art by proposing and providing solutions for the lack of clearly labeled, annotated data in downstream tasks, but with the aim of discovering how and why different things respond differently to that task. The novel and inventive concepts disclosed in this application can be used to automate the process of discovering appropriate labels for these differences (automatic discovery) and to train models so that, given a specific set of parameters for a downstream task, predictions can be generated about which of those different groups something will fall into.
[0059] As described above, this application discloses two methods to solve the problem of automatically discovering tags that can be linked together. In the first method, as referenced... Figure 2A , Figure 3A and Figure 3B The described approach employs a contrastive learning method. Embedded vectors are mapped to a continuous vector space, and then unsupervised clustering analysis is performed to discover potentially distinct groupings and any additional, unforeseen groupings. This first approach can be improved by adding more signals to the data used to create the continuous vector space to which the embedded vectors are mapped, providing a more complex vector space that is sensitive to more factors and will result in more diverse groupings.
[0060] In the second method, refer to the above. Figure 2B The first method described above can be further supplemented by adding classification or regression blocks for different types of downstream tasks. In this embodiment, the ML-based embedder and the classification or regression block, which can be implemented as a neural network (NN), can be trained together using end-to-end learning, while maintaining the automatic discovery of labels so advantageously achieved by this system and method.
[0061] Based on the above description, it will be apparent that various techniques can be used to implement these concepts without departing from the scope of the concepts described herein. Furthermore, although these concepts have been described with specific reference to certain embodiments, those skilled in the art will recognize that changes in form and detail may be made without departing from the scope of these concepts. Therefore, the described embodiments are to be considered illustrative rather than restrictive in all respects. It should also be understood that this application is not limited to the specific embodiments described herein, and many rearrangements, modifications, and substitutions are possible without departing from the scope of this disclosure.
Claims
1. A system for performing embedded adaptive content evaluation based on a machine learning model, comprising: Processing hardware; and The system memory stores software code and at least one machine learning model trained using contrastive learning based on a similarity metric to map each of a plurality of video segments to a corresponding embedding vector in a continuous vector space, wherein the similarity metric is a video coding metric, and the machine learning model is trained to label a pair of video segments as similar based on both performing well on a particular coding scheme, and to label a pair of video segments as dissimilar based on both performing well on different coding schemes; The processing hardware is configured to execute software code, thereby: It can receive input that includes multiple video clips; Using at least one machine learning model, each of the multiple video segments is mapped to a corresponding embedding vector in a continuous vector space to provide embedding vectors that correspond to multiple mappings respectively for the multiple video segments; Perform one of the classification or regression of multiple video segments using the embedding vectors of multiple mappings; Based on either classification or regression, multiple video clips are classified into video content categories among multiple video content categories relative to the similarity measure; Determine the encoding scheme corresponding to the video content category; and The encoding scheme described above is used to encode multiple video segments.
2. The system of claim 1, wherein the plurality of video content categories include live action and animation, and wherein the video content category is one of live action or animation.
3. The system of claim 1, wherein the classification comprises grouping each of at least one of the embedding vectors of the plurality of mappings into one or more clusters, each cluster corresponding to a different category of the similarity measure.
4. The system of claim 1, wherein the processing hardware is further configured to execute the software code, thereby: Based on the video content category, a preprocessing algorithm is selected for preprocessing multiple video segments; and The selected preprocessing algorithm is used to preprocess multiple video segments.
5. The system according to claim 1, wherein the at least one machine learning model comprises at least one of a one-dimensional convolutional neural network, a two-dimensional convolutional neural network, or a three-dimensional convolutional neural network.
6. The system according to claim 1, wherein the continuous vector space is multidimensional.
7. The system of claim 1, wherein the similarity measure includes one of a quantitative similarity measure or a perceptual similarity measure.
8. The system of claim 1, wherein one of the classification or regression is performed using a corresponding trained classification machine learning model or a corresponding trained regression machine learning model, and wherein at least one machine learning model and the corresponding trained classification machine learning model or trained regression machine learning model are trained independently of each other.
9. The system of claim 1, wherein one of the classification or regression is performed using a trained classification machine learning model or a trained regression machine learning model, and wherein the trained classification machine learning model or the trained regression machine learning model includes a trained neural network.
10. The system of claim 1, wherein one of the classifications or regressions is performed using a corresponding one of the classification or regression blocks of at least one machine learning model, and the at least one machine learning model including the corresponding one of the classification or regression blocks is trained using end-to-end learning.
11. A method used by a system, the system comprising processing hardware and storing software code and at least one machine learning model, the at least one machine learning model being trained using contrastive learning based on a similarity metric to map each of a plurality of video clips to a corresponding embedding vector in a continuous vector space, wherein, The similarity metric is a video coding metric, and the machine learning model is trained to label a pair of video segments as similar based on their good performance on a specific coding scheme, and to label a pair of video segments as dissimilar based on their good performance on different coding schemes. The method includes: The software code executed by the processing hardware receives input including multiple video clips; By using software code executed by the processing hardware and employing at least one machine learning model, each of a plurality of video segments is mapped to a corresponding embedding vector in a continuous vector space to provide embedding vectors that correspond to a plurality of mappings for the plurality of video segments respectively. Using software code executed by the processing hardware, one of the classification or regression operations on multiple video segments is performed using embedding vectors of multiple maps. Software code executed by processing hardware classifies multiple video clips into video content categories among multiple video content categories relative to the similarity metric, based on either classification or regression. Determine the encoding scheme corresponding to the video content category; and The encoding scheme described above is used to encode multiple video segments.
12. The method of claim 11, wherein the plurality of video content categories include live action and animation, and wherein the video content category is one of live action or animation.
13. The method of claim 11, wherein the classification comprises grouping each of at least one of the embedding vectors of the plurality of mappings into one or more clusters, each cluster corresponding to a different category of the similarity measure.
14. The method of claim 11, further comprising: Based on the video content category, select a preprocessing algorithm for preprocessing multiple video segments; as well as The selected preprocessing algorithm is used to preprocess multiple video segments.
15. The method of claim 11, wherein the at least one machine learning model comprises at least one of a one-dimensional convolutional neural network, a two-dimensional convolutional neural network, or a three-dimensional convolutional neural network.
16. The method of claim 11, wherein the continuous vector space is multidimensional.
17. The method of claim 11, wherein the similarity measure comprises one of a quantitative similarity measure or a perceptual similarity measure.
18. The method of claim 11, wherein one of the classification or regression is performed using a trained classification machine learning model or a trained regression machine learning model, and wherein the at least one machine learning model and the trained classification machine learning model or the trained regression machine learning model are trained independently of each other.
19. The method of claim 11, wherein one of the classification or regression is performed using a trained classification machine learning model or a trained regression machine learning model, and wherein the trained classification machine learning model or the trained regression machine learning model includes a trained neural network.
20. The method of claim 11, wherein one of the classification or regression is performed using a corresponding one of the classification or regression blocks of the at least one machine learning model, and wherein end-to-end learning is used to train the at least one machine learning model comprising a corresponding one of the classification or regression blocks and a trained neural network.