Type classification for video compression
Through machine learning models, the video content is type classification and dynamic encoding parameters are adjusted, which solves the problem of poor video compression effect in the existing technology, and achieves efficient video compression and resource saving.
Patent Information
- Application Number
- CN202510212344.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-26
- Filing Date
- 2025-02-25
- Publication Date
- 2025-08-26
AI Technical Summary
Existing video compression technology is difficult to automatically adjust the encoding parameters according to the type of video content, resulting in poor compression effect and it is difficult for users to select appropriate parameters to achieve the best compression effect.
The machine learning model is used to classify the video content, dynamically adjust the encoding parameters to adapt to different types of video content, and real-time updates and training are combined with feedback loops to optimize the encoder's encoding process.
Dynamic encoding according to the video content type is realized, video compression efficiency and quality is improved, computing resources are saved, and the selection of encoding parameters is optimized.
Smart Images

Figure CN120547346A_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment is directed to video compression of frames of a media stream using encoding based in part on genre classification of one or more frames. Background Art
[0002] Video compression can be used to provide simplified media streams while preserving some detail in the underlying video content. Deep video compression techniques can be used for video compression. However, such deep video compression may still require many parameters to be adjusted, such as those that determine and constrain the operation of the video compression. Many of these parameters have different effects on different videos. For example, a parameter can be used to improve the quality of a compressed portion of the video or reduce the bitrate. However, such a solution may negatively impact different types of content because it may not be appropriate for the content being compressed. While one approach might be to leave parameter selection to the user of the video compression, such as by inputting a video compression configuration, most users may not be informed of the relationship between the content's video sequences and the available parameters to provide any benefits for video compression. For example, users may not be able to determine whether using a parameter with a particular content will have a positive or negative impact. Consequently, available parameters may not be enabled by default and may not be used for video compression in consumer environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 is an illustration of a system for video compression using type classification in at least one embodiment;
[0004] Figure 2 is an illustration of aspects of a machine learning (ML) model having child ML models to perform different inferences to provide a type classification of a received frame in at least one embodiment;
[0005] Figure 3 is an illustration of aspects of a machine learning (ML) model having supervised training, unsupervised training, or semi-supervised training to provide classification of a type of received frame in at least one embodiment;
[0006] Figure 4 The computer and processor aspects of a system for video compression using type classification in at least one embodiment are shown;
[0007] Figure 5 A process flow for a system using type classification for video compression in at least one embodiment is shown;
[0008] Figure 6 Yet another process flow for a system for video compression using type classification in at least one embodiment is shown; and
[0009] Figure 7 An additional process flow for a system that uses type classification for video compression in at least one embodiment is shown. DETAILED DESCRIPTION
[0010] Figure 1 1 is a diagram of a system 100 for video compression using type classification in at least one embodiment. System 100 includes at least one circuit for use as an encoder 104, which may be a video encoder, and at least one other circuit for performing inference using a machine learning (ML) model 128. For example, the inference performed by the ML model is for determining the type of received frames in an input sequence 102 for a media stream. As used herein, type may be represented by at least one feature that can objectively distinguish between different types. Thus, determining type herein may be determining at least one feature of type by the ML model. Furthermore, as used herein, type may not be subjective because its underlying features are objectively quantified and classified by the ML model. Type, as used herein with respect to the ML model, may differ from what is readily apparent to a human observer. However, in at least one embodiment, features may be associated with the type and made apparent to a human observer through the use of labeling, for example, by using labeled features associated with the labeled type in supervised training of the ML model.
[0011] The purpose of determining the type by the ML model is to enable the encoder 104 to perform efficient encoding appropriate for the type of the input sequence 102 in the media stream. As used herein, the input sequence 102 may include multiple frames that are encoded after the underlying type can be determined by the trained ML model 128. In one example, the system 100 supports training the ML model 128 to classify each input sequence 102 of the media stream (or set of scenes within the input sequence) into at least one type. Even if shown in the singular, the input sequence is received and encoded continuously. Therefore, the input sequence can include determinations of different types therein and can be encoded differently using different encoding parameters.
[0012] The system 100 can also use the encoder 104 to select default video compression parameters that reflect different encoding parameters specific to the determined type. The system 100 can also use the encoder 104 to perform video compression or encoding that lacks some or all of the default video compression parameters. Video compression parameters are also referred to herein as encoding parameters. In at least one embodiment, the encoding parameters herein can include configuration of content adaptive parameters and an overall compression gain derived from the content adaptive parameters. The overall compression gain may be higher relative to compression lacking such content adaptive parameters.
[0013] In at least one embodiment, the types categorized herein may include, but are not limited to, nature content, camera assistance content, sports content, screen content, gaming content, cartoons, automotive content, medical imaging content, machine-generated content, handheld camera content, remote desktop applications, fixed camera content, and user-generated content (UGC). Different types may benefit from retaining different information during video compression. For example, medical imaging content may benefit from retaining information in areas relevant to the study of medical problems, while automotive content may benefit from retaining information in areas surrounding it where there are cars. Thus, the surrounding or remaining areas of any content with less or no interest may be compressed more highly than the areas where the information is retained.
[0014] The type classification performed by the ML model 128 enables the encoder 104 to perform encoding using encoding parameters that ensure that the information retained can be retained, for example, in certain areas of interest 102A, and that compression is ensured to save bits in the surrounding or remaining areas 102B of the content represented in the input sequence 102. In at least one embodiment, instead of establishing types, specific features of the video codec standard can be used to explore different types. For example, the screen content of a remote desktop application can use specific encoding parameters provided by an industry standards body. In at least one embodiment, supervised learning can be used to establish the use of such encoding parameters to train the ML model to associate with a specific type. In this way, the type can be objectified based on the encoding parameters (rather than noise characteristics, different motion vector distributions, different pixel intensity levels, or different edge characteristics).
[0015] The training of the ML model 128 can be performed naively through supervised learning, where a label or token is provided in each sequence (or set of scenes). Alternatively, the training of the ML model 128 can be performed through unsupervised learning, where different video content (or regions thereof) can initially be encoded using different encoding parameters, and these encoding sets can be classified by the trained ML model 128. In another alternative, the training of the ML model 128 can be performed through semi-supervised learning, where portions of the sequence can include labels and can be classified by the ML model 128 along with other portions of the sequence. In each approach, the ML model 128 can use best fit to inform training, for example, by establishing different types of categories. Then, during testing or in a real-time environment, the input video content of the input sequence 102 can be classified according to the established categories.
[0016] In one example, different features associated with different types can be used to train the ML model 128. In one instance, the different features can be associated with one or more of noise characteristics, motion vector distributions, pixel intensity levels, or edge characteristics. For example, one or more of different noise characteristics, different motion vector distributions, different pixel intensity levels, or different edges can be different for different types. For example, the pixel intensity levels of at least some portions of natural content can be higher than the pixel intensity levels of some medical imaging content. Similarly, the edges of cartoon or screen content can be more defined than the edges of other types.
[0017] Furthermore, different noise signatures can include white noise in an image, while different motion vector distributions enable support for content with varying motion. For example, content with low relative motion might have a low median, while content with high relative motion might have a high median, rather than a low median. Furthermore, the benefit of relying on pixel intensity levels might be important for distinguishing medical imaging content, but less so for natural content. Similarly, noise signatures and motion vector distributions might differ for gaming content and automotive content, in part due to the constant motion associated with such content. Therefore, these different signatures can distinguish between different types and can be used to define them.
[0018] Once trained, the ML model 128 can use such features in an input sequence to be encoded to determine the type of the input sequence by determining the best match to the trained categories established by the ML model 128. The encoder 104 can perform encoding on a media stream comprising the input sequence 102 based in part on encoding parameters appropriate for the determined type. The encoder 104 provides an encoded media stream, also referred to herein as an output bitstream. However, as detailed herein, the encoder 104 can be configured to provide different encoded media streams in a dynamic manner. For example, an initial input sequence can be encoded based on the first determined type. Then, based at least in part on a scene cut event that introduces at least one additional input sequence or a change in the input sequence (or set of scenes), additional classification can be performed by the trained ML model 128.
[0019] For example, an initial type can be determined by the ML model and communicated to the encoder. However, the ML model may not perform inference on the additional input sequence until the encoder communicates a scene change event in the input sequence relative to the additional input sequence to the ML model. Dynamic encoding can be supported by a feedback loop from the encoder 104 to the ML model 128. In at least one embodiment, one or more processors or execution units of the processors provide one or more different circuits that can be used to execute the encoder 104 differently from the ML model 128. Based in part on the feedback in the feedback loop, further inference types can be provided to the encoder to use different encoding parameters for the additional input sequence. The ML model 128 can be trained to classify an entire input sequence or a set of scenes into different types so that the encoder can perform encoding using different encoding parameters for different input sequences. However, the ML model 128 can be trained to classify each input sequence as it is received, and in part based on scene change events to enable the encoder 104 to perform dynamic encoding.
[0020] In addition, the classification of the ML model 128 can be used to enable the encoder 104 to make mode selections or parameter selections to perform encoding specific to the types in the input sequence. For example, the mode selection or parameter selection can be associated with the available encoding parameters in the encoding parameters to provide different but specific encodings for each type or underlying feature. In addition, while the ML model 128 can perform initial classification or inference on the media stream, a feedback loop 130 can be used to indicate scene cut events from the encoder 104 back to the ML model 128. Scene cut events can be associated with the input sequence of the media stream, for example, within two different input sequences. The scene cut event can serve as a point at which the ML model 128 participates in determining a new type for a subsequent scene or a subsequent input sequence.
[0021] In at least one embodiment, using ML model 128 to classify an input sequence of video content based in part on scene change events allows for dynamic encoding as genre changes occur in the media stream. Thus, the encoded media stream may have different video sequences corresponding to input sequence 102, with different encoding parameters representing different genres due to the different encodings required for genre changes as the video content evolves and as testing and encoding occur dynamically. These different encoding parameters representing different genres may occur over time and need not be carried in output bitstream 126 at a single point in time. However, different encoding parameters representing different genres may also be provided in output bitstream 126 over a predetermined period of time, for example, if the input sequence has different genres determined by ML model 128, or if ML model 128 receives feedback and changes based in part on updates to ML model 128 or different mode selections.
[0022] The training of the ML model may combine supervised training, unsupervised training, or semi-supervised training. In one or more such trainings, at least some, all, or any available scenario of different types of input sequences may be labeled. In at least unsupervised and semi-supervised learning, different types of media streams may be part of the encoding set, and the ML model classifies different features associated with different types to enable the encoder to perform different encodings on different types of input sequences. Thus, the ML model enables the encoder to make mode selections or parameter selections appropriate to the type provided from the ML model. Furthermore, although described in the singular, the ML model may include sub-ML models, such as Figure 2 and Figure 3 The inference may result in the encoder providing different encoding parameters in the output bitstream 126.
[0023] In at least one embodiment, different inferences can also be dynamically provided as the type of input sequence 102 changes, enabling different encoding parameters to dynamically emerge over time for the video content associated with the output bitstream 126. Separately, in response to at least one encoding parameter indicated to the ML model 128 by feedback provided from the encoder 104, the encoder 104 can be caused to provide different encoding parameters in the output bitstream 126. This can reflect adjustments or updates made by or within the ML model 128. The adjustments or updates can be based in part on associating at least one characteristic of the type with the encoding parameter based on the feedback. For example, an initial type or associated characteristic can be indicated to the encoder 104 from the ML model 128. However, the selection of the encoding parameter itself is provided by the encoder 104.
[0024] Since the ML model 128 may not initially have encoding information, adjustments or updates after the initial type or features are indicated to the ML model 128 can use a sub-ML model, which then relates the encoding parameters to the features of the type. This process enables the sub-ML model to perform adjustments or updates. In addition, since the sub-model has limited features or training relative to the main ML model, such as Figure 2 and Figure 3 As described above, using a sub-ML model can reduce the size of the ML model or speed up the inference performed by the ML model based on feedback dynamically received from the video encoder to the ML model.
[0025] In at least one embodiment, an encoder 104 (e.g., a video encoder) may receive at least one input sequence 102 associated with a media stream and may provide at least one output sequence 122 that is a compressed or modified version of the input sequence 102. Furthermore, an output bitstream 126 is an encoded media stream that includes different encoding parameters representing different types and corresponds to the input sequence 102. For example, in at least one embodiment, the different encoding parameters for different types are based on a determination by an ML model 128. In another example, different types are represented in different encoding parameters by varying at least one value or parameter of one encoding parameter in a set of encoding parameters representing different encoding parameters. Therefore, different types do not necessarily require that all different encoding parameters in the set be different or have different values. Since some video content may always have only a single type, a preliminary determination of the type used throughout the video content may be made, with feedback used only to ensure that the encoding parameters for the video content remain unchanged.
[0026] Furthermore, if feedback from the encoder to the ML model 128 indicates that different features are more prominent for the video content, encoding parameters can be updated even if the overall genre remains unchanged. Thus, features can be trained separately to different sub-ML models, and feedback can be used to cause one of the sub-ML models to provide inference about different features that reflect aspects of the genre to be used as the basis for encoding subsequent input sequences 102. For example, because dynamic encoding is provided herein, the output bitstream may change over time as genre or feature changes occur in the media stream. Therefore, as used herein, the genre of the ML model 128 may refer to one or more features that, individually or collectively, represent the genre of the video content. Therefore, although the term "genre" is used herein, when describing with reference to the ML model 128, the genre and its classification may differ from what a human observer would subjectively understand. When describing with reference to the ML model 128, the genre and its classification may be specific to the specific features trained for the ML model 128 and may differ from what a human observer would subjectively understand.
[0027] Although shown in the singular, the encoding performed by the encoder 104 is for a set of input sequences or scenes indicated by the ML model 128 as being of the same type. The encoding is performed to provide an output bitstream that is an encoded media stream having different video sequences associated with different encoding parameters of different types determined by the ML model 128. In at least one embodiment, the encoder 104 can be based in part on one of the H.264 standard, the MPEG2 standard, the AVC standard, the HEVC standard, the VP9 standard, the AV1 standard, or the VVC standard. However, the encoder 104 can be any encoder standard that allows weighting of the input, such as by using a quantization parameter (QP) for mode selection.
[0028] Figure 1 As shown, in terms of video encoding, a mode selection can be made to perform inter-frame or intra-frame mode encoding and decoding. This mode selection can be performed using the mode selection module 116. The mode selection can enable the selection of parameters associated with the available encoding parameters. The result of this mode selection is to provide a specific encoding based in part on the classification from the ML model 128. The mode selection can also allow the encoder 104 to determine how many bits it is willing to sacrifice to hide and / or eliminate distortion that may be associated with certain portions of media content belonging to certain types.
[0029] In at least one embodiment, there is a trade-off between the bits used and the distortion of the encoding performed. The trade-off may be associated with the distortion, which may vary between different encoders. For example, the trade-off may be between different user presets, different target bit rates (e.g., which may affect the bit budget), and between different frames in a group of frames (GOP) representing the input sequence 102 to be encoded. However, for the type classifications herein, the trade-off may be tailored to the different types so as to preserve useful information about the type during the encoding process. In another example, the trade-off may include the possibility of some distortion occurring within a general region 102B of the input sequence 102, and ensuring that no distortion (or relatively little distortion) is applied to certain regions of interest 102A during the encoding process, as these regions are relevant to the type determined for the input sequence.
[0030] Video compression can be computationally intensive. Current state-of-the-art compression ratios offer compression ratios of 1 / 200 to 1 / 1000, but require significantly more computing resources to perform this compression. However, with artificial intelligence and machine learning (AI / ML) workloads using large amounts of images and videos, autonomous vehicles generating large amounts of video per vehicle, applications like smart cities demanding more video data, content created for entertainment requiring higher video resolutions and increased bit depths, and today's remote work video conferencing technologies, there is a recognition that video compression must be performed more efficiently. This efficient approach can rely on type determination performed by an ML model, which is then used by the encoder to perform encoding. Furthermore, it is recognized that the limitations of the human eye, as well as the aggressive quantization or decimation of features required for color space conversion and the separation of luma (brightness) and chroma, can limit the ability to provide high-quality video compression. Using ML models can enable objective type determination for the encoder.
[0031] As used herein, a system 100 employing an ML model 128 trained to classify genres in a video sequence enables certain parameters or patterns in an encoder 104 to encode an input sequence and provide an encoded media stream, also referred to herein as an output bitstream 126. Encoding supported by the ML model for genre classification allows for bit savings. In one example, the ML model for genre classification enables the encoder to select regions within frames of the input sequence to preserve quality by encoding those regions with specific encoding parameters to reduce the impact of compression on the input sequence. This allows more bits to be used in encoding regions 102A of the video sequence to achieve the desired quality required for the genre, and allows fewer bits to be used to save computational resources on other regions 102B of the input sequence 102 where the relevant genre does not require encoding of every detailed aspect thereof.
[0032] As part of the encoding parameters, a Fourier transform or other related transform can be performed on the blocks within each frame to convert the data therein into the frequency domain and allow quantization or discarding of information based on selected frequencies. In doing so, the transform coefficients at lower frequencies may be quantized less aggressively than the transform coefficients at higher frequencies. In addition, motion estimation can be used to capture and encode movement between video frames. Although all of these methods or options attempt to improve video compression, they may all serve similar goals, namely allowing the encoder to compress the video into a smaller bitstream by eliminating noise, artifacts, allowing at least more intensive motion estimation, and exploiting temporal and spatial redundancy. However, as used herein, for certain types, additional benefits can be achieved by only reserving bits for certain portions of the video sequence related to the type. For example, the aggressiveness of the transform and quantization provided by the transform and quantization (T and Q) module 108 may be different for different types.
[0033] In view of all of these benefits, the encoder can be differentiated based in part on the selection of appropriate tools to enable various aspects of the encoder to save bits. For example, the selection of appropriate tools refers to the selection of encoding parameters that enable regions within the frames of each input sequence 102 (e.g., regions provided by macroblocks (MBs)) to be selected to be more or less compressed than other regions. This approach and other such approaches can be defined within the encoder as different modes that may require more bits or fewer bits to ensure the desired quality. The RDO module 116A can be associated with the mode selection module 116 of the encoder 104 to address requirements by using an RDO metric (e.g., sum of squared errors (SSE) or sum of transform differences (SATD)) to determine the cost associated with each choice made and to enable selection based on the cost.
[0034] Further RDO metrics allow for further mode selection that benefit from evaluation using further quality metrics, including VMAF, SSIM, MS-SSIM, or PSNR. For the encoder 104, addressing temporal effects may still remain, as it may be accomplished solely at the frame level using such further RDO metrics. Distortion may be determined as a difference from the original image. In at least one embodiment, the system 100 for video compression using type classification herein includes an ML model 128 to enable improved selection of at least those quality metrics that may be the basis for mode selection provided by the RDO output 124 of the RDO module 116A. The encoder 104 may use the improved selection of at least the quality metrics to perform video compression of the video sequence 102, and in particular, to provide the appropriate type of video compression. For example, the encoder 104 (also referred to herein as a video encoder) may receive transform coefficients or parameters, such as QP. The RDO module 116A is configured to optimize an efficient representation, which may include segmentation, prediction mode, motion vector (MV), or QP, for each point or block in the frame.
[0035] In at least one embodiment, the RDO output 124 is used to select a mode provided by the RDO module 116A. The RDO also facilitates selection of available encoding parameters based in part on the type classification of the input sequence 102. In at least one embodiment, an interface can be provided between the encoder 104 and the ML model 128 to allow input from the encoder 104 to be received in the ML model 128, the input reflecting feedback from the feedback loop 130. In addition, the interface can enable output from the ML model 128 to the encoder 104, which may result in the selection of certain video compression parameters to compress the input sequence 102. For example, the video compression parameters reflect a quality metric of the RDO output 124.
[0036] In at least one embodiment, RDO may be limited to a single point for each block in each frame of the input sequence 102 and may be represented by a linear equation of R+λ*D, where λ(lambda) is a multiplier and the (R, D) pair may be used with the multiplier to minimize the combined R+D value. R may be associated with bitrate, while D may be associated with distortion as it relates to the quality of the media. RDO allows candidate solutions to be ranked using a linear equation, for example, to select one of the candidate solutions. Thus, the lambda value may be associated with a range from 1 to the minimized cost of the (R, D) set. R may be measured in bits, while D may be a quality unit, such that the equation provides a measure of distortion units for each bit of the bitrate used in the video compression process.
[0037] To achieve a predetermined bit rate for R, a certain lambda value may be used. The ML model 128 herein is capable of selecting encoding parameters, which may include R, D, and lambda values, to allow RDO to use different quality metrics for different genres. This is done to ensure that the effectiveness of the video compression performed in the video encoder is based at least in part on the genre associated with the underlying video content. Therefore, in at least one embodiment, the system 100 herein uses the ML model 128 to optimize the encoder 104 so that different quality metrics representing different video compression parameters can be used for different genres.
[0038] In at least one embodiment, Figure 1 As shown, encoder 104 is associated with at least one execution unit of a processor that performs inference using ML model 128. Encoder 104 may include an output to provide feedback from encoder 104 to ML model 128 via feedback loop 130. In an example, the output may indicate to at least one execution unit a scene change event or at least one of different encoding parameters used or available to the video encoder. For example, at least one of the different encoding parameters is an encoding parameter of a previous input sequence.
[0039] The ML model 128 can be capable of using at least one different encoding parameter of a previous video sequence to update the ML model 128 or to further train, retrain, or test (including inference) the ML model 128. Thus, the ML model 128 can enable dynamically encoding an input sequence differently, such as for a subsequent input sequence having a different type relative to a previous or initial input sequence, the subsequent input sequence being determined by the ML model 128 and based at least in part on a scene cut event. The encoder 104 can also include an input to receive different inferences for received frames of the subsequent input sequence in response to the at least one different encoding parameter provided in the feedback loop. For example, the ML model 128 can include a sub-ML model to provide different inferences for the subsequent input sequence, such as about Figure 2 and Figure 3 Further described.
[0040] In at least one embodiment, Figure 1An encoder 104 is provided that performs H.264 encoding. The encoder 104 includes hardware or software modules such as a prediction module 112, a T and Q module 108, and an entropy codec module 110. Additional modules may be present, such as an inverse module 114, a filter module 120, a motion processing module 118 (for supporting motion estimation and related aspects), and a previous or reference frame module 106. The type of video compression used herein has no effect on the decoding process of the bitstream provided from the encoder 104, including the output frames 122. For example, the decoding process may be in accordance with H.264 decoding or other decoding related to the encoding format used to provide the output bitstream 126 from the encoder 104, particularly with respect to the entropy codec module 110.
[0041] The bitstream representing a frame of the input sequence 102 to be compressed may include different MBs. In at least one embodiment, different MB sizes may be supported in the encoder 104, including but not limited to 8x8, 8x16, 16x8, 4x4, and 16x16. The MBs may correspond to displayed pixel data obtained at the location of the block. As part of video compression, the prediction module 112 may generate predicted MBs that may be used to generate residual data reflecting the quantized data. There may be multiple prediction options associated with the prediction module 112, including intra-frame prediction associated with previously encoded data from the current sequence (e.g., the input sequence 102). Another option associated with the prediction module 112 includes inter-frame prediction using encoded data from other previously encoded frames (i.e., reference frames, such as from the previous or reference frame module 106). These reference frames may appear before or after the current frame in display order and may be associated with motion compensation, such as using the motion processing module 118 of previously encoded frames (e.g., frames provided by the previous or reference frame module 106).
[0042] Yet another option associated with prediction module 112 includes using different prediction block sizes for intra-frame and inter-frame prediction options. Using different prediction block sizes for MBs can change the accuracy associated with the prediction. Another option associated with prediction module 112 includes using multiple frames during prediction, which is available in the inter-frame prediction option to provide better prediction accuracy. Another option is to skip MB data or residual data so that encoder 104 itself performs inference of the MB data based in part on the predicted MB. One or more of these options represent encoding parameters that can be applied to compress the input sequence 102 of the media stream based in part on the type selection made by ML model 128.
[0043] In at least one embodiment, intra-frame prediction may be based at least in part on spatial data within at least one frame of the input sequence 102. The MBs generated as part of intra-frame prediction may differ from the MBs of the frames of the input sequence 102. The residual data may be a residual MB generated by subtracting the predicted MB from the current MB. The residual MB may be transformed, quantized, and entropy coded in the provided modules 108 and 110 according to a mode selected by a mode selection module 116, which may be associated with an RDO module 116A to perform, for example, RDO. Furthermore, in the encoder 104, the quantized data may be rescaled and inverse transformed in an inverse module 114. The output of the inverse module 114 may be filtered and combined with the predicted MB in a prediction module 112. Motion estimation from the motion processing module 118 may be included. The result may be a reconstructed MB or a decoded frame, which is provided to the previous or reference frame module 106 for further prediction. In at least one embodiment, additional coding parameters may be expressed using one or more of inter-frame prediction or intra-frame prediction, which may be based in part on the type selection made by the ML model 128 to compress the input sequence 102 of the media stream.
[0044] Figure 2 2 is a diagram of aspects 200 of a machine learning (ML) model having, in at least one embodiment, child ML models 1-N 210 to perform different inferences to provide type classification for received frames. In one example, the ML model 128 can include child ML models 1-N 210 to perform different inferences on an input sequence. The different inferences can be responsive to at least one of the different encoding parameters indicated by feedback to the ML model 128 via the feedback loop 130 from the encoder 104. This enables reducing the size of the ML model 128 or speeding up inference in the ML model 128 based on the feedback. Furthermore, because the feedback is dynamically received from the video encoder 104 to the ML model 128, the ML model 128 can be updated or further training or testing can be performed on the ML model 128.
[0045] Figure 2It is also shown that the ML model 128 can be executed on a different processor infrastructure 260B than the encoder 260A. In addition, the ML model 128 can be controlled by an application 250 for which or on whose behalf the encoding is performed. The application 250 can provide a control input to indicate that the ML model 128 is to perform inference on the input sequence 202. Separately, the video encoder 104 can be controlled by its respective processing infrastructure 260A to perform encoding of the media stream based in part on the capabilities associated with the processing infrastructure. For example, the capabilities can relate to the encoding standards enabled for the encoder 104, including the H.264 standard, the MPEG2 standard, the AVC standard, the HEVC standard, the VP9 standard, the AV1 standard, or the VVC standard. In addition, the application 250 and the processing infrastructure 260A, 260B can share a memory 270 (which can be part of the system 100) to enable inference and enable encoding of the media stream.
[0046] In at least one embodiment, feedback on a scene cut event can be used in conjunction with at least one different encoding parameter. While the type 212 may initially be determined by the ML model 128 and indicated to the encoder 104 for use by the encoder 104, feedback on the scene cut event or one or more of the at least one different encoding parameters can be provided to the ML model 128 to determine a different type for at least one subsequent input sequence 102. In one example, the type 212 is indicated as one or more values or other parameters that can be used by the encoder to select certain encoding parameters. However, certain encoding parameters may not be known to the ML model 128. Therefore, feedback in the feedback loop 130 can be used to inform the ML model 128 of encoding parameters that are generally available or used in conjunction with the indicated type in the video encoder. For example, the feedback in the feedback loop 130 can indicate a scene cut event to at least one execution unit and subsequently enable the encoded media stream output from the video encoder to include different encoding parameters that are dynamically provided for the encoded media stream based at least in part on the scene cut event.
[0047] Thereafter, the ML model 128 may be updated to improve accuracy by training (including retraining) the ML model 128, which training associates the different encoding parameters initially used with the types initially indicated to the encoder 104. Furthermore, this process enables the coded media stream as the output bitstream 126 to include different encoding parameters of different types, as determined by the ML model 128 and dynamically provided for the input media stream having the input sequence 102 shown. Furthermore, because the process is performed in an ongoing or dynamic manner, the coded media stream as the output bitstream 126 may change based at least in part on scene change events and may include different encoding parameters as the types in the content of the input sequence 102 change.
[0048] The ML model 128 may include a type feature dataset 204 to retain different types of features available to the main ML model 208. In at least one embodiment, the main ML model 208 performs an initial determination of the type of the input sequence 102. Thereafter, for subsequent input sequences, additional determinations may be dynamically performed using the main ML model 208 or using one or more sub-ML models 1-N 210. Thus, in at least one embodiment, the ML model 128 defaults to the main ML model 208 contained therein. Furthermore, like the ML model 128 with the type feature dataset 204, the sub-ML models 1-N 210 may be associated with their own dataset, which may be part of the type feature dataset 204. However, the main ML model 208 may use or access all features of the type feature dataset 204. Furthermore, while a dataset may retain features, it may do so only to enable training, retraining, or updating of any of the main ML model 208 and sub-ML models 1-N 210. In one example, different encoding parameters provided to the ML model 128 via feedback may be provided to the type feature dataset 204 for retraining or updating the ML model 128.
[0049] The ML model 128 includes a feature normalization module 206, which preprocesses features to provide normalized features or to classify marginal features into new or established categories. Therefore, the type feature dataset 204 and the feature normalization module 206 can be used to train and test the ML model 128. Furthermore, the main ML model 208, the child ML models 1-N 210, the type feature dataset 204, and the feature normalization module 206 can be executed by different circuits, such as memories, caches, buffers, processors, or execution units within the processors. Therefore, retraining or updating the ML model 128 can be applied to the main ML model 208 or any of the child ML models 1-N 210.
[0050] In at least one embodiment, the encoder 104 and the ML model 128 can be provided with a processed sequence 202 (which can be a downsampled or filtered version of the input sequence 102) rather than the input sequence 102 to allow for the application of video compression using type classification, as described throughout this document. In one non-limiting example, the processed sequence 202 can allow for color format conversion to be provided to the input sequence 102. In at least one embodiment, the processed sequence 202 can allow for certain aspects corresponding to features used by the ML model 128 to be enhanced or suppressed to assist in testing the ML model 128. Thus, the ML model 128 can receive the processed sequence 202 while the encoder 104 receives the input sequence 102. This can reduce the workload of the ML model 128 in classifying the input sequence 102, as classification can be performed on a downsampled version of the processed sequence 202, while encoding is performed using the input sequence 102 as provided.
[0051] Figure 3 300 is an illustration of aspects of a machine learning (ML) model with supervised, unsupervised, or semi-supervised training for providing a type classification for a received frame. Initially, supervised learning allows the ML model 128 to receive a set of labeled training data and be trained to recognize patterns in that data. Supervised learning can be provided for one or more features in the type feature dataset 204 by labels associated with such features. Any of the main ML model or child ML models that use such labeled features and have labeled categories can be considered a supervised ML model 302. The supervised ML model 302 can then be trained, retrained, or updated using further labeled features from a feedback loop.
[0052] Additionally, unsupervised learning can be provided for one or more features in the type feature dataset 204 by using existing features and allowing the ML model 128 to determine patterns in the features without labels or instructions. In contrast to supervised learning with supervised ML model 302, unsupervised learning allows for the use of existing features to determine types within unsupervised categories. Because features are implicitly associated with different types, it is assumed that the unsupervised categories provide different types of classification for encoding parameters to be used with the input sequence 102. For example, an input sequence with features classified as belonging to a certain unsupervised category of the trained ML model will result in encoding the input sequence using encoding parameters for that unsupervised category. Such unsupervised categories may not be labeled. Any primary or child ML model that uses such unlabeled features and has both unlabeled and unsupervised categories can be considered an unsupervised ML model 304. Parameters received in the feedback loop are then encoded and associated with one or more features in the unsupervised category, allowing for further training, retraining, or updating of the unsupervised ML model 304.
[0053] Furthermore, semi-supervised learning can be provided for one or more features in the type feature dataset 204 by combining features from supervised and unsupervised learning. Thus, in supervised learning, a relatively small number of features can be labeled compared to the number of features used in supervised learning. In unsupervised learning, a relatively large number of features can be used even if they are unlabeled. This is intended to allow for a broader classification than that provided by supervised learning and to provide categories even when a large number of labeled features are unavailable. Semi-supervised learning allows the use of existing features to determine types within such combined categories to provide a semi-supervised ML model 306. Because these features are a combination of labeled and latent features associated with a particular type, the combined categories provide different type classifications for encoding parameters to be used with the input sequence 102. For example, an input sequence having features classified as belonging to one of the combined categories of the trained ML model will result in encoding parameters for that combined category being used to encode the input sequence. For example, tokenization can be provided for such combined categories based in part on the labeled features. Any of the main ML models or sub-ML models that use such combined features, have labeled and unlabeled features, and provide combined categories can be considered a semi-supervised ML model 306. The parameters received in the feedback loop are then encoded and associated with one or more features in the unsupervised category, allowing for further training, retraining, or updating of the semi-supervised ML model 306 as labeled or unlabeled features.
[0054] In at least one embodiment, Figure 2 and Figure 3 In each of the , type feature datasets 204 include features such as different types of noise features, different types of motion vector distributions, different types of different pixel intensity levels, and different types of different edge features. In addition, a downsampled or filtered sequence of the input sequence 102 that highlights one or more of these features can be used to enable the processed sequence 202. This process can be used to train or test the ML model 128. In addition, Figure 2 Each sub-ML model 1-N 210 in the system 100 can be trained for each feature. Thus, each sub-ML model 1-N 210 can have its own sub-type feature dataset 204 associated with only one of the features. Thus, if a feature is determined to be prominent in the encoding of the input sequence 102 based on feedback, the sub-ML model corresponding to that feature can be used with subsequent input sequences. Thus, the system 100 supports reasoning performed using an ML model on processed versions or sequences 202 of one or more received frames. In addition, the system 100 can also support reasoning performed using an ML model on one or more sub-regions 102A of one or more received frames.
[0055] In at least one embodiment, the processed sequence 202 may also be used as a feature for one or more supervised, unsupervised, or semi-supervised ML models or sub-ML models 302-306. Because some features may be better classified than others, a learning type that is appropriate for the clarity of the classification process may be beneficial. Thus, in one embodiment, when the encoder 104 receives the input sequence 102, the ML model 128 may receive different types of processed sequences 202 depending on the sub-model 302-306 used. Thus, Figure 2 and Figure 3 The sub-ML model in (representing ML model 128) can perform different inferences on the received frames in the input sequence 102 in response to at least one of the different encoding parameters provided to ML model 128 or as feedback.
[0056] also, Figure 2 and Figure 3 The sub-ML models representing the ML model 128 in the training data may have different associated memory or processing capacities. Because one or more sub-models may be trained using downsampled features and other sub-models may be trained using whole features, there may be different types of feature datasets of different sizes and held in different memory capacities. Furthermore, the main ML model 208 and at least some of the sub-ML models 1-N 210 using whole features may require more processing power than sub-ML models using downsampled features. At least one execution unit may execute the ML model 128 using the main ML model 208 or one of the sub-ML models 1-N 210 in response to at least one of the different encoding parameters indicated to the ML model 128 via feedback. Furthermore, at least one execution unit may execute the ML model 128 based in part on a threshold capacity of at least one of the different associated memory or processing capacities. In one example, once training is complete, the main ML model 208 or the sub-ML models 1-N 210 may be stored in memory and loaded into at least one execution unit to perform inference based in part on the different encoding parameters to be provided in the output bitstream 126.
[0057] Figure 4 The computer and processor aspects of a system for video compression using type classification in at least one embodiment are shown 400. For example, each illustrated processor 402 may include one or more processing or execution units 408 that may use an encoder and use type classification from an ML model to perform any or all aspects of the video compression system 100. The system 100 may include an interface that may be between the encoder and the ML model to allow feedback loops and communication of types between these two aspects of the system.
[0058] The processing or execution unit 408 may include a plurality of circuits to support one or more aspects described herein for the encoder 104, the ML model 128, and the interface between the two aspects. In at least one embodiment, the processor 402 may include a CPU, a GPU, a DPU, which may be associated with a multi-tenant environment to perform one or more of the encoder 104, the ML model 128, and the interface between the two aspects described herein. In addition, with respect to Figure 4 4 and 402, the GPU may obviously be located in a different graphics / video card 412. Thus, even though described in the singular, the graphics / video card 412 may include multiple cards and each card may include multiple GPUs.
[0059] In accordance with at least one embodiment, the computer and processor aspect 400 may be executed by one or more processors 402, including a system on a chip (SOC) or some combination of processors, which may include an execution unit to execute instructions. In at least one embodiment, the computer and processor aspect 400 may include, but is not limited to, components, such as the processor 402, to use an execution unit 408, which includes logic for executing algorithms for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, the computer and processor aspect 400 may include a processor, such as Processor series, Xeon TM 、 XScale TM and / or StrongARM TM 、 Core TM or Nervana TM The computer and processor aspects 400 may be configured to execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0060] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0061] In at least one embodiment, the computer and processor aspects 400 may include, but are not limited to, a processor 402, which may include, but are not limited to, one or more execution units 408 to execute the computer program according to the present disclosure. Figures 1 to 3 and Figures 5 to 7 In at least one embodiment, the computer and processor aspect 400 is a single-processor desktop or server system, but in another embodiment, the computer and processor aspect 400 can be a multi-processor system.
[0062] In at least one embodiment, the processor 402 may include, but is not limited to, a Complex Instruction Set Computer ("CISC") microprocessor, a Reduced Instruction Set Computing ("RISC") microprocessor, a Very Long Instruction Word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as, for example, a digital signal processor. In at least one embodiment, the processor 402 may be coupled to a processor bus 410 that may transmit data signals between the processor 402 and other components in the computer and processor aspect 400.
[0063] In at least one embodiment, processor 402 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 404. In at least one embodiment, processor 402 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may be external to processor 402. Other embodiments may include a combination of internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, register file 406 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.
[0064] In at least one embodiment, execution unit 408 (including, but not limited to, logic to perform integer and floating-point operations) also resides in processor 402. In at least one embodiment, processor 402 may also include a microcode ("ucode") read-only memory ("ROM") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 408 may include logic to process packed instruction set 409.
[0065] In at least one embodiment, by including a packed instruction set 409 and associated circuitry for executing the instructions in the instruction set of a general-purpose processor, operations used by many multimedia applications can be performed using packed data in processor 402. In at least one embodiment, many multimedia applications can be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on packed data, which can eliminate the need to transfer smaller units of data across the processor's data bus to perform one or more operations one data element at a time.
[0066] In at least one embodiment, execution unit 408 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer and processor aspects 400 may include, but are not limited to, memory 420. In at least one embodiment, memory 420 may be a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or another memory device. In at least one embodiment, memory 420 may store instructions 419 and / or data 421 represented by data signals, which may be executed by processor 402.
[0067] In at least one embodiment, the system logic chip may be coupled to the processor bus 410 and the memory 420. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub ("MCH") 416, and the processor 402 may communicate with the MCH 416 via the processor bus 410. In at least one embodiment, the MCH 416 may provide a high-bandwidth memory path 418 to the memory 420 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 416 may direct data signals between the processor 402, the memory 420, and other components in the computer and processor aspect 400, and bridge data signals between the processor bus 410, the memory 420, and the system I / O interface 422. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 416 may be coupled to the memory 420 via the high-bandwidth memory path 418, and the graphics / video card 412 may be coupled to the MCH 416 via an accelerated graphics port ("AGP") interconnect 414. In at least one embodiment, graphics / video card 412 may be coupled to one or more processors 402 via the PCIe interconnect standard. Similarly, network controller 424 may also be coupled to one or more processors 402 via the PCIe interconnect standard.
[0068] In at least one embodiment, the computer and processor aspects 400 may use the system I / O interface 422 as a proprietary hub interface bus to couple the MCH 416 to the I / O controller hub ("ICH") 430. In at least one embodiment, the ICH 430 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 420, chipset, and processor 402. Examples may include, but are not limited to, an audio controller 429, a firmware hub ("flash BIOS") 428, a wireless transceiver 426, a data storage device 424, a legacy I / O controller 423 including a user input and keyboard interface 425, a serial expansion port 427 (e.g., a universal serial bus ("USB") port), and a network controller 434. In at least one embodiment, the data storage device 424 may include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0069] In at least one embodiment, Figure 4 A computer and processor aspect 400 is shown, comprising interconnected hardware devices or "chips," while in other embodiments, Figure 4 An exemplary SoC is shown. In at least one embodiment, Figure 4The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer and processor aspect 400 are interconnected using a Compute Express Link (CXL) interconnect.
[0070] Thus, the at least one execution unit 408 can be circuitry of the at least one processor 402 that is associated with a video encoder. The association can be such that the at least one execution unit 408 of the at least one processor 402 can execute the video encoder. The association can be such that the at least one execution unit 408 of the at least one processor 402 can load and run or execute instructions to execute the video encoder. However, the association can be such that the at least one execution unit 408 of the at least one processor 402 can be hardwired to execute the video encoder.
[0071] Furthermore, at least one execution unit 408 may be circuitry of at least one processor 402 associated with the ML model. The association may be such that the at least one execution unit 408 of the at least one processor 402 can execute the ML model. The association may be such that the at least one execution unit 408 of the at least one processor 402 can load and run or execute instructions to execute the ML model. However, the association may be such that the at least one execution unit 408 of the at least one processor 402 may be hardwired to execute the ML model. Furthermore, to support the dataset, other circuitry may be present, including a cache 404 that may be associated with the execution unit 408. However, to execute the ML model, a trained ML model may be loaded into the execution unit 408 and run or executed therefrom. Furthermore, different execution units may be present that provide an interface between the execution unit executing the ML model and the encoder.
[0072] The ML model may be configured to determine a type associated with a received media stream frame based in part on training the ML model using features associated with the different types. The training may be supervised, unsupervised, or semi-supervised. The at least one execution unit 408 of the at least one processor 402 executing the encoder may also enable the encoder to encode the media stream based in part on the determined type. The encoded media stream provided by the encoder may include different video sequences associated with different encoding parameters of different types determined by the ML model.
[0073] Furthermore, at least one execution unit 408 of at least one processor 402 causes features used in the ML model executed therein to include one or more of different noise features of different types, different motion vector distributions of different types, different pixel intensity levels of different types, or different edge features of different types. For example, these features may be used to impart training and testing of the ML model to indicate the type of input sequence to the encoder.
[0074] At least one execution unit 408 of at least one processor 402 executing the ML model may include an input to receive feedback from at least one different execution unit 408 executing the video encoder. The feedback may be an indication of a scene cut event to the ML model. The scene cut event may be a basis for the ML model to perform inference to determine a type of a subsequent input sequence. The type may be the same as a previous type of a previous input sequence. The type may include characteristics that differ from the previous type of the previous input sequence, based in part on the feedback, which may include encoding parameters used by the encoder for the previous input sequence. Thereafter, based in part on the type indicated to the encoder, an output bitstream having an encoded media stream may be enabled using different encoding parameters of different types, the encoding parameters determined by the ML model and dynamically provided at least in part based on the scene cut event.
[0075] At least one execution unit 408 of at least one processor 402 executing the ML model may include an input for receiving feedback from a video encoder to indicate at least one different encoding parameter to the at least one execution unit. The ML model may include a sub-ML model to perform different inferences on received frames in the input sequence in response to the at least one different encoding parameter. For example, adjustments or updates may need to be made by or within the ML model to associate at least one characteristic of the type with the encoding parameter based in part on the feedback. For example, an initial type or associated characteristic may be indicated from the ML model to the encoder. However, the selection of the encoding parameter itself is provided by the encoder. Because the ML model may not initially have this information, adjustments or updates made after the initial type or characteristic is indicated to the ML model may use a sub-ML model to subsequently associate the encoding parameter with the characteristic of the type. This process can reduce the size of the ML model or speed up the inference performed by the ML model based on the feedback dynamically received from the video encoder to the ML model.
[0076] At least one execution unit 408 can be circuitry of at least one processor 402 to be associated with a video encoder configured to encode a media stream based in part on a type associated with the media stream determined using an ML model. The type determined from a received frame of the media stream can be based in part on training the ML model using features associated with different types. The output bitstream of the encoder can be an encoded media stream that can include different video sequences associated with different encoding parameters of different types determined by the ML model. Because video content can be continuously encoded and transmitted to a decoder, it will be appreciated that, in at least one embodiment, the output bitstream can include different video sequences over a period of time (rather than at any instant in time).
[0077] Furthermore, at least one execution unit 408 of at least one processor 402 to be associated with the video encoder may include an output for providing feedback from the video encoder to different execution units executing the ML model. The output may indicate a scene change event or at least one different encoding parameter to the ML model. Different encoding parameters may be dynamically provided for the video stream based at least in part on the scene change event, and different encoding parameters may be provided for the video content over time. The at least one execution unit 408 to be associated with the video encoder may also include an input for receiving different inferences for a received frame. The different inferences may be responsive to at least one different encoding parameter provided to the ML model via a feedback loop. The ML model may provide the different inferences using a sub-ML model included therein.
[0078] In at least one embodiment, at least one execution unit 408 of at least one processor 402 may be configured to train an ML model using features associated with different types of media streams. Once trained, the ML model enables a video encoder to encode a media stream based in part on the type determined by the ML model for the media stream. Furthermore, once trained, the ML model enables the video encoder to provide an encoded media stream comprising different video sequences associated with different encoding parameters of different types determined by the ML model.
[0079] Figure 5A process flow or method 500 for a system for video compression using type classification in at least one embodiment is shown. The method 500 includes receiving 502 a media stream, which may include an input sequence as described herein. The method 500 includes performing 504 inference using an ML model to determine a type associated with a frame of the received media stream. This may be based in part on training the ML model using features associated with different types. Validation 506 may be performed on the type determined by performing 504. The validation may be used to ensure that classification is achieved at a determined threshold to allow the ML model to perform inference. In at least one embodiment, the method 500 includes encoding 508 the media stream using a video encoder based in part on the determined type. The encoded media stream may be provided 510 as part of the encoding, wherein the encoded media stream includes different video sequences associated with different encoding parameters of different types determined by the ML model.
[0080] Figure 5 The method 500 may include further steps or may include sub-steps, wherein the features include one or more of different noise features of different types, different motion vector distributions of different types, different pixel intensity levels of different types, or different edge features of different types. Figure 5 The method 500 may include further steps or may include sub-steps, wherein the imparted training is supervised training, unsupervised training, or semi-supervised training.
[0081] Figure 6 Yet another process flow or method 600 of a system for video compression using type classification is shown in at least one embodiment. Figure 6 The method 600 may be used with Figure 5 For example, the method 600 includes enabling 602 a feedback loop from a video encoder to at least one execution unit that performs inference using an ML model, which may be Figure 5 In one example, the feedback loop may be enabled based in part on feedback sent from the video encoder. Thus, while the feedback loop may physically exist, no useful information may be provided by the feedback until the feedback is provided to enable execution of the ML model of step 504. However, in at least one embodiment, the enabling 602 step may be performed for subsequent input sequences of the media stream of step 502.
[0082] Method 600 includes determining or verifying 604 that feedback has been received. For example, an ML model may be associated with an interface to expect certain types of information, which may be predetermined categories of information. When such information is received, it may be processed as feedback received at step 604. Method 600 includes determining 606 a scene change event from the feedback in the feedback loop. Method 600 includes enabling 608 dynamically providing different encoding for the media stream based at least in part on the scene change event. For example, initial encoding parameters may be provided for an encoded media stream, and then different types of adjusted, updated, or completely different encoding parameters may be provided for subsequent input sequences in the media stream.
[0083] In at least one embodiment, method 600 may include determining 606 at least one different encoding parameter provided in conjunction with the feedback, either separately from or in addition to the scene cut event provided in the feedback loop. Method 600 includes enabling 608 dynamically providing different encodings for the media stream based at least in part on at least one of the scene cut event and / or the different encoding parameters. For example, initial encoding parameters may be provided for an encoded media stream, and then adjusted, updated, or completely different encoding parameters may be provided for subsequent input sequences in the media stream. However, in at least one embodiment, enabling 608 the step of dynamically providing encodings may use a sub-ML model of the ML model to perform different inferences on the received frame in response to at least one of the different encoding parameters provided in the feedback.
[0084] Figure 6 Method 600 may include further steps, or may include sub-steps, in which the child ML models have different associated memory or processing capacities. Method 600 may include further steps, or may include sub-steps, in which a threshold capacity for at least one of the different associated memory or processing capacities may be determined. Then, in response to at least one of the different encoding parameters indicated to the ML model, and based in part on the threshold capacity, a different inference may be performed using one of the child ML models in method 600.
[0085] Figure 7 A further process flow or method 700 of a system for video compression using type classification in at least one embodiment is shown. Figure 7 The method 700 can be used with Figure 5 Method 500 or Figure 6 For example, the method 700 can be used with the method 600 of administering Figure 5 The step 504 of the method 500 is associated with the training of the ML model used. Figure 7The method 700 in the embodiment includes determining 702 features associated with different types for different video content. As described throughout this document, the features may include one or more of different noise features of different types, different motion vector distributions of different types, different pixel intensity levels of different types, or different edge features of different types.
[0086] Figure 7 The method 700 in
[0065] includes determining or validating 704 to provide one or more ML models. For example, determining or validating 704 can be applied to create a child ML model or to cause adjustments or updates to an existing main ML model or child ML model. For example, when feedback indicates encoding parameters of a type that may not be associated with a previously classified best fit, there may be a new category associated with the new child ML model that can be provided to provide additional classification. Figure 7 The method 700 in includes training 706 an ML model using features associated with different types, the ML model may include sub-ML models and a default main ML model. Figure 7 The method 700 in includes using 708 the trained ML model to enable a video encoder to encode a media stream based in part on a type determined by the ML model. This can support Figure 5 The method 700 further includes enabling 710 the video encoder to provide an encoded media stream having different encoding parameters of different types determined by the ML model in a manner similar to Figure 5 The same as step 510.
[0087] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be apparent to those skilled in the art that the concepts of the present invention may be practiced without one or more of these specific details.
[0088] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0089] Unless otherwise noted or clearly contradicted by the context, the use of the terms "a" and "an" and "the" and similar referents in the context of describing the disclosed embodiments (particularly in the context of the appended claims) should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include," "have," "include," and "contain" should be interpreted as open terms (meaning "including but not limited to"), unless otherwise noted. The term "connected" (which refers to a physical connection when unmodified) should be interpreted as partially or completely contained within, attached to, or connected together, even if there are some intervening objects. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, and each individual value is incorporated into the specification as if it were separately described herein. In at least one embodiment, unless otherwise noted or contradicted by the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set including one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.
[0090] Unless otherwise expressly indicated or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C" are understood in the context to generally refer to an item, term, or the like, which may be A or B or C, or any non-empty subset of the set of A, B, and C. For example, in the illustrative example of a set having three members, the conjunctions "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctions are not generally intended to imply that certain embodiments require the presence of each of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise expressly indicated or contradicted by context, the term "plurality" indicates plurality (e.g., "a plurality of items" indicates a plurality of items). In at least one embodiment, the number of items in the plurality of items is at least two, but may be more if expressly indicated or indicated by context. Further, unless stated otherwise or clear from context, the phrase "based on" means "based at least in part on" rather than "based solely on."
[0091] Unless otherwise indicated herein or clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is collectively executed by hardware or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program that includes a plurality of instructions that can be executed by one or more processors.
[0092] In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) having executable instructions stored thereon that, when executed by one or more processors of a computer system (i.e., as a result of being executed), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lack the entire code, but rather the plurality of non-transitory computer-readable storage media collectively store the entire code. In at least one embodiment, the executable instructions are implemented so that different instructions are executed by different processors, e.g., a non-transitory computer-readable storage medium stores the instructions, and a main central processing unit ("CPU") executes some instructions while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of the instructions.
[0093] In at least one embodiment, an arithmetic logic unit (ALU) is a set of combinational logic circuits that takes one or more inputs to produce a result. In at least one embodiment, a processor uses an ALU to implement mathematical operations such as addition, subtraction, or multiplication. In at least one embodiment, the ALU is used to implement logical operations such as logical AND / OR or XOR. In at least one embodiment, the ALU is stateless and is made of physical switching elements (such as semiconductor transistors arranged to form logic gates). In at least one embodiment, the ALU can be operated internally as a stateful logic circuit with an associated clock. In at least one embodiment, the ALU can be constructed as an asynchronous logic circuit whose internal state is not maintained in an associated register set. In at least one embodiment, a processor uses an ALU to combine operands stored in one or more registers of the processor and produce an output that can be stored by the processor in another register or memory location.
[0094] In at least one embodiment, as a result of processing an instruction retrieved by the processor, the processor presents one or more inputs or operands to an arithmetic logic unit, causing the arithmetic logic unit to generate a result based at least in part on the instruction code provided to the input of the arithmetic logic unit. In at least one embodiment, the instruction code provided by the processor to the ALU is based at least in part on the instruction executed by the processor. In at least one embodiment, combinatorial logic in the ALU processes the inputs and generates an output, which is placed on a bus within the processor. In at least one embodiment, the processor selects a destination register, memory location, output device, or output storage location on the output bus, thereby clocking the processor so that the result generated by the ALU is sent to the desired location.
[0095] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the operations to be performed. Furthermore, the computer system that implements at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices that operate differently, such that the distributed computer system performs the operations described herein, and such that no single device performs all of the operations.
[0096] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate embodiments of the present disclosure and does not impose a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0097] In the description and claims, the terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. Rather, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0098] Unless otherwise expressly stated, it is understood that throughout the specification, terms such as "processing," "computing," "calculating," "determining," etc., refer to the movement and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical quantities (e.g., electronic quantities) in the registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers, or other such information storage, transmission, or display devices of the computing system.
[0099] In a similar manner, the term "processor" may refer to any device or portion of a device that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As non-limiting examples, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process may refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.
[0100] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In at least one embodiment, the process of providing, outputting, transmitting, sending, or presenting analog or digital data may be accomplished by transmitting the data as an input or output parameter to a function call, a parameter to an application programming interface, or an inter-process communication mechanism.
[0101] Although the description herein sets forth example implementations of the described technology, other architectures may be used to implement the described functionality and are intended to fall within the scope of this disclosure. In addition, although specific responsibilities are defined above for descriptive purposes, the various functions and responsibilities may be allocated and divided in different ways depending on the circumstances.
[0102] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological movements, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or movements described. Rather, the specific features and movements are disclosed as example forms of implementing the claims.
Claims
1. A system comprising: at least one execution unit configured to perform inference using a machine learning (ML) model to determine a type associated with a frame of a received media stream based at least in part on using ML model features associated with different types; as well as A video encoder is provided for encoding the media stream based at least in part on the determined type. 2 . The system of claim 1 , wherein the encoded media stream output from the video encoder comprises at least two video sequences associated with different types. 3 . The system of claim 1 , wherein the encoded media stream output from the video encoder comprises different video sequences associated with different encoding parameters representing different types.
4. The system of claim 1 , wherein the features comprise one or more of different noise features of the different types, different motion vector distributions of the different types, different pixel intensity levels of the different types, or different edge features of the different types.
5. The system of claim 1 , wherein the ML model is trained using supervised training, unsupervised training, or semi-supervised training.
6. The system of claim 1 , further comprising a feedback loop from the video encoder to the at least one execution unit for indicating a scene cut event to the at least one execution unit, wherein the encoded media stream output from the video encoder includes different encoding parameters dynamically provided for the encoded media stream based at least in part on the scene cut event.
7. The system of claim 1 , further comprising a feedback loop from the video encoder to the at least one execution unit for indicating to the at least one execution unit at least one of different encoding parameters used by the video encoder, wherein the ML model is composed of sub-ML models to perform different inferences on the received frames in response to at least one of the different encoding parameters.
8. The system of claim 7 , wherein the sub-ML models have different associated memory or processing capacities, and wherein the at least one execution unit is configured to, in response to at least one of the different encoding parameters indicated to the ML model, utilize one of the sub-ML models based in part on a threshold capacity of at least one of the different associated memory or processing capacities.
9. The system of claim 1 , wherein inference using the ML model is performed on processed versions of one or more received frames.
10. The system of claim 1, wherein inference using the ML model is performed using one or more sub-regions of one or more received frames.
11. The system of claim 1 , wherein the ML model is controlled by an application to perform the inference based in part on input from the application, and wherein the video encoder is controlled by a processing infrastructure to perform encoding of the media stream based in part on capabilities associated with the processing infrastructure.
12. The system of claim 11, wherein the application and the processing infrastructure share memory of the system to enable the inference and to enable encoding of the media stream.
13. At least one execution unit, the at least one execution unit being associated with a video encoder, that performs inference using a machine learning (ML) model to determine a type associated with a frame of a received media stream based in part on using ML model features associated with different types, and enables the video encoder to encode the media stream based in part on the determined type.
14. The at least one execution unit of claim 13, wherein the characteristics include one or more of different noise characteristics of the different types, different motion vector distributions of the different types, different pixel intensity levels of the different types, and different edge characteristics of the different types.
15. The at least one execution unit of claim 13, wherein the ML model is trained using supervised training, unsupervised training, or semi-supervised training.
16. The at least one execution unit of claim 13, further comprising an input for receiving feedback from the video encoder, the feedback being used to indicate a scene cut event to the at least one execution unit, wherein An encoded media stream output from the video encoder includes different encoding parameters dynamically provided for the encoded media stream based at least in part on the scene change event.
17. The at least one execution unit of claim 13 , further comprising an input for receiving feedback from the video encoder, for indicating to the at least one execution unit at least one of different encoding parameters used by the video encoder, wherein the ML model is composed of sub-ML models to perform different inferences on the received frames in response to at least one of the different encoding parameters.
18. A video encoder for encoding a media stream based at least in part on a type associated with the media stream, the type being inferred using a machine learning (ML) model executed on at least one execution unit, the type being determined based at least in part on received frames of the media stream using ML model features associated with different types.
19. The video encoder of claim 18, further comprising: an output for providing feedback of the video encoder to the at least one execution unit, the output being configured to indicate to the at least one execution unit at least one of a scene cut event or different encoding parameters used or available to the video encoder, wherein the different encoding parameters are dynamically provided for the coded media stream based at least in part on the scene cut event; and An input for receiving different inferences for the received frame in response to at least one of the different encoding parameters, wherein the ML model is composed of sub-ML models to provide the different inferences.
20. At least one execution unit for training a machine learning (ML) model using features associated with different types of media streams, wherein: Once trained, the ML model is used to enable a video encoder to encode the media stream based in part on the type inferred by the ML model for the media stream, and to enable the video encoder to provide an encoded media stream based in part on the determined type inferred by the ML model.
21. The at least one execution unit of claim 20, wherein the characteristics comprise one or more of different noise characteristics of the different types, different motion vector distributions of the different types, different pixel intensity levels of the different types, or different edge characteristics of the different types.
22. A method for a video encoder, the method comprising: executing a machine learning (ML) model to infer a type associated with a frame of the received media stream based at least in part on using ML model features associated with different types; as well as The media stream is encoded using the video encoder based at least in part on the determined type.
23. The method of claim 22, wherein the characteristics include one or more of different noise characteristics of the different types, different motion vector distributions of the different types, different pixel intensity levels of the different types, or different edge characteristics of the different types.
24. The method of claim 22, further comprising: enabling a feedback loop from the video encoder to at least one execution unit executing the ML model; as well as A scene cut event is determined based on feedback in the feedback loop, wherein a different encoding is dynamically provided for the media stream based at least in part on the scene cut event.
25. The method of claim 22, further comprising: enabling a feedback loop from the video encoder to at least one execution unit executing the ML model; determining at least one of different encoding parameters used or available to the video encoder and provided to the at least one execution unit in the feedback loop; as well as A sub-ML model of the ML model is used to perform different inference on the received frame in response to at least one of the different encoding parameters.
26. The method of claim 22, wherein the sub-ML models have different associated memory or processing capacities, and wherein the method further comprises: determining a threshold capacity for at least one of the different associative memory or processing capacities; as well as Using one of the sub-ML models to perform different inferences based in part on the threshold capacity in response to at least one of different encoding parameters indicated to the ML model from the video encoder.