Patch-based Video Coding for Machines

The patch-based encoding system efficiently encodes regions of interest in video frames using atlases and metadata, reducing bitrate and pixel rate while maintaining quality and flexibility for machine learning applications.

JP7704365B2Active Publication Date: 2025-07-08INTEL CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022553027
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-16
Filing Date
2021-04-15
Publication Date
2025-07-08
Estimated Expiration
2041-04-15

AI Technical Summary

Technical Problem

Existing video compression/decompression systems for machine learning applications are inefficient in coding only regions of interest, leading to high bitrate and pixel rate requirements, and lack flexibility in handling variations in optimal visual features.

Method used

A patch-based encoding and decoding system that detects and encodes regions of interest (ROI) into atlases, using metadata to map these regions within the video frames, allowing for flexible encoding and decoding without full frame representation at high resolution, and supports scalable video coding.

Benefits of technology

Reduces bitrate and pixel rate requirements while maintaining video quality in regions of interest, enhancing coding efficiency and flexibility for machine learning tasks, and avoiding obsolescence with evolving machine learning technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704365000002
    Figure 0007704365000002
  • Figure 0007704365000003
    Figure 0007704365000003
  • Figure 0007704365000004
    Figure 0007704365000004
Patent Text Reader

Abstract

Devices and techniques related to implementing patch-based video coding for machines are described, including detecting a region of interest in a frame of video, extracting the detected region of interest into one or more atlases that are not present in the frame at a resolution equal to or greater than the resolution of the region of interest, and encoding the one or more atlases into a bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Priority Claim] This application claims priority to U.S. Provisional Patent Application No. 63 / 011,179, filed on April 16, 2020, entitled "PATCH BASED VIDEO CODING FOR MACHINES", the entire content of which is incorporated herein by reference for all purposes.

Background Art

[0002] In some contexts, a video compression / decompression (codec) system may be employed such that the reconstructed video is used by machines rather than for human viewing. For example, the reconstructed video may be used in the context of machine learning and other applications. Currently, the MPEG Video Coding for Machines (VCM) group is researching methods for calculating visual features from images or videos and the possibility of standardizing the compression of these visual features using compact features. The draft standards for MPEG Video Point Cloud Coding (V-PCC) and Immersive Video (MIV) describe the selection of patches from multiple views, the placement of patches into an atlas, and the encoding of the atlas as a video using conventional video codecs, and further define metadata for describing the patches and a method for signaling their placement into the atlas.

Brief Description of the Drawings

[0003] The subject matter described herein is shown by way of example and is not limited to the accompanying drawings. For the sake of brevity and clarity, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Further, reference numerals may be repeated between the drawings to indicate corresponding or similar elements where appropriate. The following are shown in the figures.

[0004]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Mode for Carrying Out the Invention

[0005] Next, with reference to the accompanying drawings, one or more embodiments or implementations will be described. It should be understood that the description of specific configurations and arrangements is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be adopted without departing from the spirit and scope of the description. It will be apparent to those skilled in the art that the techniques and / or arrangements described herein can also be used in various other systems and applications other than those described herein.

[0006] The following description describes various implementation forms that can be realized in an architecture such as, for example, a system-on-chip (SoC) architecture. However, the implementation forms of the technologies and / or arrangements described in this specification are not limited to a specific architecture and / or computing system, and may be implemented by any architecture and / or computing system for a similar purpose. By way of example, various architectures that use, for example, multiple integrated circuit (IC) chips and / or packages, and / or various computing devices and / or consumer electronics (CE) devices such as set-top boxes, smartphones, etc. may implement the technologies and / or arrangements described in this specification. Further, the following description may describe a number of specific details such as logic implementation forms, types and interrelationships of system components, options for logic partitioning / integration, etc., but the claimed subject matter may be implemented without such specific details. In other examples, some topics such as, for example, control structures and complete software instruction sequences may not be shown in detail so as not to obscure the topics disclosed herein.

[0007] The subject matter disclosed herein may be implemented in hardware, firmware, software, or any combination thereof. The subject matter disclosed herein may also be implemented as instructions stored on a machine-readable medium readable and executable by one or more processors. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.).

[0008] References to "one implementation", "an implementation", "exemplary implementations", etc. in this specification indicate that the described implementations may include specific features, structures, or characteristics, but not all implementations necessarily include the specific features, structures, or characteristics. Also, such phrases do not necessarily refer to the same implementation. Further, when a specific feature, structure, or characteristic is described in relation to one implementation, it is considered within the knowledge of those skilled in the art to achieve such feature, structure, or characteristic in relation to other implementations, whether or not explicitly described in this specification.

[0009] The terms "substantially", "close", "approximately", "near", and "about" generally refer to being within + / - 10% of a target value. For example, unless otherwise specified in the explicit context in which they are used, the terms "substantially equal", "nearly equal", and "approximately equal" mean that there are only incidental variations between those so described. In the art, such variations are typically within + / - 10% of a given target value. Unless otherwise specified, the use of ordinal adjectives such as "first", "second", and "third" to describe common objects only indicates that different instances of similar objects are being referred to, and is not intended to mean that the objects so described must be in a given order in time, space, ranking, or other ways.

[0010] In this specification, methods, devices, apparatuses, computing platforms, and articles related to machine video coding are described such that the reconstructed video is used by a machine rather than being viewed by a human.

[0011] As described, in some contexts, a video compression / decompression (codec) system may be employed such that the reconstructed video is used by a machine rather than for human viewing. For example, the reconstructed video may be used in the context of machine learning and other applications. As used herein, terms such as machine learning, machine learning operation, etc. indicate the use of a machine learning application to determine an output from a portion of a video. Such machine learning applications include, for example, human face recognition, object recognition, image recognition, etc. Such applications may be deployed in a variety of contexts such as surveillance, video classification, image tagging, etc. In some embodiments, a video analysis system requires the use of video compression to transmit the captured video to an analysis engine. The techniques described herein improve the coding efficiency, accuracy, and compression speed of videos for machine learning rather than for human viewing such that the transmission bandwidth and pixel rate are advantageously reduced.

[0012] In some embodiments, rectangular regions of interest in a video are detected and placed on patches. Although rectangular regions of interest have been illustrated and described, any suitable shape may be used. As used herein, the term region of interest refers to the region to which a machine learning operation is applied (whether the result is a positive recognition or not). The patches are packed together into a video atlas (or multiple atlases), and the video atlas may also include a downsampled representation of the entire input video frame. The video atlas is encoded by a video coder such as an AVC (Advanced Video Coding)-compliant encoder, an HEVC (High Efficiency Video Coding)-compliant encoder, a VVC (Versatile Video Coding)-compliant encoder, etc. Further, metadata is signaled to indicate the size and location of each region of interest of the original video and the atlas video. For example, the metadata maps the regions of interest between one or more atlases and the frames of the video. Such techniques are advantageous in coding only the regions of interest of the video rather than the full video to reduce the bitrate and pixel rate required for the video, maintain the video quality within the regions of interest, and avoid a degradation in the accuracy of machine learning tasks that use the reconstructed video.

[0013] As described above, one or more atlases lack a full frame of video at a resolution greater than that of the region of interest. Further, one or more bitstreams obtained from the encoding of the (one or more) atlases lack a representation of a full frame of video at a resolution greater than that of the coded region of interest. In some embodiments, the full frame of video is not present at all (e.g., the bitstream does not contain any representation of the full frame of video). In some embodiments, the full frame of video is represented at a resolution lower than that of the region of interest. In particular, the techniques described herein do not rely on predefined visual features. Instead, the actual region of interest is coded. Thus, the techniques described herein are advantageously flexible with respect to variations in optimal visual features. Further, the techniques used to identify the region of interest may be selected by the encoder and, advantageously, need not be standardized. As a result, aspects included in any standardization of the techniques described herein are avoided from becoming obsolete as state-of-the-art machine learning evolves.

[0014] FIG. 1 is a block diagram of an exemplary patch - based encoder 100, an exemplary patch - based decoder 110, and an exemplary machine - learning system 120 arranged according to at least some implementations of the present disclosure. As shown, the patch - based encoder 100 includes a region of interest (ROI) detector and extractor 111, an atlas builder 112, a video encoder 113, a metadata encoder 114, and a multiplexer 115. The patch - based decoder 110 includes a demultiplexer 121, a video decoder 122, a metadata decoder 123, and an optional image reconstructor 124. As shown, the machine - learning system 120 performs a learning task based on received data. The patch - based encoder 100, the patch - based decoder 110, the machine - learning system 120, and any other encoder, decoder, machine - learning system, or other device can each be implemented in one or more of such devices including any suitable form - factor device, or a server computer, a cloud - computing environment, a personal computer, a laptop computer, a tablet, a phablet, a smartphone, a game console, a wearable device, a display device, an all - in - one device, a two - in - one device, etc. In some embodiments, the patch - based decoder and the machine - learning system are implemented on the same device.

[0015] In the patch-based encoder 100, video 102 is input to the ROI detector and extractor 111. Such input video 102 may include any suitable video frame, video picture, sequence of video frames, group of pictures, groups of pictures, video data, etc. of any suitable resolution. For example, the video may be Video Graphics Array (VGA), high definition (HD), full HD (e.g., 1080p), 4K resolution video, 8K resolution video, etc., and the video may include any number of video frames, sequences of video frames, pictures, groups of pictures, etc. In some embodiments, the input video 102 is an immersive video having multiple views of a scene. The techniques herein are described with respect to frames and pictures, blocks, and sub-blocks of various sizes (used interchangeably) for clarity of presentation. As used herein, a block or coding unit may be of any size and shape that includes a plurality of pixel samples (usually square or rectangular) within any suitable color space such as YUV. Further, a block or coding unit may have sub-blocks or prediction units and may be characterized as a block depending on the context. Also, a block, sub-block, or coding unit may optionally be divided into transform blocks or transform units for the purpose of residual transformation. As used herein, the term size indicates the size of such coding units, transform units, etc., and does not necessarily include the unit itself. The term coding unit or transform unit may indicate its size. Such a frame may also be characterized as a picture, video picture, sequence of pictures, video sequence, etc., and such a coding unit or block may be characterized as a maximum coding unit, coding unit, coding block, macroblock, subunit, sub-block, etc.

[0016] The ROI detector and extractor 111 of the patch-based encoder 100 receives the video 102 and detects and extracts patches or regions of interest using any suitable technique. For example, the ROI may be detected using deep neural networks including feature extraction, machine learning, convolutional neural networks, etc. For example, the region of interest may include an area of a frame or picture of the video that is considered of interest to the machine learning system 120. In some embodiments, the region of interest is an area detected to include a face, and the machine learning system 120 may perform face recognition on the detected face. In particular, face detection is a relatively simple task compared to the more difficult task of face recognition. The patch or region of interest is input to the atlas constructor 112 of the patch-based encoder 100 that forms one or more atlases using the patches. As used herein, the term atlas refers to an image or frame that includes several image patches or regions of interest. For example, the atlas may be formed by clipping regions of interest from frames of the video and packing them into the atlas. Any number of atlases may be formed using such techniques having each atlas including one or more of the regions of interest. The atlas constructor 112 or the ROI detector and extractor 111 (or its ROI extractor block) may optionally perform image scaling to reduce the resolution of the individual patches.

[0017] The atlas builder 112 outputs a video atlas (e.g., a video frame at the original resolution or a subsampled resolution) for encoding using the video encoder 113 of the patch-based encoder 100, which may perform any suitable encoding such as standards-compliant encoding. For example, the video encoder 113 of the patch-based encoder 100 may generate a standards-compliant bitstream of the bitstream 105 using HEVC, AVC, VVC, etc. Also, the atlas builder 112 outputs metadata that describes patches (e.g., regions of interest) and the atlas, encoded by the metadata encoder 114 of the patch-based encoder 100. The outputs of the video encoder 113 and the metadata encoder 114 are multiplexed together into one or more bitstreams 105 via the multiplexer 115 of the patch-based encoder 100. The resulting bitstream 105 is transmitted by the patch-based encoder 100 to the patch-based decoder 110 for final decoding by the patch-based decoder 110, etc.

[0018] The patch-based decoder 110 receives a bitstream 105 from the patch-based encoder 100 or another device or system. In the patch-based decoder 110, the bitstream 105 is demultiplexed into a video bitstream and a metadata bitstream component, and is decoded by a video decoder 122 and a metadata decoder 123, respectively. Optionally, the image reconstructor 124 can reconstruct an image or frame of the same size as the input video 102. Such reconstruction may include scaling individual patches and placing the patches in the same locations where the patches were placed within the image or frame input to the patch-based encoder 100 as the input video 102. As shown, either the reconstructed image or patches (e.g., regions of interest) are input to a machine learning system 120 to perform a machine learning task to generate a machine learning (ML) output 107. The machine learning task may be any suitable machine learning task such as object detection, object recognition, face detection, face recognition, semantic segmentation, etc. The machine learning output 107 may include any data structure indicative of a machine learning task such as pixel segmentation data, pixel label data, object recognition identifiers, face recognition identifiers, etc.

[0019] FIG. 2 is a block diagram of another exemplary patch-based encoder 200, another exemplary patch-based decoder 210, and another exemplary machine learning system 220 arranged in accordance with at least some implementations of the present disclosure. As shown, the patch-based encoder 200 includes a region of interest (ROI) detector and extractor 211, an atlas builder 212, a feature encoder 213, a metadata encoder 214, and a multiplexer 215. The patch-based decoder 210 includes a demultiplexer 221, a feature decoder 222, and a metadata decoder 223. As shown, the machine learning system 220 performs a learning task based on received data. In particular, FIG. 2 shows an alternative system in which the visual feature encoder and decoder replace the video encoder and decoder described with respect to FIG. 1. The system of FIG. 2 enables the techniques discussed herein to be used in combination with feature-based encoding techniques. For example, the components of the systems of FIGS. 1 and 2 can be combined in any suitable manner.

[0020] In the patch-based encoder 200, the input video 102 is input to the ROI detector and extractor 211, and the ROI detector and extractor 211 detect and extract patches using, for example, any suitable technique described herein with respect to the ROI detector and extractor 111. The patches or regions of interest are input to the atlas constructor 212 of the patch-based encoder 200 that forms one or more atlases using the patches. The atlas constructor or ROI extractor 212 may optionally perform image scaling to reduce the resolution of the individual patches. The atlas constructor 212 outputs features for encoding using the feature encoder 213 of the patch-based encoder 200 that can implement any suitable encoding technique. The atlas constructor 212 also outputs metadata that describes the features and atlases encoded by the metadata encoder 214 of the patch-based encoder 200. The outputs of the video encoder and the metadata encoder (e.g., bitstreams) are multiplexed together into one or more bitstreams 205 via the multiplexer 215 of the patch-based encoder 200. The resulting bitstream 205 is transmitted by the patch-based encoder 200 to the patch-based decoder 210 for final decoding by the patch-based decoder 210 and the like.

[0021] The patch-based decoder 210 receives one or more bitstreams 205 from the patch-based encoder 200 or another device or system. In the patch-based decoder 210, the one or more bitstreams 205 are demultiplexed into a feature bitstream and a metadata bitstream component that are respectively decoded by the feature decoder 222 and the metadata decoder 223. As shown, the reconstructed features and the decoded metadata are input to the machine learning system 120 to perform a machine learning task based on the decoded features and their positions in the input video as indicated by the metadata. The machine learning task may be any suitable machine learning task such as object detection, object recognition, semantic segmentation, and the like.

[0022] As described above, the metadata to be encoded and decoded includes the number of patches (e.g., regions of interest), the number of atlases, the size of each atlas, the size of each patch (e.g., region of interest), and information on the correspondence between video frames and patches (e.g., regions of interest). For example, the metadata may provide, for a patch (e.g., region of interest), the correspondence between the position (and size) of the patch in the input video and the position (and size) of the patch in one of any number of atlases. For each patch, the information may include the position and size (top, left, width, height) of the original image patch as well as the position and size (top, left, width, height) of the atlas patch.

[0023] In some embodiments, the size of each patch may be explicitly signaled. However, instead of explicitly signaling the size of the patches within an atlas, the scaling factor used for both the horizontal and vertical dimensions, or either a separate horizontal scaling factor and a vertical scaling factor that are signaled, may be signaled. In some embodiments, a scaling factor for each atlas is signaled and the patches are arranged within the atlas corresponding to those scaling factors. The scaling factor selected by the encoder can be based on, for example, the expected importance of a particular patch for a machine learning task, because the patch represents an object detected with high confidence or because the patch contains significant high-frequency video content. In some embodiments, higher-confidence values or higher-frequency content videos have smaller scaling factors relative to lower-confidence values or lower-frequency video content. In some embodiments, the scaling factor is signaled as a log2 number (e.g., a downscaling factor of 8 can be signaled as log2(8) = 3 bits).

[0024] Table 1 below shows examples of the syntax for each patch (e.g., region of interest). For example, such syntax may be coded as metadata for each patch (e.g., region of interest) within one or more atlases. [Table 1] Table 1 - Examples of Syntax

[0025] In an exemplary syntax, information is signaled with respect to a rectangular bounding box around the patch. However, in some embodiments, if the patch does not occupy all sample positions of the rectangular bounding box, the patches may overlap within the atlas. As shown in Table 1, in some embodiments, for each patch (e.g., region of interest), the upper left position (patch_source_pos_x, patch_source_pos_y) within the atlas, the size (patch_source_pos_width, patch_source_pos_height) within the atlas, and the upper left position (patch_cur_pic_pos_x, patch_cur_pic_pos_y) within the frame are provided. Further, for the scaling information for sizing the patch (e.g., region of interest), a flag (patch_scale_flag) indicating whether scaling is used is first provided, and if scaling information is provided, it is provided via a scaling factor, which may be signaled as a log2 number (patch_log2_scale_x, patch_log2_scale_y).

[0026] Any suitable technique can be used to determine the region of interest (e.g., by the ROI detector and extractor 111 of the patch-based encoder 100, or the ROI detector and extractor 211 of the patch-based encoder 200). In some embodiments, the detection and extraction techniques used depend on the intended machine learning task. For example, in a use case where the machine learning task is face recognition, the ROI detector may perform face detection. In particular, face detection is a much simpler operation than face recognition. Further, face recognition / matching requires access to a library of faces, which is not required for face detection. For example, the ROI detector and extractor 111 of the patch-based encoder 100, or the ROI detector and extractor 211 of the patch-based encoder 200, may perform detection and extraction that supports the machine learning task used after decoding. In some embodiments, the ROI detector and extractor 111 or the ROI detector and extractor 211 performs object detection to support object recognition or semantic segmentation after decoding.

[0027] Figure 3 shows an exemplary input image or frame 300 including exemplary detected regions of interest 301, 302, 303, 304 arranged according to at least some implementations of the present disclosure. For example, the input image or frame 300 may be a frame of the input video 102. As shown, the ROI detector and extractor 111 of the patch-based encoder 100, or the ROI detector and extractor 211 of the patch-based encoder 200, may detect and extract regions of interest 301, 302, 303, 304 including indicating bounding boxes around each of the regions of interest 301, 302, 303, 304. In the example of Figure 3, each of the detected regions of interest 301, 302, 303, 304 is a face, and face detection can be used to determine the ROI and bounding box of each face around the regions of interest 301, 302, 303, 304. Further, as shown, each of the regions of interest 301, 302, 303, 304 is placed at a specific location within the input image or frame 300 as indicated by their bounding boxes.

[0028] FIG. 4 shows an exemplary atlas 400 formed from exemplary regions of interest of the input image or frame of FIG. 3 arranged according to at least some implementations of the present disclosure. As shown, patches (e.g., regions of interest) are determined as defined by the bounding boxes of each detected region of interest 301, 302, 303, 304 that includes objects of interest, faces of interest, etc., and are arranged in the atlas 400. For example, each detected region of interest 301, 302, 303, 304 is placed into a rectangular patch, and all of the patches are arranged in the atlas 400 as shown. For example, the atlas constructor 111 of the patch-based encoder 100 or the atlas constructor 211 of the patch-based encoder 200 may generate the atlas 400 of FIG. 4 based on the detected faces (and corresponding patches), detected objects, detected features, etc. of the input image or frame 300 shown in FIG. 3.

[0029] As described, in some embodiments, an atlas video that includes the atlas 400 and other atlases at the same time instance, or atlases at multiple time instances, is encoded using a video encoder such as the video encoder 113. Thus, it is desirable to have consistency of the content within the atlas video over consecutive pictures so that inter-frame prediction can efficiently code the video. Further, video encoders such as HEVC-compliant video encoders and AVC-compliant video encoders require that all coded pictures within a video sequence have the same picture size (e.g., width and height). In embodiments where the video encoder of the patch-based encoder 100 uses an HEVC or AVC codec (or another codec that requires the video sequence to have the same picture size), the atlas maintains the same picture size for the entire sequence until a newly intra-coded picture (e.g., an I picture or an IDR picture) is used. In some embodiments, the first and second atlases of a video sequence have the same size. In some embodiments, the first and second atlases of a video sequence have the same size in response to being encoded by one of the HEVC or AVC-compliant encoders. In some embodiments, detected and extracted patches from the same location within an input video frame are placed at the same location and have the same size (and resolution) within the atlas to enhance the temporal correlation within the video atlas for encoding.

[0030] In the VVC standard, a reference picture size change feature is provided that allows individual coded pictures within the same sequence to have different sizes. In VVC, the actual width and height of each coded picture are signaled together with the maximum picture height and width of the sequence. In some embodiments, the use of the VVC reference picture size change feature is used, for example, to add additional patches in the middle of an atlas sequence when a new object enters the video, which changes the size of the atlas without requiring modification of the size and position of other patches within the atlas from previous pictures and without requiring an intra-coded picture. Such techniques can improve inter-frame prediction and coding efficiency when a new object enters the video.

[0031] FIG. 5 shows an exemplary atlas size change 500 in the context of reference picture size change arranged according to at least some implementations of the present disclosure. As shown, in some embodiments, only regions of interest 301, 302, 303 are present in the image or frame 501 of the input video. For frame 501 (time or timestamp t), a corresponding atlas 511 is formed that includes regions of interest 301, 302, 303 as patches therein. Metadata 521 is generated for atlas 511 to provide a mapping of regions of interest 301, 302, 303 within atlas 511 and frame 501 (shown not to scale for clarity of presentation) as described herein. Also, metadata 521 may indicate, for example, the size of atlas 511 according to a standard codec.

[0032] In a subsequent frame 502 (in the case of time or time stamp t+x), the region of interest 304 is detected for the first time. In the case of frame 502, the corresponding atlas 512 is formed to include regions of interest 301, 302, 303, 304 as patches therein, and the formation of atlas 512 includes a sizing operation 523 for changing one or both of the height and width of the atlas. Metadata 522 is generated for atlas 512 to provide a mapping of regions of interest 301, 302, 303 in atlas 511 and frame 502. Further, metadata 522 may indicate the size of atlas 512 and / or the change in size with respect to atlas 511, for example, according to a standardized codec. As described, the video codec may provide sized features to all coded pictures or frames in the same sequence to have different sizes. Thereby, the region of interest 304 may be efficiently added to atlas 512, while the regions of interest 301, 302, 303 in atlas 512 may be coded using inter-coding techniques compared to previous regions of interest (having similar or the same content).

[0033] Further, as described above, better coding efficiency for inter-frame prediction can be achieved if the placement of patches into the atlas is consistent for consecutive pictures (e.g., the atlas). However, if the ROI represents a moving object, the ROI size is likely to change from picture to picture (e.g., over multiple frames). In some embodiments, especially in latency-tolerant systems, look-ahead may be performed to find the ROI of a particular object in multiple pictures. In some embodiments, the overall ROI including all positions occupied by individual picture ROIs is found and the overall ROI is used in the ROI extraction block.

[0034] In some embodiments, lookahead is not performed (e.g., lookahead may not be permitted due to latency issues). In some embodiments, the encoder forms a larger ROI for an object in the first picture by placing a buffer around each object. Thereafter, in subsequent pictures, if the object size does not become larger than the larger ROI selected, there is no need to change the ROI size. The ROI detector may also choose not to move the position of the ROI in order to maintain background consistency. In some embodiments, a buffer of a particular fraction of the ROI within the first frame (e.g., 20% - 50% additional buffer area) is added around the ROI, the patches extracted for the ROI are not moved, and the ROI size does not change across multiple atlases (from the buffered patch size including the ROI and the buffer), so that the ROI is expected to change size within the buffered patch size but not expand outside it. Such techniques may provide enhanced temporal correlation and coding in some contexts.

[0035] FIG. 6 shows an exemplary ROI size change 600 in the context of lookahead analysis or additional context of buffer regions without lookahead analysis, arranged according to at least some implementations of the present disclosure. As shown, in some embodiments, an ROI 611 with a relatively small bounding box (BB1) is detected within a frame (e.g., the frame at time or timestamp t). For example, the bounding box of the ROI 611 may be formed immediately around the detected object, or detected face, etc., such that no buffer or a very small buffer (e.g., 1 - 5 pixels) is provided around the detected object, or detected face, etc.

[0036] As shown, in some embodiments, the look-ahead analysis 605 is performed at any rate, such as per frame, every other frame, etc., or for any number of subsequent frames in time (e.g., frames at times or timestamps t+1, t+2, ···) with respect to a single frame of a specific look-ahead from the current frame. In the look-ahead frame 602, detections are performed again to determine the bounding box around the corresponding region of interest 612 such that the regions of interest 611, 612 are expected to contain the same object, face, etc. based on shared or similar features or shared or similar frame positions. Note that the region of interest 612 is larger than the region of interest 611. In some embodiments, based on the look-ahead analysis 605 that detects a larger region of interest 611 within the subsequent frame 602 in time, the region of interest 611 is expanded to the size of the region of interest 612 as shown with respect to the sizing operation 606. Next, with respect to frame 601, a region of interest 631 that includes the region of interest 611 and has the size of the region of interest 612 is inserted into the atlas 621. Similarly, for all frames from the current frame 601 to the subsequent frame 602 in time (e.g., with any number of frames between frames 601, 602), regions of interest having the size of the region of interest 612 (or the maximum region of interest size of any frame evaluated as part of the look-ahead analysis 605) are inserted into the corresponding atlas.

[0037] As shown, the atlas 621 may include the region of interest 631 and other regions of interest including regions of interest 632, 633, which may also be sized to the maximum size of the corresponding regions of interest detected using the look-ahead analysis 605 as described. Similarly, the atlas 622 may include the region of interest 612 (not sized in this example as it has the maximum region of interest size) and other regions of interest including regions of interest 634, 635, which may be sized as needed to the maximum size of the corresponding regions of interest detected using the look-ahead analysis 605. For example, the maximum region of interest size may be detected in any frame evaluated using the look-ahead analysis 605.

[0038] As described, the inter-frame prediction coding efficiency can be advantageously increased by consistently placing patches (e.g., regions of interest) in the atlas for successive atlases. To achieve such an arrangement having regions of interest that represent moving objects and thus are likely to change in size, lookahead analysis 605 is used to find an overall ROI or maximum ROI size that includes all the positions occupied by the individual picture ROIs, and the overall ROI or maximum ROI size is used within the ROI extraction block for the atlas of the corresponding frame.

[0039] In other embodiments, the look-ahead analysis 605 is not performed. For example, the look-ahead analysis 605 may not be permitted due to latency issues. In some embodiments, the sizing operation 606 may be performed without the look-ahead analysis 605 to change the size of the region of interest 611 by any suitable amount (e.g., when an object or face is first detected) to a larger size. For example, the size of the bounding box or the detected region of interest can be increased by 20% to 50%. Subsequently, a larger-sized region of interest (here represented by the sizes of region of interest 612 and region of interest 631 for clarity of presentation) is inserted into the atlas 621. Similar operations may be performed on the region of interest 632. Using such padded or expanded regions of interest increases the likelihood that the regions of interest maintain consistency through encoding. For example, in subsequent pictures or frames such as frame 602, if the object size does not become larger than the larger ROI selected, as shown with respect to region of interest 612 and bounding box BB2, there is no need to change the ROI size. The ROI detector may also choose not to move the position of the ROI to maintain background consistency. In some embodiments, a buffer of a particular fraction of the ROI within the first frame (e.g., an additional buffer region of 20% - 50%) is added around the ROI, the patches extracted for the ROI are not moved, and the ROI size does not change across multiple atlases (from the buffered patch size including the ROI and the buffer), such that the ROI is expected to change size within the buffered patch size but not expand outside of it. Such techniques may provide enhanced temporal correlation and coding in some contexts.

[0040] In some embodiments, the machine learning task application may be interested in accessing not only the ROI but also the full video. In some embodiments, the entire input image (or frame) is included within the patch, and the patch is downscaled to reduce the bitrate and pixel rate.

[0041] FIG. 7 shows an exemplary atlas 700 that includes extracted region-of-interest patches 301, 302, 303, 304 and a low-resolution full input image or frame 711 arranged according to at least some implementations of the present disclosure. As shown, the ROI detector and extractor 111 of the patch-based encoder 100, or the ROI detector and extractor 211 of the patch-based encoder 200, may detect and extract the regions of interest 301, 302, 303, 304 by showing bounding boxes around each of the regions of interest 301, 302, 303, 304, and includes each of the patches or regions of interest 301, 302, 303, 304 (as defined by the bounding boxes around each of the regions of interest 301, 302, 303, 304), as well as the corresponding video input image to the atlas 700 (e.g., the image from which the regions of interest 301, 302, 303, 304 were extracted). In the example of FIG. 7, each of the regions of interest 301, 302, 303, 304 is a face, but any type of ROI may be included.

[0042] In particular, to improve coding efficiency, the input frame 310 is downscaled to a downscale frame 711 having a resolution lower than the resolution of the regions of interest 301, 302, 303, 304 using any suitable technique such as downsampling as shown by the downscaling operation 720. In some embodiments, the regions of interest 301, 302, 303, 304 have the same resolution as the input frame 310, and the downscaled frame 711 has a resolution lower than the resolution of the input frame 310. In some embodiments, the regions of interest 301, 302, 303, 304 have a resolution lower than the input frame 310, and the downscaled frame 711 has a resolution even lower than the regions of interest 301, 302, 303, 304. For example, the regions of interest 301, 302, 303, 304 may be provided at a higher resolution (relative to the downscaled frame 711) to improve post-decoding machine learning. As shown, metadata 712 is also generated and encoded for the atlas 700 along with metadata 712 that provides a mapping between the sizes and positions of the regions of interest 301, 302, 303, 304 within the atlas 700 and the input frame 310, as well as an indication of the presence of the downscaled frame 711 and its scaling factor.

[0043] In some embodiments, scalable coding is used to code patches (e.g., regions of interest) that represent objects with respect to patches or atlases that include low-resolution full images. For example, the low-resolution full image may be encoded as a base layer of a scalable video encoder, and each patch may be encoded as an enhancement layer of the scalable video encoder using an enhancement layer that references the base layer for encoding.

[0044] FIG. 8 shows an exemplary scalable video coding 800 that includes coding regions of interest 301, 302, 303, 304 for the same region of a downscaled frame 711 arranged according to at least some implementations of the present disclosure. As shown with respect to region of interest 301, each region of interest may be extracted as a high-resolution or higher-resolution layer 812, and the corresponding region of the downscaled frame 711 may be extracted as a low-resolution or lower-resolution layer 813. Further, the low-resolution or lower-resolution layer 813 provides a base layer (BL) in the scalable video coding 800.

[0045] In the context of scalable video coding 800, on the encoding side, the low-resolution or lower-resolution layer 813 is encoded via a BL encoder 815 to generate a BL bitstream 821. Further, a local decoding loop is applied via a BL decoder 817, and the reconstructed version of the resulting low-resolution or lower-resolution layer 818 is differenced from the high-resolution or higher-resolution layer 812 to generate an enhancement layer (EL) 814. Thereafter, the enhancement layer 814 is encoded using an EL encoder 816 to generate an EL bitstream 822. The BL bitstream 821 and the EL bitstream 822 may be included in the bitstream 105 as described herein. For example, portions of an atlas including the downscaled frame 711 and the region of interest 301 may be coded using such techniques.

[0046] On the decoding side, the BL bitstream 821 is decoded using the BL decoder 817 to generate a reconstructed version of the low-resolution or lower-resolution layer 818. Also, the EL bitstream 822 is decoded using an EL decoder to generate the enhancement layer 814. Then, the enhancement layer 814 and the reconstructed version of the low-resolution or lower-resolution layer 818 are added to generate a reconstructed version of the region of interest 301. The same operation may be performed for any number of regions of interest 301, 302, 303, 304 and the corresponding regions of the downscaled frame 711. Such techniques may utilize the presence of the downscaled frame 711 to improve coding efficiency.

[0047] In some embodiments, the regions of interest 301, 302, 303, 304 overlap each other within the input image. In some embodiments, in the decoder, reconstruction of the input image at the original resolution is performed, for example, when it is preferred for a machine learning task. In some embodiments, the input image may be reconstructed at a lower resolution. Further, when the machine learning task can operate directly on the patches, the pixel rate required to perform the machine learning task is reduced. In some embodiments, only the patches are encoded and transmitted to reduce the bitrate and pixel rate. For example, the reconstructed original-sized image may have gaps for regions not included within the patches. In some embodiments, the gaps can be filled using a single color such as gray. In some embodiments, an atlas containing only ROI patches (see FIG. 4) is encoded, reconstructed (at the decoder), formed in the reconstructed image, the regions between the ROI patches are left unencoded, reconstructed in the regions between the ROI patches as described above, and provided in a single color.

[0048] FIG. 9 shows an exemplary reconstructed image or frame 900 that includes only the reconstructed regions of interest 901, 902, 903, 904 having the remaining reconstructed image regions as monochromatic 911, arranged according to at least some implementations of the present disclosure. As shown, video decoder 122 decodes the atlas, and image reconstructor 124 may reconstruct an image or frame 900 that includes patches corresponding to the detected and extracted regions of interest 901, 902, 903, 904, or ROI patches, such that the area between such regions of interest 901, 902, 903, 904 is left as a single color, such as gray. Such techniques may save bitrate and computation when machine learning tasks performed by machine learning system 120 operate only on regions of interest 901, 902, 903, 904 and do not require a fully reconstructed image corresponding to, for example, input image or frame 300.

[0049] In some embodiments, the full-resolution image includes an overlap, and the sample values of the decoded patches can be combined as an average or weighted average, or a single patch can be prioritized. The priority may be based on the signaled priority, the order in which the patch metadata is signaled, or the scaling factor of the patch. In some embodiments, when a downscaled version of the entire image (e.g., lower resolution) is included as a patch, unscaled patches (e.g., at the higher original resolution) representing the ROI are prioritized.

[0050] FIG. 10 shows an exemplary reconstructed image or frame 1000 having regions of interest 1001, 1002, 1003, 1004 of the original size or resolution, arranged according to at least some implementations of the present disclosure, where the complete reconstructed image 1010 is scaled and at a lower resolution. As shown, the reconstructed image 1000 is reconstructed using the complete reconstructed image 1010 (e.g., from the portions other than the regions of interest 1001, 1002, 1003, 1004 to the full frame size) that is downscaled and included as patches in the atlas, and other patches in the atlas (e.g., the regions of interest 1001, 1002, 1003, 1004) that are prioritized during the reconstruction process. In particular, the reconstructed patches or regions of interest 1001, 1002, 1003, 1004 have a higher resolution than the remaining portions of the complete reconstructed image 1010. For example, the detected and extracted patches (e.g., the regions of interest 301, 302, 303, 304) may be included in the atlas at the original resolution, and the full image may be downsampled or downscaled and included in the atlas as described with respect to FIG. 7. During reconstruction, the atlas is decoded, the frame is reconstructed, and the resulting reconstructed frame 1000 is based on the downsample or downscale version of the full image included in the atlas and the reconstructed ROI patches at the original resolution. In some embodiments, such decoding techniques may include scalable video coding as described with respect to FIG. 8.

[0051] In some embodiments, the patches may be divided among multiple atlases.

[0052] FIG. 11 shows exemplary first and second atlases 1102, 1103 having a first atlas 1102 that includes a low - resolution version of an input image 711 and a second atlas 1103 that includes full - resolution regions of interest 301, 302, 303, 304, arranged according to at least some implementations of the present disclosure. As shown, including the regions of interest 301, 302, 303, 304 means that the down - scaled full image is included in the second atlas 1103 (atlas 1), and the down - scaled full image is included in the first atlas 1102 (atlas 1). In some embodiments, separate atlases such as the first atlas 1102 and the second atlas 1103 are coded as separate video bitstreams and offer the option of choosing to transfer them within a media recognition network element for convenience, or of choosing to decode a subset of the atlases at a decoder. In the example of FIG. 11, a system such as a patch - based decoder 110 may choose to decode the down - scaled full image as provided in the first atlas 1102 or only the regions of interest 301, 302, 303, 304 as provided in the second atlas 1103. As shown in FIG. 11, metadata 1112 indicates the content of the first atlas 1102 as a down - scaled version of the full frame, indicates scaling factors as needed, and indicates the content of the second atlas 1103 as regions of interest, and is provided to locate and size the regions of interest within the full video frame as described herein.

[0053] Furthermore, separation into multiple atlases can provide enhanced privacy. For example, faces may be removed from the input image and / or faces may be placed in a separate atlas.

[0054] FIG. 12 shows an exemplary atlas 1200 with regions of interest patches 1201, 1202, 1203, 1204 removed, arranged according to at least some implementations of the present disclosure. The atlas 1200 of FIG. 12 may be generated, for example, by replacing the regions of interest 301, 302, 303, 304 with a single color. For example, the regions of interest 301, 302, 303, 304 (e.g., faces) may be combined together in a single separate atlas, or each of the regions of interest 301, 302, 303, 304 (e.g., faces) may be placed in its own atlas. Such separation into separate atlases provides that the video content is represented by separately encoded video sequences and enables separate handling (e.g., for encryption for privacy or access). In some embodiments, the codec standard can be defined such that a particular operating mode is selectable by the application. Alternatively, the patches may be separated into separate tiles for the video codec. For example, the regions of interest 301, 302, 303, 304 may be provided within one or more separate atlases as discussed herein. In encoding, the bitstream corresponding to such one or more separate atlases may be encrypted for privacy reasons. In decoding, such an encrypted bitstream may be decrypted before decoding, and such decoding is performed using any suitable one or more of the techniques described herein.

[0055] In some embodiments, on the encoder side, preprocessing of the patches such as edge enhancement or blurring is supported. For example, ROI patches (e.g., faces or other objects) may be blurred or removed due to privacy concerns. In some embodiments, the syntax is defined to signal the type of preprocessing performed for each patch or for each atlas so that it can be reversed at decoding time.

[0056] In some embodiments, a video encoder that encodes an atlas may select to indicate that a particular picture is marked as a long-term reference picture when the arrangement of patches within the atlas changes, or changes significantly, such as by the addition of new patches. In some embodiments, high-level syntax elements signal that the encoder uses a long-term reference picture via an indication. In some embodiments, signaling, such as a syntax flag, is used to indicate a significant change in the arrangement of patches within the atlas. In such a system, a machine learning task on the decoder side can be advantageously simplified by only decoding the long-term reference picture when operations such as classification need to be performed only when significant changes occur. For example, simpler machine learning operations may be performed on other pictures.

[0057] FIG. 13 is a flowchart showing an exemplary process 1300 for encoding and / or decoding video for machine learning, arranged in accordance with at least some implementations of the present disclosure. Process 1300 may include one or more operations 1301-1307, as shown in FIG. 13. Process 1300 may form at least a portion of an encoding process (i.e., operations 1301-1304) and / or a decoding process (i.e., operations 1305-1307) in the context of machine learning. Further, process 1300 is described herein with reference to system 1400 of FIG. 14.

[0058] FIG. 14 is an explanatory diagram of an exemplary system 1400 for encoding and / or decoding video for machine learning, arranged according to at least some implementations of the present disclosure. As shown in FIG. 14, system 1400 may include a central processor 1401, a graphics processor 1402, and a memory 1403. Also, as shown, central processor 1401 may implement one or more of an encoder 1411 (e.g., all or part of patch-based encoder 100 and / or patch-based encoder 200), a decoder 1412 (e.g., all or part of patch-based encoder 110 and / or patch-based encoder 210), and a machine learning module 1413 (e.g., to implement one or both of machine learning system 120 and machine learning system 220). In an example of system 1400, memory 1403 may store bitstream data, region of interest data, downsampled frame data, metadata, or any other data described herein.

[0059] As shown, in some examples, one or more or part of encoder 1411, decoder 1412, and machine learning module 1413 are implemented via central processor 1401. In other examples, one or more or part of encoder 1411, decoder 1412, and machine learning module 1413 are implemented via graphics processor 1402, a video processing unit, a video processing pipeline, a video or image signal processor, etc. In some examples, one or more or part of encoder 1411, decoder 1412, and machine learning module 1413 are implemented in hardware as a system on chip (SoC). In some examples, one or more or part of encoder 1411, decoder 1412, and machine learning module 1413 are implemented in hardware via an FPGA.

[0060] The graphics processor 1402 can include any number and type of image or graphics processing devices that can provide operations as described herein. Such operations can be implemented via software or hardware or a combination thereof. For example, the graphics processor 1402 can include circuitry dedicated to the manipulation and / or analysis of video data obtained from the memory 1403. The central processor 1401 can provide control and other high-level functions for the system 1400 and / or can include any number and type of processing units or modules that can provide any operations as described herein. The memory 1403 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, the memory 1403 can be implemented by cache memory. In one embodiment, one or more or a portion of the encoder 1411, decoder 1412, and machine learning module 1413 are implemented via the execution units (EUs) of the graphics processor 1402. The EUs can include programmable logic or circuitry, such as one or more logic cores that can provide a wide range of programmable logic functions. In one embodiment, one or more or a portion of the encoder 1411, decoder 1412, and machine learning module 1413 are implemented via dedicated hardware, such as fixed function circuitry. The fixed function circuitry can include dedicated logic or circuitry and can provide a set of fixed function entry points that can be mapped to dedicated logic for a fixed purpose or function.

[0061] Returning to the description of FIG. 13, process 1300 starts at operation 1301, and any number of regions of interest for machine learning operations are detected in the full frame of the video. The regions of interest can be detected using any suitable one or more techniques related to machine learning operations executed on the decoder side. For example, the regions of interest can be detected using object detection techniques, face detection techniques, etc. In some embodiments, detecting or generating a first region of interest among the regions of interest includes performing look-ahead analysis to detect corresponding subsequent first regions of interest in a plurality of frames that temporally follow the full frame, and sizing the first region of interest to include the first region of interest and all subsequent first regions of interest. In some embodiments, detecting or generating a first region of interest among the regions of interest includes determining a detected region around an object in the first region of interest and expanding the detected region to the first region of interest to provide a buffer around the region.

[0062] The process continues to operation 1302, where one or more atlases including regions of interest at a first resolution are formed. For example, each of the regions of interest detected in operation 1301 may be extracted from the frame and inserted into one atlas or multiple atlases. In some embodiments, the full frame representation of the video in which the regions of interest are detected is not included in one or more atlases. In some embodiments, a low-resolution version of the full frame of the video is included in one or more atlases. In some embodiments, process 1300 includes downscaling the full frame of the video to a second resolution lower than the first resolution and including the downscaled frames of the video at the second resolution in one or more atlases for encoding into one or more bitstreams, as described with respect to operation 1304.

[0063] The process continues with operation 1303, where metadata corresponding to one or more atlases is generated, indicating the size and position of each region of interest within the full frame of the video. As described herein, the metadata provides a mapping between one or more atlases and the full frame of the video, such that the regions of interest can be sized appropriately and positioned within the full frame of the video. The metadata can include any suitable mapping data, such as the position of each region of interest within the atlas, the size of each region of interest within the atlas, the position of each region of interest within the full frame, and the size of each region of interest within the full frame, or coded data that can determine such values. In some embodiments, the metadata includes the top left position and scaling factor of a first region within the full frame for a first region of interest among the regions of interest.

[0064] The process continues with operation 1304, where one or more atlases and the metadata are encoded into one or more bitstreams that lack a representation of the full frame of the video at a first resolution or a resolution higher than the first resolution. For example, such coding at a high resolution without such extraction of regions of interest and coding of the full frame at a high resolution provides the regions of interest at a high resolution for eventual use in a machine learning operation while saving computational time and bitstream size. In some embodiments, a low resolution full frame is provided, and in process 1300, the encoding of a first region of interest among the plurality of regions of interest includes scalable video coding based on the first region of interest and the corresponding region of the full frame of the video at a resolution lower than the first resolution.

[0065] In some embodiments, a first region of interest among a plurality of regions of interest is within a first atlas, and process 1300 further includes detecting a second region of interest in a subsequent frame of the video and resizing the first atlas to add the second region of interest to the resized first atlas, such that encoding one or more atlases and metadata includes encoding the resized first atlas. In some embodiments, a first region of interest among the regions of interest includes a facial representation, and process 1300 further includes separating the first region of interest into the first atlas and encrypting a first bitstream corresponding to the first atlas.

[0066] In some embodiments, operations 1304-1305 are performed by an encoder or encoder system, and operations 1304-1305 are performed by a decoder or decoder system that is separate from the decoder or decoder system. For example, one or more bitstreams may be transmitted from an encoder or encoder system to a decoder or decoder system (via any number of intermediate devices).

[0067] Processing continues at operation 1305, where one or more bitstreams are received that include one or more atlases representing a plurality of regions of interest at a first resolution and metadata for locating the regions of interest within the full frame of the video, and the one or more bitstreams lack a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution. As described above, coding regions of interest at a high resolution without full-frame coding provides bit savings and efficiency while providing high-quality regions of interest for machine learning operations on the decoding side. The bitstream may include any data described herein. In some embodiments, the metadata includes the upper left position and a scaling factor of a first region within the full frame for a first region of interest among the plurality of regions of interest.

[0068] The process continues with operation 1306, where one or more bitstreams are decoded to generate a plurality of regions of interest. The one or more bitstreams can be decoded using any suitable one or more techniques. In some embodiments, the one or more bitstreams are compliant with a standard, and the decoding used also complies with a standard codec such as AVC, HEVC, VVC, etc. In some embodiments, process 1300 further includes decoding one or more bitstreams to generate a full frame at a second resolution lower than the first resolution. In some embodiments, decoding a first region of interest among the plurality of regions of interest includes scalable video decoding based on the first region of interest and the corresponding downscaled region of the full frame of the video.

[0069] The process continues with operation 1307, where machine learning is applied to one or more of the regions of interest, based in part on one or more corresponding positions of one or more of the regions of interest within the full frame, to generate a machine learning output. The machine learning operation can be performed using any suitable one or more techniques. In some embodiments, the region of interest includes a face detection region, and applying machine learning includes applying face recognition to one or more of the regions of interest. For privacy purposes, in some embodiments, a first region of interest among the regions of interest includes a facial representation, and process 1300 further includes decrypting at least a portion of one or more bitstreams including the representation of the first region of interest.

[0070] Process 1300 or a portion thereof may be repeated any number of times, either serially or in parallel, for any number of frames such as a video, time instance, etc. Process 1300 may be implemented by any suitable device, system, apparatus, or platform described herein. In one embodiment, process 1300 is implemented by a system or apparatus having a memory storing any data structure described herein and a processor executing any of operations 1301 - 1307.

[0071] Here, the techniques, encoders, and decoders being discussed, as well as systems and devices for implementing them, will be discussed. For example, any encoder (encoder system), decoder (decoder system), or bitstream extractor described herein can be implemented via the system shown in FIG. 15 and / or the device implemented in FIG. 16. In particular, the techniques, encoders, and decoders described can be implemented via any suitable device or platform described herein, such as a personal computer, laptop computer, tablet, phablet, smartphone, digital camera, game console, wearable device, display device, all-in-one device, two-in-one device, etc.

[0072] The various components of the systems described herein can be implemented in software, firmware, and / or hardware, and / or any combination thereof. For example, the various components of the devices or systems described herein can be provided, at least in part, by the hardware of a computing system on chip (SoC) such as found in a computing system such as a smartphone. One of ordinary skill in the art may recognize that the systems described herein may include additional components not shown in the corresponding figures. For example, the systems described herein may include additional components not shown for clarity.

[0073] The implementations of the exemplary processes described herein may include the execution of all the operations shown in the illustrated order, but the present disclosure is not limited thereto. In various examples, the implementations of the exemplary processes herein may include only a subset of the illustrated operations, operations executed in an order different from that illustrated, or additional operations.

[0074] Furthermore, any one or more of the operations described herein may be performed in response to instructions provided by one or more computer program products. Such program products may include, for example, a signal transmission medium that provides instructions which, when executed by a processor, can provide the functions described herein. The computer program products may be provided in any form of one or more machine-readable media. Thus, for example, a processor including one or more graphics processing units or processor cores may perform one or more of the exemplary process blocks herein in response to program code and / or instructions or instruction sets transmitted to the processor by one or more machine-readable media. In general, a machine-readable medium can transmit software in the form of program code and / or instructions or instruction sets that can cause any of the devices and / or systems described herein, at least a part of the device or system, or any other module or component described herein, to be implemented.

[0075] As used in any implementation described herein, the term "module" refers to any combination of software logic, firmware logic, hardware logic, and / or circuits configured to provide the functions described herein. The software may be embodied as a software package, code and / or instruction set or instructions, and the "hardware" used in any implementation described herein may include, for example, hard-wired circuits, programmable circuits, state machine circuits, fixed function circuits, execution unit circuits, and / or firmware that stores instructions executed by programmable circuits, alone or in any combination. The module may be embodied, collectively or individually, as a circuit that forms part of a larger system, such as an integrated circuit (IC), a system on chip (SoC), etc.

[0076] FIG. 15 is an explanatory diagram of an exemplary system 1500 arranged according to at least some implementations of the present disclosure. In various implementations, system 1500 may be a mobile device system, but system 1500 is not limited to this context. For example, system 1500 may be incorporated into a personal computer (PC), laptop computer, ultra-laptop computer, tablet, touch pad, portable computer, handheld computer, palmtop computer, personal digital assistant (PDA), mobile phone, combination of mobile phone / PDA, television, smart device (e.g., smartphone, smart tablet or smart TV), mobile Internet device (MID), messaging device, data communication device, camera (e.g., point-and-shoot camera, super-zoom camera, digital single-lens reflex (DSLR) camera), surveillance camera, surveillance system including a camera, etc.

[0077] In various implementations, system 1500 includes a platform 1502 coupled to a display 1520. Platform 1502 may receive content from a content device such as content service device 1530 or content delivery device 1540, or from another content source such as image sensor 1519. For example, platform 1502 may receive image data as described herein from image sensor 1519 or any other content source. A navigation controller 1550 including one or more navigation functions may be used, for example, to interact with platform 1502 and / or display 1520. Each of these components is described in more detail below.

[0078] In various implementation forms, the platform 1502 may include any combination of a chipset 1505, a processor 1510, a memory 1512, an antenna 1513, a storage 1514, a graphics subsystem 1515, an application 1516, an image signal processor 1517, and / or a radio 1518. The chipset 1505 may provide intercommunication between the processor 1510, the memory 1512, the storage 1514, the graphics subsystem 1515, the application 1516, the image signal processor 1517, and / or the radio 1518. For example, the chipset 1505 may include a storage adapter (not shown) that can provide intercommunication with the storage 1514.

[0079] The processor 1510 may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor, an x86 instruction set compatible processor, a multi-core, or any other microprocessor or central processing unit (CPU). In various implementation forms, the processor 1510 may be a dual-core processor, a dual-core mobile processor, etc.

[0080] The memory 1512 may be implemented as a volatile memory device such as, but not limited to, random access memory (RAM), dynamic random access memory (DRAM), or static RAM (SRAM).

[0081] The storage 1514 may be implemented as a non-volatile memory device such as, but not limited to, a magnetic disk drive, an optical disk drive, a tape drive, an internal storage device, an attached storage device, a flash memory, a battery-backed SDRAM (synchronous DRAM), and / or a network-accessible storage device. In various implementation forms, the storage 1514 may include technologies for improving the protection of storage performance for valuable digital media, for example, when multiple hard drives are included.

[0082] The image signal processor 1517 may be implemented as a dedicated digital signal processor or the like used for image processing. In some examples, the image signal processor 1517 may be implemented based on a single instruction multiple data or multiple instruction multiple data architecture or the like. In some examples, the image signal processor 1517 may be characterized as a media processor. As described in this specification, the image signal processor 1517 may be implemented based on a system-on-chip architecture and / or based on a multi-core architecture.

[0083] The graphics subsystem 1515 may perform processing of images such as still images or videos for display. The graphics subsystem 1515 may be, for example, a graphics processing unit (GPU) or a vision processing unit (VPU). The graphics subsystem 1515 and the display 1520 may be communicably coupled using an analog or digital interface. For example, the interface may be any of a high-definition multimedia interface, DisplayPort, wireless HDMI (registered trademark), and / or a technology compliant with wireless HD. The graphics subsystem 1515 may be integrated with the processor 1510 or the chipset 1505. In some implementations, the graphics subsystem 1515 may be a stand-alone device communicably coupled to the chipset 1505.

[0084] The graphics and / or video processing techniques described in this specification may be implemented in various hardware architectures. For example, the graphics and / or video functions may be integrated within a chipset. Alternatively, separate graphics and / or video processors may be used. As yet another implementation, the graphics and / or video functions may be provided by a general-purpose processor including a multi-core processor. In a further embodiment, the functions may be implemented in a household appliance.

[0085] Wireless device 1518 may include one or more wireless devices that can transmit and receive signals using various suitable wireless communication technologies. Such technologies may include communication via one or more wireless networks. Exemplary wireless networks include, but are not limited to, wireless local area network (WLAN), wireless personal area network (WPAN), wireless metropolitan area network (WMAN), cellular network, and satellite network. When communicating via such a network, wireless device 1518 may operate in accordance with one or more applicable standards in any version.

[0086] In various implementations, display 1520 may include any television-type monitor or display. Display 1520 may include, for example, a computer display screen, a touch screen display, a video monitor, a television-like device, and / or a television. Display 1520 may be digital and / or analog. In various implementations, display 1520 may be a holographic display. Also, display 1520 may be a transparent surface that can receive visual projections. Such projections may convey various forms of information, images, and / or objects. For example, such a projection may be a visual overlay for a mobile augmented reality (MAR) application. Platform 1502 may display user interface 1522 on display 1520 under the control of one or more software applications 1516.

[0087] In various implementations, the content service device 1530 is hosted by any domestic, international, and / or independent service and can thus be accessible to the platform 1502 via, for example, the Internet. The content service device 1530 may be coupled to the platform 1502 and / or the display 1520. The platform 1502 and / or the content service device 1530 may be coupled to the network 1560 to communicate (e.g., transmit and / or receive) media information with the network 1560. The content delivery device 1540 may also be coupled to the platform 1502 and / or the display 1520.

[0088] The image sensor 1519 may include any suitable image sensor capable of providing image data based on a scene. For example, the image sensor 1519 may include a semiconductor charge-coupled device (CCD)-based sensor, a complementary metal-oxide semiconductor (CMOS)-based sensor, an N-type metal-oxide semiconductor (NMOS)-based sensor, and the like. For example, the image sensor 1519 may include any device capable of detecting information about a scene and generating image data.

[0089] In various implementations, the content service device 1530 may include a cable television box, a personal computer, a network, a telephone, an Internet-enabled device or appliance capable of delivering digital information and / or content, and any other similar device capable of communicating content unidirectionally or bidirectionally between a content provider and the platform 1502 and / or the display 1520, either via the network 1560 or directly. It will be understood that content may be communicated unidirectionally and / or bidirectionally between any of the components within the system 1500 and a content provider via the network 1560. Examples of content may include any media information, such as, for example, video, music, medical, and game information.

[0090] The content service device 1530 may receive content such as cable TV programs including media information, digital information, and / or other content. Examples of content providers may include any cable or satellite TV or radio or Internet content provider. The examples provided are not meant to limit the implementation forms according to the present disclosure in any way.

[0091] In various implementation forms, the platform 1502 may receive a control signal from a navigation controller 1550 having one or more navigation functions. The navigation functions of the navigation controller 1550 can be used, for example, to interact with the user interface 1522. In various embodiments, the navigation controller 1550 can be a pointing device that can be a computer hardware component (specifically, a human interface device) that enables a user to input spatial (e.g., continuous and multi-dimensional) data into a computer. Many systems such as graphical user interfaces (GUIs), as well as TVs and monitors, enable a user to control a computer or TV and provide data using physical gestures.

[0092] The movement of the navigation function of the navigation controller 1550 may be replicated on a display (e.g., display 1520) by the movement of a pointer, cursor, focus ring, or other visual indicator displayed on the display. For example, under the control of the software application 1516, the navigation function disposed on the navigation controller 1550 may be mapped to a virtual navigation function displayed on, for example, the user interface 1522. In various embodiments, the navigation controller 1550 may not be a separate component and may be integrated into the platform 1502 and / or the display 1520. However, the present disclosure is not limited to the elements or contexts illustrated or described herein.

[0093] In various implementations, a driver (not shown) may include, for example, technology that enables a user to touch a button after initial startup to immediately turn the platform 1502 on and off like a television when activated. The program logic may enable the platform 1502 to stream content to the media adapter or other content service device 1530 or content delivery device 1540 even when the platform is "off". Further, the chipset 1505 may include, for example, hardware and / or software support for 5.1 surround sound and / or high-definition 7.1 surround sound. The driver may include a graphics driver for an integrated graphics platform. In various embodiments, the graphics driver may include a Peripheral Component Interconnect (PCI) Express graphics card.

[0094] In various implementations, any one or more of the components shown in system 1500 may be integrated. For example, platform 1502 and content service device 1530 may be integrated, or platform 1502 and content delivery device 1540 may be integrated, or platform 1502, content service device 1530, and content delivery device 1540 may be integrated. In various embodiments, platform 1502 and display 1520 may be an integrated unit. For example, display 1520 and content service device 1530 may be integrated, or display 1520 and content delivery device 1540 may be integrated. These examples are not intended to limit the present disclosure.

[0095] In various embodiments, system 1500 may be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, system 1500 may include components and interfaces suitable for communicating via a wireless shared medium such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, etc. Examples of wireless shared media may include a portion of the wireless spectrum such as the RF spectrum. When implemented as a wired system, system 1500 may include components and interfaces suitable for communicating via a wired communication medium such as an input / output (I / O) adapter, a physical connector for connecting the I / O adapter to a corresponding wired communication medium, a network interface card (NIC), a disk controller, a video controller, an audio controller, etc. Examples of wired communication media may include wires, cables, metal leads, printed circuit boards (PCBs), backplanes, switch fabrics, semiconductor materials, twisted pair wires, coaxial cables, optical fibers, etc.

[0096] Platform 1502 can establish one or more logical or physical channels for communicating information. The information can include media information and control information. Media information can refer to any data representing content for users. Examples of content can include, for example, data from voice conversations, video conferences, streaming videos, email messages, voicemail messages, alphanumeric symbols, graphics, images, videos, text, etc. Data from a voice conversation can be, for example, speech information, silence periods, background noise, comfort noise, tones, etc. Control information can refer to any data representing commands, instructions, or control words intended for an automation system. For example, control information may be used to route media information through the system or to instruct nodes to process media information in a predetermined manner. However, embodiments are not limited to the elements or contexts illustrated or described in FIG. 15.

[0097] As described above, system 1500 can be embodied in various physical styles or form factors. FIG. 16 shows an exemplary small form factor device 1600 arranged according to at least some implementations of the present disclosure. In some examples, system 1500 may be implemented via device 1600. In other examples, other systems, components, or modules described herein, or portions thereof, may be implemented via device 1600. In various embodiments, for example, device 1600 may be implemented as a mobile computing device having wireless capabilities. A mobile computing device can refer to any device having a processing system and a mobile power source, such as, for example, one or more batteries.

[0098] Examples of mobile computing devices can include personal computers (PCs), laptop computers, ultra-laptop computers, tablets, touch pads, portable computers, handheld computers, palmtop computers, personal digital assistants (PDAs), mobile phones, combinations of mobile phones / PDAs, smart devices (e.g., smartphones, smart tablets or smart mobile TVs), mobile Internet devices (MIDs), messaging devices, data communication devices, cameras (e.g., point-and-shoot cameras, superzoom cameras, digital single-lens reflex (DSLR) cameras), and the like.

[0099] Examples of mobile computing devices can also include computers arranged to be implemented by an automobile or a robot, or arranged to be worn by a person, such as wrist computers, finger computers, ring computers, glasses computers, belt clip computers, armband computers, shoe computers, clothing computers, and other wearable computers. In various embodiments, for example, the mobile computing device may be implemented as a smartphone that can execute computer applications as well as voice communication and / or data communication. Some embodiments may be described by way of example using a mobile computing device implemented as a smartphone, but it will be understood that other embodiments may also be implemented using other wireless mobile computing devices. The embodiments are not limited in this context.

[0100] As shown in FIG. 16, the device 1600 may include a housing having a front surface 1601 and a back surface 1602. The device 1600 includes a display 1604, an input / output (I / O) device 1606, a color camera 1621, a color camera 1622, and an integrated antenna 1608. In some embodiments, the color camera 1621 and the color camera 1622 achieve a planar image as described herein. In some embodiments, the device 1600 does not include the color cameras 1621 and 1622, and the device 1600 obtains input image data (e.g., any input image data described herein) from another device. The device 1600 may also include a navigation function 1612. The I / O device 1606 may include any suitable I / O device for inputting information into the mobile computing device. Examples of the I / O device 1606 may include an alphanumeric keyboard, a numeric keypad, a touchpad, input keys, buttons, switches, a microphone, a speaker, a voice recognition device, and software, etc. Information may also be input into the device 1600 via a microphone (not shown), or may be digitized by a voice recognition device. As shown in the figure, the device 1600 may include color cameras 1621, 1622 and a flash 1610 integrated on the back surface 1602 (or other locations) of the device 1600. In other examples, the color cameras 1621, 1622, and the flash 1610 may be integrated on the front surface 1601 of the device 1600, or both a front and a back set of cameras may be provided. The color cameras 1621, 1622 and the flash 1610 may be components of a camera module for generating color image data with IR texture correction, and this color image data is output to the display 1604 and / or processed into an image or streaming video that is remotely communicated from the device 1600 via, for example, the antenna 1608.

[0101] Various embodiments may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements can include processors, microprocessors, circuits, circuit elements (such as transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chip sets, and the like. Examples of software can include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. The determination of whether an embodiment is implemented using hardware elements and / or software elements can vary according to any number of factors such as desired computational speed, power level, heat tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.

[0102] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium representing various logic within a processor, which, when read by a machine, causes the machine to create the logic for performing the techniques described herein. Such representations, known as IP cores, are stored on a tangible machine-readable medium and can be supplied to various customers or manufacturing facilities for loading onto a manufacturing machine that actually fabricates the logic or processor.

[0103] The specific features described in this specification have been described with reference to various implementations, but this description is not intended to be construed in a limiting sense. Accordingly, various modifications of the implementations described herein, as well as other implementations that will be apparent to those skilled in the art to which this disclosure pertains, are considered to be within the spirit and scope of this disclosure.

[0104] The following relates to further embodiments.

[0105] In one or more first embodiments, a method for coding a video for machine learning includes detecting a plurality of regions of interest for machine learning operations within a full frame of the video, forming one or more atlases including the regions of interest at a first resolution, generating metadata corresponding to the one or more atlases and indicating the size and position of each of the regions of interest within the full frame of the video, and encoding the one or more atlases and the metadata into one or more bitstreams, wherein the one or more bitstreams lack a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution.

[0106] In one or more second embodiments, in addition to the first embodiment, the method further includes downscaling the full frame of the video to a second resolution lower than the first resolution and including the downscaled frame of the video at the second resolution in the one or more atlases for encoding into the one or more bitstreams.

[0107] In one or more third embodiments, in addition to the first or second embodiment, the encoding of a first region of interest among the plurality of regions of interest includes scalable video coding based on the first region of interest and the corresponding region of the full frame of the video at a resolution lower than the first resolution.

[0108] In one or more fourth embodiments, in addition to any of the first to third embodiments, the metadata includes, for a first region of interest among the plurality of regions of interest, the upper left position and the scaling factor of a first region within the full frame.

[0109] In one or more fifth embodiments, in addition to any of the first to fourth embodiments, the first region of interest among the plurality of regions of interest is within a first atlas, and the method further includes detecting a second region of interest within a subsequent frame of the video and resizing the first atlas to add the second region of interest to the resized first atlas, and the step of encoding one or more atlases and metadata includes encoding the resized first atlas.

[0110] In one or more sixth embodiments, in addition to any of the first to fifth embodiments, the step of detecting a first region of interest among the plurality of regions of interest includes performing a look-ahead analysis to detect corresponding subsequent first regions of interest within a plurality of temporally subsequent frames with respect to the full frame and sizing the first region of interest to include the first region of interest and all subsequent first regions of interest.

[0111] In one or more seventh embodiments, in addition to any of the first to sixth embodiments, the step of detecting a first region of interest among the plurality of regions of interest includes determining a detected region around an object within the first region of interest and expanding the detected region to the first region of interest to provide a buffer around the region.

[0112] In one or more eighth embodiments, in addition to any of the first to seventh embodiments, the first region of interest among the plurality of regions of interest includes a facial representation, and the method further includes separating the first region of interest into a first atlas and encrypting a first bitstream corresponding to the first atlas.

[0113] In one or more ninth embodiments, a method for performing machine learning on a received video includes receiving one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for locating regions of interest within a full frame of the video, the one or more bitstreams lacking a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution; decoding the one or more bitstreams to generate the plurality of regions of interest; and applying machine learning to one or more of the regions of interest based at least in part on corresponding one or more positions of the one or more regions of interest within the full frame to generate a machine learning output.

[0114] In one or more tenth embodiments, in addition to any of the first through ninth embodiments, the method further includes decoding the one or more bitstreams to generate a full frame at a second resolution lower than the first resolution, wherein decoding of a first region of interest among the plurality of regions of interest includes scalable video decoding based on the first region of interest and a corresponding downscaled region of the full frame of the video.

[0115] In one or more eleventh embodiments, in addition to the ninth or tenth embodiment, the metadata includes an upper left position of a first region within the full frame and a scaling factor for a first region of interest among the plurality of regions of interest.

[0116] In one or more twelfth embodiments, in addition to any of the ninth through eleventh embodiments, the first region of interest among the plurality of regions of interest includes a representation of a face, and the method further includes decrypting at least a portion of one or more bitstreams including the representation of the first region of interest.

[0117] In one or more 13th embodiments, in addition to any of the 9th to 12th embodiments, the region of interest includes a face detection region, and the step of applying machine learning includes the step of applying face recognition to one or more of the regions of interest.

[0118] In one or more 14th embodiments, the device or system includes a memory and one or more processors for executing the method according to any one of the above embodiments.

[0119] In one or more 15th embodiments, at least one machine-readable medium includes a plurality of instructions that, in response to being executed on a computing device, cause the computing device to execute the method according to any one of the above embodiments.

[0120] In one or more 16th embodiments, the apparatus includes means for executing the method according to any one of the above embodiments.

[0121] It should be recognized that the embodiments are not limited to the described embodiments and can be implemented with modifications and changes without departing from the scope of the appended claims. For example, the above embodiments may include a specific combination of features. However, the above embodiments are not limited thereto, and in various implementations, the above embodiments may include ensuring only a subset of such features, ensuring a different order of such features, ensuring different combinations of such features, and / or ensuring additional features other than those explicitly listed features. Therefore, the scope of the embodiments should be determined with reference to the appended claims, together with the full scope of equivalents to which such claims are entitled. (Other possible items) (Item 1) a memory for storing at least a part of the input video, a processor circuit coupled to the memory, A system comprising: wherein the processor circuit is Detect a plurality of regions of interest for machine learning operations in the full frame of the input video, Form one or more atlases including the regions of interest at a first resolution, Generate metadata corresponding to the one or more atlases and indicating the size and position of each of the regions of interest within the full frame of the video, Encode the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams lacking a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution, System. (Item 2) The processor circuit, Downscale the full frame of the video to a second resolution lower than the first resolution, Include the downscaled frame of the video at the second resolution in the one or more atlases for encoding into the one or more bitstreams, The system according to item 1. (Item 3) The processor circuit for encoding a first region of interest among the plurality of regions of interest includes the processor circuit that performs scalable video encoding based on the first region of interest and a corresponding region of the full frame of the video at a resolution lower than the first resolution, the system according to item 1. (Item 4) The metadata includes, for a first region of interest among the plurality of regions of interest, the upper left position and a scaling factor of the first region within the full frame, the system according to item 1. (Item 5) A first region of interest among the plurality of regions of interest is within a first atlas, and the processor circuit, Detect a second region of interest in a subsequent frame of the video, Resizing the first atlas, adding the second region of interest to the resized first atlas, and the processor circuit that encodes the one or more atlases and the metadata includes the processor circuit that encodes the resized first atlas. The system according to any one of items 1 to 4. (Item 6) The processor circuit that detects a first region of interest among the plurality of regions of interest Performs a look-ahead analysis to detect corresponding subsequent first regions of interest in a plurality of temporally subsequent frames with respect to the full frame. The processor circuit that sizes the first region of interest to include the first region of interest and all subsequent first regions of interest The system according to any one of items 1 to 4, including. (Item 7) The processor circuit that detects a first region of interest among the plurality of regions of interest Determines a detected region around an object within the first region of interest. The processor circuit that expands the detected region to the first region of interest to provide a buffer around the region. The system according to any one of items 1 to 4, including. (Item 8) The first region of interest among the plurality of regions of interest includes a facial expression, and the processor circuit separates the first region of interest into a first atlas and encrypts a first bitstream corresponding to the first atlas. The system according to any one of items 1 to 4. (Item 9) At least one machine-readable medium that, in response to being executed on a computing device, causes the computing device to Detect a plurality of regions of interest for machine learning operations in a full frame of video. Form one or more atlases including the regions of interest at a first resolution. Generate metadata corresponding to the one or more atlases and indicating the size and position of each of the regions of interest within the full frame of the video. Encode the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams lacking a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution. At least one machine-readable medium comprising a plurality of instructions for coding a video for machine learning. (Item 10) In response to being executed on the computing device, cause the computing device to Downscale the full frame of the video to a second resolution lower than the first resolution. Include the downscaled frame of the video at the second resolution in the one or more atlases for encoding into the one or more bitstreams, and encoding of a first region of interest among the plurality of regions of interest includes scalable video encoding based on the first region of interest and a corresponding region of the full frame of the video at a resolution lower than the first resolution. The machine-readable medium of item 9, further comprising instructions for coding a video for machine learning. (Item 11) When a first region of interest among the plurality of regions of interest is within a first atlas, the machine-readable medium, in response to being executed on the computing device, causes the computing device to Detect a second region of interest in a subsequent frame of the video. Resize the first atlas to add the second region of interest to the resized first atlas, and encoding of the one or more atlases and the metadata includes encoding of the resized first atlas. The machine-readable medium of item 9 or 10, further comprising instructions for coding a video for machine learning. (Item 12) Detecting a first region of interest among the plurality of regions of interest includes performing a look-ahead analysis to detect corresponding subsequent first regions of interest in a plurality of subsequent frames temporally following the full frame; and sizing the first region of interest to include the first region of interest and all subsequent first regions of interest The machine-readable medium according to item 9 or 10, comprising: (Item 13) means for detecting a plurality of regions of interest for machine learning operations within a full frame of a video; means for forming one or more atlases including the regions of interest at a first resolution; means for generating metadata corresponding to the one or more atlases and indicating the size and position of each of the regions of interest within the full frame of the video; means for encoding the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams lacking a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution A system comprising: (Item 14) means for downscaling the full frame of the video to a second resolution lower than the first resolution; means for including the downscaled frame of the video at the second resolution in the one or more atlases for encoding into the one or more bitstreams, the means for encoding the first region of interest among the plurality of regions of interest including means for performing scalable video encoding based on the first region of interest and a corresponding region of the full frame of the video at a resolution lower than the first resolution The system according to item 13, further comprising: (Item 15) A first region of interest among the plurality of regions of interest is within a first atlas, and the system means for detecting a second region of interest within subsequent frames of the video; means for resizing the first atlas and adding the second region of interest to the resized first atlas, wherein the means for encoding the one or more atlases and the metadata includes means for encoding the resized first atlas; The system according to item 13 or 14, further comprising (Item 16) A memory storing at least a portion of one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for locating the regions of interest within the full frame of the video, wherein the one or more bitstreams lack a representation of the full frame of the video at the first resolution or a resolution higher than the first resolution; A processor circuit coupled to the memory, decoding the one or more bitstreams to generate the plurality of regions of interest; applying machine learning to one or more of the regions of interest, at least in part based on the one or more corresponding locations of the regions of interest within the full frame, to generate a machine learning output; A processor circuit Comprising a system. (Item 17) Wherein the processor circuit decodes the one or more bitstreams to generate the full frame at a second resolution lower than the first resolution, and the processor circuit that decodes a first region of interest among the plurality of regions of interest performs scalable video decoding based on the first region of interest and the corresponding downscaled region of the full frame of the video; The system according to item 16. (Item 18) The system according to item 16, wherein the metadata includes the upper left position and the scaling factor of the first region in the full frame for the first region of interest among the plurality of regions of interest. (Item 19) The system according to any one of items 16 to 18, wherein the first region of interest among the plurality of regions of interest includes a facial expression, and the processor circuit decodes at least a part of the one or more bitstreams including the expression of the first region of interest. (Item 20) The system according to any one of items 16 to 18, wherein the region of interest includes a face detection region, and the processor circuit that applies the machine learning includes a processor circuit that applies face recognition to the one or more of the regions of interest. (Item 21) At least one machine-readable medium including a plurality of instructions, in response to being executed on a computing device, causing the computing device to receive one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for locating the regions of interest within the full frame of the video, the one or more bitstreams lacking a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution, decode the one or more bitstreams to generate the plurality of regions of interest, apply machine learning to one or more of the regions of interest based at least in part on the corresponding one or more positions of the one or more of the regions of interest within the full frame to generate a machine learning output Thereby, at least one machine-readable medium including a plurality of instructions for executing machine learning. (Item 22) In response to being executed on the computing device, causing the computing device to Decoding the one or more bitstreams to generate the full frame at a second resolution lower than the first resolution, wherein decoding of a first region of interest among the plurality of regions of interest includes scalable video decoding based on the first region of interest and corresponding downscaled regions of the full frame of the video The machine-readable medium of item 21, further comprising instructions for causing machine learning to be performed thereby (Item 23) The machine-readable medium according to item 21 or 22, wherein a first region of interest among the plurality of regions of interest includes a facial expression, and the machine-readable medium further includes instructions for causing the computing device to perform machine learning by decoding at least a portion of the one or more bitstreams including the expression of the first region of interest in response to being executed on the computing device (Item 24) Means for receiving one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for identifying the locations of the regions of interest within the full frame of the video, wherein the one or more bitstreams lack a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution Means for decoding the one or more bitstreams to generate the plurality of regions of interest Means for applying machine learning to one or more of the regions of interest based at least in part on the one or more corresponding locations of the regions of interest within the full frame to generate a machine learning output A system comprising (Item 25) Means for decoding the one or more bitstreams to generate the full frame at a second resolution lower than the first resolution, wherein decoding of a first region of interest among the plurality of regions of interest includes scalable video decoding based on the first region of interest and corresponding downscaled regions of the full frame of the video The system according to item 24, further comprising

Claims

1. A memory for storing at least a part of an input video, and A processor circuit coupled to the memory, A system comprising: The processor circuit Detects a plurality of regions of interest for machine learning processing in a full frame of the input video, Forms one or more atlases including the plurality of regions of interest at a first resolution, Generates metadata corresponding to the one or more atlases and indicating the size and position of each of the plurality of regions of interest within the full frame of the input video, Encodes the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams not including a representation of the full frame of the input video at the first resolution or at a resolution higher than the first resolution, System.

2. The processor circuit Downscales the full frame of the input video to a second resolution lower than the first resolution, Includes the downscaled full frame of the input video at the second resolution in the one or more atlases for encoding into the one or more bitstreams, The system according to claim 1.

3. The processor circuit encoding a first region of interest among the plurality of regions of interest includes the processor circuit performing scalable video encoding based on the first region of interest and a corresponding region of the full frame of the input video at a resolution lower than the first resolution. The system according to claim 1 or 2.

4. The system according to any one of claims 1 to 3, wherein the metadata includes an upper left position and a scaling factor of the first region of interest within the full frame for the first region of interest among the plurality of regions of interest.

5. A first region of interest among the plurality of regions of interest is within a first atlas, and the processor circuit Detects a second region of interest in a subsequent frame of the input video, Executes resizing the first atlas and adding the second region of interest to the resized first atlas, The processor circuit encoding the one or more atlases and the metadata includes the processor circuit encoding the resized first atlas. The system according to any one of claims 1 to 4. **Claim 6** The processor circuit detecting a first region of interest among the plurality of regions of interest means that the processor circuit performs a look-ahead analysis to detect corresponding subsequent first regions of interest in a plurality of frames temporally subsequent to the full frame, sizing the first region of interest to include the first region of interest and all subsequent first regions of interest The system according to any one of claims 1 to 5, including. **Claim 7** The processor circuit detecting a first region of interest among the plurality of regions of interest means that the processor circuit determines a detected region around an object within the first region of interest, extending the detected region to the first region of interest to provide a buffer around the detected region The system according to any one of claims 1 to 6, including. **Claim 8** The first region of interest among the plurality of regions of interest includes a facial representation, and the processor circuit separates the first region of interest into a first atlas and encrypts a first bitstream corresponding to the first atlas. The system according to any one of claims 1 to 7. **Claim 9** A computer program for coding a video for machine learning, causing a computing device to detect a plurality of regions of interest for machine learning processing in a full frame of a video; form one or more atlases including the plurality of regions of interest at a first resolution; generate metadata corresponding to the one or more atlases and indicating the size and position of each of the plurality of regions of interest within the full frame of the video; encoding the one or more atlases and the metadata into one or more bitstreams, the one or more bitstreams not including a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution; procedure A computer program that causes to execute. **Claim 10** Further to the computing device A procedure for downscaling the full frame of the video to a second resolution lower than the first resolution; A procedure for including the downscaled full frame of the video at the second resolution in the one or more atlases in order to encode the one or more bitstreams, wherein the procedure for encoding a first region of interest among the plurality of regions of interest includes a procedure for performing scalable video encoding based on the first region of interest and a corresponding region of the full frame of the video at a resolution lower than the first resolution. The computer program according to claim 9.

11. A first region of interest among the plurality of regions of interest is within a first atlas, and in the computing device A procedure for detecting a second region of interest in a subsequent frame of the video; A procedure for resizing the first atlas and adding the second region of interest to the resized first atlas is further executed, The procedure for encoding the one or more atlases and the metadata includes a procedure for encoding the resized first atlas. The computer program according to claim 9 or 10.

12. The procedure for detecting a first region of interest among the plurality of regions of interest includes A procedure for performing a look-ahead analysis to detect a corresponding subsequent first region of interest in a plurality of frames temporally subsequent to the full frame; A procedure for sizing the first region of interest to include the first region of interest and all subsequent first regions of interest The computer program according to any one of claims 9 to 11.

13. Means for detecting a plurality of regions of interest for machine learning processing within the full frame of the video; Means for forming one or more atlases including the plurality of regions of interest at a first resolution; Means for generating metadata corresponding to the one or more atlases and indicating the size and position of each of the plurality of regions of interest within the full frame of the video; Means for encoding the one or more atlases and the metadata into one or more bitstreams, wherein the one or more bitstreams do not include a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution A system comprising the same

14. Means for downscaling the full frame of the video to a second resolution lower than the first resolution Means for including the downscaled full frame of the video at the second resolution in the one or more atlases for encoding into the one or more bitstreams The means for encoding a first region of interest among the plurality of regions of interest includes means for performing scalable video encoding based on the first region of interest and a corresponding region of the full frame of the video at a resolution lower than the first resolution The system according to claim 13

15. The first region of interest among the plurality of regions of interest is within a first atlas, and the system further includes Means for detecting a second region of interest in subsequent frames of the video Means for resizing the first atlas and adding the second region of interest to the resized first atlas The means for encoding the one or more atlases and the metadata includes means for encoding the resized first atlas, according to the system of claim 13 or 14

16. A memory storing at least a portion of one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for locating the plurality of regions of interest within a full frame of a video, wherein the one or more bitstreams do not include a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution A processor circuit coupled to the memory, the processor circuit Decodes the one or more bitstreams to generate the plurality of regions of interest A system that applies machine learning to one or more of the plurality of regions of interest, based in part on corresponding one or more positions within one or more of the full frames of the plurality of regions of interest.

17. The processor circuit decodes the one or more bitstreams to generate the full frame at a second resolution lower than the first resolution, and the processor circuit decoding a first region of interest among the plurality of regions of interest includes the processor circuit performing scalable video decoding based on the first region of interest and a corresponding downscaled region of the full frame of the video. The system according to claim 16.

18. The system according to claim 16 or 17, wherein the metadata includes, for a first region of interest among the plurality of regions of interest, a top-left position and a scaling factor of the first region of interest within the full frame.

19. The system according to any one of claims 16 to 18, wherein a first region of interest among the plurality of regions of interest includes a facial representation, and the processor circuit decodes at least a part of the one or more bitstreams including the representation of the first region of interest.

20. The system according to any one of claims 16 to 19, wherein the plurality of regions of interest includes a face detection region, and the processor circuit applying the machine learning includes the processor circuit applying face recognition to the one or more of the plurality of regions of interest.

21. A computing device receiving one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for locating the plurality of regions of interest within a full frame of a video, wherein the one or more bitstreams do not include a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution, decoding the one or more bitstreams to generate the plurality of regions of interest A computer program for executing machine learning by causing a procedure for applying machine learning to one or more of the plurality of regions of interest to be executed based at least in part on corresponding one or more positions within one or more of the full frames of the plurality of regions of interest.

22. The computing device further executes machine learning by causing a procedure for decoding the one or more bitstreams to generate the full frame at a second resolution lower than the first resolution. The procedure for decoding a first region of interest among the plurality of regions of interest includes a procedure for performing scalable video decoding based on the first region of interest and a corresponding downscaled region of the full frame of the video. The computer program according to claim 21.

23. When the first region of interest among the plurality of regions of interest includes a facial expression, the computing device further executes machine learning by causing a procedure for decrypting at least a part of the one or more bitstreams including the expression of the first region of interest. The computer program according to claim 21 or 22.

24. Means for receiving one or more bitstreams including one or more atlases representing a plurality of regions of interest at a first resolution and metadata for identifying the positions of the plurality of regions of interest within the full frame of the video, wherein the one or more bitstreams do not include a representation of the full frame of the video at the first resolution or at a resolution higher than the first resolution. Means for decoding the one or more bitstreams to generate the plurality of regions of interest. Means for applying machine learning to one or more of the plurality of regions of interest based at least in part on corresponding one or more positions within one or more of the full frames of the plurality of regions of interest. A system comprising.

25. means for decoding the one or more bitstreams to generate the full frame at a second resolution lower than the first resolution, wherein decoding a first region of interest of the plurality of regions of interest includes performing scalable video decoding based on the first region of interest and a corresponding downscaled region of the full frame of the video, the system of claim 24. **Claim 26** At least one machine-readable recording medium storing a computer program according to any one of claims 9 to 12 or claims 21 to 23.

Citation Information

Patent Citations

  • Tracking video reproducing apparatus

    JP2006033793A

  • Moving image encoder, moving image encoding method, moving image encoding program, moving image decoder, moving image decoding method, moving image decoding program and moving image processor

    JP2014060512A

  • Data pruning for video compression using example-based super-resolution

    US20120288015A1

  • Enhanced siamese trackers

    US20180129934A1

  • Progressive compressed domain computer vision and deep learning systems

    US20190246130A1