Dash signaling of display attenuation maps in adaptive streaming services
Display attenuation maps in adaptive streaming protocols dynamically adjust display settings to reduce energy consumption in multimedia devices, addressing the inefficiencies in existing technologies and optimizing energy usage.
Patent Information
- Application Number
- PCT/EP2025/050862
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2025-01-15
- Publication Date
- 2025-07-24
AI Technical Summary
Existing multimedia streaming technologies do not effectively optimize display energy consumption in devices like smartphones and televisions, despite the increasing energy demands due to higher display resolutions and dynamic range imaging, leading to significant power usage.
Implementing display attenuation maps within adaptive streaming protocols, such as DASH, to dynamically adjust backlight or pixel illumination levels based on image content, reducing energy consumption without impairing image quality.
This approach allows for efficient energy management in display devices by optimizing the trade-off between quality of experience and energy reduction, using pixel-wise attenuation maps and flexible signaling mechanisms.
Smart Images

Figure EP2025050862_24072025_PF_FP_ABST
Abstract
Description
DASH SIGNALING OF DISPLAY ATTENUATION MAPS IN ADAPTIVE STREAMING SERVICESCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of European Provisional Patent Application No. 24305102.6 filed January 16, 2024, the contents of which are hereby incorporated by reference hereinTECHNICAL FIELD
[0002] The present embodiments generally relate to systems and methods for multimedia processing and distribution, and in particular systems and methods for streaming multimedia.BACKGROUND
[0003] Reducing energy consumption of electronic devices has become a requirement not only for electronic devices manufacturers but also to limit, as much as possible, the environmental impact and to contribute to the emergence of a sustainable display industry. The increase in display resolution from SD to HD to 4K and soon to 8K and beyond, as well as the introduction of high dynamic range imaging, has brought about a corresponding increase in energy requirements of display devices. This is not consistent with the global need to reduce energy consumption. Indeed, displays are the most important source of energy consumption, whether it be for battery-powered devices (e.g., smartphones) or in the global video distribution chain.SUMMARY
[0004] In some aspects, the present disclosure is directed to implementations of systems and methods for flexible signaling of display attenuation maps and relative attenuation map information in an adaptive streaming protocol manifest or media presentation descriptor (MPD) file, including association of the attenuation maps with related video data in the MPD. Such implementations may enable a streaming client to identify a set of available display attenuation maps for the video being streamed and to efficientlyselect the most appropriate attenuation map(s) that optimize the display energy consumption of playback devices in streaming services.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, wherein like reference numerals in the figures indicate like elements, and wherein:
[0006] FIG. 1 is a block diagram of a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (DASH) delivery process model, according to some implementations;
[0007] FIG. 2 is an illustration of a DASH media presentation description (MPD) manifest, according to some implementations;
[0008] FIGs. 3A-3C are a representation of an example XML schema for an attenuation map descriptor, according to some implementations;
[0009] FIGs. 4A-4B are a representation of an example MPD manifest, according to some implementations;
[0010] FIGs. 5A-5B are a representation of another example MPD manifest, according to some implementations;
[0011] FIGs. 6A-6B are a representation of yet another example MPD manifest, according to some implementations;
[0012] FIGs. 7A-7C are a representation of yet another example MPD manifest, according to some implementations;
[0013] FIG. 8 is a flow chart of a method for processing streaming video with an associated display attenuation map, according to some implementations;
[0014] FIG. 9 illustrates a block diagram of an embodiment of video encoder in which various aspects of the embodiments may be implemented;
[0015] FIG. 10 illustrates a block diagram of an embodiment of video encoder in which various aspects of the embodiments may be implemented; and
[0016] FIG. 11 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented.DETAILED DESCRIPTION
[0017] The present aspects, although describing principles related to particular drafts of VVC (Versatile Video Coding) or to HEVC (High Efficiency Video Coding) specifications, are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0018] The present disclosure is directed to signaling mechanisms for video and multimedia streaming to enable a device to control the energy usage of its display and the rendered video quality in the context of an adaptative streaming protocol (e.g., MPEG-DASH, specified in ISO / IEC 23009-1 ; HTTP Live Streaming (HLS); or similar protocols) on the screen of a desktop or laptop computer, a smartphone, a tablet, a set-top box, or a television connected to the internet. Such devices may consume significant amounts of energy, but by providing brightness or dimming metadata or light attenuation maps as part of a streamed video or multi-media signal, the decoder of the display may adjust one or more backlights or pixel illumination levels to reduce energy consumption without significant impairment of image quality.
[0019] For example, Organic Light Emitting Diode (OLED) displays are getting more and more popular because of numerous advantages compared to non-emissive displays such as Thin-Film Transistor Liquid Crystal Displays (TFT-LCDs). Rather than using a uniform backlight, OLED displays are composed of LEDs as image pixels. OLEDs power consumption is therefore highly correlated to the image content and can be readily estimated by considering the luminance level of the displayed image pixels. Although OLED displays consume energy in a more controllable and efficient manner than many other displaytechnologies, they are still the most important source of energy consumption in a video transmission chain.
[0020] The ISO / IEC 23001-11 specification ("Energy-Efficient Media Consumption: Green Metadata”) promulgated by the International Organization for Standardization in 2023, incorporated by reference herein, specifies metadata, referred to as green metadata, that facilitates the reduction of energy usage during media consumption (i.e., decoding and display operations), and specifically for reducing the display power consumption. The metadata described in the specification are particularly well tailored to display technologies based on non-emissive pixels and embedding backlight illumination such as LCD, as they are designed to attain display energy reductions by using display adaptation techniques that generate dynamically, on the emitter side, some RGB-component statistics and quality indicators metrics about the consumed video content. The metadata can be used to perform RGB picture components rescaling to set the best compromise between backlight / voltage reduction and picture quality, reducing voltage, and therefore allowing to reduce the energy consumption. However, it is far from being optimal, as these metadata convey global information and do not convey any information that would help the use of a pixel-wise attenuation map, as such a map is of no use for non-emissive pixel types of displays.
[0021] In some implementations, display green metadata may be provided as a timed metadata stream, such as an MPEG-DASH timed metadata representation (e.g., with the ©codecs attribute set to “dipi”, as described in ISO / IEC 23001-11). The stream associated with this representation can be retrieved and used by the decoder to perform post-processing to all the available and associated media representations and reduce energy consumption.
[0022] A display attenuation map may comprise a video created by a video content analyzer through a pixel-wise energy-aware frame algorithm. Such maps are typically in 2D, though 3D attenuation maps may be utilized in some implementations, such as for 3D environments, pixel clouds, augmented reality or virtual reality, etc. The frame algorithm may have smoothness and scalable properties which may be dynamically adjustable in some implementations. The related attenuation map information contains thecharacteristics of the attenuation map (e.g., display target model, width, length, energy reduction rate, etc.) and the operation to be used to apply it to the original video before rendering the video on the display. Applying the information conveyed in the display attenuation map video to the analyzed video intelligently and locally reduces locally the brightness of the frames to optimize the tradeoff between quality of experience (QoE) and energy reduction.
[0023] The attenuation map is designed so that, when applied to an input image, it produces a modified image that requires less energy for display than the input image. One implementation is to scale down the luminance according to a selected energy reduction rate. More complex implementations take other parameters into account such as the similarity between the modified image and the input image, or the contrast sensitivity function of the human vision, or a smoothness characteristic that allows to downscale the attenuation map without introducing heavy artifacts when upscaling it on the decoder side, etc.
[0024] Applying an attenuation map to an image may comprise combining them together according to a selected type of operation. The values of the samples of the attenuation map and the type of operation are closely related together. Indeed, when the operation is an addition, the attenuation map comprises sample values having negative values, when the operation is a subtraction, the attenuation map comprises sample values having negative values, when the operation is a multiplication, the attenuation map comprises floating point sample values in a range between zero (pixel becomes black) and one (no attenuation). The sample values may be referred to variously as attenuation values, attenuation factors, attenuation vectors, or by any other similar description. The attenuation may be applied to luma values of pixels, chroma values of pixels, or both luma and chroma values of pixels.
[0025] To convey the attenuation map information, in some implementations, a user data supplemental enhancement information (SEI) message may be utilized, but their genericity makes them inefficient for many use cases. In other implementations, a more use-specific SEI message may be utilized, such as those described in PCT application no. PCT / EP2023 / 083362 on November 28, 2023, which is incorporated herein by reference.. This message, conveyed within the codec signaling, aims at guidingthe use of the pixel-wise display attenuation map. The dimming maps may be transmitted as auxiliary data up to the receiver side. In other implementations, display attenuation maps may be carried in ISO Based Media File Format (ISOBMFF) media containers. However, many of these implementations may require changes to the core syntax structures or underlying protocol. Instead, implementations of the systems and methods discussed herein provide for signaling of display attenuation map information in adaptation or representation sets within manifest files.
[0026] Specifically, the present disclosure provides implementation details of systems and methods that enrich dynamic adaptive streaming protocols such as dynamic adaptive streaming over HTTP (DASH) or HTTP live streaming (HLS) signaling to enable delivery of a display attenuation map-coded video and the relative information from standard HTTP servers to HTTP clients. The additional signaling enables a media player to select a video media representation, and one or several complementary display attenuation maps (sometimes referred to as dimming maps) including relative attenuation map information, to modify the video and reduce the energy consumption while displaying it on the screen of the receiver device or the screen connected to it. A display attenuation map representation can be associated with one or several video media representation(s) (e.g., at different bit rates or bandwidths, or with varying resolution, color depth, frame rate, etc.). In various implementations, these systems and methods may include signaling of the codec used for the Adaptation Sets carrying display attenuation map information as a restricted video codec, the definition and use of a new DASH descriptor to signal attenuation map information at either the Adaptation Set or Representation levels, and instructions for grouping Adaptation Sets carrying display attenuation maps with the Adaptation Set of the video using Preselections to define some experiences.
[0027] The dynamic adaptative streaming over HTTP (i.e., DASH, ISO / IEC 23009-1) protocol enables delivery of continuous media content from standard HTTP servers to HTTP clients and enables caching of content by standard HTTP caches. It allows the highest visual quality while accounting for various factors such as network bandwidth, display resolution and viewing conditions.
[0028] FIG. 1 is a block diagram of an embodiment of a DASH delivery process model 100. A DASH media presentation preparation function 110 may interact with a Media Presentation Description (MPD) delivery function 120 and a DASH segment delivery function 130. Each function 110-130 may be provided by the same or a different computing device (or a virtual device, such as a cloud computing device, media provider cloud, etc.). For example, functions 110-130 may be provided by or implemented on an HTTP server executed by a computing device. An HTTP cache 140 may store media segments and / or MPD metadata for delivery to a DASH client 150, such as a desktop computer, laptop computer, smart television, appliance, tablet computer, smart phone, or other such computing device or media consumption device. HTTP cache 140 may be part of such computing device or media consumption device, or may be implemented in a gateway, router, access point, network appliance, or other such device.
[0029] DASH media presentation preparation function 110 may be configured to prepare a media presentation for viewing by a DASH client 150, either contemporaneously (e.g., for live content) or at a prior time (e.g., for video on demand or pre-recorded content). For example, the DASH media presentation preparation function 310 may receive data regarding media content from a content delivery network and may prepare an MPD, sometimes referred to as a manifest, to describe the media content.
[0030] Referring briefly to FIG. 2, illustrated is an example of a DASH manifest or MPD 200. The MPD 200 may comprise one or more URLs for keys, initialization vectors, ciphers, segments, and / or message authentication codes. The MPD 200 may list the URLs as static addresses, or as functions that may be used to determine associated URLs. The MPD 200 may be in any suitable format, such as Extensible Markup Language (XML). An MPD 200 may comprise information for one or more periods 210. A period may comprise timing data and may represent a content period during which a consistent set of encoded versions of the media content is available (e.g., a set of available bitrates, languages, captions, subtitles, etc. that do not change). Each period description 210 may comprise one or more adaptation set descriptions 220. An adaptation set may represent a set of interchangeable encoded versions of one orseveral media content components. For example, a first adaptation set may comprise a main video component, a second adaptation set may comprise a main audio component, a third adaptation set may comprise captions, etc. An adaption set may also comprise multiplexed content, such as combined video and audio. An adaptation set description 220 may define one or more representations 230. A representation may describe a deliverable encoded version of one or more media content components, such as an ISO base media file format (ISO-BMFF) version, a Moving Picture Experts Group (MPEG) version two transport system (MPEG-2 TS) version, etc. A representation may describe, for example, any needed codecs, encryption, and / or other data needed to present the media content. A DASH client 150 may dynamically switch between representations based on network conditions, device capability, user choice, etc., which may be referred to as adaptive streaming (e.g., switching to a more compressed or lower resolution version as a result of network congestion, etc.). Each representation set 230 may comprise one or more segments 240, sometimes referred to as chunks. Each segment 240 may comprise the media content data, may be associated with a URL, and may be retrieved by the DASH client 150 using a standard method, such as an HTTP GET request. Each segment 240 may contain a pre-defined byte size (e.g., 1 ,000 bytes) and / or an interval of playback time (e.g., 2, 5, or 15 seconds) of the media content (segments do not need to be the same length or duration). A segment may comprise the minimal individually addressable units of data that can be downloaded using URLs advertised via the MPD. The periods, adaptation sets, representations, and / or segments may be described in terms of attributes and elements, which may be modified to affect the presentation of the media content by the DASH client 150. Upon preparing the MPD, the DASH media presentation preparation function 110 may deliver the MPD to the MPD delivery function 120.
[0031] Returning to FIG. 1 , the client 150 may request the MPD from the MPD delivery function 120, e.g., via an HTTP GET request. The MPD delivery function 120 may respond with the MPD directly or via an HTTP cache 140 (e.g., an edge cache or other device storing a copy of the MPD) Based on the address data in the MPD, the DASH client 150 may request appropriate segments from the DASHsegment delivery function 130 (e.g., a segment in a particular resolution, encoding type, bitrate, etc.). While only one MPD delivery function, DASH segment delivery function, and cache are illustrated, in many implementations, these functions may be duplicated across or provided by many different computing devices (including virtual computing devices or cloud devices). Accordingly, segments and / or MPDs may be retrieved from a plurality of different devices, and / or from a plurality of URLs and / or physical locations. The DASH client 150 may present the retrieved segments based on the instructions in the MPD.
[0032] The systems and methods discussed herein are directed, in one aspect, to signaling display attenuation map information via a DASH MPD manifest or similar adaptive streaming protocol manifest. The display attenuation map representations may be mapped to relevant video streams in the MPD to enable streaming clients to efficiently select the most appropriate attenuation map based on the network and battery conditions at any given point in time. In some implementations, display attenuation map video frames may be stored in ISO base media files as media samples of a media track using extensions to ISO / IEC 14496-12. In some implementations, the display attenuation map tracks may be signaled in the MPD or manifest file and associated to one or several representations of a target video, such that the decoder can apply the map on these representations before displaying or rendering them on a display. By utilizing the MPD or manifest value for signaling the display attenuation map, service providers may avoid having to use SEI messages or auxiliary data within the codec bitstream. This may reduce bandwidth requirements for the stream itself, and also allow for caching of the display attenuation map with the manifest (e.g., at geographically proximate network edge caches, which may reduce latency and total network bandwidth requirements).
[0033] In some implementations, display attenuation maps may be signaled in the MPD as representations of an adaptation set. An adaptation set that includes representations of a display attenuation map is referred to here as a Display Attenuation Map Adaptation Set. A display attenuationmap adaptation set may include a Role descriptor, as defined in ISO / IEC 23009-1 , with a @schemeldUri attribute set to "urn:mpeg:dash:role:2011” and the ©value attribute set to "ami” in some implementations.
[0034] A display attenuation map adaptation set may include one or more representations of a coded display attenuation map sequence. Each representation may be encoded using a different coding configuration and / or may result in a different energy reduction level when applied to the associated video. The association between the representations of the display attenuation map adaptation set and representation in the related video adaptation set may be signaled by including an ©association Id identifier and ©associationType attribute in the representations of the display attenuation map adaptation set. The ©attociation Id attribute may be assigned a value identical to that of the @id attribute for the corresponding representation in the video adaptation set and the ©associationType attribute may be set to the value “amit” (indicating that the adaptation set is for a display attenuation map for the associated video adaptation set).
[0035] The ©codecs attribute for a display attenuation map adaptation set, or representations of this adaptation set if ©codecs is not signaled for the AdaptationSet element, may be set based on the respective codec used for encoding the attenuation map. The value of ©codecs may be set to 'resv.gmat.XXXX', where XXXX corresponds to the four-character code (4CC) of the video codec from the original_format field in the RestrictedSchemelnfoBox of the sample entry (e.g., 'avcT or 'hvd ') of the corresponding ISOBMFF track.
[0036] In another embodiment, a number of display attenuation map adaptation sets may be associated with the same video representation. For example, each display attenuation map adaptation set may be related to a particular region within the video frames. Such a region may be spatial (e.g., a coordinate range such as from 0,0 to 250,400, a window within the frame such as the upper left, etc.) or temporal (e.g., from 0:00 to 5:30, for 120 seconds, etc.) or both spatial and temporal (e.g., with a specified coordinate range and time range). In such embodiments, the display attenuation map adaptation sets may be grouped with the associated video adaptation set using a green video preselection.
[0037] A green video preselection is signaled using a preselection element, as defined in ISO / IEC 23009-1 , with an identifier list for the @preselectionComponents attribute including the identifier of the main adaptation set (the video adaptation Set) for the volumetric media followed by the identifiers of the associated display attenuation map adaptation sets. The ©codecs attribute for the preselection may be set based on the codec used by the main adaptation set (i.e., the video adaptation set).
[0038] To signal the characteristics of a display attenuation map carried by a display attenuation map adaptation set, an Attenuation Map descriptor may be utilized. This descriptor may comprise an EssentialProperty or SupplementalProperty descriptor with the ©schemeldUri set to a unique URI (e.g., "urn:mpeg:mpegl:green:2023:ami"). An AttenuationMap descriptor can be present at the adaptation level and / or at the representation level for each representation in a display attenuation map adaptation set.
[0039] In some implementations, the ©value of the AttenuationMap descriptor may be omitted. The AttenuationMap descriptor may include an AMI element whose attributes and sub-elements describe the display attenuation map represented by the adaptation set, and its characteristics. The XML elements and attributes for this descriptor may be defined in a separate namespace "urn:mpeg:mpegl:green:2023”. The namespace designator "green:” is used to refer to this name space in this document. Table 1 is a table listing example elements and attributes of an AttenuationMap descriptor, according to some implementations.Table 1 : Elements and attributes of the Attenuation Map descriptor
[0040] In some implementations, the data types for various elements and attributes are defined in an XML schema shown in FIGs. 3A-3C. In another implementation, the AttenuationMap descriptor is a SupplementalProperty descriptor with the @schemeldUri attribute set to a unique URI (e.g., "urn:mpeg:mpegl:green:2023:ami"). In other implementations, the sizes of the defined attributes for the various elements of the attribute may be different based on the allowed range of values for the attribute.
[0041] FIGs. 4A-4B are a representation of an example MPD manifest, according to some implementations. The example demonstrates a DASH MPD with a presentation that includes a video adaptation set with three representations and a display attenuation map adaptation set with one representation that is associated with the first video representation. The attenuation map carried by the display attenuation map adaptation set results in a 20% reduction in energy consumption when applied to the associated video Representation. The application of display attenuation map representation also results in a 5% reduction of the video quality in terms of peak signal to noise ratio (PSNR).
[0042] FIGs. 5A-5B are a representation of another example MPD manifest, according to some implementations. This example demonstrates a DASH MPD with a presentation that includes a video adaptation set with three representations and a display attenuation map adaptation set with two representations that are associated with the first video representation. The attenuation map carried by the first representation of the display attenuation map adaptation set results in a 20% reduction in energy consumption when applied to the video representation, while the second representation of the display attenuation map adaptation set results in a 40% reduction in energy consumption when applied to the same video representation.
[0043] FIGs. 6A-6B are a representation of yet another example MPD manifest, according to some implementations. This example is similar to the example of FIGs. 5A-5B, but use one instance of the descriptor at the adaptation set level containing common information for all representations and additional instances for each representation with representation-specific information.
[0044] FIGs. 7A-7C are a representation of yet another example MPD manifest, according to some implementations. This example is similar to the example of FIGs. 6A-6B, but use a preselection to indicate the grouping of the video adaptation set and associated display attenuation map adaptation sets. In other implementations, variations of the above example manifests may be utilized.
[0045] As discussed above, a DASH client, such as client 150, may be guided by the information provided in the MPD. FIG. 8 is a flow chart of a method 800 for processing streaming video with anassociated display attenuation map, using the signaling presented in the MPD and the above-discussed schema, according to some implementations.
[0046] At 805, the client may retrieve the manifest. In some embodiments, the client first issues an HTTP request and downloads the MPD file from the content server. In other implementations, the client may retrieve the MPD file from a cache. At 810, in some implementations, the client may parse the MPD file to generate a corresponding in-memory representation of the XML elements in the MPD file.
[0047] To identify available display attenuation maps in a period, at 815, the streaming client scans the AdaptationSet elements to find Adaptation Sets with a Role descriptor with an ©value is set to ‘ami’ and includes an AttenuationMap descriptor element. For each Representation in the display attenuation map adaptation set, at 820, the client may also identify the associated Representation in the video Adaptation Set using the ©association^. In some implementations, 815-820 may be repeated for each additional adaptation set in the manifest.
[0048] At 825, in some implementations, the client starts selects one of the Representations of the video Adaptation Set based on its capabilities and the network conditions. For example, the client may monitor network conditions or perform measurements (e.g., transmitting short pings or requests and measuring a roundtrip time or packet loss rate, etc.) and select a Representation for which the network conditions meet a minimum requirement (e.g., based on throughput).
[0049] At 830, the client may download the Initialization Segment for the selected Representation, and may download the Initialization Segment for all Representations from the display attenuation map adaptation set that are associated with the selected video Representation. At 835, the client may begin sequentially downloading media segments from the video Representation and decoding and displaying the media segments.
[0050] During display of the media segments, in some implementations at 840, the client may constantly or periodically monitor the energy state and / or consumption of the device (e.g., the energy conditions such as on battery or on AC power; a charge state; the remaining power in the device's battery; a rate ofchange of the battery level; etc.) and based on these measurements or conditions, a remaining playback time of the media (where the media is pre-recorded), and the information signaled in the AttenuationMap descriptor, may determine whether to reduce energy consumption. If so, at 845, the client device may select one of the display attenuation map Representations.
[0051] Returning to 835, the client may subsequently download with each Media Segment from the video Representation a corresponding Media Segment from the display attenuation map Representation. The downloaded display attenuation map Media Segments are decoded and the decoded frames are applied to the corresponding decoded frames from the video before rendering.
[0052] In some implementations, 835-845 may be performed repeatedly, such as where energy consumption levels change or are restored. For example, after reducing energy consumption due to a low battery level, if the device is plugged in or the battery charged sufficiently, then the client may select a new display attenuation map representation with a higher energy consumption (lower attenuation, or higher QoE) or even an unattenuated representation.
[0053] Various methods and other aspects described in this application can be used to modify modules, for example, the motion compensation (970, 1075), motion estimation 975), entropy coding, intra (960, 1060) and / or decoding modules (945, 1030), of a video encoder 900 and decoder 1000 as shown in FIG. 9 and FIG. 10 . Moreover, the present aspects are not limited to WC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including WC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0054] Various numeric values are used in the present application, for example. The specific values are for example purposes and the aspects described are not limited to these specific values.
[0055] FIG. 9 illustrates an encoder 900. Variations of this encoder 900 are contemplated, but the encoder 900 is described below for purposes of clarity without describing all expected variations.
[0056] Before being encoded, the video sequence may go through pre-encoding processing 901), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one of the color components). Metadata can be associated with the pre-processing, and attached to the bitstream.
[0057] In the encoder 900, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned 902 and processed in units of, for example, CUs. Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction 960. In an inter mode, motion estimation 975 and compensation 970 are performed. The encoder decides 905 which one of the intra mode or inter mode to use for encoding the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. Prediction residuals are calculated, for example, by subtracting 910 the predicted block from the original image block.
[0058] The prediction residuals are then transformed 925 and quantized 930. The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded 945 to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes.
[0059] The encoder decodes an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized 940 and inverse transformed 950 to decode prediction residuals. Combining 955 the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters 965 are applied to the reconstructed picture to perform, for example, deblocking / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer 980.
[0060] FIG. 10 illustrates a block diagram of a video decoder 1000. In the decoder 1000, a bitstream is decoded by the decoder elements as described below. Video decoder 1000 generally performs adecoding pass reciprocal to the encoding pass as described in FIG. 9 . The encoder 900 also generally performs video decoding as part of encoding video data.
[0061] In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 900. The bitstream is first entropy decoded 1030 to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide 1035 the picture according to the decoded picture partitioning information. The transform coefficients are de-quantized 1040 and inverse transformed 1050 to decode the prediction residuals. Combining 1055 the decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained 1070 from intra prediction 1060 or motion-compensated prediction (i.e., inter prediction 1075. In-loop filters 1065 are applied to the reconstructed image. The filtered image is stored at a reference picture buffer 1080.
[0062] The decoded picture can further go through post-decoding processing 1085, for example, an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing 901 . The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0063] FIG. 11 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 1100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1100, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1100 are distributed across multiple ICs and / or discretecomponents. In various embodiments, the system 1100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 1100 is configured to implement one or more of the aspects described in this application.
[0064] The system 1100 includes at least one processor 1110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 1110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 1100 includes at least one memory 1120 (e.g., a volatile memory device, and / or a nonvolatile memory device). System 1100 includes a storage device 1140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 1140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as nonlimiting examples.
[0065] System 1100 includes an encoder / decoder module 1130 configured, for example, to process data to provide an encoded video / 3D object or decoded video / 3D object, and the encoder / decoder module 1130 may include its own processor and memory. The encoder / decoder module 1130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 1130 may be implemented as a separate element of system 1100 or may be incorporated within processor 1110 as a combination of hardware and software as known to those skilled in the art.
[0066] Program code to be loaded onto processor 1110 or encoder / decoder 1130 to perform the various aspects described in this application may be stored in storage device 1140 and subsequently loaded onto memory 1120 for execution by processor 1110. In accordance with various embodiments, one or more of processor 1110, memory 1120, storage device 1140, and encoder / decoder module 1130 may storeone or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video / 3D object, the decoded video / 3D object or portions of the decoded video / 3D object, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0067] In several embodiments, memory inside of the processor 1110 and / or the encoder / decoder module 1130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 1110 or the encoder / decoder module 1130 is used for one or more of these functions. The external memory may be the memory 1120 and / or the storage device 1140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for coding and decoding operations, such as for instance MPEG-2, HEVC, or VVC.
[0068] The input to the elements of system 1100 may be provided through various input devices as indicated in block 1105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a composite input terminal, (ill) a USB input terminal, and / or (iv) an HDMI input terminal.
[0069] In various embodiments, the input devices of block 1105 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (I) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (ill) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RFportion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0070] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 1110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 1110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1110, and encoder / decoder 1130 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.
[0071] Various elements of system 1100 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 1115, for example, an internal bus as known in the art, including theI2C bus, wiring, and printed circuit boards.
[0072] The system 1100 includes communication interface 1150 that enables communication with other devices via communication channel 1190. The communication interface 1150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 1190. The communication interface 1150 may include, but is not limited to, a modem or network card and the communication channel 1190 may be implemented, for example, within a wired and / or a wireless medium.
[0073] Data is streamed to the system 1100, in various embodiments, using a Wi-Fi network such as IEEE 802.11 . The Wi-Fi signal of these embodiments is received over the communications channel 1190 and the communications interface 1150 which are adapted for Wi-Fi communications. The communications channel 1190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 1100 using a set-top box that delivers the data over the HDMI connection of the input block 1105. Still other embodiments provide streamed data to the system 1100 using the RF connection of the input block 1105.
[0074] The system 1100 may provide an output signal to various output devices, including a display 1165, speakers 1175, and other peripheral devices 1185. The other peripheral devices 1185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 1100. In various embodiments, control signals are communicated between the system 1100 and the display 1165, speakers 1175, or other peripheral devices 1185 using signaling such as AV.Link, CEC, or other communications protocols that enable device- to-device control with or without user intervention. The output devices may be communicatively coupled to system 1100 via dedicated connections through respective interfaces 1160, 1170, and 1180. Alternatively, the output devices may be connected to system 1100 using the communications channel 1190 via the communications interface 1150. The display 1165 and speakers 1175 may be integrated in a single unit with the other components of system1100 in an electronic device, for example, a television. In various embodiments, the display interface 1160 includes a display driver, for example, a timing controller (T Con) chip.
[0075] The display 1165 and speaker 1175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 1105 is part of a separate set-top box. In various embodiments in which the display 1165 and speakers 1175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0076] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first”, "second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a "first decoding” and a "second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0077] Moreover, the present aspects are not limited to DASH, HLS, or other streaming protocols, and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination. Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.
[0078] As discussed above, the present aspects, although describing principles related to particular drafts of VVC (Versatile Video Coding) or to HEVC (High Efficiency Video Coding) specifications, are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations(including WC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0079] Various implementations involve decoding. "Decoding,” as used in this application, may encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. Whether the phrase "decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0080] Various implementations involve encoding. In an analogous way to the above discussion about "decoding”, "encoding” as used in this application may encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream.
[0081] In the present application, the terms "reconstructed” and "decoded” may be used interchangeably, the terms "encoded” or "coded” may be used interchangeably, the terms "pixel” or "sample” may be used interchangeably, and the terms "image,” "picture” and "frame” may be used interchangeably. Usually, but not necessarily, the term "reconstructed” is used at the encoder side while "decoded” is used at the decoder side.
[0082] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, anintegrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants ("PDAs”), and other devices that facilitate communication of information between end-users.
[0083] Reference to "one embodiment” or "an embodiment” or "one implementation” or "an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment” or "in an embodiment” or "in one implementation” or "in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment. Additionally, this application may refer to "determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0084] Further, this application may refer to "accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0085] Additionally, this application may refer to "receiving” various pieces of information. Receiving is, as with "accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, "receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0086] It is to be appreciated that the use of any of the following “ / ”, "and / or”, and "at least one of, for example, in the cases of “A / B”, "A and / or B” and "at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C” and "at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0087] Also, as used herein, the word "signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a quantization matrix for de-quantization. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways.
[0088] For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word "signal”, the word "signal” can also be used herein as a noun.
[0089] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the describedimplementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor- readable medium.
Claims
What is Claimed:1 . A method for utilizing attenuation maps for adaptive streaming, comprising: receiving, by a client device comprising a video decoder and a display, an adaptive streaming manifest comprising an identification of one or more adaptation sets, wherein a first adaptation set of the one or more adaptation sets comprises identifications of a plurality of representations of an item of streaming media content and one or more display attenuation maps; retrieving, by the client device, one or more media segments of the item of streaming media content corresponding to one of the plurality of representations; retrieving, by the client device, one or more media segments of a first display attenuation map of the one or more display attenuation maps corresponding to one of the plurality of representations; and displaying, by the client device via the display, the one or more media segments of the item of streaming media content, adjusted based on the one or more media segments of the first display attenuation map.
2. The method of claim 1 , wherein the one or more media segments of the first display attenuation map comprise pixel attenuation factors.
3. The method of claim 1 , wherein displaying the one or more media segments of the item of streaming media content further comprises, for each pixel of a media segment of the item of streaming media content, combining a luma value, a chrome value, or both a luma and a chroma value of the pixel with a corresponding pixel attenuation factor of a corresponding media segment of the display attenuation map.
4. The method of claim 1 , wherein the first adaptation set comprises identifications of a plurality of display attenuation maps, and further comprising selecting, by the client device, the first display attenuation map based on one or more of a device power condition, a battery charge level, and a time remaining of the item of streaming media content.
5. The method of claim 1 , wherein the client device receives the adaptive streaming manifest from a first server device.
6. The method of claim 1 , wherein the client device retrieves the one or more media segments of the item of streaming media content from a plurality of server devices.
7. The method of claim 1 , further comprising selecting, by the client device based on a network condition, a first representation of the plurality of representations.
8. The method of claim 7, wherein the one of the plurality of representations comprises the selected first representation.
9. The method of claim 1 , further comprising determining, by the client device, to reduce an energy utilization of display of the one or more media segments.
10. The method of claim 9, wherein determining to reduce an energy utilization of display of the one or more media segments is based on one or more of a device power condition, a battery charge level, and a time remaining of the item of streaming media content.
11. A client device utilizing attenuation maps for adaptive streaming, the client device comprising: a video decoder; and a display operably coupled to the video decoder, the client device configured to receive an adaptive streaming manifest comprising an identification of one or more adaptation sets, wherein a first adaptation set of the one or more adaptation sets comprises identifications of a plurality of representations of an item of streaming media content and one or more display attenuation maps; the client device further configured to retrieve one or more media segments of the item of streaming media content corresponding to one of the plurality of representations;the client device further configured to retrieve one or more media segments of a first display attenuation map of the one or more display attenuation maps corresponding to one of the plurality of representations; and the client device further configured to display the one or more media segments of the item of streaming media content, adjusted based on the one or more media segments of the first display attenuation map.
12. The client device of claim 11 , wherein the one or more media segments of the first display attenuation map comprise pixel attenuation factors.
13. The client device of claim 11 , wherein displaying the one or more media segments of the item of streaming media content further comprises, for each pixel of a media segment of the item of streaming media content, combining a luma value, a chrome value, or both a luma and a chroma value of the pixel with a corresponding pixel attenuation factor of a corresponding media segment of the display attenuation map.
14. The client device of claim 11 , wherein the first adaptation set comprises identifications of a plurality of display attenuation maps, and further comprising selecting, by the client device, the first display attenuation map based on one or more of a device power condition, a battery charge level, and a time remaining of the item of streaming media content.
15. The client device of claim 11 , wherein the client device receives the adaptive streaming manifest from a first server device.
16. The client device of claim 11 , wherein the client device retrieves the one or more media segments of the item of streaming media content from a plurality of server devices.
17. The client device of claim 11 , further comprising selecting, by the client device based on a network condition, a first representation of the plurality of representations.
18. The client device of claim 11 , further comprising determining, by the client device, to reduce an energy utilization of display of the one or more media segments.
19. A client device utilizing attenuation maps for adaptive streaming, the client device comprising: a video decoder; and an input / output device, the client device configured to receive an adaptive streaming manifest comprising an identification of one or more adaptation sets, wherein a first adaptation set of the one or more adaptation sets comprises identifications of a plurality of representations of an item of streaming media content and one or more display attenuation maps; the client device further configured to retrieve one or more media segments of the item of streaming media content corresponding to one of the plurality of representations; the client device further configured to retrieve one or more media segments of a first display attenuation map of the one or more display attenuation maps corresponding to one of the plurality of representations; and the client device further configured to cause the one or more media segments of the item of streaming media content to be displayed, adjusted based on the one or more media segments of the first display attenuation map.
20. The client device of claim 11 , wherein the client device causing the one or more media segments of the item of streaming media content to be displayed by sending the one or more media segments of the item of streaming media content to be displayed to an external display.
Citation Information
Patent Citations
Method and device for encoding and decoding attenuation map for energy aware images
WO2024126030A1