Adaptive transfer function for video encoding and decoding

By using an adaptive transfer function method, the transfer function is dynamically adjusted according to the characteristics of the video data, which solves the problem of low efficiency in HDR video encoding and enables efficient display on HDR devices.

CN116320394BActive Publication Date: 2025-12-09APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310395840.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-02-28
Filing Date
2015-02-25
Publication Date
2025-12-09
Estimated Expiration
2035-02-25

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively process high dynamic range (HDR) video data, resulting in low encoding efficiency and difficulty in displaying on HDR devices.

Method used

An adaptive transfer function approach is adopted, which dynamically adjusts the transfer function based on the brightness and texture features of the video data, limits the range of video data within the codec, uses fewer bits for representation during encoding, and extends to the full dynamic range of HDR devices during decoding.

Benefits of technology

It improves video encoding efficiency, simplifies the implementation of HDR technology, and ensures high-quality video display on HDR devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320394B_ABST
    Figure CN116320394B_ABST
Patent Text Reader

Abstract

The present disclosure relates to adaptive transfer functions for video encoding and decoding. The invention provides a video encoding and decoding system that implements an adaptive transfer function method within a codec for signal representation. A focused dynamic range for representing the effective dynamic range of the human visual system can be dynamically determined for each scene, sequence, frame, or region of input video. Video data can be clipped and quantized into the bit depth of the codec according to a transfer function for encoding within the codec. The transfer function can be the same as the transfer function of the input video data, or can be a transfer function internal to the codec. The encoded video data can be decoded and expanded into the dynamic range of one or more displays. The adaptive transfer function method can enable a codec to use fewer bits for internal representation of a signal, while still representing the entire dynamic range of the signal in the output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of application no. 201580010095.8, filed on February 25, 2015, entitled "Adaptive Transfer Function for Video Encoding and Decoding", which claims priority to U.S. Provisional Patent Application No. 62 / 108, 1 10, filed on January 26, 2015, entitled "Adaptive Transfer Function for Video Encoding and Decoding". TECHNICAL FIELD

[0002] The present disclosure relates generally to digital video or image processing and display. BACKGROUND

[0003] Various devices, including but not limited to personal computer systems, desktop computer systems, laptop and notebook computers, tablet or slate devices, digital cameras, digital video recorders, and mobile or smart phones, can include software and / or hardware that can implement one or more video processing methods. For example, a device can include an apparatus (e.g., an integrated circuit (IC), such as a system on a chip (SOC), or a subsystem of an IC) that can receive and process digital video input from one or more sources and output processed video frames according to one or more video processing methods. As another example, a software program can be implemented on a device that can receive and process digital video input from one or more sources and output processed video frames to one or more destinations according to one or more video processing methods.

[0004] For example, a video encoder can be implemented as an apparatus or an alternative software program in which digital video input is encoded or converted to another format, e.g., a compressed video format such as the H.264 / Advanced Video Coding (AVC) format or the H.265 High Efficiency Video Coding (HEVC) format, according to a video encoding method. As another example, a video decoder can be implemented as an apparatus or an alternative software program in which video in a compressed video format such as AVC or HEVC is received and decoded or converted to another (decompressed) format, e.g., a display format used by a display device, according to a video decoding method. The H.264 / AVC standard is published by ITU-T in a document entitled "ITU-T Recommendation H.264: Advanced video coding for generic audiovisual services." The H.265 / HEVC standard is published by ITU-T in a document entitled "ITU-T Recommendation H.265: High Efficiency Video Coding."

[0005] In many systems, an apparatus or software program can implement both a video encoder component and a video decoder component; such an apparatus or program is often referred to as a codec. Note that a codec can encode / decode both visual / image data and audio / sound data in a video stream.

[0006] Generally defined, dynamic range is the ratio between the maximum and minimum possible values of a variable quantity, such as in sound and light. In digital image and video processing, extended dynamic range or high dynamic range (HDR) imaging refers to techniques that produce a wider range of luminances in an electronic image (e.g., as displayed on a display screen or display device) than is obtained using standard digital imaging techniques, referred to as standard dynamic range or SDR imaging.

[0007] An electro-optical transfer function (EOTF) can map digital code values to light values, e.g., to luminance values. The inverse process, generally referred to as an optical-electrical transfer function (OETF), maps light values to electronic / digital values. The EOTF and OETF can be collectively referred to as transfer functions. The SI unit of luminance is the candela per square meter (cd / m 2 ). The non-SI term for this unit is "NIT". In standard dynamic range (SDR) imaging systems, a fixed transfer function, e.g., a fixed power-law gamma transfer function, is typically used for internal representation of video image content in order to simplify the encoding and decoding processes in the encoding / decoding system or codec. With the advent of high dynamic range (HDR) imaging techniques, systems, and displays, the need for more flexible transfer functions has arisen. SUMMARY

[0008] Embodiments of video encoding and decoding systems and methods are described that implement an adaptive transfer function for internal representation of video image content within a video encoding and decoding system or codec. The embodiments can dynamically determine a focused dynamic range for a current scene, sequence, frame, or region of a frame in an input video based on one or more characteristics of the image data (e.g., luminance, texture, etc.), clip the input video dynamic range to the focused range, and then appropriately map (e.g., quantize) values within the clipped range from the bit depth of the input video to the bit depth of the codec according to the transfer function used to represent the video data in the codec.

[0009] In embodiments, various transfer functions can be used to represent the input video data and the focused range video data in the codec. In some embodiments, the transfer function used to represent the video data in the codec can be the same as the transfer function used to represent the input video data. In some embodiments, a different transfer function (referred to as an internal transfer function or secondary transfer function) can be used to represent the video data in the codec than the transfer function used to represent the input video data (referred to as a primary transfer function).

[0010] In embodiments, the focused dynamic range used by the encoder for a scene, sequence, frame,

[0011] The focus range, transfer function, quantization parameters, and other format information for a region or area can be signaled to the decoder components, for example, by metadata embedded in the output bitstream. In the decoder, the encoded bitstream can be decoded and dynamically expanded to the full dynamic range of the target device (e.g., a display that supports high dynamic range (HDR)) according to the signaled focus range(s) for the scene, sequence, frame, or video area.

[0012] By dynamically adapting the transfer function to the input video data, embodiments can allow video data to be represented in the codec with fewer bits than when used to represent the input video data, while also allowing the codec to output video data that is expanded and fills the dynamic range of an HDR device such as a display that supports HDR. Embodiments of the adaptive transfer function approach can, for example, enable a video encoding and decoding system to use 10 bits or fewer for internal representation and processing of video data within the codec, while using 12 bits or more to represent the expanded dynamic range or high dynamic range of the video data when outputting the video to an HDR device such as a display that supports HDR. Thus, embodiments of the adaptive transfer function approach of a video encoding and decoding system can simplify the implementation of HDR technology, and thus its adoption, especially in the consumer space. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 An example codec or video encoding and decoding system implementing embodiments of an adaptive transfer function is shown.

[0014] Figure 2 An example encoder applying an adaptive transfer function approach to video input data and generating encoded video data according to some embodiments is shown.

[0015] Figure 3 An example decoder decoding encoded video data according to an adaptive transfer function approach and expanding the decoded video data to generate display format video data according to some embodiments is shown.

[0016] Figure 4 An example full range of input video data and an example focus range of the video data according to some embodiments is shown.

[0017] Figure 5 An example of mapping N-bit input video data within a focus range to generate C-bit video data according to some embodiments is shown.

[0018] Figure 6 An example of expanding C-bit decoded video into the full dynamic range of an HDR- enabled device to generate D-bit video data for the device according to some embodiments is shown graphically.

[0019] Figures 7A to 7C Graphically illustrates applying different focus ranges to different portions of a video sequence or video frame according to an embodiment of the adaptive transfer function method.

[0020] Figure 8 Flowchart for a video encoding method according to some embodiments that applies the adaptive transfer function method to video input data and generates encoded video data.

[0021] Figure 9 Flowchart for a video decoding method according to some embodiments that decodes encoded video data according to the adaptive transfer function method and expands the decoded video data to generate display format video data.

[0022] Figure 10 Block diagram of one embodiment of a system on a chip (SOC) that can be configured to implement various aspects of the systems and methods described herein.

[0023] Figure 11 Block diagram of one embodiment of a system that can include one or more SOCs.

[0024] Figure 12 Illustrates an exemplary computer system that can be configured to implement various aspects of the systems and methods described herein, according to some embodiments.

[0025] Figure 13 Illustrates a block diagram of a portable multifunctional device according to some embodiments.

[0026] Figure 14 Depicts a portable multifunctional device according to some embodiments.

[0027] While the application is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the drawings and detailed description thereto are not intended to limit the application to the particular form disclosed, but on the contrary, the intention is to cover all modifications, equivalents and alternatives falling within the spirit and scope of the application. As used throughout this application, the word "may" is used in a permissive sense (i.e., meaning having the potential to), rather than the mandatory sense (i.e., meaning must). Similarly, the words "include," "including," and "includes" mean "including, but not limited to." It will be readily understood that the components of the present application, as generally described and illustrated in the Figures herein, could be arranged and designed in a wide variety of different configurations.

[0028] Various units, circuits, or other components can be described as being "configured to" perform a task or tasks. In such contexts, "configured to" is often used interchangeably with "having structure that" performs the task or tasks. Similarly, various units, circuits, or other components can be described as "configured by" a structure to perform a task or tasks, which in such contexts, is often used interchangeably with "having the structure that" performs the task or tasks. Reciting that a unit, circuit, or other component is "configured to" perform one or more tasks is expressly intended not to invoke 35 U.S.C. § 112, paragraph 6, on the unit, circuit, or other component. Accordingly, "configured to" can describe structure that is not yet in existence. DETAILED DESCRIPTION

[0029] Embodiments of video encoding and decoding systems and methods are described that implement an adaptive transfer function for internal representation of video image content within a video encoding and decoding system or codec. These embodiments can allow the dynamic range of input video data to be adapted to the codec during the encoding and decoding process. In embodiments, the transfer function can be dynamically adapted to individual scenes, sequences, or frames of input video to the codec. In some embodiments, the transfer function can be dynamically adapted within regions of a frame. These embodiments can dynamically adapt the transfer function to the input video data, preserve only information within a dynamically determined human visual system effective dynamic range, referred to herein as the focus dynamic range or simply the focus range, and map the data within the focus range from the bit depth of the input video to the bit depth of the codec according to the transfer function for processing within the codec. The output of the codec can be expanded to fill the dynamic range of an output device or target device, including but not limited to an extended dynamic range device or a high dynamic range (HDR) device, such as an HDR-capable display. Figures 1 to 3 An exemplary video encoding and decoding system or codec in which embodiments of the adaptive transfer function method can be implemented is shown.

[0030] The human visual system covers a significant dynamic range overall. However, the human visual system tends to adapt and limit the dynamic range based on the current scene or image being viewed, e.g., according to the luminance (illuminance) and texture characteristics of the scene or image. Thus, while the overall dynamic range of the human visual system is quite large, the effective dynamic range of a given scene, sequence, frame, or region of a video can be quite small according to the image characteristics, including but not limited to luminance and texture. Embodiments of the video encoding and decoding systems and methods described herein can exploit this characteristic of human vision to employ similar strategies for dynamically limiting the range within the codec according to the characteristics of the current scene, sequence, frame, or region. In embodiments, according to one or more characteristics (e.g., luminance, texture, etc.) of the scene, sequence, frame, or region currently being processed, an encoder component or process can dynamically limit the range (e.g., luminance) of the input video scene, sequence, frame, or region to be within the effective dynamic range of the human visual system and the required range of the codec (e.g., bit depth). This can be done, for example, by dynamically determining the area within the current scene, sequence, frame, or region in the input video according to the sample values and focusing the bit depth of the codec (e.g., 10 bits) into that range. In some embodiments, the focusing can be performed in the encoder by clipping the input video dynamic range into that area, and then mapping (e.g., quantizing) the values in the clipped range (referred to as the focused dynamic range or focused range) from the bit depth of the input video to the bit depth of the codec appropriately according to the transfer function used to represent the video data in the codec. Figure 4 and Figure 5 An exemplary focused range determined for input video data and mapping of N-bit input video data to the bit depth of the codec according to the transfer function is shown graphically according to some embodiments. As used herein, N-bits refers to the bit depth of the input video, C-bits refers to the bit depth used to represent the video data within the codec, and D-bits refers to the bit depth of the target device (e.g., display device) for the decoded video.

[0031] In implementations, various transfer functions can be used to represent the input video data and the focus range video data in the codec. Examples of transfer functions that can be used to represent the video data in implementations can include, but are not limited to, a gamma-based power law transfer function, a logarithm-based transfer function, and a transfer function based on human visual perception, such as the perceptual quantizer (PQ) transfer function proposed by Dolby Laboratories, Inc. In some implementations, the transfer function used to represent the video data in the codec can be the same as the transfer function used to represent the input video data. However, in other implementations, the video data in the codec can be represented using a different transfer function (referred to as an internal or secondary transfer function) than the transfer function used to represent the input video data (referred to as a primary transfer function). This can allow, for example, the video data (which can also be referred to as a video signal) to be represented with higher precision within the codec than can be accomplished using the primary transfer function.

[0032] In implementations, the focus range, transfer function, quantization parameter, and other format information for a scene, sequence, frame, or region of the encoder can be signaled to the decoder components, for example, by metadata embedded in the output bitstream. At the decoder, the encoded bitstream can be decoded and dynamically expanded to the full dynamic range of a target device, such as an HDR device (e.g., an HDR-capable display), according to the signaled focus range or ranges for the scene, sequence, frame, or video region. Figure 6 Expanding decoded video data into the full dynamic range of an HDR device is shown graphically according to some implementations.

[0033] By dynamically adapting the transfer function to the input video data, implementations can allow the video data to be represented in the codec with fewer bits than when used to represent the input video data, while also allowing the codec to output the video data so that it is expanded and fills the dynamic range of an HDR device, such as an HDR-capable display. Implementations of the adaptive transfer function approach can, for example, enable a video encoding and decoding system to use 10 bits or fewer for the internal representation and processing of video data within the codec, while using, for example, 12 bits or more to represent the expanded dynamic range or high dynamic range of the video data when outputting the video to an HDR device, such as an HDR-capable display. Thus, implementations of the adaptive transfer function approach of a video encoding and decoding system can simplify the implementation of HDR technology, and therefore its adoption, especially in the consumer space.

[0034] The adaptive transfer function approach described herein can be applied to all color components of a video signal, or alternatively can be applied to one or both of the luminance component and the chrominance components of a video signal separately.

[0035] Figure 1 An exemplary video encoding and decoding system or codec 100 implementing embodiments of the adaptive transfer function is shown. In embodiments, an adaptive transfer function 110 component or module of the codec 100 can convert N-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) video input data 102 to C-bit (e.g., 10-bit or fewer bits) video input data 112 and to an encoder 120 component or module of the codec 100 in accordance with the adaptive transfer function method. In some embodiments, to convert the N-bit video input data 102 to the C-bit video input data 112, the adaptive transfer function 110 can dynamically determine an area within a current scene, sequence, frame, or region in the N-bit video input data 102 in accordance with sample values and focus C-bits into the area. In some embodiments, this focusing can be done in the adaptive transfer function 110 by clipping the N-bit video input data 102 to a determined focus range, and then mapping (e.g., quantizing) the N-bit values in the focus range to the codec's bit depth appropriately according to a transfer function used to represent the video data in the encoder 120 to generate the C-bit video input data 112 for the encoder 120.

[0036] In some embodiments, the adaptive transfer function 110 can also generate format metadata 114 as output to the encoder 120. The format metadata 114 can describe the adaptive transfer function method as being applied to the input video data 102. For example, the format metadata 114 can indicate the determined focus range, parameters used to map the video data to the encoder bit depth, and can also contain information about the transfer function applied to the focus range.

[0037] In some embodiments, the encoder 120 can then encode the C-bit video input data 112 to generate an encoded stream 122 as output. In some embodiments, the encoder 120 can encode the C-bit video input data 112 according to a compressed video format such as the H.264 / Advanced Video Coding (AVC) format or the H.265 / High Efficiency Video Coding (HEVC) format. However, other encoding formats can also be used. In some embodiments, the encoder 120 can embed the format metadata 114 into the output stream 122, for example, so that the format metadata 114 can be provided to a decoder 130 component. The output encoded stream 122 can be stored to a memory, or alternatively can be sent directly to a decoder 130 component or module of the codec 100. Figure 2 An exemplary encoder component of the video encoding and decoding system or codec 100 is shown in more detail.

[0038] Decoder 130 can read or receive encoded stream 122 and decode it to generate a C-bit decoded stream 132 as output to an inverse adaptive transfer function 140 component or module of codec 100. In some embodiments, the adaptive transfer function method described as being applied by adaptive transfer function 110 to format metadata 134 of input video data 102 can be extracted from input stream 122 and passed to inverse adaptive transfer function 140. Inverse adaptive transfer function 140 can convert C-bit decoded stream 132 according to format metadata 134 and display information 144 to generate D-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) HDR output 142 to one or more displays 190 or other devices. Figure 3 An example decoder component of video encoding and decoding system or codec 100 is shown in more detail.

[0039] Embodiments of video encoding and decoding system or codec 100 implementing the adaptive transfer function method as described herein can be implemented, for example, in a device or system that includes one or more image capture devices and / or one or more display devices. An image capture device can be any device that includes an optical or light sensor capable of capturing digital images or video. Image capture devices can include, but are not limited to, video cameras and still image cameras, as well as image capture devices that can capture both video and single images. An image capture device can be a stand-alone device or can be a camera integrated into other devices, including but not limited to a smart phone, a mobile phone, a PDA, a tablet or tablet device, a multi-function device, a computing device, a laptop computer, a notebook computer, a netbook computer, a desktop computer, and the like. Note that an image capture device can include a small form factor camera suitable for small devices such as mobile phones, PDAs, and tablet devices. A display or display device can include a display screen or panel integrated into other devices, including but not limited to a smart phone, a mobile phone, a PDA, a tablet or tablet device, a multi-function device, a computing device, a laptop computer, a notebook computer, a netbook computer, a desktop computer, and the like. A display device can also include a video monitor, a projector, or generally any device that can display or project digital images and / or digital video. A display or display device can use LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies can also be used.

[0040] Figures 10 to 14Non-limiting examples of devices in which embodiments can be implemented are shown. A device or system that includes an image capture device and / or a display device can include hardware and / or software for implementing at least some of the functionality for processing video data as described herein. In some embodiments, a portion of the functionality as described herein can be implemented on one device, while another portion can be implemented on another device. For example, in some embodiments, a device that includes an image capture device can implement a sensor pipeline that processes and compresses (i.e., encodes) images or video captured via a light sensor, while another device that includes a display panel or display screen can implement a display pipeline that receives and processes the compressed images (i.e., decodes) for display to the display or screen. In some embodiments, at least some of the functionality as described herein can be implemented by one or more components or modules of a system on a chip (SOC) that can be used in a device, including but not limited to a multi-function device, a smart phone, a tablet or tablet device, and other portable computing devices such as laptop computers, notebook computers, and netbook computers. Figure 10 An example SOC is shown, and Figure 11 An example device implementing the SOC is shown. Figure 12 An example computer system that can implement the methods and apparatus described herein is shown. Figure 13 And Figure 14 An example multi-function device that can implement the methods and apparatus described herein is shown.

[0041] Embodiments are now generally described herein as processing video. However, embodiments can also be applied to processing a single image or still image in addition to or instead of processing video frames or video sequences. Thus, it should be understood that when “video,” “video frame,” “frame,” and the like are used herein, these terms can generally refer to a captured digital image.

[0042] Figure 2 An example encoder 200 that applies an adaptive transfer function method to video input data 202 and generates encoded video data 232 as output, in accordance with some embodiments, is shown. As Figure 1As shown, the encoder 200 can be, for example, a component or module of the codec 100. In embodiments, the adaptive transfer function component or module 210 of the encoder 200 can convert N-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) video input data 202 to C-bit (e.g., 10-bit or fewer) data 212 according to a transfer function method, and output the C-bit data 212 to a processing pipeline 220 component of the encoder 200. In some embodiments, to convert the N-bit video input data 202 to C-bit video data 212, the adaptive transfer function module 210 can dynamically determine an area within a current scene, sequence, frame, or region in the N-bit video input data 202 in terms of sample values and focus C-bits into that area. In some embodiments, this focusing can be done in the adaptive transfer function module 210 by clipping the dynamic range of the N-bit video input data 202 into a determined focus range, and then mapping (e.g., quantizing) the N-bit values in the focus range to C-bits appropriately according to a transfer function used to represent video data in the encoder 200. In some embodiments, the adaptive transfer function module 210 can also generate format metadata 214 as output to the encoder 220. The format metadata 214 can describe the adaptive transfer function method as being applied to the input video data 202. For example, the format metadata 214 can indicate the determined focus range parameters used to map the video data to the encoder bit depth, and can also contain information about the transfer function applied to the focus range.

[0043] In some embodiments, the encoder 200 can then encode the C-bit video input data 212 generated by the adaptive transfer function module 210 to generate an encoded stream 232 as output. In some embodiments, the encoder 120 can encode the C-bit video input data 112 according to a compressed video format such as the H.264 / Advanced Video Coding (AVC) format or the H.265 / High Efficiency Video Coding (HEVC) format. However, other encoding formats can also be used.

[0044] In some embodiments, input video frames are subdivided into the encoder 200 and processed therein according to blocks of picture elements, referred to as pixels or primitives. For example, 16x16 pixel blocks, referred to as macroblocks, can be used for H.264 encoding. HEVC encoding uses blocks referred to as coding tree units (CTUs) that can vary in size from 16x16 pixels to 64x64 pixels. CTUs can be divided into coding units (CUs), and can be further subdivided into prediction units (PUs) that can vary in size from 64x64 pixels down to 4x4 pixels. In some embodiments, the video input data 212 can be separated into luma and chroma components, and the luma and chroma components can be processed separately at one or more components or stages of the encoder.

[0045] In Figure 2 In the example encoder 200 shown, the encoder 200 includes a processing pipeline 220 component and an entropy encoding 230 component. The processing pipeline 220 can include multiple components or stages that process video input data 212. The processing pipeline 220 may, for example, implement intra- and inter- estimation 222, mode decision 224, motion compensation and reconstruction 226 operations as one or more stages or components.

[0046] The operation of an example processing pipeline 220 is described below at a high level, and is not intended to be limiting. In some embodiments, intra- and inter- estimation 222 can determine previously encoded pixel blocks to be used in the process of encoding a block input to the pipeline. In some video encoding techniques such as H.264 encoding, each input block can be encoded using a currently intra-encoded pixel block. The process of determining these blocks can be referred to as intra-estimation or simply intra-estimation. In some video encoding techniques such as H.264 and H.265 encoding, blocks can also be encoded using pixel blocks from one or more previously reconstructed frames (referred to as reference frames, shown in FIG. 2 as reference data 240). In some embodiments, reconstructed and encoded frames output by the encoder pipeline can be decoded and stored to the reference data 240 for use as reference frames. In some embodiments, the reference frames stored in the reference data 240 for use in reconstructing a current frame can include one or more reconstructed frames that occur temporally prior to the current frame in the video being processed, and / or one or more reconstructed frames that occur temporally later than the current frame in the video being processed. The process of finding matching pixel blocks in reference frames can be referred to as inter-estimation, or more generally as motion estimation. In some embodiments, mode decision 224 can receive output for a given block from inter- and intra- estimation 222, and determine the best prediction mode (e.g., inter-prediction mode or intra-prediction mode) and corresponding motion vectors for the block. This information is passed to motion compensation and reconstruction 226. Figure 2

[0047] The operation of motion compensation and reconstruction 226 can depend on the best mode received from mode decision 224. If the best mode is inter-prediction, the motion compensation component obtains reference frame blocks corresponding to the motion vectors, and combines these blocks into a predicted block. The motion compensation component then applies weighted prediction to the predicted block to generate a final block prediction that is passed to the reconstruction component. In weighted prediction, values from the reference data can be weighted according to one or more weighting parameters, and shifted by an offset value to generate prediction data that will be used to encode the current block. If the best mode is intra-prediction, intra-prediction is performed using one or more neighboring blocks to generate a predicted block for the current block being processed at this stage of the pipeline.

[0048] ​The reconstruction component performs a block (e.g., macroblock) reconstruction operation of the current block from the motion compensated output. The reconstruction operation can include, for example, forward transform and quantization (FTQ) operations. Motion compensation and reconstruction 226 can output the transformed and quantized data to an entropy encoding 230 component of the encoder 200.

[0049] The entropy encoding 230 component can apply, for example, entropy encoding techniques to compress the transformed and quantized data output by the pipeline 220 to generate an encoded output stream 232. Example entropy encoding techniques that can be used can include, but are not limited to, Huffman coding techniques, CAVLC (Context Adaptive Variable Length Coding) and CABAC (Context Adaptive Binary Arithmetic Coding). In some embodiments, the encoder 200 can embed the format metadata 214 into the encoded output stream 232 so that the format metadata 214 can be provided to a decoder. The output encoded stream 232 can be stored to a memory or, alternatively, can be transmitted directly to a decoder component.

[0050] Reference data 240 can also be output by the pipeline 220 and stored to a memory. The reference data 240 can include, but is not limited to, one or more previously encoded frames (referred to as reference frames) that can be accessed, for example, in the motion estimation and motion compensation and reconstruction operations of the pipeline 220.

[0051] In some embodiments, the encoder 200 can include a reformatting 250 component that can be configured to reformat reference frame data obtained from the reference data 240 for use by the pipeline 220 components when processing a current frame. The reformatting 250 can involve, for example, converting the reference frame data from a focus range / transfer function used to encode the reference frame to a focus range / transfer function to be used to encode the current frame. For example, the focus range / transfer function mapping for the reference frame luminance can be from 0.05 cd / m 2 to 1000 cd / m 2 . For the current frame, the focus range can be extended due to brighter image content; for example, the focus range can be extended or increased to 2000 cd / m 2 . Thus, the focus range / transfer function mapping for the current frame luminance can be from 0.05 cd / m 2 to 2000 cd / m 2 . In order to use the reference frame data in the pipeline 220 (e.g., in the motion estimation operation), the reference frame data can be reformatted according to the focus range / transfer function used to encode the current frame by the reformatting 250 component.

[0052] In a given example, the reformatting 250 component can convert the reference frame data from 0.05 cd / m 2 to 1000 cd / m 2The range within which it ranges from 0.05 cd / m 2 Range reconstructed to 2000 cd / m for the current frame. 2 Range. In this example, since the reference frame only contains values ​​from 0.05 cd / m 2 Up to 1000 cd / m 2 The data, therefore, contains some code words in the C-bit representation of the reformatted reference frame data (e.g., indicating greater than 1000 cd / m). 2 The coded words of the values ​​may not be used in the reformatted reference data. However, multiple values ​​are mapped to the current focus range for prediction or other operations.

[0053] As indicated by the arrows returning from reference data 240 to adaptive transfer function module 210, in some implementations, adaptive transfer function module 210 may use adaptive transfer function information from one or more previously processed frames to determine the focus range / transfer function of the current frame to be processed in pipeline 220.

[0054] Figure 3 An exemplary decoder 300 is shown, according to some embodiments, which decodes encoded video data 302 and expands decoded video data 322 to generate display format video data 332 using an adaptive transfer function method. Figure 1 As shown, decoder 300 may be, for example, a component or module of codec 100. In Figure 3 In the exemplary decoder 300 shown, decoder 300 includes an entropy decoding 310 component, an inverse vectorization and transformation 320 component, and an inverse adaptive transfer function module 330. The entropy decoding 310 component may, for example, apply entropy decoding techniques to decompress data generated by the encoder (e.g., Figure 2 The encoder 200 (shown) outputs an encoded stream 302. The inverse vectorization and transformation unit 320 performs inverse vectorization and transformation operations on the data output by entropy decoding 310 to generate a C-bit decoded stream 322 as the output of the inverse adaptive transfer function module 330 to decoder 300. In some embodiments, the adaptive transfer function method is described as applying format metadata 324 to the current scene, sequence, frame, or region, which can be extracted from the input stream 302 and passed to the inverse adaptive transfer function module 330. The inverse adaptive transfer function module 330 can extend the C-bit decoded stream 322 according to the format metadata 324 and display information 392 to generate a D-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) HDR output 332 for one or more displays 390 or other devices.

[0055] Figures 4 to 6An exemplary focus range for N-bit input video data is shown graphically, according to some embodiments, and the input video data is mapped according to a transfer function to C-bits available within a codec, and the decoded video is expanded to the full dynamic range of an HDR device to generate D-bit video data for that device.

[0056] Figure 4 An exemplary full dynamic range for N-bit input video data is shown, and an exemplary focus range determined for the input video data, according to some embodiments. In Figure 4 the vertical axis represents N-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) code values in the input video data. The horizontal axis represents the dynamic range of luminance in the input video data, in this example, 0 cd / m 2 - 10000 cd / m 2 , where cd / m 2 (candela per square meter) is the SI unit of luminance. The non-SI term for this unit is "NIT". This curve represents an exemplary transfer function for the input video data. The focus range (2000 cd / m 2 - 4000 cd / m 2 in this example) represents the human visual system effective dynamic range for the current scene, sequence, frame, or region, as determined according to one or more characteristics (e.g., luminance, texture, etc.) of the respective video data. As can be seen, in this example, the focus range is represented by a ratio of N-bit code values (about 1 / 8). Note that different focus ranges can be determined for different scenes, sequences, frames, or regions within a video stream, as Figures 7A to 7C indicated.

[0057] Figure 5 An example is shown of mapping N-bit input video data within the determined focus range to generate C-bit video data within a codec, according to some embodiments. In Figure 5 the vertical axis represents C-bit (e.g., 12-bit, 10-bit, 8-bit, etc.) code values available in the codec. The horizontal axis shows the focus range (2000 cd / m 2 - 4000 cd / m 2The curve represents an exemplary transfer function used to represent focus range data within the codec. In some implementations, the transfer function used to represent video data in the codec may be the same as the transfer function used to represent input video data. However, in other implementations, a different transfer function (referred to as an internal or second-level transfer function) may be used to represent video data in the codec than the transfer function used to represent input video data (referred to as the first-level transfer function). This allows, for example, to represent the focus range of the video signal with higher precision within the codec compared to using the first-level transfer function.

[0058] Figure 6 An example is shown, according to some implementation schemes, of extending C-bit decoded video into the dynamic range of an HDR-enabled device to generate D-bit video data for that device. Figure 6 In the diagram, the vertical axis represents the data from the decoder (e.g., ...). Figure 3 The D-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) code values ​​in the video data output by the decoder 300 are shown. The horizontal axis represents the dynamic range of brightness supported by the display device to which the decoder outputs the D-bit video data; in this example, the dynamic range shown is 0 cd / m². 2 -10000cd / m 2 The curve represents an exemplary transfer function of the display device. After the encoded video signal is decoded by the decoder, the focus range used for the signal's internal representation in the codec is remapped from the C-bit representation to a wider dynamic range and the D-bit representation of the display device. It should be noted that different focus ranges can be encoded for different scenes, sequences, frames, or regions within the video stream, such as... Figures 7A to 7C As shown. In some embodiments, format metadata may be provided to a decoder in or with the encoded video stream, indicating the focus range used to encode each scene, sequence, frame, or region. In some embodiments, the metadata may also include one or more parameters for mapping video data to the bit depth of the codec. In some embodiments, the metadata may also include information about the internal transfer function used to represent the focus range data in the codec. This metadata can be used by the decoder to remap the focus range data from a C-bit representation to a wider dynamic range and a D-bit representation of the display device.

[0059] Figures 7A to 7C The diagram graphically illustrates how different focus ranges are applied to different parts of a video sequence or video frame according to an implementation of the adaptive transfer function method.

[0060] Figure 7ADifferent focus ranges are applied to different scenes or sequences of an input video according to some embodiments of the adaptive transfer function method are shown graphically. In this example, focus range 710A is dynamically determined and used for encoding frames 700A-700E, and focus range 710B is dynamically determined and used for encoding frames 700F-700G. Frames 700A-700E and frames 700F-700H may, for example, represent different scenes within a video, or may represent different sequences of frames within or across multiple scenes in a video.

[0061] Figure 7B Different focus ranges are applied to each frame of an input video according to some embodiments of the adaptive transfer function method are shown graphically. In this example, focus range 710C is dynamically determined and used for encoding frame 70OK, focus range 710D is dynamically determined and used for encoding frame 70OL, and focus range 710E is dynamically determined and used for encoding frame 70OK.

[0062] Figure 7C Different focus ranges are applied to different regions of a video frame according to some embodiments of the adaptive transfer function method are shown graphically. As shown in Figure 7A and Figure 7B In some embodiments, the adaptive transfer function method can be applied to a frame, scene, or sequence. However, in some embodiments, in addition or as an alternative, the adaptive transfer function method can be applied to two or more regions 702 within a frame, as shown in Figure 7CThe focus range 710F is dynamically determined and used to encode region 702A of frame 700L, the focus range 710G is dynamically determined and used to encode region 702B of frame 700L, and the focus range 710H is dynamically determined and used to encode region 702C of frame 700L. In various embodiments, the regions 702 to be encoded according to the adaptive transfer function method can be rectangular or can have other shapes, including but not limited to arbitrarily determined irregular shapes. In some embodiments, there can be a fixed number of regions 702 to be encoded according to the adaptive transfer function method in each frame. In some embodiments, the number of regions 702 to be encoded according to the adaptive transfer function method and / or the shape of the regions 702 can be determined for each frame, scene, or sequence. In some embodiments, the format metadata generated by the encoder and passed to the decoder can indicate the focus range, quantization parameter, and other information used to encode each region 702 in one or more frames 700. In embodiments in which a secondary transfer function or internal transfer function is used within the codec, this format metadata can contain information for converting from the secondary transfer function to the primary transfer function. In embodiments in which the number of regions 702 can vary, an indication of the actual number of regions 702 can be included in the metadata passed to the decoder. In some embodiments, the coordinates and shape information of the regions 702 can also be included in the metadata passed to the decoder.

[0063] In video content, the various parameters used in the adaptive transfer function method (e.g., focus range, quantization parameter, etc.) can be similar from one region or frame to the next in the input video being processed, especially within a scene. Thus, in some embodiments, the adaptive transfer function method can provide for intra (spatial) prediction and / or inter (temporal) prediction of one or more of these adaptive transfer function parameters in the encoder, for example, by utilizing the weighted prediction process in codecs such as the AVC and HEVC codecs. In inter (temporal) prediction, previously processed reference data from one or more temporally past or future frames can be used to predict one or more parameters for content in the current frame. In intra (spatial) prediction, data from one or more neighboring blocks or regions within a frame can be used to predict one or more parameters for content in the current block or region.

[0064] Figure 8 A high level flowchart of a video encoding method to apply the adaptive transfer function method to video input data and generate encoded video data according to some embodiments. As Figure 8As indicated at 800, N-bit (e.g., 12-bit, 14-bit, 16-bit, etc.) video data can be received for encoding. In some implementations, frames of the video stream are processed sequentially by an encoder. In some implementations, the input video frames are subdivided into encoders and processed by them based on pixel blocks (e.g., macroblocks, CUs, PUs, or CTUs).

[0065] like Figure 8 As indicated at 802, the focus range of the input video data can be determined. In some implementations, the focus range of each scene, sequence, or frame of the video input to the encoder can be dynamically determined. In some implementations, the focus range of each of two or more regions within a frame can be dynamically determined. In some implementations, the focus range can be determined based on one or more features (e.g., brightness, texture, etc.) of the current scene, sequence, frame, or region being processed. The focus range represents the effective dynamic range of the human visual system for image data (e.g., brightness) in the current scene, sequence, frame, or region. For example, if the brightness dynamic range in the input video data is 0 cd / m², the focus range can be determined based on these features. 2 -10000cd / m 2 Therefore, an exemplary focusing range based on specific features of various scenes, sequences, frames, or regions can be 2000 cd / m². 2 -4000cd / m 2 0cd / m 2 -1000cd / m 2 1000cd / m 2 -2500cd / m 2 etc. Figure 4 An exemplary focus range determined for N-bit input video data according to some implementation schemes is illustrated graphically.

[0066] like Figure 8 As noted at 804, N-bit video data within the focus range can be mapped to C-bit video data according to a transfer function. In some implementations, the input video data can be cropped according to the determined focus range, and then the cropped data values ​​can be appropriately mapped (e.g., quantized) to the C bits available in the encoder according to a transfer function used to represent the video data in the encoder. Figure 5According to some embodiments, an exemplary mapping of N-bit input video data to C-bit available in the encoder is illustrated graphically within an exemplary focus range according to a transfer function. In these embodiments, various transfer functions can be used to represent the N-bit input video data and C-bit video data in the encoder. Examples of transfer functions that can be used to represent video data in these embodiments include, but are not limited to, gamma-based power-law transfer functions, logarithmic transfer functions, and transfer functions based on human visual perception such as the PQ transfer function. In some embodiments, the transfer function used to represent the C-bit video data in the encoder may be the same as the transfer function used to represent the N-bit input video data. However, in other embodiments, the input video data can be represented according to a first-level transfer function, and different (second-level) transfer functions can be used to represent the video data in the encoder. The second-level transfer function can, for example, represent the video signal with higher precision within the encoder compared to the precision that may be achieved using a first-level transfer function.

[0067] like Figure 8 As shown in element 806, the adaptive transfer function processing executed at elements 800 to 804 will iterate as long as there is still N bits of input video data to be processed (e.g., scene, sequence, frame, or region). Figure 8 As indicated at 810, each unit of C-bit video data (e.g., scene, sequence, frame, or region) output by the adaptive transfer function method at elements 800 to 804 is input to and processed by the encoder. As previously mentioned, in some embodiments, the N-bit input video data may be subdivided into and processed by the adaptive transfer function module or component, and then passed to the encoder as C-bit video data according to pixel blocks (e.g., macroblocks, CUs, PUs, or CTUs).

[0068] like Figure 8 As noted at 810, one or more components of the encoder can process C-bit video data to generate encoded (and compressed) video data output (e.g., CAVLC output or CABAC output). In some embodiments, the encoder can encode the C-bit video input data according to a compressed video format such as H.264 / AVC or H.265 / HEVC. However, other encoding formats may also be used. In some embodiments, focus range, transfer function, quantization parameters, and other format information used to encode each scene, sequence, frame, or region can be embedded as metadata into the encoded output stream or otherwise signaled to one or more decoders. Figure 2The encoder 200 illustrates an exemplary encoder for processing C-bit video data to generate an encoded output stream. In some embodiments, the encoded output stream, including metadata, may be written to memory, for example, via direct memory access (DMA). In some embodiments, as an alternative or supplement to writing the output stream and metadata to memory, the encoded output stream and metadata may be sent directly to at least one decoder. The decoder may be implemented on the same or different devices and apparatuses as the encoder.

[0069] Figure 9 This is a high-level flowchart of a video decoding method according to some implementation schemes. The method decodes and expands encoded video data using an adaptive transfer function approach to generate video data in a display format. For example... Figure 9 As noted at 900, the decoder can obtain encoded data (e.g., data encoded and compressed using CAVLC or CABAC). The encoded data can be obtained, for example, read from memory, received from the encoder, or otherwise. Figure 9 As noted at 902, the decoder can decode the encoded data to generate decoded C-bit video data and format metadata. In some implementations, the decoder can decode the encoded data according to a compressed video format such as H.264 / AVC or H.265 / HEVC. However, other encoding / decoding formats may also be used. The format metadata extracted from the encoded data may include, for example, focus range, transfer function, quantization parameters, and other format information used for encoding each scene, sequence, frame, or region.

[0070] like Figure 9 As shown in element 904, as long as encoded data exists, the decoding performed at 900 and 902 can continue. Figure 9 As noted at 910, each unit of decoded data (e.g., each scene, sequence, frame, or region) decoded by the decoder can be output and processed according to the inverse adaptive transfer function method. At least a portion of the format metadata extracted from the encoded data can also be output to the inverse adaptive transfer function method.

[0071] like Figure 9 As noted at point 910, the inverse adaptive transfer function method dynamically extends decoded C-bit video data from the focus range to generate full dynamic range D-bit video data for output to target devices such as HDR-enabled displays. For a given scene, sequence, frame, or region, the extension can be dynamically performed based on the appropriate format metadata extracted from the encoded data by the decoder. Figure 6Graphically illustrates extending decoded video data into the full dynamic range of an HDR device according to some embodiments

[0072] Primary transfer function and secondary transfer function

[0073] Examples of transfer functions that can be used to represent video data in embodiments of the adaptive transfer function method as described herein can include, but are not limited to, gamma-based power-law transfer functions, logarithmic (log-based) transfer functions, and transfer functions based on human visual perception, such as the perceptual quantizer (PQ) transfer function proposed by Dolby Laboratories, Inc.

[0074] In some embodiments, the transfer function used to represent video data in the codec can be the same as the transfer function used to represent the input video data (referred to as the primary transfer function). In these embodiments, the input video data can be clipped according to the focus range determined for the scene, sequence, frame, or region, and then mapped (e.g., quantized) into the available bits of the codec according to the primary transfer function. The focus range, clipping, and mapping (e.g., quantization) parameters can be signaled to the decoder, for example, as metadata in the output encoded stream, so that the inverse adaptive transfer function method can be performed on the decoder to generate the full range video data for the device (e.g., an HDR-capable display).

[0075] However, in some embodiments, a different transfer function (referred to as the internal transfer function or secondary transfer function) can be used to represent the video data within the codec than the primary transfer function. In these embodiments, the input video data can be clipped to the focus range determined for the scene, sequence, frame, or region, and then mapped, scaled, or quantized into the available bits of the codec according to the secondary transfer function. The secondary transfer function can allow, for example, a higher precision in representing the video signal within the codec than can be possible with the primary transfer function. In these embodiments, in addition to the focus range, clipping, and mapping (e.g., quantization) parameters, the decoder can also be signaled information about how to convert the video data from the secondary transfer function to the primary transfer function, for example, as metadata in the output stream. This information can include, but is not limited to, the type of secondary transfer function (e.g., power-law gamma transfer function, log transfer function, PQ transfer function, etc.) and one or more control parameters of the transfer function. In some embodiments, information describing the primary transfer function can also be signaled.

[0076] In some implementations of the adaptive transfer function method, the internal or second-level transfer function used within the codec can be a discontinuous transfer function representation, wherein one or more portions of the original transfer function or first-level transfer function of the internal representation can be retained, while the remainder of the internal representation uses a different transfer function. In these implementations, for example, a table can be used to precisely describe the area in the output stream, and signals can be sent to notify one or more regions of the retained first-level transfer function. Signals can also be sent to notify information about different transfer functions.

[0077] like Figures 7A to 7C As shown, implementations of an adaptive transfer function method can be performed intraframe at the frame level or region level. Therefore, adjacent or neighboring encoded and decoded video frames or regions within a frame can have significantly different transfer function representations, which can negatively impact coding (e.g., motion estimation, motion compensation, and reconstruction) because the transfer function representation of the current frame or region being encoded can be significantly different from that of adjacent frames or regions available for the coding process. However, in some implementations, weighted prediction, provided in codecs using coding formats such as Advanced Video Coding (AVC) and High Efficiency Video Coding (HEVC), can be used to significantly improve coding efficiency. In these implementations, weighting parameters can be provided for frames or regions that can be used in the weighted prediction process to adjust for differences in transfer function representations. These weighting parameters can be signaled, for example, by including them in frame or block metadata. For example, in an AVC or HEVC encoder, the weighting parameters can be signaled in the slice header metadata and can be selected by changing the reference index within a macroblock in AVC or within a prediction unit (PU) in HEVC. In some implementations, weighting information may be explicitly signaled on a block-by-block basis, which may or may not be related to the slice header portion. For example, in some implementations, the first-level weights or sovereign weights may be signaled in the slice header and have a δ or difference that can be used to adjust the first-level weights signaled at the block level. In some implementations, the weighting parameters may also include color weighting information. In some implementations, these weighting parameters may be used for both intra-frame and inter-frame prediction. For example, in intra-frame prediction, as an alternative or supplement to using previously processed neighboring or adjacent data (e.g., pixels or pixel blocks) as is, additional weighting and offset parameters may be provided to adjust the prediction based on the potentially different transfer function characteristics of neighboring or adjacent samples.

[0078] In some implementations in which an internal transfer function or a secondary transfer function is used to represent data within a codec, in addition to the weighting parameters, information about the secondary transfer function can be signaled in the codec for transfer function prediction. In some implementations, the secondary transfer function information can be signaled, for example, in a slice header using a transfer function table, such that each reference data index is associated with one or more transfer function adjustment parameters. In some implementations, alternatively or in addition, the secondary transfer function information can be signaled at a block level. In implementations using a piecewise transfer function representation, multiple weighting parameters (e.g., parameters having different impact on different luminance values or levels) can be signaled in the slice header and / or at a block level.

[0079] In some implementations in which an internal or secondary transfer function is used and adjusted to represent data within a codec, the internal transfer function and adjustment for a scene, sequence, frame, or region can be determined dynamically. In some implementations, the internal transfer function can be determined based on one or more characteristics of the current video frame or one or more regions of the frame. The characteristics can include, but are not limited to, minimum luminance and peak luminance, motion, texture, color, histogram pattern, percentile concentration, etc. In some implementations, determining the internal transfer function for the current frame can also utilize past and / or future frames as reference frames. In some implementations, for example, determining the internal transfer function can utilize a window of one, two, or more reference frames occurring before and / or after the current frame in the video or video sequence. Using window-based information to determine the internal transfer function can result in changing characteristics of each frame, and can help avoid or smooth out too abrupt jumps or discontinuities in the internal transfer function adjustment, which can otherwise adversely affect encoding. A smooth, window-based adaptive transfer function approach can provide better transfer function selection for encoding purposes and for final signal reconstruction at the full transfer function range / required transfer function range. In some implementations, a window-based adaptive transfer function approach can include estimating the peak luminance and minimum luminance of the frames within the window to determine the internal transfer function. In some implementations, the luminance histogram, histogram pattern, concentration of values, and how to adjust them, and partitioning of area or regions within the window can be estimated by the window-based approach for determining the internal transfer function. In some implementations, information about the human visual system can be utilized in this selection to achieve improved performance.

[0080] In some embodiments, the decision windows used in a window-based adaptive transfer function method can not overlap, and are determined individually for each frame or region. In some embodiments, there can be overlap at the window boundaries, while the boundary frames outside the current window are considered in the estimation but not adjusted for themselves. In some embodiments, multiple windows can influence the boundary frames through an adaptive process, such as a weighted average of the internal transfer function information based on the window distance, which can help ensure smooth transitions between neighboring windows. In some embodiments, the window-based decision process for the internal transfer function can be based on a "running window" approach, in which the internal transfer function for each frame is determined based on the characteristics of the frame itself and one or more past and / or future frames and / or windows in time. In some embodiments, the internal transfer function for past and / or future frames and / or windows can also be utilized to determine the internal transfer function for the current frame. In some embodiments, a multi-pass approach can be implemented, in which the behavior and performance of the pre-selected internal transfer function in a previous pass is considered to tune the subsequent internal transfer function, enabling higher coding efficiency and better rendering of a given final transfer function target.

[0081] In some embodiments, the internal transfer function (e.g., range, type, bit depth, etc.) can be determined and adjusted based on the characteristics of the video data being encoded. In some embodiments, the compression capabilities or characteristics of the particular transfer function of the signal of interest can also be considered in selecting and / or adjusting the internal transfer function. In some embodiments, in addition or alternatively, the internal transfer function can be determined and adjusted based on one or more target displays and their characteristics or limitations. For example, if it is known that the current target display has a limited dynamic range compared to the dynamic range supported by the primary transfer function, it can be pointless to include values in the compressed signal that are outside the target display range. Instead, an internal transfer function representation that best fits the dynamic range that the signal can be provided to the display can be determined, which can allow for a better compressed representation of the signal for that particular display capability. In some embodiments, if multiple displays are to be supported by the codec, a dynamic range can be selected for the display, and an internal transfer function representation can be determined that provides a best fit to the selected dynamic range. For example, the most capable display (e.g., with the highest dynamic range) can be selected, and its dynamic range can be used to adjust the internal transfer function to generate video output for all displays. As another example, the dynamic range can be selected based on a pricing model in which one or more characteristics of the display (e.g., dynamic range) can be weighted (e.g., based on a determined or indicated ordering of the displays according to importance or other factors).

[0082] Embodiments of the adaptive transfer function method for video encoding and decoding systems or codecs as described herein can provide functionality to signal (from the encoder) and support (from the encoder and decoder) extended dynamic range or high dynamic range, while keeping the complexity (e.g., bit depth) of the codec within a reasonable range. This can be achieved by dynamically determining a focus range on a region, frame, scene, or sequence level, clipping input data to the focus range, and mapping the clipped data from the bit depth of the input video to the bit depth of the codec, while signaling appropriate parameters (e.g., focus range, quantization parameters, etc.) from the encoder to the decoder on a region, frame, scene, or sequence level, so that the dynamic range can be extended to the full range of an HDR-capable display. Additionally, in some embodiments, when different transfer function representations are used to encode different frames, appropriate parameters (e.g., weights, etc.) for performing weighting parameters can be signaled.

[0083] Exemplary devices and apparatuses

[0084] Figures 10 to 14 Non-limiting examples of devices and apparatuses in which or with which embodiments of the various digital video or image processing and display methods and apparatuses as described herein can be implemented are shown. Figure 10 An exemplary SOC is shown, and Figure 11 An exemplary device implementing a SOC is shown. Figure 12 An exemplary computer system that can implement the methods and apparatuses described herein is shown. Figure 13 And Figure 14 An exemplary multi-function device that can implement the methods and apparatuses described herein is shown.

[0085] An exemplary system on a chip (SOC)

[0086] Referring now to Figure 10which shows a block diagram of one embodiment of a system on a chip (SOC) 8000 that can be used in embodiments. The SOC 8000 is illustrated as coupled to a memory 8800. As the name implies, the components of the SOC 8000 can be integrated onto a single semiconductor substrate as an integrated circuit “chip.” In some embodiments, the components can be implemented on two or more discrete chips in a system. However, the SOC 8000 will be used as an example herein. In the illustrated embodiment, the components of the SOC 8000 include a central processing unit (CPU) complex 8020, on-chip peripheral components 8040A-8040C (more simply referred to as “peripherals”), a memory controller (MC) 8030, and a communication fabric 8010. The components 8020, 8030, 8040A-8040C can all be coupled to the communication fabric 8010. The memory controller 8030 can be coupled to the memory 8800 during use, and the peripherals 8040B can be coupled to external interfaces 8900 during use. In the illustrated embodiment, the CPU complex 8020 includes one or more processors (P) 8024 and a level two (L2) cache 8022.

[0087] The peripherals 8040A-8040B can be any collection of additional hardware functionality included in the SOC 8000. For example, the peripherals 8040A-8040B can include video peripherals such as an image signal processor configured to process image capture data from a camera or other image sensor, a display controller configured to display video data on one or more display devices, a graphics processing unit (GPU), a video encoder / decoder or codec, a scaler, a rotator, a mixer, etc. The peripherals can include audio peripherals such as a microphone, a speaker, an interface to a microphone and a speaker, an audio processor, a digital signal processor, a mixer, etc. The peripherals can include peripheral interface controllers (e.g., peripheral 8040B) for various interfaces 8900 external to the SOC 8000, including interfaces such as universal serial bus (USB) ports, peripheral component interconnect (PCI) ports (including PCI express (PCIe) ports), serial ports, parallel ports, etc. The peripherals can include networking peripherals such as a media access controller (MAC). Any collection of hardware can be included.

[0088] The CPU complex 8020 can include one or more CPU processors 8024 that serve as the CPU(s) of the SOC 8000. The CPU(s) of the system include one or more processors that execute system- main control software such as an operating system. Generally, software executed by the CPU(s) during use can control other components of the system to implement the desired functionality of the system. The processor(s) 8024 can also execute other software such as application programs. The application programs can provide functionality to a user and can rely on the operating system for lower-level device control. Thus, the processor(s) 8024 can also be referred to as application processors. The CPU complex 8020 can further include other hardware such as an L2 cache 8022 and / or interfaces to other components of the system (e.g., to the communication fabric 8010). Generally, a processor can include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor. Instructions and data operated on by the processor in response to executing instructions can generally be stored in the memory 8800, although certain instructions can be defined to also access peripherals directly by the processor. The processor(s) can encompass processor cores implemented on integrated circuits with other components as a system on a chip (SOC 8000) or other levels of integration. The processor(s) can further include discrete microprocessors, processor cores, and / or microprocessors integrated into multi-chip module implementations, processors implemented as multiple integrated circuits, etc.

[0089] The memory controller 8030 can generally include circuitry to receive memory operations from other components of the SOC 8000 and to access the memory 8800 to complete the memory operations. The memory controller 8030 can be configured to access any type of memory 8800. For example, the memory 8800 can be static random access memory (SRAM), dynamic RAM (DRAM), such as synchronous DRAM (SDRAM) including double data rate (DDR, DDR2, DDR3, etc.) DRAM. Low power / mobile versions of DDR DRAM can be supported (e.g., LPDDR, mDDR, etc.). The memory controller 8030 can include a memory operation queue to order (and possibly reorder) the operations and present them to the memory 8800. The memory controller 8030 can also include data buffers to store write data waiting to be written to memory and read data waiting to be returned to the source of the memory operation. In some embodiments, the memory controller 8030 can include a memory cache to store recently accessed memory data. For example, in SOC implementations, the memory cache can reduce power consumption in the SOC by avoiding re-accessing data from the memory 8800 if it is expected to be accessed again soon. In some cases, the memory cache can also be referred to as a system cache, which is distinct from a private cache, such as the L2 cache 8022 or a cache in the processor 8024, which only serves certain components. Further, in some embodiments, a system cache need not be located within the memory controller 8030.

[0090] In one embodiment, the memory 8800 can be packaged with the SOC 8000 in a chip-on-chip configuration or a package-on-package configuration. A multi-chip module configuration of the SOC 8000 and the memory 8800 can also be used. Such configurations can be relatively more secure (in terms of data observability) than transmissions to other components of the system, e.g., to the endpoints 16A-B. Thus, protected data can reside in the memory 8800 unencrypted, while protected data can be encrypted for exchange between the SOC 8000 and external endpoints.

[0091] The communication fabric 8010 can be any communication interconnect and protocol for communicating between components of the SOC 8000. The communication fabric 8010 can be bus-based, including a shared bus configuration, a crossbar configuration, and a hierarchical bus with bridges. The communication fabric 8010 can also be packet-based and can be hierarchical with bridges, crossbar, point-to-point, or other interconnects.

[0092] Note that the number of components of the SOC 8000 (and Figure 10The number of sub-components of the components shown (such as within CPU complex 8020) can be different in different embodiments. There can be more or fewer of each component / sub-component than shown. Figure 10 more or fewer of each component / sub-component than shown.

[0093] Figure 11 is a block diagram of one embodiment of a system 9000 that includes at least one instance of SOC 8000 coupled to external memory 8800 and one or more external peripherals 9020. A power management unit (PMU) 9010 is provided that supplies a supply voltage to SOC 8000, as well as one or more supply voltages to memory 8800 and / or peripherals 9020. In some embodiments, more than one instance of SOC 8000 can be included (as well as more than one memory 8800).

[0094] Depending on the type of system 9000, the peripherals 9020 can include any desired circuitry. For example, in one embodiment, the system 9000 can be a mobile device (e.g., a personal digital assistant (PDA), a smartphone, etc.) and the peripherals 9020 can include devices for various types of wireless communication, such as wifi, Bluetooth, cellular, global positioning system, etc. The peripherals 9020 can also include additional storage, including RAM storage, solid state storage, or disk storage. The peripherals 9020 can include user interface devices, such as a display screen including a touch display screen or a multi-touch display screen, a keyboard or other input device, a microphone, a speaker, etc. In other embodiments, the system 9000 can be any type of computing system (e.g., a desktop personal computer, a laptop, a workstation, a network set-top box, etc.).

[0095] The external storage 8800 can include any type of memory. For example, the external storage 8800 can be SRAM, dynamic RAM (DRAM) such as synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM, RAMBUS DRAM, low power versions of DDR DRAM (e.g., LPDDR, mDDR, etc.), etc. The external storage 8800 can include one or more memory modules to which memory devices can be mounted, such as single inline memory modules (SIMMs), dual inline memory modules (DIMMs), etc. Alternatively, the external storage 8800 can include one or more memory devices mounted on SOC 8000 in a chip-on-chip configuration or a package-on-package implementation.

[0096] Multifunction device example

[0097] Figure 13A block diagram of a portable multifunctional device according to some embodiments is shown. In some embodiments, the device is a portable communication device such as a mobile telephone that also contains other functions, such as PDA, camera, video capture and / or playback, and / or music player functions. Exemplary embodiments of the portable multifunctional device include, without limitation, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Other portable electronic devices, such as laptops, mobile telephones, smartphones, tablet or slate computers, with touch-sensitive surfaces (e.g., touch screen displays and / or touch pads) are, optionally, used. devices, iPod devices, and devices. Other portable electronic devices, such as laptops, mobile telephones, smartphones, tablet or slate computers, with touch-sensitive surfaces (e.g., touch screen displays and / or touch pads) are, optionally, used.

[0098] In the discussion that follows, an electronic device that includes a display and a touch-sensitive surface is described. It should be understood, however, that the electronic device can include one or more other physical user-interface devices, such as a physical keyboard, a mouse and / or a joystick.

[0099] The device typically supports a variety of applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a game application, a telephone application, a video conferencing application, an e-mail application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0100] Various applications 1126 can be executed on the electronic device 1101, including one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a game application, a telephone application, a video conferencing application, an e-mail application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0101] Device 2100 can include a memory 2102 (which can include one or more computer readable storage media), a memory controller 2122, one or more processing units (CPU(s)) 2120, a peripherals interface 2118, RF circuitry 2108, audio circuitry 2110, a speaker 2111, a touch- sensitive display system 2112, a microphone 2113, an input / output (I / O) subsystem 2106, other input control devices 2116, and external port 2124. Device 2100 can include one or more optical sensors or cameras 2164. These components can communicate over one or more communication buses or signal lines 2103.

[0102] It should be appreciated that device 2100 is only one example of a portable multifunctional device, and that device 2100 can have more or fewer components than shown, can combine two or more components, or a can have a different configuration or arrangement of the components. Figure 13 The various components shown can be implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.

[0103] Memory 2102 can include high-speed random access memory and can also include nonvolatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other nonvolatile solid-state storage devices. Access to memory 2102 by other components of device 2100, such as CPU(s) 2120 and peripherals interface 2118, can be controlled by a memory controller 2122.

[0104] Peripherals interface 2118 can be used to couple input peripherals and output peripherals to CPU(s) 2120 and memory 2102. One or more processors 2120 run or execute various software programs and / or sets of instructions stored in memory 2102 to perform various functions and processes data for device 2100.

[0105] In some embodiments, peripherals interface 2118, CPU(s) 2120, and memory controller 2122 can be implemented on a single chip, such as chip 2104. In some other embodiments, they can be implemented on separate chips.

[0106] RF (radio frequency) circuitry 2108 receives and sends RF signals, also called electromagnetic signals. RF circuitry 2108 converts electrical signals to / from electromagnetic signals and communicates with communications networks and other communications devices via the electromagnetic signals. The RF circuitry 2108 can include well-known circuitry for detecting and

[0107] Audio circuitry 2110, speaker 2111, and microphone 2113 provide an audio interface between a user and device 2100. Audio circuitry 2110 receives audio data from peripherals interface 2118, converts the audio data to an electrical signal, and transmits the electrical signal to speaker 2111. Speaker 2111 converts the electrical signal to sound waves that are audible to the human ear. Audio circuitry 2110 also receives electrical signals converted by microphone 2113 from sound waves. Audio circuitry 2110 converts the electrical signals to audio data and transmits the audio data to peripherals interface 2118 for processing. Audio data can be retrieved from and / or transmitted to memory 2102 and / or RF circuitry 2108 by peripherals interface 2118 in some embodiments. In some embodiments, audio circuitry 2110 also includes a headset jack. The headset jack provides an interface between audio circuitry 2110 and removable audio input / output peripherals, such as output-only headphones or a headset with both output (e.g., stereo

[0108] I / O subsystem 2106 couples input / output peripherals on device 2100, such as touch screen 2112 and other input control devices 2116, to peripherals interface 2118. I / O subsystem 2106 can include display controller 2156 and one or more input controllers 2160 for other input control devices 2116. The one or more input controllers 2160 receive / send electrical signals from / to other input control devices 2116. The other input control devices 2116 can include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, and so forth. In some alternative embodiments, input controller(s) 2160 can be coupled to any one or combination of the following: a keyboard, infrared port, USB port, and a pointer device such as a mouse. The one or more buttons can include an up / down button for

[0109] Touch-sensitive display 2112 provides an input interface and an output interface between the device and a user. Display controller 2156 receives and / or sends electrical signals from / to touch screen 2112. Touch screen 2112 displays visual output to the user. The visual output can include graphics, text, icons, video, and any combination thereof (collectively termed "graphics"). In some embodiments, some or all of the visual output can correspond to user-interface objects.

[0110] The touch screen 2112 has a touch-sensitive surface, sensor, or set of sensors that accepts input from the user based on haptic and / or tactile contact. The touch screen 2112 and the display controller 2156 (along with any associated modules and / or sets of instructions in memory 2102) detect contact (and any movement or breaking of the contact) on the touch screen 2112 and convert the detected contact into interaction with user-interface objects (e.g., one or more soft keys, icons, web pages, or images) that are displayed on the touch screen 2112. In an exemplary embodiment, a point of contact between the touch screen 2112 and the user corresponds to a finger of the user.

[0111] The touch screen 2112 can use LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies can be used in other embodiments. The touch screen 2112 and the display controller 2156 can detect contact and any movement or breaking of the same using any of a plurality of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more contact points with the touch screen 2112. In an exemplary embodiment, the touch screen 2112 uses projected mutual capacitance sensing technology, such as that found in the iPhone® and iPod touch® from Apple Inc. of Cupertino, California. iPod and technologies found in the iPhone® and iPod touch® from Apple Inc. of Cupertino, California.

[0112] The touch screen 2112 can have a video resolution in excess of 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user can make contact with the touch screen 2112 using any suitable object or appendage, such as a stylus, finger, and the like. In some embodiments, the user interface is designed to work primarily with finger-based contacts and gestures, which can be less precise than stylus-based input due to the larger area of contact of a finger on the touch screen. In some embodiments, the device translates the rough finger-based input into a precise pointer / cursor position or command for performing the actions desired by the user.

[0113] In some embodiments, in addition to the touch screen 2112, the device 2100 can include a touchpad (not shown) for activating or deactivating particular functions. In some embodiments, the touchpad is a touch-sensitive area for the device that is separate from the touch screen 2112 and that does not display visual output. The touchpad can be a touch-sensitive surface that is separate from the touch screen 2112 or an extension of the touch-sensitive surface formed by the touch screen.

[0114] Device 2100 also includes power system 2162 for powering the various components of device 2100. Power system 2162 can include a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)) and any other components associated with the generation, management and distribution of power for device 2100.

[0115] Device 2100 can also include one or more optical sensors or cameras 2164. Figure 13 An optical sensor coupled to optical sensor controller 2158 in I / O subsystem 2106 is shown. The optical sensor 2164 can, for example, include a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) optical transducer or optical sensor. The optical sensor 2164 receives light from the environment, projected through one or more lenses, and converts the light to data representing an image. In conjunction with an imaging module 2143 (also called a camera module), the optical sensor 2164 can capture still images or video sequences. In some embodiments, at least one optical sensor can be located on the front of device 2100, opposite the touch screen display 2112 on the front of the device. In some embodiments, the touch screen display can be used as a viewfinder for still and / or video image acquisition. In some embodiments, the optical sensor can be located on the front of the device, as an alternative to, or in addition to, the optical sensor located on the back of the device.

[0116] Device 2100 can also include one or more proximity sensors 2166. Figure 13 The proximity sensor 2166 is shown coupled to the peripherals interface 2118. Alternatively, the proximity sensor 2166 can be coupled to the input controller 2160 in the I / O subsystem 2106. In some embodiments, the proximity sensor turns off and disables the touch screen 2112 when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).

[0117] Device 2100 can also include one or more orientation sensors 2168. In some embodiments, the one or more orientation sensors include one or more accelerometers (e.g., one or more linear accelerometers and / or one or more rotational accelerometers). In some embodiments, the one or more orientation sensors include one or more gyroscopes. In some embodiments, the one or more orientation sensors include one or more magnetometers. In some embodiments, the one or more orientation sensors include one or more of a Global Positioning System (GPS), a Global Navigation Satellite System (GLONASS), and / or other global navigation system receiver. The GPS, GLONASS, and / or other global navigation system receiver can be used to obtain information about the location and orientation (e.g., portrait or landscape) of device 2100. In some embodiments, the one or more orientation sensors include any combination of orientation sensors / rotation sensors. Figure 13 The one or more orientation sensors 2168 are shown coupled to peripherals interface 2118. Alternatively, the one or more orientation sensors 2168 can be coupled to input controller 2160 in I / O subsystem 2106. In some embodiments, information is displayed on touch screen display in a portrait view or a landscape view based on analysis of data received from the one or more orientation sensors.

[0118] In some embodiments, device 2100 can also include one or more other sensors (not shown), including but not limited to an ambient light sensor and a motion detector. These sensors can be coupled to peripherals interface 2118, or can be coupled to input controller 2160 in I / O subsystem 2106. For example, in some embodiments, device 2100 can include at least one forward-facing (away from the user) light sensor and at least one rearward-facing (toward the user) light sensor that can be used to collect ambient light measurements from the environment in which device 2100 is located for use in video and image capture, processing, and display applications.

[0119] In some embodiments, software components stored in memory 2102 include operating system 2126, communication module 2128, contact / motion module 2130 (or set of instructions), graphics module 2132, text input module 2134, Global Positioning System (GPS) module 2135, and applications 2136. Furthermore, in some embodiments, memory 2102 stores device / global internal state 2157. Device / global internal state 2157 includes one or more of: active application state, indicating which applications, if any, are currently active; display state, indicating what, if anything, is currently displayed on touch screen display 2112; sensor state, including information on

[0120] Operating system 2126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates wireless communications between various hardware and software components.

[0121] Communication module 2128 facilitates communication with other devices over one or more external ports 2124 and also includes various software components for handling data received by RF circuitry 2108 and / or external port 2124. External port 2124 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted for coupling directly to other devices or indirectly over a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, or compatible with, the 30-pin connector used on iPod (Apple Inc.'s trademark) devices.

[0122] Contact / motion module 2130 can detect contact with touch screen 2112 (in conjunction with display controller 2156) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). Contact / motion module 2130 includes various software components for performing various operations related to detection of contact, such as determining if contact has occurred (e.g., detecting a finger-down event), determining if at least

[0123] Contact / motion module 2130 can detect gestures on touch-sensitive surface. Different gestures on the touch-sensitive surface have different contact patterns. Thus, a gesture can be detected by detecting a particular contact pattern. For example, detecting a one-finger tap gesture includes detecting a finger-down event followed by a finger-up event at the same location (or substantially the same location) as the finger-down event (e.g., at an icon location). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger-down event followed by one or more finger-drag events, and then followed by a finger-up event.

[0124] Graphics module 2132 includes various software components for rendering and displaying graphics on touch screen 2112 or other display, including components for changing the intensity of graphics that are displayed. As used herein, the term "graphics" includes any object that can be displayed to a user, including without limitation text, web pages, icons (such as user-interface objects including soft keys), digital images, videos, animations and the like.

[0125] In some embodiments, graphics module 2132 stores data representing graphics to be used. Each graphic can be assigned a corresponding code. Graphics module 2132 receives, from applications etc., one or more codes specifying graphics to be displayed, and then generates screen image data for a

[0126] Text input module 2134, which can be a component of graphics module 2132, provides soft keyboard for entering text in various applications.

[0127] GPS module 2135 determines the location of the device and provides this information for use in various applications (e.g., to telephone module 2138 for use in

[0128] Applications 2136 can also include one or more modules (or sets of instructions) for implementing the functionality of the device, such as the following modules (or sets of instructions):

[0129] • a telephone module 2138;

[0130] • a video conference module 2139;

[0131] • a camera module 2143 for still image and / or video image capture;

[0132] • an image management module 2144;

[0133] • a browser module 2147;

[0134] • a search module 2151;

[0135] • a video and music player module 2152, which can be made up of any

[0136] • an online video module 2155.

[0137] Examples of other applications 2136 that can be stored on the storage memory 2102 include, but are not limited to, other word processing applications, other image editing applications, a drawing application, a presentation application, a communication / social media application, a map application, a JAVA-enabled application, an encryption application, a digital rights management application, a voice recognition and / or a voice replication application.

[0138] In conjunction with RF circuitry 2108, audio circuitry 2110, speaker 2111, microphone 2113, touch screen 2112, display controller 2156, contact module 2130, graphics module 2132, and text input module 2134, telephone module 2138 can be used to enter a sequence of characters corresponding to a telephone number, access one or more telephone numbers in the address book, modify a telephone number from the address book, dial a respective telephone number, conduct a conversation, and disconnect or hang up when the conversation is completed. As described above, the wireless communication can use any of a variety of communications standards, protocols, and technologies.

[0139] In conjunction with RF circuitry 2108, audio circuitry 2110, speaker 2111, microphone 2113, touch screen 2112, display controller 2156, optical sensor 2164, optical sensor controller 2158, contact / motion module 2130, graphics module 2132, text input module 2134, and telephone module 2138, video conference module 2139 includes executable instructions to initiate, conduct, and terminate a video conference between a user and one or more other participants in accordance with user instructions.

[0140] In conjunction with touch screen 2112, display controller 2156, one or more optical sensors 2164, optical sensor controller 2158, contact / motion module 2130, graphics module 2132, and image management module 2144, camera module 2143 includes executable instructions to capture still images or video (including a video stream) and store them into memory 2102, modify characteristics of a still image or video, or delete a still image or video from memory 2102.

[0141] In conjunction with touch screen 2112, display controller 2156, contact / motion module 2130, graphics module 2132, text input module 2134, and camera module 2143, image management module 2144 includes executable instructions to arrange, modify (e.g., edit), or otherwise manipulate, label, delete, present (e.g., in a digital slide show or album), and store still and / or video images.

[0142] In conjunction with RF circuitry 2108, touch screen 2112, display system controller 2156, contact / motion module 2130, graphics module 2132, and text input module 2134, browser module 2147 includes executable instructions to browse, using user inputs, the Internet, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as linking to and displaying of attachments and other files linked to web pages.

[0143] In conjunction with touch screen 2112, display system controller 2156, contact / motion module 2130, graphics module 2132, and text input module 2134, search module 2151 includes executable instructions to search for text, music, sound, image, video, and / or other files in memory 2102 that match one or more search criteria (e.g., one or more search terms).

[0144] Incorporating touchscreen 2112, display system controller 2156, contact / motion module 2130, graphics module 2132, audio circuitry 2110, speaker 2111, RF circuitry 2108, and browser module 2147, the video and music player module 2152 includes executable instructions allowing users to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touchscreen 2112 or on an external display connected via external port 2124). In some embodiments, device 2100 may include the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0145] Incorporating a touchscreen 2112, a display system controller 2156, a touch / motion module 2130, a graphics module 2132, an audio circuit 2110, a speaker 2111, an RF circuit 2108, a text input module 2134, and a browser module 2147, the online video module 2155 includes instructions that allow users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touchscreen or on an external display connected via external port 2124), and otherwise manage online video in one or more video formats such as H.264 / AVC or H.265 / HEVC.

[0146] Each module and application identified above corresponds to a set of executable instructions for performing one or more of the functions described above and the methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (i.e., instruction sets) need not be implemented as standalone software programs, processes, or modules; therefore, various subsets of these modules can be combined or otherwise rearranged in various embodiments. In some embodiments, memory 2102 may store a subset of the modules and data structures identified above. Furthermore, memory 2102 may store additional modules and data structures not described above.

[0147] In some implementations, device 2100 is the only device that performs operations of a set of predefined functions on the device via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for the operation of device 2100, the number of physical input control devices (such as push-buttons, dial pads, etc.) on device 2100 can be reduced.

[0148] The set of predefined functions that can be uniquely performed through the touch screen and / or the touch pad includes navigation between user interfaces. In some embodiments, the touch pad, when touched by a user, navigates the device 2100 from any user interface that can be displayed on the device 2100 to a main, home, or root menu. In such embodiments, the touch pad can be referred to as a "menu button." In some other embodiments, the menu button can be a physical push button or other physical input control device, instead of a touch pad.

[0149] Figure 14 A portable multifunction device 2100 with a touch screen 2112 in accordance with some embodiments is shown. The touch screen can display one or more graphical user interfaces (UIs) 2200. In some embodiments, a user can interact with the touch screen, for example, by making postures on the graphical objects displayed on the touch screen 2112 with one or more fingers 2202 (not necessarily drawn to scale in the figure) or with one or more styluses 2203 (not necessarily drawn to scale in the figure).

[0150] The device 2100 can also include one or more physical buttons, such as "home" or menu buttons 2204. As described previously, the menu button 2204 can be used to navigate to any application 2136 in a set of applications that can be executed on the device 2100. Alternatively, the menu button can be implemented as a soft key in a GUI displayed by the touch screen 2112 in some embodiments.

[0151] In one embodiment, the device 2100 includes the touch screen 2112, the home button or the menu button 2204, a push button 2206 for the on / off control of the device and the lock control of the device, one or more volume adjustment buttons 2208, a user identity module (SIM) card slot 2210, a headset jack 2212, and a dock / charging external port 2124. The push button 2206 can be used to turn the power on / off on the device by depressing the button and holding the button in the depressed state for a predetermined time interval; to lock the device by depressing the button and releasing the button before the predefined time interval has elapsed; and / or to unlock or initiate the unlock process of the device. In an alternative embodiment, the device 2100 can also accept verbal input for activation or deactivation of some functions through the microphone 2113.

[0152] Device 2100 can also include one or more cameras 2164. Camera(s) 2164 can include, for example, a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) photogate or photosensor. Camera(s) 2164 receive light projected through one or more lenses from an environment and convert the light into data representing an image or video frame. In some embodiments, at least one camera 2164 can be located on the back of device 2100 opposite the touchscreen display 2112 on the front of the device. In some embodiments, alternatively or additionally, at least one camera 2164 can be located on the front of a device with touchscreen display 2112, for example so that a user can view other video conference participants on touchscreen display 2112 while obtaining an image of the user for a video conference. In some embodiments, at least one camera 2164 can be located on the front of device 2100 and at least one camera 2164 can be located on the back of device 2100. In some embodiments, touchscreen display 2112 can be used as a viewfinder and / or user interface for still image and / or video sequence capture applications.

[0153] Device 2100 can include video and image processing hardware and / or software, including but not limited to video encoding and / or decoding components, codecs, modules, or pipelines that can be used to capture, process, convert, compress, decompress, store, modify, transmit, display, otherwise manage and manipulate still images and / or video frames or video sequences captured or otherwise acquired (e.g., via network interface) via camera(s) 2164. In some embodiments, device 2100 can also include one or more light sensors or other sensors that can be used to collect environmental light measurements or other measurements from the environment in which device 2100 is located for video and image capture, processing, and display use.

[0154] Exemplary computer system

[0155] Figure 12 An exemplary computer system 2900 that can be configured to perform any or all of the embodiments described above is shown. In different embodiments, computer system 2900 can be any of a variety of types of devices including, but not limited to: a personal computer system, desktop computer, laptop, notebook, tablet, slate, slate, a tablet, a set-top box, a mobile device, a consumer device, an application server, a storage device, a video recording device, a peripheral device such as a switch, modem, router, or generally any type of computing or electronic device.

[0156] The various embodiments as described herein can be executed within one or more computer systems 2900 that can interact with various other devices. Note that although specific embodiments can be illustrated and described herein, it is but one implementation. Many embodiments can be made and applied without departing from the scope and spirit of claimed subject matter. Thus, the specification and drawings are to be regarded as illustrative in nature and not as restrictive. Figures 1 to 11 Any components, actions, or functions described herein can be implemented within one or more computers configured to Figure 12 The computer system 2900, as shown in the illustrated embodiment, includes one or more processors 2910 coupled to a system memory 2920 via an input / output (I / O) interface 2930. The computer system 2900 further includes a network interface 2940 coupled to the I / O interface 2930, and one or more input / output devices or components 2950, such as a cursor control 2960, a keyboard 2970, one or more displays 2980, one or more cameras 2990, and one or more sensors 2992, including but not limited to light sensors and motion detectors. In some cases, it is contemplated that embodiments can be implemented using a single instance of computer system 2900 while in other embodiments multiple such systems, or multiple nodes making up a computer system 2900, can be configured to host different portions or instances of embodiments. For example, in one embodiment some elements can be implemented via one or more nodes of computer system 2900 that are distinct from those nodes used to implement other elements.

[0157] In various embodiments, computer system 2900 can be a uniprocessor system including one processor 2910, or a multiprocessor system including several processors 2910 (e.g., two, four, eight, or other suitable number). Processors 2910 can be any suitable processors capable of executing instructions. For example, in various embodiments, processors 2910 can be general- purpose or embedded processors implementing any of a variety of instruction set architectures (ISAs), such as the x86, PowerPC, SPARC, or MIPS ISAs, or any other suitable ISA. In multiprocessor systems, each of the processors 2910 can typically, but is not required to, implement the same ISA.

[0158] The system memory 2920 can be configured to store program instructions 2922 and / or data accessible by the processor 2910. In various embodiments, the system memory 2920 can be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), nonvolatile / Flash memory, or any other type of memory. In the illustrated embodiment, program instructions 2922 can be configured to implement any of the functions described herein. Additionally, the memory 2920 can comprise any of the information structures or data structures described herein. In some embodiments, program instructions and / or data can be received, sent or stored upon different types of computer-accessible media or similar media, depending on the particular implementation. Although the computer system 2900 is described as implementing the functionality of the functional blocks of the preceding figures, any of the functionality described herein can be implemented via such a computer system.

[0159] In one embodiment, the I / O interface 2930 can be configured to coordinate I / O traffic between the processor 2910, the system memory 2920, and any peripheral devices in the device, including peripheral interface 2940 or other peripheral interfaces, such as input / output devices 2950. In some embodiments, the I / O interface 2930 can perform any necessary protocol, timing or other data transformations to convert data signals from one component (e.g., system memory 2920) into a format suitable for use by another component (e.g., the processor 2910). In some embodiments, the I / O interface 2930 can include support for devices attached through various types of peripheral buses, such as a variant of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard, for example. In some embodiments, the

[0160] The network interface 2940 can be configured to allow the computer system 2900 to exchange data with other devices attached to a network 2985, such as a server or a client device, or with a node of the computer system 2900, over the network 2985. In various embodiments, the network 2985 can comprise one or more networks including, but not limited to, a local area network (LAN), such as an Ethernet or corporate network, a wide area network (WAN), such as the Internet, a wireless data network, some other electronic data network, or some combination thereof. In various embodiments, the network interface 2940 can support communication via wired or wireless general data networks, such as any suitable type of Ethernet network, via telecommunications / telephony networks such as analog voice networks or digital fiber communications networks, via storage area networks such as Fibre Channel SANs, or via any other suitable type of network and / or protocol.

[0161] The input / output devices 2950 can in some embodiments include one or more display terminals, keyboards, keypads, touchpads, scanning devices, voice or optical recognition devices, or any other devices suitable to input or output data. Multiple input / output devices 2950 can be present in the computer system 2900, or can be distributed on various nodes of the computer system 2900. In some embodiments, similar input / output devices can be separate from the computer system 2900 and can interact with one or more nodes of the computer system 2900 through a wired or wireless connection, such as over the network interface 2940.

[0162] As shown in Figure 12 The memory 2920 can include program instructions 2922 that can be executable by the processor, possibly to implement any of the elements or actions described above. In one embodiment, the program instructions can implement the methods described above. In other embodiments, different elements and data can be included. Note that the data can include any of the data or information described above.

[0163] Those skilled in the art will appreciate that the computer system 2900 is merely illustrative and is not intended to limit the scope of embodiments. In particular, the computer system and devices can include any combination of hardware or software that can perform the indicated functions, including computers, network devices, internet appliances, personal digital assistants, wireless phones, pagers, modems, etc. The computer system 2900 can be connected to any other devices in any way that provides for a transfer of data among software applications or hardware components. Moreover, the computer system 2900 can operate in a networked environment using logical connections to one or more other devices, or as a self-contained, independent system.

[0164] Those skilled in the art will also recognize that while various items are shown as being stored in memory or on storage devices during use, these items, or portions thereof, may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, some or all of the software components may be executed in memory on another device and communicate with the illustrated computer system via inter-computer communication. Some or all of the system components or data structures may also be stored (e.g., as instructions or structured data) on a computer-accessible medium or portable article of manufacture for reading by a suitable drive, various examples of which have been described above. In some embodiments, instructions stored on a computer-accessible medium separate from computer system 2900 may be transmitted to computer system 2900 via a transmission medium or signal (such as an electrical signal, electromagnetic signal, or digital signal), the transmission medium or signal being transmitted via a communication medium (such as a network and / or wireless link). Various embodiments may further include receiving, transmitting, or storing instructions and / or data implemented according to the above description on a computer-accessible medium. Generally, computer-accessible media may include non-transitory computer-readable storage media or memory media, such as magnetic or optical media, such as discs or DVD / CD-ROMs, and volatile or non-volatile media, such as RAM (e.g., SDRAM, DDR, RDRAM, SRAM, etc.), ROM, etc. In some embodiments, computer-accessible media may include transmission media or signals, such as electrical signals, electromagnetic signals, or digital signals transmitted via communication media such as networks and / or wireless links.

[0165] In various embodiments, the methods described herein may be implemented in software, hardware, or a combination thereof. Furthermore, the block order of the methods may be changed, and various elements may be added, reordered, combined, omitted, modified, etc. Various modifications and changes will be apparent to those skilled in the art who benefit from this disclosure. The various embodiments described herein are intended to be exemplary and not restrictive. Many variations, modifications, additions, and improvements are possible. Thus, multiple instances may be provided for a component described herein as a single instance. The boundaries between various components, operations, and data storage devices are somewhat arbitrary, and specific operations are illustrated in the context of a particular exemplary configuration. Other allocations of functionality are contemplated, which may fall within the scope of the appended claims. Finally, the structure and function of discrete components presented in exemplary configurations may be implemented as combined structures or components. These and other variations, modifications, additions, and improvements may fall within the scope of the embodiments defined by the appended claims.

Claims

1. A system comprising: a video decoder configured to perform operations comprising: receiving a bitstream, the bitstream comprising (i) encoded video data corresponding to a plurality of frames of source video data encoded by an encoder and (ii) format metadata, the encoded video data representing a focused range of luminance values of the source video data, the format metadata providing information about the focused range luminance values and a plurality of transfer functions used to map a full dynamic range of the source video data to the focused range of luminance values, wherein different ones of the plurality of transfer functions are used corresponding to the plurality of frames; obtaining the encoded video data and the format metadata from the bitstream; and decoding a particular frame from the plurality of frames of the encoded video data, wherein the decoding comprises: using the information provided by the format metadata to select a particular transfer function from the plurality of transfer functions corresponding to the particular frame; and determining luminance values for the particular frame using the focused range of luminance values indicated by the format metadata and the particular transfer function.

2. The system of claim 1, wherein the plurality of transfer functions comprises (i) a first transfer function and (ii) a second transfer function, the first transfer function configured to clip source video data corresponding to a frame from the plurality of frames according to a focused range of the source video data corresponding to the frame, the second transfer function configured to map the clipped source video data corresponding to the frame to an available number of bits used to represent the encoded video data.

3. The system of claim 2, wherein the source video data and the encoded video data are represented using different numbers of bits, and wherein a first number of bits used to represent the source video data is greater than a second number of bits used to represent the encoded video data.

4. The system of claim 3, wherein the decoder is configured to decode the encoded video data obtained from the bitstream using a third number of bits different from the first number of bits and the second number of bits.

5. The system of claim 2, wherein a first number of frames from the plurality of frames are encoded using one or more portions of the first transfer function and a second number of frames from the plurality of frames are encoded using the second transfer function.

6. The system of claim 5, wherein the format metadata comprises information indicating the one or more portions of the first transfer function.

7. The system of claim 1, wherein a first transfer function from the plurality of transfer functions and a different second transfer function are used to map a first frame and a second frame from the plurality of frames, respectively.

8. The system of claim 7, wherein the first frame and the second frame are adjacent frames in the source video data.

9. A method performed by a video decoder, the method comprising: ​ ​ receiving a bitstream, the bitstream including (i) encoded video data corresponding to a plurality of frames of source video data encoded by an encoder and (ii) format metadata, the encoded video data representing a focused range of luminance values of the source video data, the format metadata providing information about the focused range of luminance values and a plurality of transfer functions used to map a full dynamic range of the source video data to the focused range of luminance values, wherein different ones of the plurality of transfer functions are used corresponding to the plurality of frames; obtaining the encoded video data and the format metadata from the bitstream; and decoding a particular frame from the plurality of frames of the encoded video data, wherein the decoding includes: using the information provided by the format metadata to select a particular transfer function from the plurality of transfer functions corresponding to the particular frame; and determining luminance values for the particular frame using the focused range of luminance values indicated by the format metadata and the particular transfer function.

10. The method of claim 9, wherein the plurality of transfer functions includes (i) a first transfer function and (ii) a second transfer function, the first transfer function configured to clip source video data corresponding to a frame from the plurality of frames according to a focused range of luminance values of the source video data corresponding to the frame, the second transfer function configured to map the clipped source video data corresponding to the frame to a number of bits available for representing the encoded video data.

11. The method of claim 10, wherein the source video data and the encoded video data are represented using different numbers of bits, and wherein a first number of bits used to represent the source video data is greater than a second number of bits used to represent the encoded video data.

12. The method of claim 11, wherein the decoder is configured to decode the encoded video data obtained from the bitstream using a third number of bits different from the first number of bits and the second number of bits.

13. The method of claim 10, wherein a first number of frames from the plurality of frames are encoded using one or more portions of the first transfer function and a second number of frames from the plurality of frames are encoded using the second transfer function.

14. The method of claim 13, wherein the format metadata includes information indicating the one or more portions of the first transfer function.

15. The method of claim 9, wherein a first transfer function from the plurality of transfer functions and a different second transfer function are used to map a first frame and a second frame, respectively, from the plurality of frames.

16. The method of claim 15, wherein the first frame and the second frame are adjacent frames in the source video data.

17. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause a video decoder to perform operations comprising: receiving a bitstream, the bitstream including (i) encoded video data corresponding to a plurality of frames of source video data encoded by an encoder and (ii) format metadata, the encoded video data representing a focused range of luminance values of the source video data, the format metadata providing information about the focused range luminance values and a plurality of transfer functions used to map a full dynamic range of the source video data to the focused range of luminance values, wherein different ones of the plurality of transfer functions are used corresponding to the plurality of frames; obtaining the encoded video data and the format metadata from the bitstream; and decoding a particular frame from the plurality of frames of the encoded video data, wherein the decoding includes: using the information provided by the format metadata to select a particular transfer function from the plurality of transfer functions corresponding to the particular frame; and using the focused range of luminance values indicated by the format metadata and the particular transfer function to determine luminance values for the particular frame.

18. The non-transitory computer readable medium of claim 17, wherein the plurality of transfer functions includes (i) a first transfer function and (ii) a second transfer function, the first transfer function configured to clip source video data corresponding to a frame from the plurality of frames according to a focused range of the source video data corresponding to the frame, the second transfer function configured to map the clipped source video data corresponding to the frame to a number of bits available for representing the encoded video data.

19. The non-transitory computer readable medium of claim 18, wherein the source video data and the encoded video data are represented using different numbers of bits, wherein a first number of bits used to represent the source video data is greater than a second number of bits used to represent the encoded video data, and wherein the decoder is configured to decode the encoded video data obtained from the bitstream using a third number of bits different from the first number of bits and the second number of bits.

20. The non-transitory computer readable medium of claim 18, wherein a first number of frames from the plurality of frames are encoded using one or more portions of a first transfer function and a second number of frames from the plurality of frames are encoded using a second transfer function, and wherein the format metadata includes information indicating the one or more portions of the first transfer function.

21. The non-transitory computer readable medium of claim 17, wherein a first transfer function from the plurality of transfer functions and a different second transfer function are used to map a first frame and a second frame, respectively, from the plurality of frames, and wherein the first frame and the second frame are adjacent frames in the source video data. ​