Video content encoding and decoding

By generating and allocating video frames to data structures at different time levels, and dynamically adjusting the encoding quality level and bit rate, the adaptability problem of video exchange between electronic devices is solved, achieving a balance between video quality and real-time performance under different conditions.

CN115866249BActive Publication Date: 2026-03-27APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently exchange video content in real time between electronic devices, adapting to differences in video complexity, communication network capabilities, and device capabilities, making it difficult to balance video quality and real-time performance.

Method used

By generating data structures representing multiple frames of a video, allocating them to different temporal layers, and dynamically adjusting the encoding quality level based on video complexity, network capabilities, and device capabilities, the frames are encoded using different bit rates and quantization parameters, generating highly adaptable data structures.

Benefits of technology

It enables dynamic adjustment of video quality and frame rate under different device and network conditions, ensuring consistency and efficiency of video content quality during real-time exchange.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115866249B_ABST
    Figure CN115866249B_ABST
Patent Text Reader

Abstract

This disclosure relates to video content encoding and decoding. In one example method, a system receives a plurality of frames of a video, and generates a data structure representing the video and representing a plurality of temporal layers. Generating the data structure includes: (i) determining a plurality of quality levels for presenting the video, wherein each of the quality levels corresponds to a different respective sampling period for sampling the frames of the video, (ii) assigning each of the frames to a respective one of the temporal layers of the data structure based on the sampling periods, and (iii) indicating in the data structure one or more relationships between (a) at least one of the frames assigned to at least one of the temporal layers of the data structure and (b) at least another one of the frames assigned to at least another one of the temporal layers of the data structure. Further, the system outputs the data structure.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] This disclosure relates generally to video content encoding and decoding. BACKGROUND

[0002] Electronic devices can be used to generate and present video content. As one example, an electronic device can generate a video of an object and transmit the video to one or more other electronic devices for presentation. In some implementations, the electronic devices can exchange video content in real-time (e.g., to facilitate a video conference or video call between them). SUMMARY

[0003] In one aspect, a method includes: receiving, by one or more processors, a plurality of frames of a video; generating, by the one or more processors, a data structure representing the video, wherein the data structure represents a plurality of temporal layers, and wherein generating the data structure includes: (i) determining a plurality of quality levels for presenting the video, wherein each of the quality levels corresponds to a different respective sampling period for sampling the frames of the video, (ii) assigning each of the frames to a respective temporal layer of the temporal layers of the data structure based on the sampling periods, and (iii) indicating, in the data structure, one or more relationships between (a) at least one of the frames assigned to at least one of the temporal layers of the data structure and (b) at least another of the frames assigned to at least another of the temporal layers of the data structure; and outputting, by the one or more processors, the data structure.

[0004] Implementations of the aspect can include one or more of the following features.

[0005] In some implementations, the data structure can include a group of pictures (GOP) structure.

[0006] In some implementations, outputting the data structure can include transmitting a bitstream having the data structure.

[0007] In some implementations, the method can further include generating the video using one or more cameras of a first mobile device, and transmitting the data structure from the first mobile device to one or more second mobile devices via a communication network.

[0008] In some implementations, the video can include visual content for a communication session between the first mobile device and the one or more second mobile devices.

[0009] In some implementations, for each frame of the frames, assigning each frame of the frames to a respective temporal layer of the temporal layers of the data structure can include determining an index number associated with the frame, identifying, from the quality levels, a particular quality level of the plurality of quality levels that corresponds to a sampling period that is divisible by the index number, and assigning the frame to one of the temporal layers based on the identified quality level.

[0010] In some implementations, the frames are assigned to the temporal layer having an index value that is the same as an index value of the identified quality level.

[0011] In some implementations, identifying the particular quality level can include identifying, from the particular quality levels, a subset of the quality levels, each quality level of the subset of the quality levels corresponding to a respective sampling period that is divisible by the index number, and selecting, from the subset, a quality level having a largest sampling period.

[0012] In some implementations, the plurality of quality levels can include a first quality level corresponding to a first sampling period and a second quality level corresponding to a second sampling period, wherein the first sampling period is a multiple of the second sampling period.

[0013] In some implementations, the plurality of quality levels can further include a third quality level corresponding to a third sampling period, wherein the second sampling period is a multiple of the third sampling period.

[0014] In some implementations, the method can further include encoding frames assigned to a first temporal layer of the temporal layers of the data structure according to a first bit rate, and encoding frames assigned to a second temporal layer of the temporal layers of the data structure according to a second bit rate, wherein the first bit rate is different than the second bit rate.

[0015] In some implementations, the method can further include encoding frames assigned to a first temporal layer of the temporal layers of the data structure according to a first quantization parameter, and encoding frames assigned to a second temporal layer of the temporal layers of the data structure according to a second quantization parameter, wherein the first quantization parameter is different than the second quantization parameter.

[0016] Other implementations relate to systems, apparatuses, and non-transitory computer- readable media having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform the operations described herein.

[0017] The details of one or more implementations are set forth in the accompanying drawings and the detailed description below. Other features and advantages will be apparent from the detailed description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is a diagram for exchanging video content between users.

[0019] Figure 2 is a diagram of an example system for encoding and decoding video content.

[0020] Figure 3A is a diagram of an example set of video frames.

[0021] Figure 3B is a diagram of an example data structure for encoding information about video frames, and a diagram showing a relationship between temporal layers of the data structure and respective quality levels or grades for presenting video content.

[0022] Figure 4 is a diagram of an example process for modifying a data structure to increase a frame rate for presenting video content according to one or more service levels or grades.

[0023] Figure 5 is a diagram of an example process for modifying a data structure to decrease a frame rate for presenting video content according to one or more service levels or grades.

[0024] Figure 6 is a flowchart of an example process for encoding data about video frames in a data structure.

[0025] Figure 7 is a flowchart of an example process for modifying a data structure.

[0026] Figure 8 is a diagram of an example process for encoding visual content.

[0027] Figure 9 is a diagram of an example device architecture for implementing the features and processes described in reference to Figures 1 to 8 . DETAILED DESCRIPTION

[0028] Generally, an electronic device can generate video content and present that video content to one or more users. As one example, a first electronic device can generate a video of an object (e.g., containing synchronized audio content and visual content about the object), encode information about the video in a data structure, and transmit the data structure to one or more second electronic devices. Each of the second electronic devices can decode the data structure to extract information about the video and present the video (or an approximation thereof) to one or more users.

[0029] Furthermore, in some implementations, electronic devices can exchange video content in real time (e.g., to facilitate video conferencing or video calls between them). As an example, two or more electronic devices can be communicatively coupled to each other (e.g., via a communication network). Moreover, each electronic device can generate a video of its user, encode information about the video, and transmit this data structure to each of the other electronic devices in real time. Simultaneously, each electronic device can also receive data structures from each of the other electronic devices, decode the data structures to extract information about the other user's video, and present the video to its user in real time. Therefore, multiple users can communicate with each other through the real-time exchange of video content.

[0030] In some implementations, electronic devices can dynamically adjust the encoding of video content based on the complexity of the video content, the capabilities of the communication network, and / or the capabilities of one or more other electronic devices. For example, this is beneficial in enabling electronic devices to exchange video content in real time, despite differences or fluctuations in the complexity of the video content, the capabilities of the communication network, and / or the capabilities of the electronic devices.

[0031] As an example, if the video content is complex (e.g., requires a large amount of data to encode the representation of the video content), electronic devices can allocate more computing resources to encode the video content (e.g., a higher bitrate), allowing the video content to be presented with higher fidelity or detail.

[0032] Furthermore, if the video content is not complex (e.g., requires a small amount of data to encode the representation of the video content), electronic devices can allocate less computing resources to encode the video content (e.g., a lower encoding bit rate), thereby converting computing resources.

[0033] As another example, if a communication network only enables the exchange of data between computer systems at low throughput, electronic devices can encode video at a low quality level (e.g., to produce a data structure with a small data size). This allows the encoded video content to be transmitted over the communication network, decoded by a receiving electronic device, and presented in real time by the receiving electronic device, without exceeding the capabilities of the communication network.

[0034] Furthermore, if the communication network enables data to be exchanged between the electronic devices according to a high throughput, the electronic devices can encode the video according to a high quality level (e.g., to produce data structures having a large data size). This enables a high quality version of the encoded video content (e.g., having a high level of fidelity and / or detail) to be transmitted over the communication network, decoded by the receiving electronic device, and presented in real-time by the receiving electronic device.

[0035] As another example, if the receiving electronic device has low technical capabilities (e.g., has a slow processor, limited memory, etc.), the electronic device can encode the video according to a low quality level (e.g., to produce data structures that have low computational requirements for decoding). This enables the encoded video content to be transmitted over the communication network, decoded by the receiving electronic device, and presented in real-time by the receiving electronic device without exceeding the capabilities of the receiving electronic device.

[0036] Furthermore, if the receiving electronic device has high technical capabilities (e.g., has a fast processor, large memory, etc.), the electronic device system can encode the video according to a high quality level (e.g., to produce data structures that have high computational requirements for decoding). This enables a high quality version of the encoded video content to be transmitted over the communication network, decoded by the receiving electronic device, and presented in real-time by the receiving electronic device.

[0037] Exemplary techniques for dynamically adjusting the encoding of video content are described in further detail below.

[0038] Figure 1 An exemplary system 100 for exchanging video content between electronic devices is shown. In some implementations, the system 100 can be configured to implement a video conference or video call between two or more users of electronic devices.

[0039] The system 100 includes several electronic devices 102a-102c and a video conference server 104 communicatively coupled to each other via a communication network 106. During example operation of the system 100, the electronic devices 102a-102c establish one or more communication sessions using the communication network 106 and exchange video content with each other (e.g., in real-time).

[0040] In some implementations, the electronic devices 102a-102c can establish a communication session with each other directly. As one example, the electronic devices 102a-102c can establish one or more communication links between each other over the communication network 106. Further, each of the electronic devices 102a-102c can generate video content and transmit the video content directly to one or more of the other electronic devices 102a-102c via the communication links. Moreover, each of the electronic devices 102a-102c can receive video content from one or more of the other electronic devices 102a-102c directly via the communication links and present the video content to a user.

[0041] In some implementations, the electronic devices 102a-102c establish a communication session with each other using the video conferencing server 104 as an intermediary. As one example, each of the electronic devices 102a-102c can establish a communication link with the video conferencing server over the communication network 106. Further, each of the electronic devices 102a-102c can generate video content and transmit the video content to the video conferencing server 104 via the communication link. The video conferencing server 104 can receive the video content from each of the electronic devices 102a-102c and route the video content to each of the other electronic devices 102a-102c (e.g., such that each electronic device receives video content generated by each of the other electronic devices). In turn, each of the electronic devices 102a-102c can present the received video content to a user.

[0042] In practice, the electronic devices 102a-102c can be any devices configured to receive, process, and transmit data. As one example, at least one of the electronic devices 102a-102c can be a computing device, such as a client computing device (e.g., a desktop or notebook computer), a server computing device (e.g., a server computer or cloud computing system), a mobile computing device (e.g., a cellular phone, a smartphone, a tablet computer, a personal digital assistant, or a notebook computer), a wearable computing device (e.g., a smartwatch, a virtual reality headset, or an augmented reality headset), or other computing device capable of receiving, processing, and transmitting data. In some implementations, at least one of the electronic devices 102a-102c can operate using one or more operating systems (e.g., Apple macOS, Apple iOS, Microsoft Windows, Linux, Unix, Google Android, etc.) and one or more architectures (e.g., x86, PowerPC, ARM, etc.).

[0043] Although Figure 1 Three example electronic devices 102a-102c are shown, but in practice, the system 100 can include any number of electronic devices configured to exchange video with each other over the communication network 106 (e.g., two, three, four, five, or more).

[0044] Furthermore, in practice, the video conference server 104 can also be any device configured to receive, process, and transmit data. As one example, the video conference server 104 can be a computing device, such as a server computing device (e.g., a server computer or a cloud computing system). In some implementations, the video conference server 104 can operate using one or more operating systems (e.g., Apple macOS, Apple iOS, Microsoft Windows, Linux, Unix, Google Android, etc.) and one or more architectures (e.g., x86, PowerPC, ARM, etc.).

[0045] The communication network 106 can be any communication network through which data can be transmitted and shared. For example, the communication network 106 can be a local area network (LAN) or a wide area network (WAN), such as the Internet. The communication network 106 can be implemented using various network interfaces, such as a wireless network interface (e.g., Wi-Fi, Bluetooth, or infrared) or a wired network interface (e.g., Ethernet or serial connection). The communication network 106 can also include a combination of more than one network, and can be implemented using one or more network interfaces.

[0046] Figure 2 is a diagram of an example system 200 for processing and displaying video content. The system 100 includes an encoder 202, a decoder 204, a Tenderer 206, and an output device 208.

[0047] In some implementations, the system 200 can be implemented at least in part using the system 100. As one example, each of the electronic devices 102a-102c can include a respective encoder 202 (e.g., for encoding video content to be exchanged with the other electronic devices). As another example, each of the electronic devices 102a-102c can include a respective decoder 204, Tenderer 206, and output device 208 (e.g., for decoding, rendering, and presenting video content received from the other electronic devices).

[0048] During example operation of system 200, encoder 202 receives a number of successive video frames 210 of a video. In some implementations, video frames 210 can be provided to encoder 202 in real-time. For example, a camera subsystem of an electronic device can continuously produce video frames 210 (e.g., during a communication session (e.g., a video conferencing session or a video telephony session)) and provide them to encoder 202 as they are captured.

[0049] In some implementations, each of video frames 210 can include two-dimensional and / or three-dimensional images (or a portion thereof), each representing a particular portion of the video. As one example, each of video frames 210 can include an image that is captured at a particular point in time (e.g., by a camera subsystem of an electronic device). Moreover, video frames 210 can be presented in order (e.g., in the order in which they were captured) to indicate dynamic visual information, such as movement or other changes of objects of the images over time.

[0050] In some implementations, video frames 210 can be captured according to a particular frame rate and / or bit rate. As one example, video frames 210 include images that are captured at a frequency of QUOTE (e.g., each image is separated in time by a time interval QUOTE and according to a bit rate of QUOTE successively.

[0051] Figure 3A An example set of video frames 210 is shown in FIG. 13. In this example, a set of eight video frames (labeled “0” through “7”) are captured over a time interval QUOTE of frame rate QUOTE of time interval QUOTE Also, a set of additional video frames can be captured after time interval QUOTE (e.g., another set of eight video frames starting at frame “8(0)” over time interval QUOTE ).

[0052] The encoder 202 generates encoded content 212 based on the video frames 210. The encoded content 212 includes information that enables an electronic device to generate video content representative of the video frames 210. As one example, the encoded content 212 can include a data structure that includes a representation of one or more of the video frames 210. Further, the data structure can include information about relationships between at least some of the video frames 210, such as an order in which the video frames 210 are to be presented in a sequence and / or one or more prediction vectors between the video frames 210.

[0053] Figure 3B An example data structure 300 for encoding information about the video frames 210 is shown. In some implementations, the data structure 300 can be referred to as a group of pictures (GOP) structure.

[0054] In general, the data structure 300 can assign each of the video frames 210 to one of a number of temporal layers according to different quality levels or “tiers.” As one example, a high quality level can correspond to a presentation of video content according to a high fidelity and / or level of detail, while a low quality level can correspond to a presentation of video content according to a low fidelity and / or level of detail.

[0055] In some implementations, each of the quality levels or tiers can be associated with a respective bit rate for transmitting and / or receiving the video content. As one example, a high quality level can be associated with a high bit rate (e.g., such that a large amount of data can be transmitted and / or received to present the video content), while a lower quality level can be associated with a low bit rate (e.g., such that a small amount of data can be transmitted and / or received to present the video content).

[0056] For example, in the example shown, Figure 3B the quality level “Tier 0” can be associated with a bit rate of QUOTE the quality level “Tier 1” can be associated with a bit rate of QUOTE the quality level “Tier 2” can be associated with a bit rate of QUOTE and the quality level “Tier 3” can be associated with a bit rate of QUOTE In practice, each of these tiers can be associated with any other bit rate, depending on the implementation.

[0057] In some implementations, a high quality level can correspond, at least in part, to a high frame rate for presenting the video content, while a low quality level can correspond, at least in part, to a low frame rate for presenting the video content. Moreover, the data structure 300 can be arranged according to a set of temporal layers that provide increasing temporal resolution for the video content. In some implementations, this technique can be referred to as temporal scaling.

[0058] For example, in the example shown, the data structure 300 can be used to generate video content according to a low quality level (“Level 0”), in which one out of every eight captured video frames is presented (e.g., corresponding to a frame rate of approximately 0.125 frames per second). Figure 3B As another example, the data structure 300 can be used to generate video content according to a higher quality level (“Level 1”), in which one out of every four captured video frames is presented (e.g., corresponding to a frame rate of approximately 0.25 frames per second). As another example, the data structure 300 can be used to generate video content according to an even higher quality level (“Level 2”), in which one out of every two captured video frames is presented (e.g., corresponding to a frame rate of approximately 0.5 frames per second). As another example, the data structure 300 can be used to generate video content according to an even higher quality level (“Level 3”), in which every captured video frame is presented (e.g., corresponding to a frame rate of approximately 1.0 frames per second).

[0059] Each of the quality levels or levels can be associated with a respective temporal layer of the data structure 300. Moreover, each of the video frames 210 can be assigned to a particular temporal layer of these temporal layers depending on whether the video frame is to present the video content with the corresponding quality level.

[0060] As one example, as shown in the diagram 302, the quality level “Level 0” can include each of the frames that have been assigned to the temporal layer “Layer 0.” Moreover, the quality level “Level 1” can include each of the frames that have been assigned to the temporal layer “Layer 1” as well as each of the frames that have been assigned to the lower quality level (e.g., “Level 0”). Furthermore, the quality level “Level 2” can include each of the frames that have been assigned to the temporal layer “Layer 2” as well as each of the frames that have been assigned to the lower quality levels (e.g., “Level 1” and “Level 0”). Moreover, the quality level “Level 3” can include each of the frames that have been assigned to the temporal layer “Layer 3” as well as each of the frames that have been assigned to the lower quality levels (e.g., “Level 2,” “Level 1,” and “Level 0”). ​​

[0061] Video frames 210 can be assigned to particular temporal layers of data structure 300 according to diagram 302. For example, video frames "0" and "8(0)" will be presented in the video content according to each of the quality levels or tiers. Thus, video frames "0" and "8(0)" can be assigned to the lowest temporal layer, "Layer 0."

[0062] As another example, video frame "4" will be presented in the video content according to each of quality levels "Tier 1" through "Tier 3" only. Thus, video frame "4" can be assigned to temporal layer "Layer 1."

[0063] As another example, video frames "2" and "6" will be presented in the video content according to each of quality levels "Tier 2" and "Tier 3" only. Thus, video frames "2" and "6" can be assigned to temporal layer "Layer 2."

[0064] As another example, video frames "1," "3," "5," and "7" will be presented in the video content according to quality level "Tier 3" only. Thus, video frames "1," "3," "5," and "7" can be assigned to temporal layer "Layer 3."

[0065] To number the video frames in the data structure, each video frame assigned to the lowest temporal layer can be assigned an index value of 0, and each subsequent video frame can be assigned the next higher integer (unless the video frame is assigned to the lowest temporal layer, in which case the lowest temporal layer will be assigned an index value of 0).

[0066] In some implementations, encoder 202 can store data representing changes or "deltas" between one video frame and another, rather than storing a complete, independent representation of the content of each of the individual video frames in data structure 300. This can be beneficial, for example, in reducing the size of data structure 300 (e.g., by reducing redundancy in the data stored therein).

[0067] Furthermore, encoder 202 can generate data structure 300 such that data structure 300 includes one or more prediction vectors indicating a relationship between two or more video frames in data structure 300. For example, a prediction vector can indicate that data stored in data structure 300 regarding a particular video frame represents a change relative to another video frame (e.g., a predicted picture or "P-frame") stored in data structure 300, rather than an independent representation of the content of the video frame (e.g., an intra-coded picture or "I-frame" or "key frame").

[0068] As an illustrative example, in the example of diagram 302, video frame "4" can be assigned to temporal layer "Layer 1." Video frame "4" can be assigned to temporal layer "Layer 1" because video frame "4" is presented in the video content according to each of quality levels "Tier 1" through "Tier 3." Thus, video frame "4" can be assigned to temporal layer "Layer 1" because video frame "4" is presented in the video content according to each of quality levels "Tier 1" through "Tier 3." Figure 3BIn the illustrated data structure 300, the prediction vector extends from video frame “2” to video frame “0” (indicated by the arrow). This indicates that the data stored in data structure 300 regarding video frame “2” represents only the changes between video frame “2” and video frame “0” (e.g., rather than an independent representation of the contents of video frame “2”). Thus, to render video frame “2”, a decoder would obtain video frame “0” and modify the contents of video frame “0” based on the data contained in data structure 300 regarding video frame “2”.

[0069] As another example, in some implementations, the data structure 300 can include a prediction vector that extends from video frame “3” to video frame “2” (indicated by another arrow). This indicates that the data stored in data structure 300 regarding video frame “3” represents only the changes between video frame “3” and video frame “2” (e.g., rather than an independent representation of the contents of video frame “3”). Thus, to render video frame “3”, a decoder would (i) obtain video frame “0”, (ii) modify the contents of video frame “0” based on the data contained in data structure 300 regarding video frame “2” to obtain video frame “2”, and (iii) modify the contents of video frame “2” based on the data contained in data structure 300 regarding video frame “3” to obtain video frame “3”. Figure 3B In the illustrated data structure 300, the prediction vector extends from video frame “3” to video frame “2” (indicated by another arrow). This indicates that the data stored in data structure 300 regarding video frame “3” represents only the changes between video frame “3” and video frame “2” (e.g., rather than an independent representation of the contents of video frame “3”). Thus, to render video frame “3”, a decoder would (i) obtain video frame “0”, (ii) modify the contents of video frame “0” based on the data contained in data structure 300 regarding video frame “2” to obtain video frame “2”, and (iii) modify the contents of video frame “2” based on the data contained in data structure 300 regarding video frame “3” to obtain video frame “3”.

[0070] A decoder can perform similar techniques to obtain each of the other video frames represented by data structure 300.

[0071] Figure 3B The illustrated example data structure 300 is shown with four temporal layers corresponding to four quality levels or grades. However, in practice, data structure 300 can include any number of temporal layers corresponding to any number of quality levels or grades (e.g., one, two, three, four, or more). Moreover, although the example frame rates are shown in Figure 3B The example frame rates are shown in the table in FIG. 3. However, in practice, the quality levels or grades can correspond to any frame rate that is greater than or equal to the frame rate of the next lower quality level or grade. Moreover, in some implementations, the frame rate of each quality level or grade can be a multiple of the frame rate of the next lower quality level or grade.

[0072] In some implementations, encoder 202 can assign video frames 210 to the temporal layers of data structure 300 based on (i) an index number associated with each of video frames 210 and (ii) a sampling period associated with each of the quality levels or grades.

[0073] As one example, each of video frames 210 can be assigned an integer index number iQUOTEd (For example, for the first video frame in data structure 300), and incremented by 1 for each consecutive video frame in data structure 300 in the sequence. As described above, in some specific implementations, an index value of 0 may be assigned to each video frame assigned to the lowest time layer, and a next higher integer may be assigned to each subsequent video frame (unless that video frame is assigned to the lowest time layer, where an index value of 0 will be assigned to that lowest time layer).

[0074] Furthermore, a sampling period can be determined for each quality level or grade. The sampling period for a quality level can correspond to the interval at which captured video frames are sampled or encoded at that quality level. In some implementations, the sampling period may also be referred to as the sampling frequency or sampling rate and / or sampling interval (depending on the context).

[0075] For example, in Figure 3A and Figure 3B In the example shown, the set of captured video frames 210 includes time intervals QUOTE Eight frames within. Based on the highest quality level, "Level 3," corresponding to the sampling period QUOTE. Each video frame in the captured video frame is sampled or encoded (e.g., each individual video frame is sampled or encoded).

[0076] Furthermore, according to the next lower quality level, "Level 2", the corresponding sampling period is QUOTE The video frames are sampled or encoded for every second video frame captured. In some implementations, video frames in "Level 2" may be subsampled from video frames in "Level 3" (e.g., according to subsampling period 2).

[0077] Furthermore, according to the next lower quality level, "Level 1", the corresponding sampling period is QUOTE The video frames are sampled or encoded every fourth video frame in the captured video frames. In some specific implementations, video frames in "Level 1" can be subsampled from video frames in "Level 2" (e.g., according to subsampling period 2).

[0078] Furthermore, according to the lowest quality level "Level 0", the corresponding sampling period is QUOTE Every eighth video frame in the captured video frames is sampled or encoded. In some specific implementations, video frames in "Level 0" can be subsampled from video frames in "Level 1" (e.g., according to subsampling period 2).

[0079] For the first frame in data structure 300, QUOTE , the encoder 202 can determine that the maximum of QUOTE , QUOTE , QUOTE , QUOTE … QUOTE is divisible by 8. Thus, video frame “0” can be assigned to temporal layer “Tier 0” corresponding to quality level “Tier 0” having a sampling period of 8.

[0080] For example, in the example shown in Figure 3A and Figure 3B , for video frame “0”, 0 can be divisible by each of QUOTE , QUOTE , QUOTE , and QUOTE (8, 4, 2, and 1, respectively). Thus, video frame “0” can be assigned to temporal layer “Tier 0” corresponding to quality level “Tier 0” having a sampling period of 8.

[0081] As another example, in the example shown in Figure 3A and Figure 3B , for video frame “4”, 4 can be divisible by each of QUOTE , , and (4, 2, and 1, respectively). Thus, video frame “4” can be assigned to temporal layer “Tier 1” corresponding to quality level “Tier 1” having a sampling period of 4.

[0082] As another example, in the example shown in Figure 3A and Figure 3B , for video frames “2” and “6”, both 2 and 6 can be divisible by each of QUOTE , and (2 and 1, respectively). Thus, video frames “2” and “6” can be assigned to temporal layer “Tier 2” corresponding to quality level “Tier 2” having a sampling period of 2.

[0083] As another example, in the example shown in Figure 3A and Figure 3BIn the illustrated example, for video frames “1,” “3,” “5,” and “7,” each of 1, 3, 5, and 7 can be encoded only once (1) integer division. Thus, video frames “1,” “3,” “5,” and “7” can be assigned to temporal layer “tier 3” corresponding to quality level “level 3” having a sampling period of 1.

[0084] In some implementations, encoder 202 can encode data in data structure 300 differently according to the temporal tiers in which the data is stored. As one example, encoder 202 can encode data assigned to one of the temporal tiers according to a particular bitrate, and encode data assigned to another of the temporal tiers according to a different bitrate. This can be beneficial, for example, because data stored in different temporal tiers can have different complexities. Thus, more complex data (e.g., having a greater degree of change between successive video frames) can be encoded using a higher bitrate, while less complex data (e.g., having a lesser degree of change between successive video frames) can be encoded using a lower bitrate, thereby appropriately allocating computational resources therebetween. In some implementations, this technique enables decoder 204 and Tenderer 206 to generate video content having a consistent quality level over time, while also maintaining a consistent bitrate for each quality level or tier.

[0085] In some implementations, data in data structure 300 can be enabled according to one or more quantization parameter (QP) values. A quantization parameter value refers to the degree to which a range of values is quantized or compressed into a single quantum value during the encoding process. For example, a large quantization parameter value can indicate that a large range of values in the data is compressed into a particular quantum value during the encoding process, which can greatly reduce the size of the data at the expense of fidelity. As another example, a small quantization parameter value can indicate that a small range of values in the data is compressed into a particular quantum value, which can more modestly reduce the size of the data with less impact on fidelity.

[0086] In some implementations, the encoder 202 can generate a model for each temporal layer that determines a relationship between (i) a bit rate at which data of that temporal layer is encoded and (ii) a quantization parameter value used to encode that data (e.g., a "rate-QP model"). The model can account for various characteristics of the data in that temporal layer, such as a number of video frames in that temporal layer, a type of data stored in the data structure 300 regarding those video frames (e.g., whether the data represents an I-frame or a P-frame), a complexity of the content (e.g., a level of visual detail in the video frames, a degree to which content of the video frames changes over time, etc.). Moreover, the encoder 202 can select a quantization parameter value for each temporal layer such that a bit rate for each quality level or level of quality is consistent with its allocated bit rate.

[0087] In some implementations, the encoder 202 can allocate a larger number of bits to encode video frames in lower layers and a smaller number of bits to encode video frames in higher layers. This can be beneficial because video frames in lower layers can consume more bits to achieve the same quality as video frames in higher layers. For example, the temporal spacing between video frames in higher layers is smaller. As a result, the content of those video frames is more likely to be similar to one another and can require a smaller number of bits to encode. In contrast, the temporal spacing between video frames in lower layers is larger. As a result, the content of those video frames is less likely to be similar to one another and can require a larger number of bits to encode.

[0088] In some implementations, the encoder 202 can modify the data structure 300 to accommodate differences or fluctuations in attributes of the video content.

[0089] As one example, if the complexity of a portion of the video content decreases, the encoder 202 can require fewer bits to encode that portion of the video content. In response, the encoder 202 can modify the arrangement of the data structure 300 such that the frame rate of one or more of the quality levels or levels of quality increases while also ensuring that the bit rate does not exceed the allocated bit rate for the quality levels or levels of quality. This is beneficial, for example, in providing a higher quality video content (e.g., a video content with a higher frame rate) to a recipient while maintaining a consistent bit rate over time.

[0090] As an illustrative example, Figure 4 The data structure 400 is shown with a first portion 402a having a number of video frames "0" through "7" each allocated to one of a number of temporal layers "Layer 0" through "Layer 3." The portion 402a of the data structure 400 can be similar to the portion shown in FIG. 3.

[0091] As referenced above with respect to FIG. 3, the data structure 400 can be modified to account for changes in the video content. For example, the data structure 400 can be modified to increase the frame rate of one or more of the quality levels or levels of quality while also ensuring that the bit rate does not exceed the allocated bit rate for the quality levels or levels of quality. Figure 3BThe arrangement corresponds to a quality level "Level 3" of a sampling period of QUOTE a quality level "Level 2" of a sampling period of QUOTE a quality level "Level 2" of a sampling period of QUOTE a quality level "Level 1" of a sampling period of QUOTE a quality level "Level 0" of a sampling period of QUOTE

[0092] In this example, portion 402b after portion 402a represents a decrease in complexity of the video content (e.g., as compared to the complexity of portion 402a). If portion 402b were to be encoded according to the same data structure as portion 402a, this would result in a bit rate for one or more of the corresponding quality levels or levels that is less than the bit rate allocated for those quality levels or levels (e.g., the encoder 202 can allocate a larger number of bits for encoding the video frames without exceeding the bit rate allocated for those quality levels or levels).

[0093] In response, beginning with the first video frame in portion 402b (e.g., video frame "8(0)"), the encoder 202 can modify the data structure 400 to increase the frame rate for at least some of the quality levels or levels. As one example, in portion 402b, the encoder 202 can assign video frames to a lower temporal layer that has a frequency and / or proportion that is greater than the frequency and / or proportion in the initial portion 402a, such that for one or more of the quality levels or levels, the video frames are presented at a higher rate in the video content.

[0094] For example, in portion 402b, the encoder 202 can assign video frames to a lower temporal layer that has a frequency and / or proportion that is greater than the frequency and / or proportion in the initial portion 402a, such that for one or more of the quality levels or levels, the video frames are presented at a higher rate in the video content. Figure 4In the illustrated example, the sampling period for the lowest quality level ("Level 0") can be reduced in portion 402b by replacing at least some of the video frames that had previously been assigned to temporal layer "Layer 1" through "Layer 3" with video frames that are instead assigned to temporal layer "Layer 0" (e.g., resulting in an increase in the frame rate for quality level "Level 0"). Similarly, the sampling period for the next lowest quality level ("Level 1") can be reduced in portion 402b by replacing at least some of the video frames that had previously been assigned to temporal layer "Layer 2" and "Layer 3" with video frames that are instead assigned to temporal layer "Layer 1" (e.g., resulting in an increase in the frame rate for quality level "Level 1"). In this example, the frame rate for quality level "Level 0" is increased by a factor of four (e.g., from being presented once every eight frames before the modification to being presented once every two frames after the modification). Further, in this example, the frame rate for quality level "Level 1" is also increased by a factor of four (e.g., from being presented once every four frames before the modification to being presented once every frame after the modification).

[0095] Further, if the complexity of a portion of the video content increases, the encoder 202 can need more bits to encode that portion of the video content, which can exceed the bit rate allocated to one or more of the quality levels or tiers. In response, the encoder 202 can modify the arrangement of the data structure 300 such that the frame rate for one or more of the quality levels or tiers is decreased and such that the bit rate does not exceed the bit rate allocated for those quality levels or tiers.

[0096] As an illustrative example, Figure 5 A data structure 500 is shown having a first portion 502a with a number of video frames "0" through "7" each assigned to one of a number of temporal layers "Layer 0" through "Layer 3". The portion 502a of the data structure 500 can be similar to the portion shown in FIG. 3.

[0097] As described with reference to Figure 3B the arrangement corresponds to a sampling period of QUOTE a quality level of "Level 3" (e.g., every video frame of the originally captured video is sampled), a sampling period of QUOTE a quality level of "Level 2" (e.g., every second video frame of the originally captured video is sampled), a sampling period of QUOTE a quality level of "Level 1" (e.g., every fourth video frame of the originally captured video is sampled), and a sampling period of QUOTE a quality level of "Level 0" (e.g., every eighth video frame of the originally captured video is sampled).

[0098] In this example, portion 502b following portion 502a represents an increase in the complexity of the video content (e.g., compared to the complexity of portion 502a). If portion 502b is encoded according to the same data structure as portion 502a, this will result in one or more of the corresponding quality levels or grades having a bit rate greater than the bit rate allocated to those quality levels or grades (e.g., encoder 202 will need to reduce the number of bits used to encode the video frame to avoid exceeding the allocated bit rate).

[0099] In response, starting with the first video frame in section 502b (e.g., video frame "8(0)"), encoder 202 may modify data structure 500 to increase the frame rate of one or more of the quality levels or grades. As an example, in the next section 502b of data structure 500, encoder 202 may assign video frames to a higher time layer that has a higher frequency and / or proportion than in the initial section 502a, such that for one or more of the quality levels or grades, the video frames are rendered at a lower rate in the video content.

[0100] For example, in Figure 5 In the example shown, the sampling period for the lowest quality level ("Level 0") can be increased in section 502b by replacing at least some of the video frames previously assigned to time layers "Layer 0" with those previously assigned to time layers "Layer 1" through "Layer 3" (e.g., causing a reduction in the frame rate of quality level "Level 0"). Similarly, the sampling period for the next lowest quality level ("Level 1") can be increased in section 502b by replacing at least some of the video frames previously assigned to time layers "Layer 0" and "Layer 1" with those previously assigned to one or more of time layers "Layer 2" or "Layer 3" (e.g., causing a reduction in the frame rate of quality level "Level 1"). In this example, the frame rate of quality level "Level 0" is reduced to half its original value (e.g., rendered every eight frames before modification, and every 16 frames after modification). Furthermore, in this example, the frame rate of quality level "Level 1" is also reduced to half its original value (e.g., rendered every four frames before modification, and every eight frames after modification).

[0101] Furthermore, in some specific implementations, one or more video frames from the temporal layer can be omitted or "discarded" from the data structure. For example, in Figure 5In the illustrated example, the sampling period for quality level "Level 2" can be increased by alternatively assigning video frames "2" and "6" (which have previously been assigned to temporal layer "Layer 2") to temporal layer "Layer 3," resulting in a decrease in the frame rate for quality level "Level 2." Additionally, the sampling period for quality level "Level 3" can be increased by omitting or discarding video frames "1," "3," "5," and "7" (which have previously been assigned to temporal layer "Layer 3"), resulting in a decrease in the frame rate for quality level "Level 3."

[0102] In some implementations, the encoder 202 can determine whether to modify the data structure periodically or in response to certain conditions. For example, each time a video frame is assigned to the lowest temporal layer (e.g., "Layer 0"), the encoder 202 can determine whether to maintain the current configuration of the data structure for encoding future video frames (e.g., maintain the frame rate for each of the corresponding quality levels or levels) or to modify the configuration of the data structure for encoding future frames (e.g., modify the frame rate for one or more of the corresponding quality levels or levels).

[0103] Referring back to Figure 2 After generating the encoded content 212 (e.g., including a data structure encoded using one or more of the techniques described herein), the encoder 202 provides the encoded content 212 to the decoder 204 for processing. In some implementations, the encoded content 212 can be transmitted to the decoder 204 via the network 106 (e.g., as described with reference to Figure 1

[0104] In some implementations, the electronic devices 102a-102c can transmit at least some of the encoded content 212 generated by their encoders 202 directly to another electronic device 102a-102c for decoding, rendering, and output. In some implementations, the electronic devices 102a-102c can transmit at least some of the encoded content 212 generated by their encoders 202 to the video conference server 104, which in turn can transmit at least some of the encoded content 212 to one or more electronic devices 102a-102c for decoding, rendering, and output.

[0105] In some implementations, at least some of the encoded content 212 can be transmitted continuously in real-time between the electronic devices 102a-102c and / or the video conference server 104 (e.g., to facilitate a real-time video conference or video call). For example, at least some of the encoded content 212 can be transmitted in real-time between the electronic devices 102a-102c and / or the video conference server 104 as a continuous stream of data (e.g., a bitstream). ​

[0106] In some implementations, the electronic devices 102a-102c can generate a single version of the encoded content 212 and transmit an instance of the encoded content 212 to other ones of the electronic devices 102a-102c and / or the video conference server 104. In some implementations, the electronic devices 102a-102c can generate multiple versions of the encoded content 212 (e.g., each version corresponding to a different quality level or grade) and transmit different versions of the encoded content 212 to other ones of the electronic devices 102a-102c and / or the video conference server 104. For example, the first electronic device 102a can selectively transmit a first set of temporal layers of the data structure to the second electronic device 102b (e.g., to provide video content having a first quality level or grade) and selectively transmit a second set of temporal layers of the data structure to the third electronic device 102c (e.g., to provide video content having a second quality level or grade).

[0107] The decoder 204 receives the encoded content 212 and extracts information about at least some of the video frames 210 included in the encoded content 212 (e.g., in the form of decoded data 214). For example, the decoder 204 can extract information about at least some of the video frames 210 from the encoded data structure of the encoded content 212, such as the content of each of the video frames 210 (e.g., the colors, textures, visual patterns, opacity, and / or other characteristics of the video frames 210). As another example, the encoder 202 can determine relationships between at least some of the video frames 210 (e.g., the order of the video frames 210, prediction vectors between the video frames 210, etc.) and use this information to recreate the video frames (or approximations thereof).

[0108] In some implementations, the decoder 204 can selectively decode a subset of the encoded content 212 and selectively refrain from decoding other portions of the encoded content 212. For example, the decoder 204 can selectively decode portions of the encoded data structure corresponding to certain temporal layers to obtain video frames (e.g., corresponding to a desired quality level or grade), while selectively refraining from decoding portions of the encoded data structure corresponding to other temporal layers (e.g., corresponding to a high quality level or grade).

[0109] As one example, with reference to Figure 3B To provide video content having the highest quality level (“Grade 3”), the decoder 204 can decode the data structure 300 to selectively extract video frames from each of the temporal layers “Tier 0” through “Tier 3.”

[0110] As another example, to provide video content having the next highest level of quality (“Level 2”), the decoder 204 can decode the data structure 300 to selectively extract video frames from each of temporal layers “Layer 0” through “Layer 2,” and refrain from extracting video frames from temporal layer “Layer 3.”

[0111] As another example, to provide video content having the next highest level of quality (“Level 1”), the decoder 204 can decode the data structure 300 to selectively extract video frames from each of temporal layers “Layer 0” through “Layer 1,” and refrain from extracting video frames from temporal layers “Layer 2” and “Layer 3.”

[0112] As another example, to provide video content having the next highest level of quality (“Level 1”), the decoder 204 can decode the data structure 300 to selectively extract video frames from each of temporal layers “Layer 0” through “Layer 1,” and refrain from extracting video frames from temporal layers “Layer 2” and “Layer 3.”

[0113] The decoder 204 provides the decoded data 214 to the Tenderer 206. The Tenderer 206 renders video content based on the decoded data 214, and presents the rendered content to a user using the output device 208. As one example, if the output device 208 is configured to present content according to two dimensions (e.g., using a flat-panel display such as a liquid crystal display or a light-emitting diode display), the Tenderer 206 can render video content according to two dimensions and according to a particular viewing angle, and instruct the output device 208 to display the content accordingly. As another example, if the output device 208 is configured to present content according to three dimensions (e.g., using a holographic display or a headset), the Tenderer 206 can render content according to three dimensions and according to a particular viewing angle, and instruct the output device 208 to display the content accordingly.

[0114] In some implementations, the data structures described herein can provide a degree of resilience against frame loss. For example, for the example data structure 300 shown in FIG. 3A, if information regarding video frame “0” is lost (e.g., due to corruption of the data structure), the decoder will need to re-request information regarding video frame “0” in order to accurately decode video frames “1” through “7” (e.g., because each of video frames “1” through “7” at least indirectly depends on video frame “0”). However, if information regarding video frame “5” is lost, the decoder will still be able to accurately decode each of the other video frames (e.g., because none of the other video frames depend on video frame “5”). Figure 3B

[0115] ​The data structure can be encoded in a manner that utilizes such resilience against frame loss. As one example, due to the arrangement of the temporal layers of the data structure, a corruption in a lower temporal layer has as high a likelihood of being perceived by a user as a corruption in a higher temporal layer (e.g., because a loss of a video frame in a lower layer is more likely to result in a loss of other video frames in a higher temporal layer due to dependencies between them). Thus, a different degree of error correction can be applied to each of the temporal layers. This can be beneficial because the implementation of error correction can come with some computational and / or network overhead. By selectively applying a different degree of error correction to each temporal layer, the data structure can be made more resilient to errors in a resource-efficient manner.

[0116] As one example, forward error correction (FEC) can be applied to at least some of the temporal layers. Further, the degree of FEC applied to each of the temporal layers can be less than the degree of FEC applied to temporal layers below it. For example, in the illustrated example, 100% FEC can be applied to temporal layer “Layer 0,” 50% FEC can be applied to temporal layer “Layer 1,” 25% FEC can be applied to temporal layer “Layer 2,” and no FEC can be applied to temporal layer “Layer 3.” Although example degrees of FEC are described above, these are merely illustrative examples. In practice, other degrees of FEC can be used in addition to or instead of the FEC degrees described above. Figure 3B

[0117] Further, in some implementations, the electronic devices 102a-102c can selectively apply error correction to some or all of the temporal layers of the data structure depending on the intended recipient of the data structure. For example, if a first intended recipient of the data structure is experiencing frame loss, the electronic devices 102a-102c can selectively apply error correction to some or all of the temporal layers of the data structure (e.g., corresponding to video frames intended for transmission to the first recipient) and provide at least a portion of the data structure to the first recipient. As another example, if a second intended recipient of the data structure is not experiencing frame loss, the electronic devices 102a-102c can selectively refrain from applying error correction to some or all of the temporal layers of the data structure (e.g., corresponding to video frames intended for transmission to the second recipient) and provide at least a portion of the data structure to the second recipient.

[0118] ​In some implementations, the source of the encoded content (e.g., electronic devices 102a-102c) may notify each of the several recipients of the encoded content (e.g., other electronic devices in electronic devices 102a-102c) of several available quality levels or grades for receiving the encoded content. Furthermore, the source of the encoded content may provide different data structures and / or portions of data structures based on each recipient's selection. As an example, if a recipient selects the highest quality level or grade, the source may encode a data structure having each of the available time layers and transmit that data structure to the recipient. As another example, if another recipient selects a lower quality level or grade, the source may encode a data structure having a subset of time layers (e.g., omitting the highest time layer) and transmit that data structure to the recipient.

[0119] In some specific implementations, the source of the encoded content can be defined as QUOTE. A set of m quality levels or grades, and n of these quality levels or grades are provided simultaneously to one or more recipients upon request. For example, the source of the encoded content will be m QUOTE A set of quality levels or grades is announced to the recipients for selection, and the encoded content is provided to each of the recipients based on their selection (e.g., by mapping the selected quality levels or grades to one or more of the quality levels or grades in the data structure described herein).

[0120] In some specific implementations, the recipient may request a quantity of quality level or grade. o, this quantity is less than or equal to the number of quality levels or grades that the source of the encoded content simultaneously provides to the receiver. n QUOTE In this case, the source of the encoded content maps each of the selected quality levels or grades to the corresponding quality level or grade of the data structure, and transmits at least a portion of the data structure to each receiver.

[0121] For example, the source of the encoded content can define a set of seven quality levels or grades, corresponding to bit rates of 512kbps, 1Mbps, 2Mbps, 3Mbps, 4Mbps, 5Mbps, and 6Mbps, respectively. Furthermore, the source of the encoded content can simultaneously provide up to four quality levels or grades to the receiver. If the receiver selects four quality levels or grades (e.g., corresponding to bit rates of 512kbps, 1Mbps, 3Mbps, and 6Mbps), the source of the encoded content dynamically maps each of the selected quality levels to the corresponding quality level or grade of the data structure and transmits at least a portion of the data structure to each receiver. For example, see reference...Figure 3B In the illustrated example, a bit rate of 512 kbps can be mapped to a "Tier 0" of the data structure, a bit rate of 1 Mbps can be mapped to a "Tier 1" of the data structure, a bit rate of 3 Mbps can be mapped to a "Tier 2" of the data structure, and a bit rate of 6 Mbps can be mapped to a "Tier 3" of the data structure.

[0122] In some implementations, a recipient can request a number o of quality levels or tiers that is greater than the number n of quality levels or tiers of the encoded content, while the recipient is provided with a number QUOTE n of quality levels or tiers. In this case, the source of the encoded content can select a subset of the o quality levels or tiers selected by the recipient (e.g., all n QUOTE In some implementations, a recipient can request a number o of quality levels or tiers that is greater than the number n of quality levels or tiers of the encoded content, while the recipient is provided with a number QUOTE

[0123] For example, the source of the encoded content can define a set of seven quality levels or tiers corresponding to bit rates of 512 kbps, 1 Mbps, 2 Mbps, 3 Mbps, 4 Mbps, 5 Mbps, and 6 Mbps, respectively. In addition, the source of the encoded content can provide up to four quality levels or tiers to a recipient. If the recipient selects five quality levels or tiers (e.g., corresponding to bit rates of 512 kbps, 1 Mbps, 2 Mbps, 4 Mbps, and 6 Mbps), the source of the encoded content dynamically maps four of the selected quality levels to respective quality levels or tiers of the data structure and transmits at least a portion of the data structure to each recipient.

[0124] In some implementations, the source of the encoded content can dynamically map the n lowest quality levels or tiers selected by the user. For example, in the above example, a bit rate of 512 kbps can be mapped to a "Tier 0" of the data structure, a bit rate of 1 Mbps can be mapped to a "Tier 1" of the data structure, a bit rate of 2 Mbps can be mapped to a "Tier 2" of the data structure, and a bit rate of 4 Mbps can be mapped to a "Tier 3" of the data structure. In addition, a bit rate of 6 Mbps can be omitted from the quality levels or tiers of the data structure.

[0125] Furthermore, the source of encoded content can dynamically change the mapping of quality levels or tiers of the data structure. As one example, to accommodate higher quality video content, the source of encoded content can dynamically increase the bit rate mapped to one or more of the quality levels or tiers (e.g., increase the bit rate mapped to “Tier 3” from 4 Mbps to 5 Mbps). As another example, to accommodate lower quality video content, the source of encoded content can dynamically decrease the bit rate mapped to one or more of the quality levels or tiers (e.g., decrease the bit rate mapped to “Tier 3” from 4 Mbps to 3 Mbps).

[0126] Although an exemplary set of quality levels or tiers is described above, these are merely illustrative examples. In practice, any set of quality levels or tiers (e.g., corresponding to any set of corresponding bit rates) can be used, depending on the particular implementation.

[0127] Exemplary Process

[0128] Figure 6 An exemplary process 600 of encoding data in a data structure regarding video frames using one or more of the techniques described herein is shown. In some implementations, process 600 can be performed, at least in part, by encoder 202 (e.g., implemented in one or more of electronic devices 102a-102c).

[0129] According to process 600, the system receives a video frame for encoding (block 602). As one example, the video frame can be a frame or image from a video generated by the electronic device (e.g., captured by a camera subsystem of the electronic device during a video conference or video telephony session).

[0130] The system determines a temporal tier for the received video frame (block 602). For example, with reference to FIGS. 3-5, the system can determine the temporal tier based on the order of the video frame in the sequence (e.g., an index number associated with the video frame) and the sampling period for each of the number of quality levels or tiers such that the sequence of video frames is presented according to the particular frame rate for each of the quality levels or tiers (e.g., as described above). Figure 5 Exemplary techniques for selecting a temporal tier for a video frame are described. As one example, the temporal tier can be selected based on the order of the video frame in the sequence (e.g., an index number associated with the video frame) and the sampling period for each of the number of quality levels or tiers such that the sequence of video frames is presented according to the particular frame rate for each of the quality levels or tiers (e.g., as described above).

[0131] Furthermore, the system determines an encoding parameter for the video frame (block 606). As one example, the system can determine the encoding parameter based on the selected temporal tier (e.g., based on a rate-QP model for that temporal tier, as described above).

[0132] Additionally, the system encodes the video frames according to the determined encoding parameters (block 608). As one example, the system can store data in a data structure that represents the content of the video frames, and assign the video frames to the determined temporal layers in the data structure.

[0133] If the video frames are assigned to the lowest temporal layer (e.g., “Layer 0”), the system determines whether to update the arrangement of the data structure (block 610). As one example, the system can determine whether to modify the arrangement of the data structure to increase the frame rate of the presented content according to one or more quality levels or tiers (e.g., as referenced Figure 4 to the quality levels or tiers) or decrease the frame rate of the presented content according to one or more quality levels or tiers (e.g., as referenced Figure 5 to the quality levels or tiers).

[0134] The system can repeat the process 600 until no video frames remain (e.g., until the end of the video).

[0135] Figure 7 An example process 700 is shown for modifying a data structure using one or more of the techniques described herein. In some implementations, the process 700 can be performed, at least in part, by the encoder 202 (e.g., implemented in one or more of the electronic devices 102a-102c).

[0136] According to the process 700, the system selects a quality level or tier for processing (block 702). In some implementations, the system can start with the second highest quality level or tier, and process each quality level or tier in order until the lowest quality level or tier.

[0137] Additionally, the system receives encoding data for the selected tier (block 704). For example, the system can determine a current encoding bitrate for the selected quality level or tier, and a bitrate that has been allocated to the quality level or tier (e.g., a target bitrate for the quality level or tier corresponding to a particular service level).

[0138] If the current encoding bitrate is equal to the allocated bitrate (e.g., within a particular error bound or tolerance), the system can maintain the frame rate for the quality level or tier (block 706). For example, the system can maintain the same arrangement for the data structure such that the video frames are assigned to each temporal layer of the data structure according to the same pattern as before.

[0139] If the current encoding bitrate is greater than the allocated bitrate (e.g., greater than a particular error bound or tolerance), the system can decrease the frame rate for the quality level or tier (block 708). For example, with reference to Figure 5Exemplary techniques for reducing the frame rate of a quality level or tier are described.

[0140] If the current encoding bitrate is less than the allocated bitrate (e.g., less than a particular error bound or tolerance), the system can determine whether increasing the frame rate of the quality level or tier is feasible. If so, the system can increase the frame rate of the quality level or tier (block 710). Otherwise, the system can maintain the frame rate of the quality level or tier (block 706). For example, referring to FIG. 7, the system can determine whether the current encoding bitrate is less than the allocated bitrate (block 702). If so, the system can determine whether increasing the frame rate of the quality level or tier is feasible (block 704). If so, the system can increase the frame rate of the quality level or tier (block 710). Otherwise, the system can maintain the frame rate of the quality level or tier (block 706). Figure 4 Exemplary techniques for increasing the frame rate of a quality level or tier are described.

[0141] In some implementations, the system can determine whether increasing the frame rate is feasible by determining a bitrate required to encode the video frames according to the higher frame rate, and determining whether the new bitrate would exceed the bitrate allocated for the quality level or tier (e.g., exceed a particular error bound or tolerance). If so, the system can determine that increasing the frame rate is not feasible. Otherwise, the system can determine that increasing the frame rate is feasible.

[0142] In some implementations, the system can determine whether increasing the frame rate is feasible by determining whether increasing the frame rate would exceed the frame rate at which the video was originally captured. If so, the system can determine that increasing the frame rate is not feasible. Otherwise, the system can determine that increasing the frame rate is feasible.

[0143] The system can repeat the process 700 until no tiers remain.

[0144] Figure 8 An exemplary process 800 for encoding visual content is shown. The process 800 can be performed, at least in part, using one or more devices (e.g., a computer system, such as the computer systems shown in FIG. 1, FIG. 2, and FIG. 3). Figure 9 The one or more computer systems shown can be used to perform the process 800.

[0145] According to the process 800, the device receives a plurality of frames of a video (block 802).

[0146] In addition, the device generates a data structure representing the video and representing a plurality of temporal layers (block 804). In some implementations, the data structure can include a group of pictures (GOP) structure. For example, exemplary data structures are shown in FIG. 4, FIG. 5, and FIG. 6. Figure 3B 、 Figure 4 and Figure 5 Exemplary data structures are shown in FIG. 4, FIG. 5, and FIG. 6.

[0147] The device generates a data structure at least in part by determining a plurality of quality levels for presenting the video, where each of the quality levels corresponds to a different respective sampling period for sampling the frames of the video. Further, the device assigns each of the frames to a respective temporal layer of the temporal layers of the data structure based on the sampling period. Moreover, the device indicates in the data structure one or more relationships between (i) at least one of the frames assigned to at least one of the temporal layers of the data structure and (ii) at least another one of the frames assigned to at least another one of the temporal layers of the data structure.

[0148] In some implementations, for each of the frames, assigning each of the frames to a respective temporal layer of the temporal layers of the data structure can include (i) determining an index number associated with the frame, (ii) identifying from the quality levels a particular quality level of the plurality of quality levels that corresponds to a sampling period that is divisible by the index number, and (iii) assigning the frame to one of the temporal layers based on the identified quality level.

[0149] In some implementations, the frames are assigned to the temporal layer that has an index value that is the same as an index value of the identified quality level.

[0150] In some implementations, identifying the particular quality level can include (i) identifying from the particular quality levels a subset of the quality levels, each of the quality levels of the subset corresponding to a respective sampling period that is divisible by the index number, and (ii) selecting from the subset the quality level that has a largest sampling period.

[0151] In some implementations, the plurality of quality levels can include a first quality level that corresponds to a first sampling period and a second quality level that corresponds to a second sampling period, where the first sampling period is a multiple of the second sampling period. In some implementations, the plurality of quality levels can also include a third quality level that corresponds to a third sampling period, where the second sampling period is a multiple of the third sampling period.

[0152] In some implementations, the frames assigned to a first temporal layer of the temporal layers of the data structure can be encoded according to a first bitrate. Further, the frames assigned to a second temporal layer of the temporal layers of the data structure can be encoded according to a second bitrate, where the first bitrate is different than the second bitrate.

[0153] In some implementations, frames assigned to a first temporal layer of the temporal layers of the data structure can be encoded according to a first quantization parameter. In addition, frames assigned to a second temporal layer of the temporal layers of the data structure can be encoded according to a second quantization parameter, where the first quantization parameter is different than the second quantization parameter.

[0154] In addition, the device outputs the data structure (block 806). In some implementations, outputting the data structure can include transmitting a bitstream that includes the data structure (or a portion thereof).

[0155] In some implementations, process 800 can also include generating the video using one or more cameras of the first mobile device, and transmitting the data structure from the first mobile device to one or more second mobile devices via a communication network. In addition, the video can include visual content for a communication session between the first mobile device and the one or more second mobile devices.

[0156] Example computer system

[0157] Figure 9 is used to implement the features and processes described with reference to Figures 1 to 8 The block diagram of an example device architecture 900 for implementing the features and processes described. For example, architecture 900 can be used to implement electronic devices 102a-102c, video conferencing server 104, and / or any of the components described with reference to Figure 2 The architecture 900 can be implemented in any device used to generate the features described with reference to Figures 1 to 8 The architecture 900 can be implemented in any device used to generate the features described with reference to

[0158] Architecture 900 can include a memory interface 902, one or more data processors 904, one or more data coprocessors 974, and a peripheral interface 906. Memory interface 902, one or more processors 904, one or more coprocessors 974, and / or peripheral interface 906 can be independent components, or can be integrated into one or more integrated circuits. One or more communication buses or signal lines can couple the various components.

[0159] The one or more processors 904 and / or the one or more coprocessors 974 can operate in conjunction with each other to perform the operations described herein. For example, the one or more processors 904 can include one or more central processing units (CPUs) configured to act as a host computer processor of the architecture 900. For example, the one or more processors 904 can be configured to perform generalization data processing tasks for the architecture 900. Additionally, at least some of the data processing tasks can be offloaded to the one or more coprocessors 974. For example, specialized data processing tasks, such as processing motion data, processing image data, encrypting data, and / or performing certain types of arithmetic operations, can be offloaded to one or more specialized coprocessors 974 for handling these tasks. In some cases, the one or more processors 904 can be relatively more powerful and / or can consume more power than the one or more coprocessors 974. For example, this can be useful because it enables the one or more processors 904 to quickly process generalization tasks while also offloading certain other tasks to the one or more coprocessors 974, which can perform those tasks more efficiently and / or effectively. In some cases, the one or more coprocessors can include one or more sensors or other components (e.g., as described herein) and can be configured to process data acquired using these sensors or components and provide the processed data to the one or more processors 904 for further analysis.

[0160] Sensors, devices, and subsystems can be coupled to the peripherals interface 906 to facilitate multiple functionalities. For example, a motion sensor 910, a light sensor 912, and a proximity sensor 914 can be coupled to the peripherals interface 906 to facilitate orientation, lighting, and proximity functions of the architecture 900. For example, in some implementations, the light sensor 912 can be utilized to adjust the brightness of a touch screen 946 based on available ambient light. In some implementations, a motion sensor 910 can be utilized to detect motion of the device. For example, the motion sensor 910 can include one or more accelerometers (e.g., to measure acceleration, which can in turn be used to measure the acceleration of the motion sensor 910 and / or architecture 900 over time) and / or one or more gyroscopes (e.g., to measure the orientation of the motion sensor 910 and / or the mobile device). In some cases, measured information acquired by the motion sensor 910 can take the form of one or more time-varying signals (e.g., a time-varying profile of acceleration and / or orientation over a period of time). Additionally, display objects or media can be presented according to a detected orientation (e.g., according to a "portrait" orientation or a "landscape" orientation). In some cases, the motion sensor 910 can be integrated directly into a coprocessor 974 that is configured to process measurements acquired by the motion sensor 910. For example, the coprocessor 974 can include one or more accelerometers, gyroscopes, and / or

[0161] Other sensors can also be connected to the peripherals interface 906, such as temperature sensors, biometric sensors, or other sensing devices to facilitate relevant functionality. For example, as shown, the architecture 900 can include a heart rate sensor 932 that measures the user's heartbeats. Similarly, these other sensors can also be integrated directly into one or more coprocessors 974 that are configured to process measurements acquired from those sensors. Figure 9

[0162] A position processor 915 (e.g., a GNSS receiver chip) can be connected to the peripherals interface 906 to provide geographic reference. An electronic magnetometer 916 (e.g., an integrated circuit chip) can also be connected to the peripherals interface 906 to provide data that can be used to determine the magnetic North direction. Thus, the electronic magnetometer 916 can be used as an electronic compass.

[0163] A camera subsystem 920 and an optical sensor 922, e.g., a charged coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) optical sensor, can be utilized to facilitate camera functions, such as recording photographs and video clips.

[0164] ​Communication functions can be facilitated through one or more communication subsystems 924. The communication subsystems 924 can include one or more wireless and / or wired communication subsystems. For example, wireless communication subsystems can include radio frequency receivers and transmitters and / or optical (e.g., infrared) receivers and transmitters. As another example, wired communication subsystems can include a port device (e.g., a universal serial bus (USB) port) or some other wired port connection that can be used to establish a wired connection to other computing devices, such as other communication devices, network access devices, personal computers, printers, display screens, or other processing devices capable of receiving or transmitting data.

[0165] The specific design and implementation of the communication subsystems 924 can depend on the communication network or networks with which the architecture 900 is intended to operate. For example, the architecture 900 can include communication subsystems designed to operate over a global system for mobile communications (GSM) network, a GPRS network, an enhanced data GSM environment (EDGE) network, an 802.x communication network (e.g., Wi-Fi, Wi-Max), a code division multiple access (CDMA) network, an NFC and Bluetooth® network, or some other ™ wireless communication subsystems that operate over a network. The wireless communication subsystems can also include host protocols so that the architecture 900 can be configured as a base station for other wireless devices. As another example, the communication subsystems can use one or more protocols, such as TCP / IP protocols, HTTP protocols, UDP protocols, and any other known protocols to allow the architecture 900 to synchronize with host devices.

[0166] An audio subsystem 926 can be coupled to a speaker 928 and one or more microphones 930 to facilitate voice-enabled functionality, such as voice recognition, voice replication, digital recording, and telephony functions.

[0167] The I / O subsystem 940 can include a touch controller 942 and / or other input controller(s) 944. The touch controller 942 can be coupled to a touch surface 946. The touch surface 946 and the touch controller 942 can, for example, detect contact and movement or break thereof using any of a plurality of touch sensitivity technologies, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch surface 946. In one implementation, the touch surface 946 can display virtual or soft buttons and virtual keyboards that a user can use as an input / output device.

[0168] Other input controller(s) 944 can be coupled to other input / control devices 948, such as one or more buttons, rocker switches, thumbwheel, infrared port, USB port, and / or a pointer device such as a stylus. The one or more buttons (not shown) can include an up / down button for volume control of the speaker 928 and / or the microphone 930.

[0169] In some implementations, the architecture 900 can present recorded audio and / or video files, such as MP3, AAC, and MPEG video files. In some implementations, the architecture 900 can include functionality of an MP3 player and can include a pin connector for connecting to other devices. Other input / output and control devices can be used.

[0170] The memory interface 902 can be coupled to the memory 950. The memory 950 can include high-speed random access memory or nonvolatile, such as one or more disk storage devices, one or more optical storage devices, or flash memory storage (e.g., NAND, NOR). The memory 950 can store operating system 952, such as Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or an embedded operating system such as VxWorks. The operating system 952 can include instructions for handling basic system services and for performing hardware dependent tasks. In some implementations, the operating system 952 can include a kernel (e.g., a UNIX kernel).

[0171] The memory 950 can also store communication instructions 954 to facilitate communicating with one or more additional devices, one or more computers or servers communicating through a one or more additional devices, including peer-to-peer communication. The communication instructions 954 can further be used to select an operating mode or communication medium for the device based on a geographic location of the device (obtained by the GPS / navigation instructions 968). The memory 950 can include graphical user interface instructions 956 to facilitate graphic user interface processing; sensor processing instructions 958 to facilitate sensor-related processing and functions; phone instructions 960 to facilitate phone-related processes and functions; electronic messaging instructions 962 to facilitate electronic-messaging related processes and functions; web browsing instructions 964 to facilitate web browsing-related processes and functions; media processing instructions 966 to facilitate media processing-related processes and functions; GPS / Navigation instructions 969 to facilitate GPS and navigation-related processes; camera instructions 970 to facilitate camera-related processes and functions; and other instructions 972 for performing some or all of the processes described herein. The memory 950 can also include instructions 974 for controlling the frequency of the clock 980.

[0172] Each of the above identified instructions and applications can correspond to a set of instructions for performing one or more functions described herein. These instructions need not be implemented as separate software programs, procedures, or modules. Memory 950 can include additional or fewer instructions. Furthermore, various functions of the device can be performed in the

[0173] The described features can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. The features can be implemented in a computer program product tangibly embodied in an information carrier, for example, in a machine-readable storage device, for execution by a programmable processor; and method steps can be performed by a programmable processor acting in response to

[0174] The described features can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one input device, at least one output device, and at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, the data storage system. Computer programs are sets of instructions that can be used directly or indirectly in a computer to perform activities or bring about a certain result. Computer programs can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and they can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0175] Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from, or input to, or output from, or combinations, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits).

[0176] To provide for interaction with a user, the features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the author and a keyboard and a pointing device such as a mouse or a trackball by which the author can provide input to the computer.

[0177] The features can be implemented in a computer system that includes a back-end component such as a data server, or that includes a middleware component such as an application server or an Internet server, or that includes a front-end component such as a client computer having a graphical user interface or an Internet browser, or any combination of them. The components of the system can be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include a LAN, a WAN and the computers and networks forming the Internet.

[0178] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0179] One or more features or steps of the disclosed implementations can be implemented using an application programming interface (API). An API can define one or more parameters that are passed between calling applications and other software code (e.g., an operating system, a library, a function) that provides a service, provides data, or performs an operation or computation.

[0180] An API can be implemented as one or more calls in program code that send or receive one or more parameters through a parameter list or other structure based on a call convention defined in an API specification document. A parameter can be a constant, a key, a data structure, an object, an object class, a variable, a data type, a pointer, an array, a list, or another call. API calls and parameters can be implemented in any programming language. A programming language can define a vocabulary and a call convention for programmers to use when accessing functions supported by an API.

[0181] In some implementations, an API call can report to an application the capabilities of a device running the application, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, and the like.

[0182] A number of implementations have been described. Nevertheless, it will be understood that various modifications can be made. Elements of one or more implementations can be combined, deleted, modified, or supplemented to form additional implementations. As another example, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve the desired results. Further, other steps can be provided, or steps can be eliminated, from the described flows, and other components can be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.

Claims

1. A method for encoding video content, the method comprising: receiving, by one or more processors, a plurality of frames of a video; generating, by the one or more processors, a data structure representing the video, wherein the data structure represents a plurality of temporal layers, each temporal layer corresponding to a different respective quality level of a plurality of quality levels for presenting the video, and wherein generating the data structure comprises: determining that a first portion of the video has a first complexity and a second portion of the video has a second complexity different from the first complexity, for the first portion of the video: determining that each of the quality levels corresponds to a different respective sampling period for sampling frames of the first portion of the video, wherein for each sampling period, a value N of the sampling period indicates that every Nth frame of the video is sampled, assigning, based on the sampling periods, each of the frames of the first portion of the video to a respective one of the temporal layers of the data structure, and for the second portion of the video: modifying the plurality of quality levels for presenting the second portion of the video, wherein modifying the plurality of quality levels comprises modifying values of at least some of the sampling periods of the quality levels, and assigning, based on the modified sampling periods, each of the frames of the second portion of the video to a respective one of the temporal layers of the data structure, and indicating, in the data structure, one or more relationships between (i) at least one of the frames assigned to at least one of the temporal layers of the data structure and (ii) at least another one of the frames assigned to at least another one of the temporal layers of the data structure; and outputting, by the one or more processors, the data structure.

2. The method of claim 1, wherein the data structure comprises a group of pictures (GOP) structure.

3. The method of claim 1, wherein outputting the data structure comprises transmitting a bitstream comprising the data structure.

4. The method of claim 1, further comprising: generating the video using one or more cameras of a first mobile device, and transmitting the data structure from the first mobile device to one or more second mobile devices via a communication network.

5. The method of claim 4, wherein the video comprises visual content for a communication session between the first mobile device and the one or more second mobile devices.

6. The method of claim 1, wherein assigning each of the frames to the respective one of the temporal layers of the data structure comprises, for each of the frames: determining an index number associated with the frame, identifying, from the quality levels, a particular quality level of the plurality of quality levels that corresponds to a sampling period that is divisible by the index number, and assigning the frame to one of the temporal layers based on the identified quality level. ​ 7. The method of claim 6, wherein the frames are assigned to temporal layers having an index value that is the same as an index value of the identified quality level.

8. The method of claim 6, wherein identifying the particular quality level comprises: identifying, from the particular quality level, a subset of the quality levels, each quality level in the subset corresponding to a respective sampling period that is divisible by the index number, and selecting, from the subset, a quality level having a largest sampling period.

9. The method of claim 1, wherein the plurality of quality levels comprises: a first quality level corresponding to a first sampling period, and a second quality level corresponding to a second sampling period, wherein the first sampling period is a multiple of the second sampling period.

10. The method of claim 9, wherein the plurality of quality levels further comprises a third quality level corresponding to a third sampling period, wherein the second sampling period is a multiple of the third sampling period.

11. The method of claim 1, further comprising: encoding frames assigned to a first temporal layer of the temporal layers of the data structure according to a first bit rate, and encoding frames assigned to a second temporal layer of the temporal layers of the data structure according to a second bit rate, wherein the first bit rate is different than the second bit rate.

12. The method of claim 1, further comprising: encoding frames assigned to a first temporal layer of the temporal layers of the data structure according to a first quantization parameter, and encoding frames assigned to a second temporal layer of the temporal layers of the data structure according to a second quantization parameter, wherein the first quantization parameter is different than the second quantization parameter.

13. A system for encoding video content, the system comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving a plurality of frames of a video; generating a data structure representing the video, wherein the data structure represents a plurality of temporal layers, each temporal layer corresponding to a different respective quality level of a plurality of quality levels for presenting the video, and wherein generating the data structure comprises: determining that a first portion of the video has a first complexity and a second portion of the video has a second complexity different than the first complexity, for the first portion of the video: determining that each quality level of the quality levels corresponds to a different respective sampling period for sampling frames of the first portion of the video, wherein for each sampling period, a value N of the sampling period indicates that every Nth frame of the video is sampled, assigning each frame of the frames of the first portion of the video to a respective temporal layer of the temporal layers of the data structure based on the sampling periods, and for the second portion of the video: ​ modifying the plurality of quality levels used to present the second portion of the video, wherein modifying the plurality of quality levels comprises modifying values of at least some of the sampling periods of the quality levels, and assigning each of the frames of the second portion of the video to a respective one of the temporal layers of the data structure based on the modified sampling periods, and indicating in the data structure one or more relationships between (i) at least one of the frames assigned to at least one of the temporal layers of the data structure and (ii) at least another one of the frames assigned to at least another one of the temporal layers of the data structure; and outputting the data structure.

14. The system of claim 13, wherein the data structure comprises a group of pictures (GOP) structure, and wherein outputting the data structure comprises transmitting a bitstream comprising the data structure.

15. The system of claim 13, the operations further comprising: generating the video using one or more cameras of a first mobile device, and transmitting the data structure from the first mobile device to one or more second mobile devices via a communication network, wherein the video comprises visual content for a communication session between the first mobile device and the one or more second mobile devices.

16. The system of claim 13, wherein assigning each of the frames to the respective one of the temporal layers of the data structure comprises, for each of the frames: determining an index number associated with the frame, identifying, from the quality levels, a particular quality level of the plurality of quality levels that corresponds to a sampling period that is divisible by the index number, and assigning the frame to one of the temporal layers based on the identified quality level, wherein the frame is assigned to the temporal layer that has an index value that is the same as an index value of the identified quality level.

17. The system of claim 16, wherein identifying the particular quality level comprises: identifying, from the quality levels, a subset of the quality levels, each quality level of the subset corresponding to a respective sampling period that is divisible by the index number, and selecting, from the subset, a quality level that has a largest sampling period.

18. The system of claim 13, wherein the plurality of quality levels comprises: a first quality level corresponding to a first sampling period, and a second quality level corresponding to a second sampling period, wherein the first sampling period is a multiple of the second sampling period.

19. The system of claim 13, the operations further comprising: encoding frames assigned to a first one of the temporal layers of the data structure according to a first bit rate, and encoding frames assigned to a second one of the temporal layers of the data structure according to a second bit rate, wherein the first bit rate is different from the second bit rate.

20. The system of claim 13, the operations further comprising: ​ encoding frames assigned to a first temporal layer of the temporal layers of the data structure according to a first quantization parameter, and encoding frames assigned to a second temporal layer of the temporal layers of the data structure according to a second quantization parameter, wherein the first quantization parameter is different from the second quantization parameter.