Identification and marking of video data units for network transfer of video data.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2023-06-08
- Publication Date
- 2026-06-02
AI Technical Summary
Existing video compression technologies face challenges in efficiently processing and delivering video data over networks due to the lack of access to payload data by network devices, leading to issues with packet delivery and decoding, especially in the context of Extended Reality (XR) media services, where unified and differentiated PDU set processing are not adequately addressed.
Incorporating identification information, such as picture order count (POC) and coding layer identifiers, into the packet header outside the payload allows network devices to determine frame status and processing requirements, enabling efficient delivery and decoding of video data without accessing the encrypted payload.
This approach enhances the delivery and decoding efficiency of video data by allowing network devices to manage frame reception and decoding processes effectively, optimizing the consumption experience for XR media services.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to U.S. Patent Application No. 18 / 329,994, filed Jun. 6, 2023, and U.S. Provisional Patent Application No. 63 / 381,902, filed Nov. 1, 2022. U.S. Patent Application No. 18 / 329,994, filed Jun. 6, 2023, claims the benefit of U.S. Provisional Patent Application No. 63 / 381,902, filed Nov. 1, 2022, and the entire contents of each application are hereby incorporated by reference.
[0002] This disclosure relates to the transfer of encoded video data.
Background Art
[0003] Digital video capabilities can be incorporated into a wide range of devices including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular phones or satellite radiotelephones, video teleconferencing devices, and the like. Digital video devices implement video compression techniques such as those defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also referred to as High Efficiency Video Coding (HEVC)), and those described in extensions of such standards.
[0004] Video compression techniques perform spatial and / or temporal predictions to reduce or eliminate redundancy inherent in video sequences. In block-based video coding, a video frame or slice can be divided into macroblocks. Each macroblock can be further divided. Macroblocks in an intracoded (I) frame or slice are encoded using spatial predictions with respect to adjacent macroblocks. Macroblocks in an intercoded (P or B) frame or slice can use spatial predictions with respect to adjacent macroblocks within the same frame or slice, or temporal predictions with respect to other reference frames.
[0005] After video data has been encoded, it can be packetized for transmission or storage. The video data can then be assembled into a video file conforming to one of various standards, such as the International Organization for Standardization (ISO) based media file formats and their extensions, such as AVC. [Overview of the project]
[0006] In general, this disclosure describes techniques relating to signaling characteristics of media data contained within a network packet in the packet header. Such characteristics may include, for example, identifiers relating to the media data (e.g., relating to a slice of a picture and / or relating to a picture). Identifiers may indicate, or relate to, whether the media data can be discarded in certain circumstances, such as in circumstances following random access techniques used to initiate reception of a bitstream containing the media data. For example, a packet may be discarded if the media data depends on media data from a previously unreceived bitstream. Such circumstances may arise if the media data is contained within a gradual decoder refresh (GDR) picture used for random access, if the media data is contained within a random access skippable leading (RASL) picture, or in other such cases.
[0007] In one embodiment, a method for receiving video data includes receiving a packet comprising a packet header and a payload containing at least a portion of the video data frames, wherein the packet header is separate from the payload; extracting a video frame identifier relating to the video data frames from the packet header; and processing the payload according to the video frame identifier.
[0008] In another embodiment, a device for receiving video data includes a memory configured to store the video data, and one or more processors implemented in the circuit, configured to receive packets, each containing a packet header and a payload containing at least a portion of the frames of video data, wherein the packet header is separate from the payload, to extract a video frame identifier relating to the frames of video data from the packet header, and to process the payload according to that video frame identifier.
[0009] In another embodiment, a device for receiving video data includes means for receiving a packet comprising a packet header and a payload containing at least a portion of the video data frame, wherein the packet header is separate from the payload; means for extracting a video frame identifier relating to the video data frame from the packet header; and means for processing the payload according to the video frame identifier.
[0010] In another embodiment, a computer-readable storage medium stores instructions, which, when executed, cause a processor to receive a packet comprising a packet header and a payload containing at least a portion of a frame of video data, wherein the packet header is separate from the payload; to extract a video frame identifier relating to a frame of video data from the packet header; and to process the payload according to that video frame identifier.
[0011] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from this description, drawings, and claims. [Brief explanation of the drawing]
[0012] [Figure 1] This is a block diagram illustrating an exemplary system that implements technology for streaming media data over a network. [Figure 2]This is a conceptual diagram illustrating an exemplary architecture for delivering Extended Reality (XR) traffic. [Figure 3] This is a conceptual diagram showing a series of progressive decoder refresh (GDR) frames of video data. [Figure 4] This is a conceptual diagram illustrating an exemplary set of identification information related to protocol data units (PDUs). [Figure 5] This is a block diagram illustrating the elements of an example video file. [Figure 6] This flowchart illustrates an exemplary method using the technology of this disclosure, including transmitting a packet containing media data. [Figure 7] This flowchart illustrates an exemplary method using the technology of this disclosure, including receiving a packet containing media data. [Modes for carrying out the invention]
[0013] In general, this disclosure describes technologies related to sending and receiving Extended Reality (XR) media data, such as augmented reality (AR), mixed reality (MR), and / or virtual reality (VR). For example, a media communication session between two network devices, such as between a server device and a client device, or between two user equipment (UE) devices, may include audio data, video data, and / or XR data. Thus, a user can participate in an XR communication session while communicating with one or more other users. An XR communication session may correspond to an XR-based telecommunication session, a game, and so on.
[0014] Various issues related to XR and media (XRM) services are currently under research. Two main issues are the use of unified protocol data unit (PDU) set packet processing and differentiated PDU set processing, which generally aim to enhance PDU set processing in Fifth Generation (5G) User Plane Functions (UPF) to optimize the XRM consumption experience. A PDU set can contain multiple PDUs, each containing data relating to a common presentation time. For example, a PDU set may contain PDUs containing data relating to video data frames and / or graphical XR data relating to computer-generated graphics. Therefore, a PDU set can correspond to a video frame, and each PDU can be a slice of that video frame or a network abstraction layer (NAL) unit.
[0015] In a 5G system (5GS), the interface between the application domain and the 5GS is based on quality of service (QoS) flows. QoS flows represent the finest granularity of QoS differentiation within a PDU session. QoS flows in 5GS can be identified using QoS flow identifiers (QFIs). User plane traffic with the same QFI within a PDU session can receive the same traffic forwarding processing.
[0016] Each PDU may correspond to a packet communicated, for example, over a computer-based network. PDUs can be forwarded using the Real-time Transport Protocol (RTP) or other protocols. RTP is typically performed via the Uniform Datagram Protocol (UDP). Therefore, packets may be delivered out of order, and UDP does not provide packet delivery guarantees. Furthermore, packet processing at the network level generally does not have access to the packet payload data (which may include video coding layer (VCL) data) because the payload data may be encrypted or otherwise inaccessible to network elements.
[0017] Therefore, according to the technology of this disclosure, certain video information can be included in the packet header outside the payload so that the packet can be processed by a network device that does not have access to the payload data. Such data may include, for example, identification information relating to a frame or portion of a frame of video data. This identification information can specifically identify a frame, for example, by using a picture order count (POC) value that indicates the display order of the frame. As another example, the frame number may indicate a coding order value of the frame, which may differ from the display order. The identification information may also (and or alternatively) include data representing coding layers, such as a time layer identifier, which generally corresponds to the number of reference frames that the frame may be able to use for prediction, and whether subsequent frames can use that frame as a reference frame.
[0018] In this way, a network device, or a network component of a user device (e.g., user equipment), can determine identifier information relating to the media data contained within the payload of each received packet. Therefore, the network device or network component can determine, for example, whether all packets of a frame have been received, whether a reference frame relating to that frame has been received, and whether subsequent frames used for reference can be decoded (depending on whether that frame has been received). In this way, the network device or network component can determine, for example, whether the video data in the packet's payload should be provided to the video decoder, whether missing reference frames should be retrieved, and whether one or more sets of video data should be discarded without sending the video data to the video decoder.
[0019] The technology disclosed herein can be applied to video files that conform to video data encapsulated according to any of the following: ISO-based media file formats, Scalable Video Coding (SVC) file formats, Advanced Video Coding (AVC) file formats, Third Generation Partnership Project (3GPP®) file formats, and / or Multiview Video Coding (MVC) file formats, or other similar video file formats.
[0020] Figure 1 is a block diagram showing an exemplary system 10 that implements technology for streaming media data over a network. In this embodiment, system 10 includes a content preparation device 20, a server device 60, and a client device 40. The client device 40 and the server device 60 are communicated together by a network 74, which may include the Internet. In some embodiments, the content preparation device 20 and the server device 60 can also be communicated together by the network 74 or another network, or they can be communicated together directly. In some embodiments, the content preparation device 20 and the server device 60 may include the same device.
[0021] In the embodiment of Figure 1, the content preparation device 20 includes an audio source 22 and a video source 24. The audio source 22 may include, for example, a microphone that generates electrical signals representing captured audio data, which will be encoded by the audio encoder 26. Alternatively, the audio source 22 may include a storage medium storing previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. The video source 24 may include a video camera, a storage medium encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data that will generate video data, which will be encoded by the video encoder 28. The content preparation device 20 is not necessarily communicatively coupled to the server device 60 in all embodiments, and multimedia content can also be stored on a separate medium read by the server device 60.
[0022] The raw audio data and video data can include analog data or digital data. The analog data can be digitized before being encoded by the audio encoder 26 and / or the video encoder 28. The audio source 22 can obtain audio data from the participant during the speech, and the video source 24 can simultaneously obtain video data of the participant during the speech. In other embodiments, the audio source 22 can include a computer-readable storage medium including stored audio data, and the video source 24 can include a computer-readable storage medium including stored video data. Thus, the techniques described in this disclosure can be applied to live, streaming, real-time audio data and video data, or can also be applied to archived, pre-recorded audio data and video data.
[0023] An audio frame corresponding to a video frame is generally an audio frame that includes audio data captured (or generated) by the audio source 22 simultaneously with the video data captured (or generated) by the video source 24 included within that video frame. For example, while a participant during a speech is generally generating audio data by speaking, the audio source 22 captures that audio data, and the video source 24 simultaneously captures, that is, while the audio source 22 is capturing the audio data, the video data of the participant during the speech. Therefore, an audio frame can correspond temporally to one or more specific video frames. Thus, an audio frame corresponding to a video frame generally means that the audio data and the video data are captured simultaneously, and the audio frame and the video frame correspond to a situation where they respectively include the simultaneously captured audio data and video data.
[0024] In some embodiments, the audio encoder 26 can encode a time stamp representing the time at which the audio data for the encoded audio frame was recorded in each encoded audio frame, and similarly, the video encoder 28 can encode a time stamp representing the time at which the video data for the encoded video frame was recorded in each encoded video frame. In such embodiments, an audio frame corresponding to a video frame can include that the audio frame includes a time stamp and the video frame includes the same time stamp. The content preparation device 20 can include an internal clock such that the audio encoder 26 and / or the video encoder 28 can generate a time stamp, or the audio source 22 and the video source 24 can include an internal clock that can be used to associate the audio data and the video data with the time stamp, respectively.
[0025] In some embodiments, the audio source 22 can send data corresponding to the time at which the audio data was recorded to the audio encoder 26, and the video source 24 can send data corresponding to the time at which the video data was recorded to the video encoder 28. In some embodiments, the audio encoder 26 can encode a sequence identifier that indicates the relative temporal order of the encoded audio data but not necessarily the absolute time at which the audio data was recorded in the encoded audio data, and similarly, the video encoder 28 can also use a sequence identifier to indicate the relative temporal order of the encoded video data. Similarly, in some embodiments, the sequence identifier can be mapped to a time stamp or correlated with a time stamp in another manner.
[0026] The audio encoder 26 generally generates a stream of encoded audio data, while the video encoder 28 generates a stream of encoded video data. Each individual stream of data (whether audio or video) may be referred to as an elementary stream. An elementary stream is a single digitally coded (and possibly compressed) component of a media presentation. For example, the coded video or audio portion of a media presentation can be an elementary stream. An elementary stream can be converted to a packetized elementary stream (PES) before being encapsulated within a video file. Within the same media presentation, a stream ID can be used to distinguish PES packets belonging to one elementary stream from others. The basic unit of data in an elementary stream is a packetized elementary stream (PES) packet. Therefore, coded video data generally corresponds to an elementary video stream. Similarly, audio data corresponds to one or more elementary streams.
[0027] In the embodiment shown in Figure 1, the encapsulation unit 30 of the content preparation device 20 receives an elementary stream containing coded video data from the video encoder 28 and an elementary stream containing coded audio data from the audio encoder 26. In some embodiments, the video encoder 28 and the audio encoder 26 may each include a packetizer for forming PES packets from the coded data. In other embodiments, the video encoder 28 and the audio encoder 26 may each interface with a corresponding packetizer for forming PES packets from the coded data. In yet another embodiment, the encapsulation unit 30 may include a packetizer for forming PES packets from the coded audio and video data.
[0028] The video encoder 28 can encode video data of multimedia content in various ways to produce various representations of the multimedia content at various bitrates having various characteristics, such as pixel resolution, frame rate, compliance with various coding standards, compliance with various profiles and / or levels of profiles relating to various coding standards, representation having one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. When used in this disclosure, a representation may include audio data, video data, text data (e.g., for closed captions), or one of the other such data. A representation may include an elementary stream, such as an audio elementary stream or a video elementary stream. Each PES packet may include a stream_id that identifies the elementary stream to which the PES packet belongs. The encapsulation unit 30 is involved in assembling the elementary stream into streamable media data.
[0029] The encapsulation unit 30 receives PES packets from the audio encoder 26 and video encoder 28 relating to the elementary stream of the media presentation, and forms corresponding Network Abstraction Layer (NAL) units from these PES packets. Coated video segments can be organized into NAL units, which provide a "network-friendly" video representation for applications such as videophone, storage, broadcast, or streaming. NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may include a core compression engine and may contain block, macroblock, and / or slice-level data. Other NAL units may be non-VCL NAL units. In some embodiments, a coded picture in a given time instance, typically presented as a primary coded picture, can be included within an access unit that may contain one or more NAL units.
[0030] Non-VCL NAL units may include, among many others, parameter set NAL units and SEI NAL units. Parameter sets may include sequence-level header information (within sequence parameter sets (SPS)) and picture-level header information that does not change frequently (within picture parameter sets (PPS)). Parameter sets (e.g., PPS and SPS) allow information that does not change frequently to be avoided by not having to be repeated for each sequence or picture, thus improving coding efficiency. Furthermore, the use of parameter sets can avoid the need for redundant transmission for error tolerance by enabling out-of-band transmission of important header information. In an example of out-of-band transmission, parameter set NAL units may be transmitted on a different channel than other NAL units such as SEI NAL units.
[0031] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding coded picture samples from VCL NAL units, but can assist processes related to decoding, display, error tolerance, and other purposes. SEI messages can be included within non-VCL NAL units. SEI messages are a normative part of some standard specifications and are therefore not necessarily required for standard-compliant decoder implementations. SEI messages can be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information can be included within SEI messages, such as scalability information SEI messages in SVC embodiments and view scalability information SEI messages in MVC. These exemplary SEI messages can, for example, convey information about the extraction and characteristics of operation points.
[0032] The server device 60 includes a Real-Time Transport Protocol (RTP) transmission unit 70 and a network interface 72. In some embodiments, the server device 60 may include multiple network interfaces. Furthermore, some or all of the features of the server device 60 may be implemented on other devices in the content distribution network, such as routers, bridges, proxy devices, switches, or other devices. In some embodiments, an intermediate device in the content distribution network may cache data for multimedia content 64 and may include components substantially compliant with the components of the server device 60. Generally, the network interface 72 is configured to send and receive data over the network 74.
[0033] The RTP transmission unit 70 is configured to deliver media data to client devices 40 via the network 74 in accordance with RTP, which is standardized in Request for Comment (RFC) 3550 by the Internet Engineering Task Force (IETF). The RTP transmission unit 70 can also implement RTP-related protocols such as the RTP Control Protocol (RTCP), Real-time Streaming Protocol (RTSP), Session Initiation Protocol (SIP), and / or Session Description Protocol (SDP). The RTP transmission unit 70 can transmit media data via a network interface 72 that can implement the Uniform Datagram Protocol (UDP) and / or the Internet Protocol (IP). Therefore, in some embodiments, the server device 60 can use the network 74 to transmit media data via RTP and RTSP over UDP.
[0034] The RTP transmission unit 70 can receive an RTSP description request from, for example, a client device 40. The RTSP description request may include data indicating what types of data are supported by the client device 40. The RTP transmission unit 70 can respond to the client device 40 with data indicating a media stream, such as media content 64, which can be sent to the client device 40, along with a corresponding network location identifier, such as a uniform resource locator (URL) or uniform resource name (URN).
[0035] Next, the RTP transmission unit 70 can receive an RTSP setup request from the client device 40. The RTSP setup request may generally indicate how the media stream should be transferred. The RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and transport specifiers such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. The RTP transmission unit 70 may reply to the RTSP setup request with an acknowledgment and data representing the port of the server device 60 that will send the RTP data and control data. The RTP transmission unit 70 can then receive an RTSP playback request to "play" the media stream, i.e., to send the media stream to the client device 40 over the network 74. The RTP transmission unit 70 can also receive an RTSP teardown request to terminate the streaming session, and in response to that request, the RTP transmission unit 70 may stop sending media data to the client device 40 for the corresponding session.
[0036] The RTP receiving unit 52 can similarly initiate a media stream by first sending an RTSP description request to the server device 60. The RTSP description request may indicate the type of data supported by the client device 40. The RTP receiving unit 52 can then receive a reply from the server device 60 specifying an available media stream, such as media content 64, that can be sent to the client device 40, along with a corresponding network location identifier, such as a Uniform Resource Locator (URL) or Uniform Resource Name (URN).
[0037] Next, the RTP receiving unit 52 can generate an RTSP setup request and send it to the server device 60. As described above, the RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and transport specifiers such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. In response, the RTP receiving unit 52 may receive from the server device 60 an acknowledgment including the port of the server device 60 that the server device 60 will use to transmit the media data and control data.
[0038] After establishing a media streaming session between the server device 60 and the client device 40, the RTP transmission unit 70 of the server device 60 can send media data (e.g., packets of media data) to the client device 40 according to that media streaming session. The server device 60 and the client device 40 can exchange control data (e.g., RTCP data) that, for example, indicates reception statistics by the client device 40, thereby enabling the server device 60 to perform congestion control or to diagnose and address transmission failures in other ways.
[0039] The network interface 54 can receive the media of the selected media presentation and provide it to the RTP receiving unit 52, which can similarly provide the media data to the decapsulation unit 50. The decapsulation unit 50 decapsulates the elements of the video file into constituent PES streams, depackets those PES streams to extract the encoded data, and can send the encoded data to either the audio decoder 46 or the video decoder 48, depending on whether the encoded data is part of an audio stream or part of a video stream, for example, as indicated by the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data, which may contain multiple views of the stream, to the video output 44.
[0040] The video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiver unit 52, and decapsulation unit 50 can each be implemented, where applicable, as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof. The video encoder 28 and video decoder 48 may each be contained within one or more encoders or decoders, and either may be integrated as part of a composite encoder / decoder (CODEC). Similarly, the audio encoder 26 and audio decoder 46 may each be contained within one or more encoders or decoders, and either may be integrated as part of a composite CODEC. The apparatus, which includes a video encoder 28, a video decoder 48, an audio encoder 26, an audio decoder 46, an encapsulation unit 30, an RTP receiving unit 52, and / or a decapsulation unit 50, may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular phone.
[0041] The client device 40, the server device 60, and / or the content preparation device 20 can be configured to operate in accordance with the techniques of this disclosure. For illustrative purposes, this disclosure describes these techniques with respect to the client device 40 and the server device 60. However, it should be understood that the content preparation device 20 can be configured to perform these techniques instead of (or in addition to) the server device 60.
[0042] The encapsulation unit 30 can form NAL units, each containing a header that identifies the program to which the NAL unit belongs, and a payload, such as audio data, video data, or data describing the transport stream or program stream to which the NAL unit corresponds. For example, in H.264 / AVC, a NAL unit contains a 1-byte header and a variable-size payload. A NAL unit containing video data in its payload may contain video data at various levels of granularity. For example, a NAL unit may contain blocks of video data, multiple blocks, slices of video data, or entire pictures of video data. The encapsulation unit 30 can receive encoded video data from the video encoder 28 in the form of PES packets of elementary streams. The encapsulation unit 30 can associate each elementary stream with a corresponding program.
[0043] The encapsulation unit 30 can also assemble an access unit from multiple NAL units. Generally, an access unit may include one or more NAL units for representing frames of video data, and, if audio data corresponding to those frames is available, such audio data. Generally, an access unit includes all NAL units for a given output time instance, e.g., all audio and video data for a given time instance. For example, if each view has a frame rate of 20 frames per second (fps), each time instance may correspond to a time interval of 0.05 seconds. During this time interval, specific frames for all views of the same access unit (same time instance) can be rendered simultaneously. In one embodiment, an access unit may include a coded picture in a given time instance that can be presented as a primary coded picture.
[0044] Therefore, an access unit may contain all audio and video frames of a common time instance, for example, all views corresponding to time X. This disclosure also refers to the encoded picture of a particular view as a “view component.” That is, a view component may contain an encoded picture (or frame) relating to a particular view at a particular time. Therefore, an access unit can be defined as containing all view components of a common time instance. The decoding order of an access unit does not necessarily have to be the same as the output order or display order.
[0045] After the encapsulation unit 30 assembles the NAL unit and / or access unit into a video file based on the received data, the encapsulation unit 30 passes the video file to the output interface 32 for output. In some embodiments, the encapsulation unit 30 may either store the video file locally or send it to a remote server via the output interface 32, rather than sending the video file directly to the client device 40. The output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium such as an optical drive or magnetic media drive (e.g., a floppy drive), a universal serial bus (USB) port, a network interface, or other output interface. The output interface 32 outputs the video file to a computer-readable medium such as a transmission signal, magnetic media, optical media, memory, a flash drive, or other computer-readable medium.
[0046] The network interface 54 can receive NAL units or access units via the network 74 and provide those NAL units or access units to the decapsulation unit 50 via the RTP receiving unit 52. The decapsulation unit 50 decapsulates the elements of the video file into constituent PES streams, depackets those PES streams to extract the encoded data, and can send the encoded data to either the audio decoder 46 or the video decoder 48, depending on whether the encoded data is part of an audio stream or part of a video stream, for example, as indicated by the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data, which may contain multiple views of the stream, to the video output 44.
[0047] Figure 2 is a conceptual diagram illustrating an exemplary architecture for Extended Reality (XR) traffic delivery. Figure 2 shows the Application / Service Layer 100, User Plane Function (UPF) 102, Access Network (AN) 110, and Client Device 120. Client Device 120 may correspond to Client Device 40 in Figure 1 and may generally include components similar to those of Client Device 40. The UPF 102 and AN 110 may correspond to network devices in Network 74 in Figure 1.
[0048] In this embodiment, UPF102 includes a packet detection unit 104 and a packet detection rule 106, AN110 includes a Quality of Service (QoS) to AN resource mapping unit 112 and a wireless interface 114, and client device 120 includes a QoS rule 122, a QoS to AN resource mapping unit 124 and a wireless interface 126.
[0049] The client device 120 can send and receive media data, for example, in the form of audio data, video data, and / or Extended Reality (XR) data. Data sent from the client device 120 to another client device or another device, such as the server device 60 in Figure 1, is transmitted via the uplink stream, while data received by the client device 120 is received via the downlink stream. In the downlink (DL) Extended Reality and Media (XRM) service stream, the UPF 102 can classify incoming data packets based on the packet filter set of packet detection rule 106. The UPF 102 can communicate the classification of user plane traffic belonging to a QoS flow through QoS flow identifier (QFI) marking. The QoS-to-AN resource mapping unit 112 of the AN 110 binds QoS flows to AN resources (i.e., data radio bearers). For an uplink (UL) XRM service stream, the client device 120 (which may be a user equipment (UE)) performs a PDU set identification procedure similar to that of UPF102. The client device 120 can then perform marking of the identified PDU set to the AN.
[0050] 3GPP TR23.700-60 reports a candidate solution for identifying PDU and PDU set boundaries by matching RTP / SRTP headers, header extensions, and payloads. New parameters such as PDU set sequence number, PDU set identifier, and PDU set type are proposed for identifying PDUs and PDU sets. The candidate solution also proposes marking the priority, dependency, or importance of PDU sets by matching relevant parameters in the RTP header extension or payload, such as the Video Network Abstraction Layer (NAL) type, Temporal ID (TID), and Layer ID (LID). The most important PDUs or PDU sets can be assigned to QoS flows with higher QoS requirements, while less important PDUs or PDU sets can be discarded during network congestion or if dependent PDUs or PDU sets are not fully delivered. It is also proposed that an AN can discard a PDU set if some PDUs in that set are not received.
[0051] PDUs can be mapped to video data packets. Encoders or application functions (AFs) can be configured to recognize properties of video data packets, such as packet type and dependencies, and data packet identifiers, such as picture order counter (POC) values, related to video packets. These identifiers are typically embedded within the media data packet and are not exposed in the RTP header or header extension. The link between a specific video data packet and its corresponding properties can be lost during data encapsulation, filtering, and mapping to QoS flows, especially in unordered transmissions where the packer order may be shuffled. Adding attribute or property fields for each packet to the RTP header or header extension can increase overhead costs, as many RTP packets may share the same properties. Adding video data packet identifiers to the RTP header or header extension as an index to a list of data packet properties may be beneficial.
[0052] PDU sets may also have marked priority values. In some cases, priority between PDU sets can be identified and the priority or importance of PDU sets can be marked using picture type, temporal identifier (TID), or layer identifier (LID). However, the complexity of deriving the final overall priority in relation to multiple priorities may be added to the intermediate device. Furthermore, an encoder may assign the same TID to all frames, and the importance of those frames with the same TID may also differ. Generally, it is reasonable to mark intra-coded pictures as having the highest priority and pictures not used for reference as having the lowest priority. However, there may be cases where a picture is intra-coded but not used for reference (e.g., all in intra mode). For inspection purposes, a single indication may be needed to show the picture priority at the codec level.
[0053] PDU sets may also contain marked dependency information. Picture dependencies are managed by reference picture management in most video coding schemes. Reference pictures are marked as "not used for reference" and stored in the decoded picture buffer (DPB) for interpretation until they are removed from the DPB.
[0054] ITU-T H.266 / Versatile Video Coding (VVC) specifies two reference picture lists (RPLs), called List 0 and List 1. Predefined candidate RPLs are signaled in the Sequence Parameter Set (SPS). Indices referencing default candidate RPLs are signaled in the picture header (PH) if all slices of the picture have the same RPL, and otherwise in the slice header (SH). New RPLs can also be directly signaled in the picture header (PH) and slice header (SH). Reference pictures can be short-term reference pictures, long-term reference pictures, or inter-layer reference pictures, and when used for inter-layer prediction, reference pictures are marked with a picture order count (POC) and layer ID. The reference picture used for inter-prediction of the current picture is the current picture's active reference picture.
[0055] ITU-T H.265 / High Efficiency Video Coding (HEVC) specifies reference pictures in a reference picture set (RPS). The RPS can be signaled in a sequence parameter set (SPS) or slice header (SH).
[0056] ITU-T H.264 / Advanced Video Coding (AVC) specifies reference pictures based on two marking mechanisms: an implicit sliding window process and an explicit memory management control operation process. Typically, up to 16 unique reference pictures are allowed, given the decoded picture buffer size.
[0057] AOMedia Video1 (AV1) allows up to seven reference frames, and frame marking is specified in the frame header open bitstream unit (OBU). Reference picture indications can be presented in the frame header OBU or in additional headers for each tile.
[0058] Due to the various reference picture management designs in different codecs, it is complex for UPF102 to derive dependencies for each PDU set in a codec-independent manner. This disclosure recognizes that it is beneficial to explicitly indicate picture dependencies in specific NAL units or SEI messages to support existing codecs.
[0059] Furthermore, while the marking of reference pictures is based on POC values representing the output order of the pictures, the frames or PDU sets within the bitstream are encoded in a specific order, meaning that POC values may not be consecutive with respect to adjacent frames within the bitstream. Mapping POC values to the sequence numbers (SN) of PDU sets using existing methods is not straightforward.
[0060] H.266 supports subpicture partitioning, which allows each subpicture to be decoded independently of other subpictures within the same frame. Even if all subpictures share the same reference picture, a frame may contain a mixture of intra-coded and inter-coded subpictures. Additional attributes regarding subpicture identification and marking may be required to facilitate proper PDU processing. For example, if subpictures are independently coded and some PDUs in a PDU set are missing, the remaining PDUs can be delivered continuously. If subpictures are not independently coded and some PDUs are missing, the remaining PDUs can be discarded.
[0061] AV1 supports tile lists, which contain tile data associated with each frame, allowing each tile to be decoded independently. The tile list allows the decoder to process a subset of tiles and display the corresponding portion of that frame, without needing to completely decode all the tiles in the frame.
[0062] Figure 3 is a conceptual diagram showing a series of progressive decoder refresh (GDR) frames of video data. Generally, when performing random access (i.e., starting streaming from a point in the video other than the beginning of the video), the stream is accessed starting from the stream access point. The stream access point can be fully intra-predictive coded, thereby allowing the entire frame to be decoded and playback to begin from that stream access point. Such frames are sometimes referred to as instantaneous decoder refresh (IDR) pictures. However, fully intra-predictive coded frames have a relatively high bitrate.
[0063] Therefore, instead of having a single IDR stream access point, a Gradual Decoder Refresh (GDR) stream access point may be used. Generally, a GDR stream access point contains a series of frames, some of which are intrapredictively coded, while others are interpredictively coded. When performing random access from a GDR stream access point, the interpredictively coded portion may be decodeable or undecodeable depending on which side of the intrapredictive portion the interpredictive portion occurs on.
[0064] GDR allows for smoothing the bitrate of a bitstream by distributing intra-coded slices or blocks across multiple pictures, as opposed to an encoder intra-coding the entire picture. This enables a significant reduction in end-to-end latency, which is particularly important for ultra-low latency applications. When initiating a decoding process involving GDR picture decoding, some areas of the picture cannot be decoded accurately. After decoding several additional pictures, known as a recovery period, the entire picture with respect to that recovery point, along with all subsequent pictures in the output order, will be accurately decoded. Figure 3 shows an example of such a GDR picture recovery period, where clean and intra-predicted areas are areas that can be accurately decoded, and dirty areas are areas that cannot be accurately decoded with respect to random access. As a result, PDUs in dirty areas that cannot be accurately decoded can be discarded when random access occurs in the associated GDR picture. Identification and marking of GDR picture PDUs and PDU sets are not addressed in these candidate solutions.
[0065] More specifically, the embodiment in Figure 3 shows GDR frames 140A to 140D. GDR frame 140A includes an intra-predictive coding region 142A and an undecodeable inter-predictive coding region 144A. GDR frame 140B includes a decodeable inter-predictive coding region 146B, an intra-predictive coding region 142B, and an undecodeable inter-predictive coding region 144B. GDR frame 140C includes a decodeable inter-predictive coding region 146C, an intra-predictive coding region 142C, and an undecodeable inter-predictive coding region 144C. GDR frame 140D includes a decodeable inter-predictive coding region 146D and an intra-predictive coding region 142D.
[0066] The undecodeable inter-predictive coding regions 144A-144C are not decodeable when performing random access starting from GDR frame 140A because these regions may reference reference frames that precede GDR frame 140A in the coding order. The decodeable inter-predictive regions 146B-146D are decodeable because they can be predicted only from reference frames starting from GDR frame 140A, and only from either the decodeable inter-predictive region or the intra-predictive region.
[0067] In H.266 / VVC, if PictureOutputFlag is equal to 0, the picture is not output. For example, if a picture is a Random Access Skippable Reading (RASL) picture and the NoOutputBeforeRecoveryFlag of the associated Intra Random Access Point (IRAP) picture is equal to 1, the picture is not output. If a picture is a Gradual Decoding Refresh (GDR) picture with a NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with a NoOutputBeforeRecoveryFlag equal to 1, the picture is not output. If the ph_pic_output_flag of a picture is equal to 0, the picture is not output. The value of NoOutputBeforeRecoveryFlag can be set by external means or by the position of the associated picture in the bitstream. A PDU set can be discarded if it is not output and is not used as a reference for subsequent picture prediction.
[0068] Figure 4 is a conceptual diagram illustrating an exemplary set of identification information for protocol data units (PDUs). This disclosure describes a technique that allows the use of a high level of syntax design in the video codec to facilitate the identification and marking of PDUs and sets of PDUs. Video frame identifiers (e.g., POC values in AVC / HEVC / VVC, display_frame_id or current_frame_id in AV1) can be added to the RTP header extension and further marked in sets of PDUs or PDUs. An application function (AF) can use its video frame identifier to display video frame properties, such as frame type, priority, and dependencies, to UPF102 (Figure 2) via a Policy Control Function (PCF) and a Session Management Function (SMF). UPF102 can classify video packets for QoS marking by examining the frame identifier of a data packet and linking that frame identifier to the corresponding properties for QoS marking. Figure 4 shows one example of using video frame IDs as an index to link PDU sets to a table or list containing associated video data properties.
[0069] In some embodiments, a slice or tile identifier (e.g., slice segment address in HEVC, slice_address in VVC, tile count in AV1) indicating a specific slice / tile within a frame can be signaled in the RTP header extension along with the frame identifier. The slice / tile identifier can be mapped to a PDU attribute field. The AF can use both the frame identifier and the slice / tile identifier to indicate the properties of the video slice / tile to the UPF102. The UPF102 can classify the video packet for QoS marking by examining the frame identifier and slice / tile identifier of the data packet and linking these values to the corresponding properties for QoS marking.
[0070] In some embodiments, the frame marking RTP header extension may include priority marking information. For example, the frame marking RTP header extension may include data indicating whether the corresponding frame is intra-predictively coded (I-frame) or a discardable frame (D), a time ID (TID) for that frame, and / or a layer ID (LID). Any or all of this data may represent the priority marking of a PDU set. Such priority attribute data can be signaled in the network abstraction layer (NAL) unit header, a specific NAL unit, or in a supplemental additional information (SEI) message relating to PDU set priority marking checks.
[0071] In some embodiments, the NAL unit priority indicator, nuh_priority, can be signaled in the NAL unit header or header extension. This priority indicator can be a 3-bit code that specifies the priority of the associated NAL unit, taking values from 1 to 7, where 1 represents the highest priority and 7 represents the lowest priority. Table 1 below shows one embodiment of how the NAL unit priority indicator is represented in the NAL unit header extension. In Table 1, "[added:""]" represents an addition to the existing NAL unit syntax, for example, ITU-T H.266, while "[removed:""]" represents a deletion from the existing NAL unit syntax.
[0072] [Table 1]
[0073] The semantics for the additional syntax elements in the examples in Table 1 can be as follows: A nuh_extension_flag equal to 0 indicates that the syntax elements nuh_priority and muh_reserved_zero_5bits are not present in the NAL unit header syntax structure. A nuh_extension_flag equal to 1 indicates that the syntax elements nuh_priority and muh_reserved_zero_5bits may be present in the NAL unit header syntax structure.
[0074] `nuh_priority` specifies the NAL unit priority. A `nuh_priority` equal to 1 represents the highest priority, and 7 represents the lowest priority.
[0075] As another example, the picture priority instruction, aud_priority, can be signaled in the AU delimiter (AUD), as shown in Table 2:
[0076] [Table 2]
[0077] The semantics for the additional syntax elements in the examples in Table 2 can be as follows: aud_priority specifies the priority of the access unit (AU) including the AU delimiter. an aud_priority equal to 1 represents the highest priority, and 7 represents the lowest priority.
[0078] In yet another embodiment, picture priority indications can be signaled in the picture header (PH), as shown in Table 3:
[0079] [Table 3]
[0080] The semantics for the additional syntax elements in the examples in Table 3 can be as follows: `ph_priority` specifies the priority of the current picture. A `ph_priority` equal to 1 represents the highest priority, and 7 represents the lowest priority. The lowest priority means that the current picture will never be used as a reference picture.
[0081] In yet another embodiment, the picture priority instruction can specify priority among pictures that share the same TID and LID as the current picture. In this case, the PDU set priority marking can be derived from ph_priority, TID, and LID.
[0082] In some embodiments, slice priority indicators can be signaled in the slice header (SH) to indicate priority between slices within the same picture, or between slices of a picture that share the same TID and LID.
[0083] In some embodiments, picture priority or slice priority can be signaled in the SEI message to indicate picture priority, or to indicate slice priority using a slice address.
[0084] In some embodiments, frame priority instructions can be signaled in specific OBUs related to the AV1 codec, such as a frame header OBU or a metadata OBU. Tile priority instructions can be signaled in a tile group OBU or a tile list OBU. Table 4 shows one embodiment of frame header OBU syntax with frame priority instructions.
[0085] [Table 4]
[0086] The frame_priority syntax element can signal or map to other protocols, such as the RTP header extension or the GPRS Tunneling Protocol User Plane (GTP-U) extension header.
[0087] Deriving PDU set dependencies from a list of reference pictures is generally extremely complex. Similarly, mapping POC values to PDU set sequence numbers (SNs) incurs additional costs. Therefore, according to the techniques of this disclosure, the distance between the current picture and the active reference picture in the bitstream (e.g., POC distance) can be marked in a specific NAL unit or SEI message. All active reference pictures precede the current picture in the bitstream. Depending on the PDU set boundaries, this distance can be in units of access units (AUs) or picture units (PUs).
[0088] Table 5 below is an abstract syntax structure showing a list of active reference pictures in the bitstream related to the current picture. This list includes short-term active reference pictures, long-term active reference pictures, and cross-layer active reference pictures.
[0089] [Table 5]
[0090] The semantics for the syntax elements in the examples in Table 5 can be as follows: num_st_act_ref_pos specifies the number of the short-term active reference picture location in the syntax structure.
[0091] st_act_ref_delta_pos specifies the distance between the current picture and the active short-term referenced picture, in units of access units (AUs).
[0092] For a single-layer bitstream, one AU contains one picture. Assuming the sequence number (SN) of the current PDU set is N, the SN of the PDU set containing the i-th short-lived active reference picture is (N - st_act_ref_delta_pos[i]).
[0093] In another embodiment, st_act_ref_delta_pos can also specify the distance between the current PU and the active short-term reference picture, on a picture-by-picture basis.
[0094] num_lt_act_ref_pos specifies the number of the long-term active reference picture location in the syntax structure.
[0095] lt_act_ref_delta_pos specifies the distance between the current picture and the active short-term reference picture, in AU units.
[0096] For a single-layer bitstream, one AU contains one picture. Assuming the SN of the current picture or PDU set is N, the SN of the PDU set containing the i-th long-term active reference picture is (N - st_act_ref_delta_pos[i]).
[0097] In some embodiments, lt_act_ref_delta_pos can also specify the distance between the current picture and the active long-term reference picture, on a picture-by-picture basis.
[0098] num_il_act_ref_pos specifies the number of the inter-layer active reference picture location in the syntax structure.
[0099] lt_act_ref_delta_pos specifies the layer difference between the current picture and the active inter-layer reference picture.
[0100] Assuming the current picture or PDU set has a SN of N, and each PDU set contains a picture, the SN of the PDU set containing the i-th interlayer active reference picture is (N - il_act_ref_delta_pos[i]).
[0101] In some embodiments, the data lengths of st_act_ref_delta_pos, lt_act_ref_delta_pos, and il_act_ref_delta_pos can be extended to accommodate the number of layers if each PDU set contains pictures instead of AUs.
[0102] Table 6 below shows one example of a simplified active reference picture list structure:
[0103] [Table 6]
[0104] The semantics of the syntax elements in the examples in Table 6 can be as follows: num_act_ref_pos specifies the index of the active reference picture location in the syntax structure.
[0105] `act_ref_delta_pos` specifies the distance between the current picture and the active referenced picture in Access Units (AUs). Assuming the SN of a PDU set is N, the SN of the PDU set containing the i-th active referenced picture is (N - `act_ref_delta_pos[i]`).
[0106] num_il_act_ref_pos specifies the number of the inter-layer active reference picture location in the syntax structure.
[0107] Il_act_ref_delta_pos specifies the distance between the current picture and the active reference picture in picture units (PUs). Assuming the SN of a PDU set is N, the SN of the PDU set containing the i-th interlayer active reference picture is (N - il_act_ref_delta_pos[i]).
[0108] The proposed active reference picture position syntax structure can be carried in AU delimiter (AUD) or SEI messages.
[0109] In some embodiments, the reference picture distance in the syntax structure described above can be replaced by a sequence number indicating the position of the AU or PU within the bitstream. Such a sequence number is incremented by one each time a new AU or PU is added to the bitstream. This sequence number can be mapped to the sequence number of the PDU set to facilitate the derivation of frame dependencies.
[0110] In some embodiments, the position of a reference picture in the decoding order can be signaled in a specific metadata type OBU (e.g., metadata_itut_t35) or a new metadata type OBU in AV1 to indicate the dependency of the current picture in the bitstream.
[0111] In some embodiments, the proposed active reference picturelist syntax element can be signaled to or mapped to other protocols, such as RTP or GTP-U extension headers.
[0112] In some embodiments, the active reference picture list structure can be carried in the active reference picture location list SEI message. A persistence flag is proposed in the SEI message to indicate whether the active reference picture list is applicable only to the current picture or to the current picture and all subsequent pictures in the current layer. This persistence can be canceled when the proposed cancellation flag is set or when a new coded layer video sequence (CLVS) for the current layer begins.
[0113] Independent PDU marking can be signaled to indicate whether the current PDU can be decoded without using other PDUs in the same PDU set. Independent PDUs can be transferred even if other PDUs in the same PDU set are missing. An independent PDU can be a slice containing independent subpictures, or a slice consisting of a motion-constrained tile set (MCTS) in the H265 bitstream.
[0114] In some embodiments, all slice boundaries within the CLVS (when signaled in SPS) or the associated picture (when signaled in PPS, PH, or AUD) can be treated as picture boundaries, and syntax elements can be signaled in a specific video codec NAL unit to indicate that there is no loop filtering beyond those slice boundaries. Each slice can be decoded independently, without other slice samples within the same picture.
[0115] In some embodiments, a syntax element can be signaled in the slice header (SH) to indicate whether the current slice can be decoded independently without using predictions from other samples in the same frame.
[0116] Since each tile in AV1 can be decoded independently, if a PDU is an AV1 tile and a tile list OBU exists, independent PDU markings can be set for each PDU.
[0117] For HEVC bitstreams, if the NoRaslOutputFlag of the associated intra-random access point (IRAP) picture is equal to 1, the PDU set of random access skippable reading (RASL) pictures can be marked as discarded pictures in frame marking or PDU set marking.
[0118] For VVC bitstreams, if the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1, then a PDU set of RASL pictures with a pps_mixed_nalu_types_in_pic_flag equal to 0 can be marked as a discarded picture in frame marking or PDU set marking.
[0119] Since detecting discardable pictures is not straightforward, syntax elements can be signaled in specific NAL units such as AUD, PH, or SH, or in SEI messages, to indicate whether the relevant picture can be discarded for transmission or decryption purposes.
[0120] For VVC bitstreams, a RASL picture with pps_mixed_nalu_types_in_pic_flag equal to 1 may contain both one or more RASL subpictures and one or more random access decodable leading (RADL) subpictures. RADL subpictures can be used as active reference subpictures that should not be discarded. RASL subpictures, or slices of RASL pictures, can be discarded, while the associated PDUs can be marked as discardable PDUs.
[0121] In the case of a VVC bitstream, if a GDR picture has a NoOutputBeforeRecoveryFlag equal to 1, or if it is a recovery picture of a GDR picture with a NoOutputBeforeRecoveryFlag equal to 1, slices of the GDR picture that cannot be accurately decoded may be discarded. Corresponding slices that cannot be accurately decoded may be marked in the RTP header extension so that UPF102 can discard the associated PDU during congestion time.
[0122] In the case of VVC, the picture header syntax element ph_non_ref_pic_flag indicates that the current picture will never be used as a reference picture. Such syntax elements can be used for discardable marking. In the case of the AV1 codec, discardable or non-reference indications can be signaled in the frame header OBU uncompressed header or tile group OBU to indicate whether the associated frame will not be used as a reference picture for subsequent pictures and can be discarded without affecting the decoding process. Table 7 below shows one example of frame header OBU uncompressed header syntax with such a discardable frame indication, indicated by the tag "[added:""]".
[0123] [Table 7]
[0124] Since the discarding of a slice or frame can be decided on the fly, an indication of whether or not the associated NAL unit can be discarded can be signaled in the SEI message, in a rewritable field of a specific NAL unit, or in the metadata type OBU.
[0125] In some embodiments, a slice can be marked as a discardable slice if any of the following conditions are true: the current picture is a RASL picture and the NAL unit type of the current slice is RASL; the current picture is a GDR picture and the current slice is a P slice or a B slice (unidirectional interprediction or bidirectional interprediction); or the current picture is a recovery picture of a GDR picture and the current slice is a P slice or a B slice that follows a preceding I slice in the same picture.
[0126] Figure 5 is a block diagram showing the elements of an exemplary video file 150. As described above, video files conforming to the ISO-based media file format and its extensions store data within a set of objects referred to as "boxes." In the embodiment of Figure 5, video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a segment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. While Figure 5 represents one embodiment of a video file, it should be understood that other media files may contain other types of media data (e.g., audio data, timed text data, etc.) that are structured similarly to the data in video file 150, according to the ISO-based media file format and its extensions.
[0127] The File Type (FTYP) box 152 generally represents the file type for the video file 150. The File Type box 152 may contain data that identifies a specification describing the best use of the video file 150. The File Type box 152 may also be placed before the MOOV box 154, the Movie Fragment box 164, and / or the MFRA box 166.
[0128] In the embodiment shown in Figure 5, the MOOV box 154 includes a movie header (MVHD) box 156, a track (TRAK) box 158, and one or more movie extends (MVEX) boxes 160. Generally, the MVHD box 156 can describe the general characteristics of the video file 150. For example, the MVHD box 156 may include data describing when the video file 150 was first created, data describing when the video file 150 was last modified, data describing the timescale of the video file 150, data describing the duration of playback of the video file 150, or other data that describes the video file 150 in general.
[0129] The TRAK box 158 may contain data relating to the track of the video file 150. The TRAK box 158 may also contain a track header (TKHD) box that describes the characteristics of the track corresponding to the TRAK box 158. In some embodiments, the TRAK box 158 may contain a coded video picture, while in other embodiments, the coded video picture for that track may be contained within a movie fragment 164 that can be referenced by the data in the TRAK box 158 and / or the sidx box 162.
[0130] In some embodiments, the video file 150 may contain two or more tracks. Therefore, the MOOV box 154 may contain a number of TRAK boxes equal to the number of tracks in the video file 150. The TRAK boxes 158 can describe the characteristics of the corresponding tracks in the video file 150. For example, the TRAK box 158 can describe temporal and / or spatial information about the corresponding track. If the encapsulation unit 30 (Figure 1) includes a parameter set track in a video file such as the video file 150, a TRAK box similar to the TRAK box 158 in the MOOV box 154 can describe the characteristics of the parameter set track. The encapsulation unit 30 can signal the presence of a sequence-level SEI message in the TRAK box describing the parameter set track.
[0131] The MVEX box 160 can, for example, describe the characteristics of the corresponding movie fragment 164 to signal that the video file 150 contains, in addition to the video data contained within the MOOV box 154, a movie fragment 164 if present. In the context of streaming video data, coded video pictures can be contained within the movie fragment 164 rather than within the MOOV box 154. Therefore, all coded video samples can be contained within the movie fragment 164 rather than within the MOOV box 154.
[0132] A MOOV box 154 may contain a number of MVEX boxes 160 equal to the number of movie fragments 164 in the video file 150. Each MVEX box 160 can describe one corresponding characteristic of one of the movie fragments 164. For example, each MVEX box may contain a movie extends header box (MEHD) that describes the duration of one of the corresponding movie fragments 164.
[0133] As described above, the encapsulation unit 30 can store a sequence dataset within a video sample that does not contain the actual coded video data. A video sample can generally correspond to an access unit, which is a representation of a coded picture at a particular time instance. In the context of AVC, a coded picture includes one or more VCL NAL units containing information for constructing all the pixels of the access unit, and other related non-VCL NAL units, such as SEI messages. Thus, the encapsulation unit 30 can include a sequence dataset, which may contain sequence-level SEI messages, within one of the movie fragments 164. The encapsulation unit 30 can further signal the presence of the sequence dataset and / or sequence-level SEI messages within one of the movie fragments 164 in one of the MVEX boxes 160 corresponding to that movie fragment 164.
[0134] The SIDX box 162 is an optional element of the video file 150. That is, video files conforming to the 3GPP file format, or any other such file format, do not necessarily include the SIDX box 162. According to an embodiment of the 3GPP file format, the SIDX box can be used to identify subsegments of a segment (e.g., a segment contained within the video file 150). The 3GPP file format defines a subsegment as "a self-contained set of one or more consecutive movie fragment boxes having corresponding media data boxes, where the media data box containing the data referenced by the movie fragment box must follow that movie fragment box and precede the next movie fragment box containing information about the same track." The 3GPP file format also indicates that the SIDX box "contains a sequence of references to subsegments of the (sub)segment documented by that box. The referenced subsegments are consecutive within the presentation time. Similarly, the bytes referenced by the segment index box are always consecutive within the segment. The referenced size gives a count of the number of bytes in the reference material."
[0135] The SIDX box 162 generally provides information representing one or more subsegments of a segment contained within the video file 150. For example, such information may include the playback time at which the subsegment begins and / or ends, a byte offset relative to the subsegment, whether the subsegment contains a stream access point (SAP) (e.g., starting from an SAP), the type of SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), and the position of the SAP within the subsegment (in terms of playback time and / or byte offset).
[0136] A movie fragment 164 may contain one or more coded video pictures. In some embodiments, a movie fragment 164 may contain one or more groups of pictures (GOPs), each GOP may contain several coded video pictures, e.g., frames or pictures. Furthermore, as described above, in some embodiments, a movie fragment 164 may contain a sequence dataset. Each movie fragment 164 may contain a movie fragment header box (MFHD, not shown in Figure 5). The MFHD box can describe the characteristics of the corresponding movie fragment, such as the sequence number for that movie fragment. The movie fragments 164 may be included in the video file 150 in order of their sequence numbers.
[0137] The MFRA box 166 can describe random access points within movie fragments 164 of the video file 150. This can help perform trick modes, such as performing a seek to a specific temporal position (i.e., playback time) within the segment encapsulated by the video file 150. In some embodiments, the MFRA box 166 is generally optional and does not need to be included in the video file. Similarly, a client device such as the client device 40 does not necessarily need to refer to the MFRA box 166 in order to accurately decode and display the video data of the video file 150. The MFRA box 166 may include a number of track fragment random access (TFRA) boxes (not shown) equal to the number of tracks in the video file 150, or, in some embodiments, a number equal to the number of media tracks (e.g., non-hint tracks) in the video file 150.
[0138] In some embodiments, the movie fragment 164 may include one or more stream access points (SAPs), such as IDR pictures. Similarly, the MFRA box 166 can provide indication of the location of those SAPs within the video file 150. Thus, a temporal subsequence of the video file 150 can be formed from the SAPs of the video file 150. This temporal subsequence may also include other pictures, such as P-frames and / or B-frames, that depend on the SAPs. Frames and / or slices of a temporal subsequence can be placed within segments so that frames / slice of the temporal subsequence can be properly decoded, depending on other frames / slice of that subsequence. For example, in a hierarchical arrangement of data, data used for predictions about other data may also be included within that temporal subsequence.
[0139] Figure 6 is a flowchart illustrating an exemplary method of transmitting a packet containing media data according to the technology of this disclosure. First, a server device, such as the server device 60 in Figure 1, can receive media data from, for example, a content preparation device 20 (Figure 1) (200). The media data may be encapsulated media data or encoded media data. The media data may include at least a portion of a frame / picture of video data, for example, a slice or tile.
[0140] The server device 60 may also receive data (202) indicating whether the media data is discardable, and an identifier (204) relating to the media data. The identifier may be, for example, a frame number, a picture sequence count (POC) value, and / or other such identifiers. The identifier may further indicate, for example, a time identifier (TID), a layer identifier (LID), etc. The identifier may also indicate a particular slice, tile, or other portion of the frame or picture that the media data represents.
[0141] Next, the server device 60 can encapsulate the media data within the packet (206). Alternatively, in some embodiments, the received media data may already be encapsulated within the network packet. In either case, following the method shown in Figure 6, the server device 60 can add data representing an identifier to the RTP header extension of the packet (208). The identifier itself may indicate whether the packet's media data is discardable, or the server device 60 may further add data to the RTP header extension indicating whether the packet's media data is discardable. For example, media data can be discarded if, for example, the reference media data precedes a random access point accessed by the client device and is therefore not sent to the client device. The server device 60 can then send the packet to the client device.
[0142] Figure 7 is a flowchart illustrating an exemplary method of the present invention, including receiving a packet containing media data. In this embodiment, for example, the client device 40 in Figure 1 can first receive a packet containing media data (250). That is, the packet may include a packet header and a payload separate from the packet. The payload may correspond to application layer data of the packet, for example, media data formatted according to an ISO-based media file format, as described above with respect to Figure 5. For example, the payload may include XR data, audio data, and / or video data of a PDU set.
[0143] The packet header may include an RTP header extension according to the technology of this disclosure. Therefore, the RTP header extension may include, among other data, an identifier relating to the media data of the packet. The client device 40 can extract the identifier relating to the media data from the RTP header extension (252). The client device 40 can then use the identifier to determine whether the media data is discardable (254). For example, the client device 40 may determine that the identifier includes a POC value relating to the picture of the media data, a slice or tile identifier relating to the picture, and a TID and / or LID relating to the picture. The client device 40 can further determine whether the bitstream containing the media data has been accessed randomly from a stream access point other than the beginning of the bitstream. If the bitstream has been accessed randomly, the client device 40 can further determine whether the stream access point follows one or more reference pictures relating to the media data of the packet as if no reference pictures had been received. If no reference pictures have been received, the client device 40 can determine that the media data is discardable.
[0144] In response to determining that the media data is discardable (branch "yes" at 256), the client device 40 may discard the media data without sending it to, for example, the video decoder 48 (Figure 1) (258). In response to determining that the media data is not discardable (branch "no" at 256), the client device 40 may transfer the media data to the video decoder 48 (260).
[0145] Although the method is described with respect to the client device 40 in Figure 1, the method in Figure 7 can also be performed by, for example, the client device 120 in Figure 2. Similarly, the same method can be performed by the UPF 102 in Figure 2. In embodiments in which the UPF 102 or another device performs the method in Figure 7, if media data is sent to a video decoder, it can be assumed that the video decoder forms part of the client device 120, and thereby sending media data to the video decoder includes sending packets to the client device 120.
[0146] Thus, the method shown in Figure 7 represents one embodiment of a method that includes receiving a packet comprising a packet header and a payload containing at least a portion of the video data frames, wherein the packet header is separate from the payload; extracting a video frame identifier relating to the video data frames from the packet header; and processing the payload according to the video frame identifier.
[0147] Various embodiments of the technology disclosed herein are summarized in the following clauses: Clause 1: A method for receiving video data, comprising receiving a packet comprising a packet header and a payload comprising at least a portion of the frames of the video data, wherein the packet header is separate from the payload; extracting a video frame identifier relating to the frames of the video data from the packet header; and processing the payload in accordance with the video frame identifier.
[0148] Clause 2: The method of Clause 1, wherein the video frame identifier includes a picture order count (POC) value.
[0149] Clause 3: The method of Clause 1, wherein the video frame identifier includes a display frame identifier or a current frame identifier.
[0150] Clause 4: At least a portion of the frame includes a slice of the frame, and the video frame identifier includes a slice identifier relating to that slice, in any way described in Clauses 1 through 3.
[0151] Clause 5: At least a portion of the frame includes a frame tile, and the video frame identifier includes a tile identifier relating to that slice, in any way of Clauses 1 through 4.
[0152] Clause 6: Any method of Clauses 1 to 5, further comprising using a video frame identifier to determine one or more of the following: frame type, priority of a frame, or dependency information of a frame.
[0153] Clause 7: At least a portion of the video data frames include PDUs of a set of protocol data units (PDUs) in any of the manner described in Clauses 1 through 6.
[0154] Clause 8: Any method of Clauses 1 to 7, further comprising extracting from the packet header one or more of the following: a network abstraction layer (NAL) unit type for at least a portion of the frame, a time identifier (TID) for at least a portion of the frame, a layer identifier (LID) for at least a portion of the frame, data indicating whether at least a portion of the frame is intra-predictively coded, or data indicating whether at least a portion of the frame is discardable.
[0155] Clause 9: Any method of Clauses 1 to 8, further comprising processing a Network Abstraction Layer (NAL) unit header relating to a frame, wherein the NAL unit header includes data indicating a priority value relating to the frame.
[0156] Clause 10: Further comprising processing an Access Unit Delimiter (AUD) relating to an Access Unit corresponding to a frame, wherein the AUD includes data representing a priority value relating to the Access Unit, in any manner from Clauses 1 to 9.
[0157] Clause 11: Further comprising processing the picture header relating to the frame, in any way of Clauses 1 through 10, wherein the picture header contains data representing a priority value relating to the frame.
[0158] Clause 12: Any method of Clauses 1 to 8, further comprising processing an open bitstream unit (OBU) relating to at least a portion of the frame, wherein the OBU includes data representing a priority value relating to at least a portion of the frame.
[0159] Clause 13: The method of Clause 12, wherein the OBU includes either a frame header OBU or a metadata OBU.
[0160] Clause 14: Any method of Clauses 1 to 13, further comprising receiving data indicating the picture distance between a frame and a reference frame in the coding order relating to the frame.
[0161] Clause 15: The method of Clause 14, wherein receiving data indicating picture distance includes receiving a Network Abstraction Layer (NAL) unit containing such data, or a Supplemental Additional Information (SEI) message containing such data.
[0162] Clause 16: Any method of Clauses 14 and 15, including receiving data indicating the picture distance between a frame and a reference frame relating to that frame, or receiving data indicating the picture distance between a frame and each active reference frame in the coding order.
[0163] Clause 17: Any method of Clauses 1 to 16, further comprising receiving information indicating whether a portion of a frame of video data can be decoded without using the other portions of the frame.
[0164] Clause 18: The method of Clause 17, which indicates whether a portion of a frame contains a protocol data unit (PDU), the frame corresponds to a set of PDUs containing the PDU, and the information can decode the PDU without using other PDUs in the set.
[0165] Clause 19: Any method of Clauses 1 to 18, further comprising receiving information indicating whether or not loop filtering should be performed across one or more boundaries between a portion of a frame and one or more other portions of the frame.
[0166] Clause 20: Any method of Clauses 1-19, further comprising receiving information indicating whether a portion of a frame is coded independently without relying on predictions from other portions of the frame.
[0167] Clause 21: Any method of Clauses 1 through 20, further including receiving data indicating that the frame is a discardable frame.
[0168] Clause 22: The method of Clause 21, which includes receiving a frame in a Network Abstraction Layer (NAL) unit, an Access Unit Delimiter, a Picture Header, a Slice Header, a Supplemental Additional Information (SEI) message, or a Frame Header Open Bitstream Unit (OBU), which receives data indicating that the frame is a discardable frame.
[0169] Clause 23: Any method of Clauses 1 to 22, wherein processing the payload includes determining that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is independently coded, and providing at least a portion of the frame to the video decoder in response to determining that at least a portion of the frame is independently coded.
[0170] Clause 24: Any method of Clauses 1 to 22, wherein processing the payload includes determining that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is coded for a reference frame; determining that the GDR frame is the ordinal first frame retrieved for a bitstream containing video data such that no reference frame has been retrieved; and discarding at least a portion of the video frame in response to the fact that no reference frame has been retrieved.
[0171] Clause 25: A device for extracting media data, comprising one or more means for performing any of the methods described in Clauses 1 to 24.
[0172] Clause 26: A device of Clause 25, in which one or more means include one or more processors implemented within the circuit.
[0173] Clause 27: The device of Clause 25, further comprising memory configured to store video data.
[0174] Clause 28: A device according to Clause 25, wherein the device includes at least one of an integrated circuit, a microprocessor, and a wireless communication device.
[0175] Clause 29: A device for receiving media data, comprising means for receiving a packet, the packet header and the payload comprising at least a portion of a frame of video data, wherein the packet header is separate from the payload; means for extracting a video frame identifier relating to a frame of video data from the packet header; and means for processing the payload according to the video frame identifier.
[0176] Clause 30: A method for receiving video data, comprising receiving a packet comprising a packet header and a payload comprising at least a portion of the frames of the video data, wherein the packet header is separate from the payload; extracting a video frame identifier relating to the frames of the video data from the packet header; and processing the payload in accordance with the video frame identifier.
[0177] Clause 31: The method of Clause 30, wherein the video frame identifier includes a picture order count (POC) value.
[0178] Clause 32: The method of Clause 30, wherein the video frame identifier includes a display frame identifier or a current frame identifier.
[0179] Clause 33: The method of Clause 30, wherein at least a portion of the frame includes a slice of the frame, and the video frame identifier includes a slice identifier relating to that slice.
[0180] Clause 34: The method of Clause 30, wherein at least a portion of the frame includes a frame tile, and the video frame identifier includes a tile identifier relating to that slice.
[0181] Clause 35: The method of Clause 30, further comprising using a video frame identifier to determine one or more of the following: frame type, priority of a frame, or dependency information of a frame.
[0182] Clause 36: The method of Clause 30, wherein at least a portion of the frames of video data includes PDUs of a set of protocol data units (PDUs).
[0183] Clause 37: The method of Clause 30, further comprising extracting from the packet header one or more of the following: a network abstraction layer (NAL) unit type relating to at least a portion of the frame, a time identifier (TID) relating to at least a portion of the frame, a layer identifier (LID) relating to at least a portion of the frame, data indicating whether at least a portion of the frame is intra-predictively coded, or data indicating whether at least a portion of the frame is discardable.
[0184] Clause 38: The method of Clause 30, further comprising processing a Network Abstraction Layer (NAL) unit header relating to a frame, wherein the NAL unit header includes data indicating a priority value relating to the frame.
[0185] Clause 39: The method of Clause 30, further comprising processing an Access Unit Delimiter (AUD) relating to an Access Unit corresponding to a frame, wherein the AUD includes data representing a priority value relating to the Access Unit.
[0186] Clause 40: The method of Clause 30, further comprising processing a picture header relating to a frame, wherein the picture header contains data representing a priority value relating to the frame.
[0187] Clause 41: The method of Clause 30, further comprising processing an open bitstream unit (OBU) relating to at least a portion of a frame, wherein the OBU includes data representing a priority value relating to at least a portion of the frame.
[0188] Clause 42: The method of Clause 41, wherein the OBU includes either a frame header OBU or a metadata OBU.
[0189] Clause 43: The method of Clause 30, further comprising receiving data indicating the picture distance between a frame and a reference frame relating to that frame.
[0190] Clause 44: The method of Clause 43, wherein receiving data indicating picture distance includes receiving a Network Abstraction Layer (NAL) unit containing such data, or a Supplemental Additional Information (SEI) message containing such data.
[0191] Clause 45: The method of Clause 43, which includes receiving data indicating the picture distance between a frame and a reference frame relating to that frame, and receiving data indicating the picture distance between a frame and each active reference frame.
[0192] Clause 46: The method of Clause 30, further comprising receiving information indicating whether a portion of a frame of video data can be decoded without using the other portions of the frame.
[0193] Clause 47: The method of Clause 46 for indicating whether a portion of a frame contains a protocol data unit (PDU), that the frame corresponds to a set of PDUs containing the PDU, and that the information can decode the PDU without using other PDUs in the set.
[0194] Clause 48: The method of Clause 30, further comprising receiving information indicating whether loop filtering should be performed across one or more boundaries between a portion of a frame and one or more other portions of the frame.
[0195] Clause 49: The method of Clause 30, further comprising receiving information indicating whether a portion of a frame is coded independently without relying on predictions from other portions of the frame.
[0196] Clause 50: The method of Clause 30, further including receiving data indicating that the frame is a discardable frame.
[0197] Clause 51: The method of Clause 50, which includes receiving a frame in a Network Abstraction Layer (NAL) unit, an Access Unit Delimiter, a Picture Header, a Slice Header, a Supplemental Additional Information (SEI) message, or a Frame Header Open Bitstream Unit (OBU), thereby receiving data indicating that the frame is a discardable frame.
[0198] Clause 52: The method of Clause 30, wherein processing the payload includes determining that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is independently coded, and providing at least a portion of the frame to a video decoder in response to determining that at least a portion of the frame is independently coded.
[0199] Clause 53: The method of Clause 30, wherein processing the payload includes determining that a frame is a gradual decoder refresh (GDR) frame and that at least a portion of the frame is coded for a reference frame; determining that the GDR frame is the first frame retrieved with respect to a bitstream containing video data such that no reference frame has been retrieved; and discarding at least a portion of the video frame in response to the fact that no reference frame has been retrieved.
[0200] Clause 54: A method for receiving video data, comprising receiving a packet comprising a packet header and a payload comprising at least a portion of the frames of video data, wherein the packet header is separate from the payload; extracting a video frame identifier relating to the frames of video data from the packet header; and processing the payload in accordance with the video frame identifier.
[0201] Clause 55: The method of Clause 54, wherein the video frame identifier includes a picture order count (POC) value.
[0202] Clause 56: The method of Clause 54, wherein the video frame identifier includes a display frame identifier or a current frame identifier.
[0203] Clause 57: The method of Clause 54, wherein at least a portion of the frame includes a slice of the frame, and the video frame identifier includes a slice identifier relating to that slice.
[0204] Clause 58: The method of Clause 54, wherein at least a portion of the frame includes a frame tile, and the video frame identifier includes a tile identifier relating to that tile.
[0205] Clause 59: The method of Clause 54, further comprising using a video frame identifier to determine one or more of the following: frame type, priority of a frame, or dependency information of a frame.
[0206] Clause 60: The method of Clause 54, wherein at least a portion of the frames of video data includes PDUs of a set of protocol data units (PDUs).
[0207] Clause 61: The method of Clause 54, further comprising extracting from the packet header one or more of the following: a network abstraction layer (NAL) unit type relating to at least a portion of the frame, a time identifier (TID) relating to at least a portion of the frame, a layer identifier (LID) relating to at least a portion of the frame, data indicating whether at least a portion of the frame is intra-predictively coded, or data indicating whether at least a portion of the frame is discardable.
[0208] Clause 62: The method of Clause 54, further comprising processing a Network Abstraction Layer (NAL) unit header relating to a frame, wherein the NAL unit header includes data indicating a priority value relating to the frame.
[0209] Clause 63: The method of Clause 54, further comprising processing an Access Unit Delimiter (AUD) relating to an Access Unit corresponding to a frame, wherein the AUD includes data representing a priority value relating to the Access Unit.
[0210] Clause 64: The method of Clause 54, further comprising processing a picture header relating to a frame, wherein the picture header includes data representing a priority value relating to the frame.
[0211] Clause 65: The method of Clause 54, further comprising processing an open bitstream unit (OBU) relating to at least a portion of a frame, wherein the OBU includes data representing a priority value relating to at least a portion of the frame.
[0212] Clause 66: The method of Clause 65, wherein the OBU includes either a frame header OBU or a metadata OBU.
[0213] Clause 67: The method of Clause 54, further comprising receiving data indicating the picture distance between a frame and a reference frame relating to that frame.
[0214] Clause 68: The method of Clause 67, wherein receiving data indicating picture distance includes receiving a Network Abstraction Layer (NAL) unit containing such data, or a Supplemental Additional Information (SEI) message containing such data.
[0215] Clause 69: The method of Clause 67, which includes receiving data indicating the picture distance between a frame and a reference frame relating to that frame, and receiving data indicating the picture distance between a frame and each active reference frame.
[0216] Clause 70: The method of Clause 54, further comprising receiving information indicating whether a portion of a frame of video data can be decoded without using the other portions of the frame.
[0217] Clause 71: The method of Clause 70 for indicating whether a portion of a frame contains a protocol data unit (PDU), that the frame corresponds to a set of PDUs containing the PDU, and that the information can decode the PDU without using other PDUs in the set.
[0218] Clause 72: The method of Clause 54, further comprising receiving information indicating whether loop filtering should be performed across one or more boundaries between a portion of a frame and one or more other portions of the frame.
[0219] Clause 73: The method of Clause 54, further comprising receiving information indicating whether a portion of a frame is coded independently without relying on predictions from other portions of the frame.
[0220] Clause 74: The method of Clause 54, further including receiving data indicating that the frame is a discardable frame.
[0221] Clause 75: The method of Clause 74, which includes receiving a frame in a Network Abstraction Layer (NAL) unit, an Access Unit Delimiter, a Picture Header, a Slice Header, a Supplemental Additional Information (SEI) message, or a Frame Header Open Bitstream Unit (OBU), thereby receiving data indicating that the frame is a discardable frame.
[0222] Clause 76: The method of Clause 54, wherein processing the payload includes determining that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is independently coded, and providing at least a portion of the frame to a video decoder in response to determining that at least a portion of the frame is independently coded.
[0223] Clause 77: The method of Clause 54, wherein processing the payload includes determining that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is coded for a reference frame; determining that the GDR frame is the first frame retrieved with respect to a bitstream containing video data such that no reference frame has been retrieved; and discarding at least a portion of the video frame in response to the fact that no reference frame has been retrieved.
[0224] Clause 78: A device for retrieving media data, comprising a memory and a processing system comprising one or more processors implemented in the circuit, wherein the processing system is configured to receive a packet comprising a packet header and a payload comprising at least a portion of a frame of video data, the packet header being separate from the payload, to extract a video frame identifier relating to a frame of video data from the packet header, and to process the payload according to the video frame identifier.
[0225] Clause 79: A device under Clause 78 whose video frame identifier includes at least one of the following: picture order count (POC) value, display frame identifier, or current frame identifier.
[0226] Clause 80: A device under Clause 78 in which at least a portion of a frame includes a slice or tile of a frame, and the video frame identifier includes an identifier relating to that slice or tile of a frame.
[0227] Clause 81: A device under Clause 78, whose processing system is configured to determine one or more of the following using a video frame identifier: frame type, frame priority, or frame dependency information.
[0228] Clause 82: A device of Clause 78 in which at least a portion of the frames of video data includes PDUs of a set of protocol data units (PDUs).
[0229] Clause 83: A device under Clause 78, further configured to extract from a packet header one or more of the following: a network abstraction layer (NAL) unit type relating to at least a portion of a frame; a time identifier (TID) relating to at least a portion of a frame; a layer identifier (LID) relating to at least a portion of a frame; data indicating whether at least a portion of a frame is intra-predictively coded; or data indicating whether at least a portion of a frame is discardable.
[0230] Clause 84: The processing system is further configured to process a network abstraction layer (NAL) unit header relating to a frame, the device of Clause 78, wherein the NAL unit header contains data indicating a priority value relating to the frame.
[0231] Clause 85: The processing system is further configured to process an Access Unit Delimiter (AUD) relating to an Access Unit corresponding to a frame, the AUD containing data representing a priority value relating to the Access Unit, as per the device of Clause 78.
[0232] Clause 86: The device of Clause 78, further configured to process picture headers relating to frames, wherein the picture headers contain data representing priority values relating to frames.
[0233] Clause 87: The processing system is further configured to process open bitstream units (OBUs) relating to at least a portion of a frame, the devices of Clause 78, wherein the OBUs contain data representing priority values relating to at least a portion of the frame.
[0234] Clause 88: The device of Clause 78, further configured to receive data indicating the picture distance between a frame and a reference frame relating to that frame, wherein this data is contained within a Network Abstraction Layer (NAL) unit or a Supplementary Additional Information (SEI) message.
[0235] Clause 89: The device of Clause 78, further configured to receive information indicating whether a portion of a frame of video data can be decoded without using the other portion of the frame.
[0236] Clause 90: The device of Clause 78, further configured to receive information indicating whether loop filtering should be performed across one or more boundaries between a portion of a frame and one or more other portions of the frame.
[0237] Clause 91: A device of Clause 78, configured to process a payload, wherein the processing system determines that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is independently coded, and in response to determining that at least a portion of the frame is independently coded, provides at least a portion of the frame to a video decoder.
[0238] Clause 92: A device of Clause 78, configured to process a payload, the processing system to determine that a frame is a progressive decoder refresh (GDR) frame and that at least a portion of the frame is coded with respect to a reference frame, and that the GDR frame is the first frame retrieved with respect to a bitstream containing video data such that no reference frame has been retrieved, and in response to the absence of a reference frame, discard at least a portion of the video frame.
[0239] Clause 93: A device for receiving video data, comprising means for receiving a packet, the packet header and the payload comprising at least a portion of a frame of video data, wherein the packet header is separate from the payload; means for extracting a video frame identifier relating to a frame of video data from the packet header; and means for processing the payload according to the video frame identifier.
[0240] Clause 94: A computer-readable storage medium storing instructions, wherein, when the instructions are executed, causes a processor to receive a packet comprising a packet header and a payload comprising at least a portion of a frame of video data, wherein the packet header is separate from the payload; to extract a video frame identifier relating to a frame of video data from the packet header; and to process the payload according to the video frame identifier.
[0241] In one or more embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. Examples of computer-readable mediums include tangible computer-readable storage media such as data storage media, or communication media including any medium that facilitates the transfer of computer programs from one location to another, for example, according to a communication protocol. Thus, computer-readable media may generally correspond to (1) non-transient tangible computer-readable storage media, or (2) communication media such as signals or carrier waves. Data storage media may be any available medium accessible by one or more computers or one or more processors for retrieving instructions, codes, and / or data structures for implementation of the technology described herein. Computer program products may include computer-readable media.
[0242] For example, but not limited to, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures and that are accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then those coaxial cables, fiber optic cables, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but instead refer to non-temporary, tangible storage media. As used herein, "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where a "disk" typically reproduces data magnetically, while a "disc" reproduces data optically using a laser. Any combination of these should also be included within the scope of computer-readable media.
[0243] Instructions can be executed by one or more processors, such as digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term “processor” as used herein may refer to any of the above-described structures or any other structure suitable for implementing the technologies described herein. Furthermore, in some embodiments, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated within a composite codec. These technologies can also be fully implemented in one or more circuits or logic elements.
[0244] The technology disclosed herein can be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described herein to highlight the functional aspects of devices configured to perform the disclosed technology, but these do not necessarily require implementation by different hardware units. Rather, as described above, various units can be combined within a codec hardware unit, or provided by a set of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0245] Various embodiments have been described. These embodiments and other embodiments are within the scope of the following claims.
Claims
1. A method for receiving video data, Receiving a packet comprising a packet header and a payload containing at least a portion of a frame of video data, wherein the packet header is separate from the payload; Extracting video frame identifiers relating to at least a portion of the video data frames from the packet header, A method comprising processing the payload according to the video frame identifier.
2. The received packet is a first real-time transport protocol (RTP) packet, The packet header is the first RTP packet header, The aforementioned payload is the first payload, The aforementioned portion of the frame is the first portion of the frame, The first RTP packet header further includes a first RTP header extension, at least the first portion of the frame corresponds to a first protocol data unit (PDU), and the frame of the video data corresponds to a set of protocol data units (PDUs) including the first PDU. The aforementioned method, Receiving a second RTP packet, wherein the second RTP packet comprises a second RTP packet header and a second payload containing at least a second portion of the frame of the video data, the second RTP packet header being separate from the second payload, the second RTP packet header further comprising a second RTP header extension, and the at least second portion of the frame corresponding to a second PDU of the PDU set. From the second RTP header extension, extract the video frame identifier for at least the second portion of the video data frame, For the PDU set, the Quality of Service (QoS) flow identifier (QFI) is determined according to the video frame identifier, It further includes, Processing the payload according to the video frame identifier includes performing QoS processing on the first RTP packet and the second RTP packet according to the QFI relating to the PDU set. The method according to claim 1.
3. The method according to claim 1, wherein the video frame identifier includes a picture order count (POC) value, or the video frame identifier includes a display frame identifier or a current frame identifier.
4. At least a portion of the frame A slice of the aforementioned frame, wherein the video frame identifier includes a slice identifier relating to the aforementioned slice, or A tile of the frame, wherein the video frame identifier includes a tile identifier relating to the tile, The method according to claim 1, including the method described in claim 1.
5. The method according to claim 1, further comprising determining one or more of the following using the video frame identifier: a frame type relating to the frame, a priority relating to the frame, or dependency information relating to the frame.
6. The method according to claim 1, wherein at least a portion of the frames of the video data includes a PDU of a set of protocol data units (PDUs).
7. The method according to claim 1, further comprising extracting from the packet header one or more of the following: a network abstraction layer (NAL) unit type relating to at least a portion of the frame, a time identifier (TID) relating to at least a portion of the frame, a layer identifier (LID) relating to at least a portion of the frame, data indicating whether or not at least a portion of the frame is intra-predictively coded, or data indicating whether or not at least a portion of the frame is discardable.
8. Processing the network abstraction layer (NAL) unit header relating to the frame, wherein the NAL unit header includes data indicating a priority value relating to the frame, or Processing an access unit delimiter (AUD) relating to an access unit corresponding to the frame, wherein the AUD includes data representing a priority value relating to the access unit, or Processing the picture header relating to the frame, wherein the picture header includes data representing a priority value relating to the frame, or Processing an open bitstream unit (OBU) relating to at least a portion of the frame, wherein the OBU includes data representing a priority value relating to at least a portion of the frame. Optionally, the OBU may include either a frame header OBU or a metadata OBU. The method according to claim 1, further comprising:
9. The process further includes receiving data indicating the picture distance between the frame and a reference frame relating to the frame, and optionally receiving data indicating the picture distance. Receiving a Network Abstraction Layer (NAL) unit containing the aforementioned data, or a Supplemental Additional Information (SEI) message containing the aforementioned data, or This includes receiving data indicating the picture distance between the frame and each active reference frame. The method according to claim 1.
10. Information indicating whether a portion of a frame of the video data can be decoded without using the other portion of the frame, optionally comprising: the portion of the frame including a protocol data unit (PDU), the frame corresponding to a set of PDUs including the PDU, and the information indicating whether the PDU can be decoded without other PDUs in the set; or Information indicating whether loop filtering should be performed across one or more boundaries between the portion of the frame and one or more other portions of the frame, or Information indicating whether the portion of the frame is coded independently without using predictions from other parts of the frame, The method according to claim 1, further comprising receiving.
11. The process further includes receiving data indicating that the frame is a discardable frame, and optionally, receiving data indicating that the frame is a discardable frame includes receiving the frame in a Network Abstraction Layer (NAL) unit, an Access Unit Delimiter, a Picture Header, a Slice Header, a Supplemental Additional Information (SEI) message, or a Frame Header Open Bitstream Unit (OBU). The method according to claim 1.
12. Processing the aforementioned payload It is determined that the frame is a gradual decoder refresh (GDR) frame, and that at least a portion of the frame is independently coded. In response to determining that at least a portion of the frame is independently coded, the at least portion of the frame is provided to a video decoder, or, Processing the aforementioned payload It is determined that the frame is a gradual decoder refresh (GDR) frame, and that at least a portion of the frame is coded with respect to a reference frame. The GDR frame is determined to be the first frame extracted with respect to the bitstream containing the video data, such that the reference frame has not been extracted. In response to the fact that the aforementioned reference frame has not been retrieved, the following is performed: The method according to claim 1. The method according to claim 1.
13. A device for extracting media data, Memory and A processing system comprising one or more processors implemented within the circuit, wherein the processing system Receiving a packet comprising a packet header and a payload containing at least a portion of a frame of video data, wherein the packet header is separate from the payload; Extracting video frame identifiers relating to at least a portion of the video data frames from the packet header, Processing the payload according to the video frame identifier, A device configured to perform the following actions.
14. The device according to claim 13, further configured to carry out the method described in any one of claims 2 to 12.
15. A computer-readable storage medium in which instructions are stored, wherein when an instruction is executed, it causes a processor to perform the method according to any one of claims 1 to 12.