Signaling media timing information from media application to network element
By sending signaling to notify the frame rate of media data in 5G network, the power consumption and frame loss problems caused by changes in the frame rate of client devices are solved, and frame rate adaptation and battery life are improved.
Patent Information
- Application Number
- CN202480009604.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-02
- Filing Date
- 2024-01-19
- Publication Date
- 2025-08-26
AI Technical Summary
In 5G networks, changes in frame rate of client devices may cause network components to be unable to adjust in time, resulting in increased power consumption and frame loss, and it is difficult for the prior art to effectively manage frame rate adaptation.
By sending signaling notification frame rates in network devices to inform media data, allowing network components to adjust configurations to adapt to the current frame rate, such as signaling notification frame rates or delay information in RTP packets, the client device can deactivate and activate hardware components at the appropriate time to reduce power consumption.
It realizes effective frame rate adaptation in 5G networks, reduces power consumption and ensures accurate frame reception, and improves battery life and processing efficiency.
Smart Images

Figure CN120548698A_ABST
Abstract
Description
[0001] This application claims priority to U.S. application No. 18 / 163,622, filed February 2, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to the storage and transmission of encoded media data. Background Art
[0003] Digital video capabilities may be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, video teleconferencing equipment, etc. Digital video devices implement video compression techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions to these standards, to more efficiently transmit and receive digital video information.
[0004] Video compression techniques use spatial and / or temporal prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, video frames or slices can be partitioned into macroblocks. Each macroblock can be further partitioned. Macroblocks in intra-coded (I) frames or slices are encoded using spatial prediction relative to neighboring macroblocks. Macroblocks in inter-coded (P or B) frames or slices can use spatial prediction relative to neighboring macroblocks in the same frame or slice or temporal prediction relative to other reference frames.
[0005] After the video data is encoded, it can be packaged for transmission or storage. The video data can be assembled into a video file that conforms to any of a variety of standards, such as the International Organization for Standardization (ISO) base media file format and its extensions (such as AVC). Summary of the Invention
[0006] In general, this disclosure describes techniques for exchanging media data over a network. In a 5G network, for example, a client device (user equipment, or "UE") sends reception statistics to a base station (gNB). Reception statistics are often associated with the frame rate of video data, so signaling of reception statistics is based on a frame rate cadence. In some cases, the client device may wish to perform frame rate adaptation (e.g., based on network conditions). Modifying the frame rate in this way may alter the reception statistics reported. However, the gNB cannot determine when the frame rate changes. This disclosure describes techniques for signaling the frame rate of media data to the gNB so that the gNB can be configured to correctly receive the reception statistics.
[0007] Modifying the frame rate can also modify the time at which video data frames are received. For example, the client device can deactivate one or more hardware components related to receiving the frame's data during the time between frames. Such hardware components may include, for example, antennas, processing circuitry, and so on. That is, after receiving a frame of video data at a particular frame rate, if the frame rate is constant, the client device can deactivate the hardware component for a period of time determined by the frame rate, and then reactivate the hardware component when the next frame is to be delivered. Furthermore, if a subsequent frame is received early, the gNB or other transmitting device can also determine not to transmit the subsequent frame until the client device expects to receive it. In this way, the client device can reduce power consumption when streaming media data over the network, which can improve the battery life of battery-powered devices such as smartphones.
[0008] In one example, a method for exchanging media data via a network includes: receiving, by a network device, data representing an expected time between a first frame of media data and a second frame of media data from a media application; receiving, by the network device, the first frame of media data at a first time; waiting, by the network device, to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and processing, by the network device, the second frame of media data at the second time.
[0009] In another example, a device for exchanging media data via a network includes: a memory configured to store the media data; and one or more processors implemented in a circuit and configured to: retrieve data representing an expected time between a first frame of media data and a second frame of media data from a media application; receive the first frame of media data at a first time; wait to process the second frame of media data until a second time equal to or greater than the first time plus the expected time; and process the second frame of media data at a second time.
[0010] In another example, a computer-readable storage medium has instructions stored thereon that, when executed, cause a processor of a network device to: receive data from a media application representing an expected time between a first frame of media data and a second frame of media data; receive the first frame of media data at a first time; wait to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and process the second frame of media data at the second time.
[0011] In another example, a device for exchanging media data via a network includes: a component for receiving data representing an expected time between a first frame of media data and a second frame of media data from a media application; a component for receiving the first frame of media data at a first time; a component for waiting to process the second frame of media data until a second time equal to or greater than the first time plus the expected time; and a component for processing the second frame of media data at the second time.
[0012] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a block diagram illustrating an example system that implements techniques for streaming media data over a network.
[0014] Figure 2 is a conceptual diagram illustrating an example of signaling data representing a frame rate or other time between packets in packets including media data, according to techniques of this disclosure.
[0015] Figure 3 is a block diagram illustrating elements of an example video file.
[0016] Figure 4 is a flow chart illustrating an example method of exchanging media data via a network, according to the techniques of this disclosure.
[0017] Figure 5 is a flow chart illustrating another example method of exchanging media data via a network, in accordance with the techniques of this disclosure. DETAILED DESCRIPTION
[0018] In general, this disclosure describes techniques for exchanging media data over a network. The network can be a 5G network or other radio access network (RAN). In a 5G network, the physical downlink control channel (PDCCH) is used to transmit scheduling information to user equipment (UE) client devices. For example, for extended reality (XR) media data, such as augmented reality (AR), mixed reality (MR), virtual reality (VR), video data, audio data, etc., the PDCCH monitoring cadence can be aligned with the media data traffic cadence in the downlink channel. For example, for media data at 60 frames per second (FPS), a discontinuous reception (DRX) cycle can be used to monitor the downlink data channel at 16 ms, 17 ms, and 17 ms if the cycle repeats in this manner. PDCCH skipping can occur during inter-periods, during which the downlink channel is not monitored and no data is transmitted.
[0019] DRX is initiated in this manner when both the client device (e.g., user equipment (UE)) and the source device (e.g., base station (gNB)) can determine the frame rate of the corresponding media data. In some cases, the frame rate may change without network components being aware of the frame rate change. That is, the frame rate may change at the application layer of the network stack, but information indicating the frame rate change may not be available to lower layers of the network stack. Some video codecs, such as ITU-T H.265 / High Efficiency Video Coding (HEVC) and ITU-T H.266 / Virtual Video Coding (VVC), can support frame rates significantly higher than 60 FPS. Some applications, such as Web Real-Time Communication (WebRTC), include techniques for performing frame rate adaptation based on network conditions. For example, when available bandwidth increases, higher frame rate data may be exchanged, while when available bandwidth decreases, lower frame rate data may be exchanged. This disclosure describes techniques for providing information indicating the current frame rate of media data to lower-layer components (e.g., networking components of a network device) to adapt the network configuration to the current frame rate.
[0020] Providing information indicating the current frame rate in this manner allows network components to adjust network configurations based on the current frame rate. For example, the PDCCH / downlink transmission / monitoring cadence can be modified to accommodate the current frame rate. Consequently, each device involved in network communication of media data can adapt the transmission / monitoring cadence so that frames are not lost, while also allowing for battery conservation and improved processing efficiency when frames are not being transmitted.
[0021] Figure 1 is a block diagram illustrating an example system 10 implementing techniques for streaming media data over a network. In this example, system 10 includes content preparation device 20, server device 60, and client device 40. Client device 40 and server device 60 are communicatively coupled via network 74, which may include the Internet. In some examples, content preparation device 20 and server device 60 may also be coupled via network 74 or another network, or may be directly communicatively coupled. In some examples, content preparation device 20 and server device 60 may comprise the same device.
[0022] exist Figure 1 In the example shown, content preparation device 20 includes an audio source 22 and a video source 24. Audio source 22 may include, for example, a microphone that generates an electrical signal representing captured audio data to be encoded by audio encoder 26. Alternatively, audio source 22 may include a storage medium storing previously recorded audio data, an audio data generator (such as a computerized synthesizer), or any other source of audio data. Video source 24 may include a camera that generates video data to be encoded by video encoder 28, a storage medium encoded with previously recorded video data, a video data generation unit (such as a computer graphics source), or any other source of video data. In all examples, content preparation device 20 need not be communicatively coupled to server device 60, but rather may store multimedia content on a separate medium that is read by server device 60.
[0023] The original audio data and video data may include analog data or digital data. The analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may obtain audio data from a speaking participant while the speaking participant is speaking, and video source 24 may simultaneously obtain video data of the speaking participant. In other examples, audio source 22 may include a computer-readable storage medium containing stored audio data, and video source 24 may include a computer-readable storage medium containing stored video data. In this manner, the techniques described in this disclosure may be applied to live, streaming, real-time audio and video data, or to archived, pre-recorded audio and video data.
[0024] An audio frame corresponding to a video frame is typically an audio frame containing audio data captured (or generated) by audio source 22 concurrently with the video data captured (or generated) by video source 24 contained within the video frame. For example, when a speaking participant typically generates audio data by speaking, audio source 22 captures the audio data, and video source 24 captures video data of the speaking participant at the same time (i.e., while audio source 22 is capturing the audio data). Thus, an audio frame may temporally correspond to one or more specific video frames. Accordingly, an audio frame corresponding to a video frame typically corresponds to a scenario where audio data and video data were captured at the same time; in this case, the audio frame and video frame, respectively, include audio data and video data captured at the same time.
[0025] In some examples, audio encoder 26 may encode a timestamp in each encoded audio frame indicating the time at which the audio data of the encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp in each encoded video frame indicating the time at which the video data of the encoded video frame was recorded. In such an example, the audio frames corresponding to the video frames may include an audio frame including a timestamp and a video frame including the same timestamp. Content preparation device 20 may include an internal clock, and audio encoder 26 and / or video encoder 28 may generate timestamps based on the internal clock, or audio source 22 and video source 24 may use the internal clock to associate audio data and video data with timestamps, respectively.
[0026] In some examples, audio source 22 may send data to audio encoder 26 corresponding to the time at which the audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to the time at which the video data was recorded. In some examples, audio encoder 26 may encode a sequence identifier in the encoded audio data to indicate the relative temporal ordering of the encoded audio data, but not necessarily the absolute time at which the audio data was recorded. Similarly, video encoder 28 may also use a sequence identifier to indicate the relative temporal ordering of the encoded video data. Similarly, in some examples, the sequence identifier may be mapped to a timestamp or otherwise associated with a timestamp.
[0027] The audio encoder 26 typically produces an encoded audio data stream, while the video encoder 28 produces an encoded video data stream. Each individual data stream (whether audio or video) can be referred to as an elementary stream. An elementary stream is a single digitally decoded (possibly compressed) component of a media presentation. For example, the decoded video portion or the decoded audio portion of a media presentation can be an elementary stream. An elementary stream can be converted into a packetized elementary stream (PES) before being encapsulated into a video file. Within the same media presentation, a stream ID can be used to distinguish PES packets belonging to one elementary stream from PES packets belonging to another elementary stream. The basic data unit of an elementary stream is a packetized elementary stream (PES) packet. Thus, decoded video data typically corresponds to an elementary video stream. Similarly, audio data corresponds to one or more corresponding elementary streams.
[0028] exist Figure 1 In the example of FIG, encapsulation unit 30 of content preparation device 20 receives an elementary stream including decoded video data from video encoder 28 and an elementary stream including decoded audio data from audio encoder 26. In some examples, video encoder 28 and audio encoder 26 may each include a packetizer for forming PES packets from the encoded data. In other examples, video encoder 28 and audio encoder 26 may each interface with a respective packetizer for forming PES packets from the encoded data. In yet another example, encapsulation unit 30 may include a packetizer for forming PES packets from the encoded audio data and the encoded video data.
[0029] The video encoder 28 can encode the video data of the multimedia content in various ways, thereby producing different representations of the multimedia content at various bit rates and with various characteristics, such as pixel resolution, frame rate, compliance with various decoding standards, compliance with various profiles and / or profile levels of various decoding standards, representations with one or more views (e.g., for two-dimensional or three-dimensional playback), or other such characteristics. As used in this disclosure, a representation can include audio data, video data, text data (e.g., for closed captioning), or one of other such data. A representation can include elementary streams, such as an audio elementary stream or a video elementary stream. Each PES packet can include a stream_id that identifies the elementary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling the elementary streams into streamable media data.
[0030] Encapsulation unit 30 receives PES packets of elementary streams of a media presentation from audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. Decoded video segments can be organized into NAL units, which provide a "network-friendly" video representation for applications such as video telephony, storage, broadcast, or streaming. NAL units can be classified as Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units can contain the core compression engine and can include block-level, macroblock-level, and / or slice-level data. Other NAL units can be non-VCL NAL units. In some examples, a coded picture in a time instance (typically presented as a primary coded picture) can be contained in an access unit, which can include one or more NAL units.
[0031] Non-VCL NAL units can include parameter set NAL units and SEI NAL units, among others. Parameter sets can contain sequence-level header information (in a sequence parameter set (SPS)) and infrequently changing picture-level header information (in a picture parameter set (PPS)). Utilizing parameter sets (e.g., PPS and SPS), infrequently changing information does not need to be repeated for each sequence or picture; thus, decoding efficiency can be improved. Furthermore, the use of parameter sets enables out-of-band transmission of important header information, thereby avoiding the need for redundant transmission for error resilience. In the out-of-band transmission example, parameter set NAL units can be sent on a different channel than other NAL units (such as SEI NAL units).
[0032] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding coded picture samples from VCL NAL units, but may aid processes related to decoding, display, error resilience, and other purposes. SEI messages may be included in non-VCL NAL units. SEI messages are a normative part of some standard specifications and therefore are not always mandatory for standard-compliant decoder implementations. SEI messages may be sequence-level SEI messages or picture-level SEI messages. Some sequence-level information may be included in SEI messages, such as the scalability information SEI messages in the example of SVC and the view scalability information SEI messages in MVC. These example SEI messages may convey information about, for example, the extraction of operation points and the characteristics of the operation points.
[0033] Server device 60 includes a Real-time Transport Protocol (RTP) transmitter 70 and a network interface 72. In some examples, server device 60 may include multiple network interfaces. Furthermore, any or all features of server device 60 may be implemented on other devices in a content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediary devices in the content delivery network may cache data for multimedia content 64 and include components substantially identical to those of server device 60. Generally, network interface 72 is configured to send and receive data via network 74.
[0034] RTP sending unit 70 is configured to deliver media data to client device 40 via network 74 in accordance with RTP, which is standardized in Request for Comment (RFC) 3550 of the Internet Engineering Task Force (IETF). RTP sending unit 70 may also implement RTP-related protocols, such as RTP Control Protocol (RTCP), Real-time Streaming Protocol (RTSP), Session Initiation Protocol (SIP), and / or Session Description Protocol (SDP). RTP sending unit 70 may send media data via network interface 72, which may implement Uniform Datagram Protocol (UDP) and / or Internet Protocol (IP). Thus, in some examples, server device 60 may use network 74 to send media data via RTP and RTSP over UDP.
[0035] The RTP sending unit 70 may receive an RTSP describe request from, for example, the client device 40. The RTSP describe request may include data indicating what data types are supported by the client device 40. The RTP sending unit 70 may respond to the client device 40 with data indicating media streams that can be sent to the client device 40, such as the media content 64, and a corresponding network location identifier, such as a uniform resource locator (URL) or a uniform resource name (URN).
[0036] The RTP sending unit 70 may then receive an RTSP setup request from the client device 40. The RTSP setup request may generally indicate how the media stream is to be transmitted. The RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. The RTP sending unit 70 may reply to the RTSP setup request with an acknowledgment and data indicating the port on which the server device 60 will send the RTP data and control data. The RTP sending unit 70 may then receive an RTSP play request to cause the media stream to be "played" (i.e., sent to the client device 40 via the network 74). The RTP sending unit 70 may also receive an RTSP teardown request to end a streaming session, in response to which the RTP sending unit 70 may stop sending media data for the corresponding session to the client device 40.
[0037] Likewise, the RTP receiving unit 52 may initiate a media stream by initially sending an RTSP describe request to the server device 60. The RTSP describe request may indicate the types of data supported by the client device 40. The RTP receiving unit 52 may then receive a reply from the server device 60 specifying available media streams (such as media content 64) that can be sent to the client device 40, along with a corresponding network location identifier (such as a uniform resource locator (URL) or uniform resource name (URN)).
[0038] The RTP receiving unit 52 may then generate an RTSP setup request and send the RTSP setup request to the server device 60. As described above, the RTSP setup request may include a network location identifier for the requested media data (e.g., media content 64) and a transport specifier, such as a local port for receiving RTP data and control data (e.g., RTCP data) on the client device 40. In response, the RTP receiving unit 52 may receive an acknowledgment from the server device 60, including the port of the server device 60 that the server device 60 will use to send the media data and control data.
[0039] As part of establishing a media stream, the RTP receiving unit 52 can request a specific frame rate for the media data of the media stream. The frame rate can correspond to visual media data, such as video data, extended reality (XR) data, augmented reality (AR) data, mixed reality (MR) data, and / or virtual reality (VR) data. During the media session, the RTP receiving unit 52 can request an updated frame rate, for example, an increased frame rate if available bandwidth increases, or a decreased frame rate if available bandwidth decreases.
[0040] After a media session is established between server device 60 and client device 40, RTP sending unit 70 of server device 60 may send media data (e.g., packets of media data) to client device 40 according to the media session. Server device 60 and client device 40 may exchange control data (e.g., RTCP data) indicating, for example, reception statistics of client device 40, so that server device 60 may perform congestion control or otherwise diagnose and resolve transmission failures.
[0041] The network interface 54 can receive the media of the selected media presentation and provide it to the RTP receiving unit 52, which in turn can provide the media data to the decapsulation unit 50. The decapsulation unit 50 can decapsulate the elements of the video file into constituent PES streams, depacketize the PES streams to retrieve the encoded data, and send the encoded data to the audio decoder 46 or the video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, as indicated by the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data (which may include multiple views of the stream) to the video output 44.
[0042] According to the techniques of this disclosure, the RTP receiving unit 52 can request a specific frame rate from the RTP sending unit 70 of the server device 60. The RTP sending unit 70 can, in turn, signal the frame rate in the RTP packet so that the network interface 72 and the network interface 54 can extract and determine the frame rate from the RTP packet. The frame rate can be signaled in the RTP header or in the payload of the packet as a payload portion accessible to lower-level components (such as the network interface 72 and the network interface 54).
[0043] The network interface 72 may extract frame rate information from the current packet in order to determine scheduling information for subsequent packets of the media session. If the network interface 72 receives the subsequent packet before the scheduled time for sending the subsequent packet, the network interface 72 may buffer the subsequent packet until the scheduled time and then send the subsequent packet at the scheduled time. Therefore, in some examples, the RTP sending unit 70 of the server device 60 may represent an application that indicates a delay between two frames or a frame rate of two frames. Alternatively, the encapsulation unit 30 of the content preparation device 20 may signal the delay to the output interface 32, and the output interface 32 may signal the delay or frame rate in a packet header of the packet sent to the server device 60 or the client device 40.
[0044] In some examples, client device 40 may include components similar to those of content preparation device 20 and / or server device 60, such that client device 40 can participate in instant messaging sessions with other client devices. That is, each client device can send and receive media data according to the techniques of this disclosure. When sending media data, client device 40 may perform functions belonging to content preparation device 20 and server device 60.
[0045] The gNB may include components similar to those of server device 60 and may receive packets from upstream network equipment, such as a router or content preparation device 20. When the gNB determines the RTP / SRTP packet burst (i.e., a group of RTP packets) for the current frame, it can determine when the RTP packet burst for the next frame will arrive based on the signaled time difference and network jitter. Thus, signaling the delay between frames can support arbitrary frame rate adaptation. Components of the radio access network (RAN) can leverage this information in dynamic scheduling scenarios. For example, downlink control information (DCI) can provide a duration equal to the indicated delay between frames, during which client device 40 can skip PDCCH monitoring.
[0046] As another example, the frame rate itself can be signaled via a signaling protocol, signaling packet, or in the packet itself, such as an RTP, SRTP, or RTCP packet. In some examples, if the frame rate is not highly dynamic, content preparation device 20 and / or server device 60 can signal statistics related to the frame rate, such as jitter (e.g., using Session Description Protocol (SDP)). This indication can be sent to an application function (AF) or a network exposure function (NEF).
[0047] In some examples, if the frame rate is highly dynamic, content preparation device 20 or server device 60 may signal the frame rate in the packet itself. The packet may carry the instantaneous frame rate value, which may be added to a field of the RTP or RTCP payload or header, or added to the SRTP header.
[0048] In some examples, devices within network 74 may signal various supported frame rates to at least one application executed by server device 60 and the at least one application executed (e.g., via an application function (AF) or a network exposure function (NEF)). The supported frame rates may be signaled as an indexed list, allowing client device 40 to select one of the supported frame rates and send the index corresponding to the selected frame rate to server device 60 or content preparation device 20. In some examples, a frame rate may be negotiated between a head-mounted device (HMD) (such as XR / AR / MR / VR glasses or other XR-enabled devices) and an application providing XR data to the HMD. This negotiation may include signaling the supported frame rates to the network. The application may be configured to select from the available frame rates based on, for example, device capabilities, network conditions, available bandwidth, etc., and indicate this selection using SDP or RTP / SRTP packets. This indication may be used by the network (e.g., server device 60) for network configuration and resource allocation.
[0049] In this manner, content preparation device 20 and server device 60 represent examples of devices for exchanging media data via a network, the devices comprising: a memory configured to store media data; and one or more processors embodied in circuitry and configured to: retrieve data representing an expected time between a first frame of media data and a second frame of media data from a media application; receive the first frame of media data at a first time; wait to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and process the second frame of media data at a second time.
[0050] The data representing the expected time may directly indicate the expected time. For example, the data representing the expected time may be a value equal to the expected time itself. Alternatively, the data representing the expected time may be an indirect representation of the expected time, such as a frame rate value or an offset value.
[0051] The network interface 72 may retrieve the signaled frame rate or the signaled delay value from the RTP data received from the RTP sending unit 70. Alternatively, the RTP sending unit 70 may signal available frame rates and receive data from the client device 40 including an index indicating a selected frame rate from the available frame rates. The signaled frame rate represents the expected time between the media data of the current packet (e.g., a frame) and the media data of the subsequent packet (the subsequent frame). In some examples, the signaled frame rate may represent the reciprocal of the expected time between the media data of the current packet (e.g., a frame) and the media data of the subsequent packet (the subsequent frame), and thus, the unit may be the number of frames per a certain time unit (e.g., frames per second). The network interface 72 may buffer the data of the subsequent frame in a memory until the scheduled time for transmission, and then process (transmit) the subsequent frame at the scheduled time.
[0052] Similarly, network interface 54 may extract frame rate information from the current packet in order to determine scheduling information for subsequent packets of the media session. Specifically, network interface 54 may be configured to disable one or more hardware components associated with monitoring the downlink channel for receiving data from network interface 72 until the scheduled time for the subsequent packet to be received. Additionally or alternatively, network interface 54 may observe a delay period until the scheduled time for the subsequent packet to be received. In some examples, network interface 54 may delay processing of the subsequent packet until the scheduled time.
[0053] In this manner, client device 40 represents another example of a device for exchanging media data via a network, the device comprising: a memory configured to store media data; and one or more processors embodied in circuitry and configured to: retrieve data representing an expected time between a first frame of media data and a second frame of media data from a media application; receive the first frame of media data at a first time; wait to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and process the second frame of media data at a second time.
[0054] Specifically, the network interface 54 may retrieve a signaled frame rate from the RTP data received from the server device 60 (specifically, the RTP sending unit 70 of the server device 60). The signaled frame rate represents the expected time between the media data of the current packet (e.g., a frame) and the media data of the subsequent packet (the subsequent frame). In some examples, the signaled frame rate represents the inverse of the expected time between the media data of the current packet (e.g., a frame) and the media data of the subsequent packet (the subsequent frame), and thus may be expressed in frames per a certain time unit (e.g., frames per second). The network interface 54 may disable one or more hardware components associated with monitoring the downlink channel from the server device 60, thereby waiting to process the subsequent frame until the scheduled time for transmission. At the scheduled time, the network interface 54 may activate the one or more hardware components to receive the subsequent frame from the server device 60. Additionally or alternatively, the network interface 54 may observe a delay period until the scheduled time for receipt of the subsequent packet. In some examples, the network interface 54 may delay processing the subsequent packet until the scheduled time.
[0055] Each of the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiving unit 52, and decapsulation unit 50 can be implemented as any of a variety of suitable processing circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuits, software, hardware, firmware, or any combination thereof, as applicable. Each of the video encoder 28 and video decoder 48 can be included in one or more encoders or decoders, which can be integrated as part of a combined video encoder / decoder (CODEC). Similarly, each of the audio encoder 26 and audio decoder 46 can be included in one or more encoders or decoders, which can be integrated as part of a combined CODEC. An apparatus including the video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, encapsulation unit 30, RTP receiving unit 52, and / or decapsulation unit 50 can include an integrated circuit, a microprocessor, and / or a wireless communication device (such as a cellular phone).
[0056] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate according to the techniques of this disclosure. For purposes of example, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may be configured to perform these techniques instead of (or in addition to) server device 60.
[0057] Encapsulation unit 30 may form NAL units that include a header identifying the program to which the NAL unit belongs and a payload (e.g., audio data, video data, or data describing the transport stream or program stream to which the NAL unit corresponds). For example, in H.264 / AVC, a NAL unit includes a 1-byte header and payloads of varying sizes. A NAL unit that includes video data in its payload may include video data at various levels of granularity. For example, a NAL unit may include a block of video data, multiple blocks, a slice of video data, or an entire picture of video data. Encapsulation unit 30 may receive encoded video data in the form of PES packets for elementary streams from video encoder 28. Encapsulation unit 30 may associate each elementary stream with a corresponding program.
[0058] Encapsulation unit 30 may also assemble an access unit from multiple NAL units. In general, an access unit may include one or more NAL units representing a frame of video data and, when such audio data is available, audio data corresponding to the frame. An access unit typically includes all NAL units for an output time instance, e.g., all audio data and video data for a time instance. For example, if each view has a frame rate of 20 frames per second (fps), then each time instance may correspond to a time interval of 0.05 seconds. During this time interval, a particular frame of all views of the same access unit (the same time instance) may be presented simultaneously. In one example, in a time instance, an access unit may include a decoded picture, which may be presented as a primary decoded picture.
[0059] Accordingly, an access unit can include all audio frames and video frames for a common time instance (e.g., all views corresponding to time X). This disclosure also refers to the coded pictures of a particular view as "view components." That is, a view component can include the coded pictures (or frames) of a particular view at a specific time. Accordingly, an access unit can be defined as including all view components for a common time instance. The decoding order of access units does not necessarily have to be the same as the output order or display order.
[0060] After encapsulation unit 30 assembles the NAL units and / or access units into a video file based on the received data, encapsulation unit 30 passes the video file to output interface 32 for output. In some examples, encapsulation unit 30 may store the video file locally or send the video file to a remote server via output interface 32, rather than sending the video file directly to client device 40. Output interface 32 may include, for example, a transmitter, a transceiver, a device for writing data to a computer-readable medium (such as an optical drive, a magnetic media drive (e.g., a floppy drive)), a universal serial bus (USB) port, a network interface, or other output interface. Output interface 32 outputs the video file to a computer-readable medium, such as a transmission signal, magnetic media, optical media, memory, a flash drive, or other computer-readable medium.
[0061] The network interface 54 can receive NAL units or access units via the network 74 and provide the NAL units or access units to the decapsulation unit 50 via the RTP receiving unit 52. The decapsulation unit 50 can decapsulate the elements of the video file into constituent PES streams, depacketize the PES streams to retrieve the encoded data, and send the encoded data to the audio decoder 46 or the video decoder 48, depending on whether the encoded data is part of an audio stream or a video stream, for example, as indicated by the PES packet header of the stream. The audio decoder 46 decodes the encoded audio data and sends the decoded audio data to the audio output 42, while the video decoder 48 decodes the encoded video data and sends the decoded video data (which may include multiple views of the stream) to the video output 44.
[0062] Figure 2 is a conceptual diagram illustrating an example of signaling data indicating a frame rate or other time between packets in a packet including media data according to the techniques of this disclosure. Figure 2 In the example of , media stream 110 includes packet 100A and packet 100B. Packet 100A includes timestamp value 106A, delta time (ΔT) value 104A, and media data 102A, and packet 100B includes timestamp value 106B, ΔT value 104B, and media data 102B.
[0063] In some examples, packet 100A is the first packet carrying a larger media data (e.g., a video frame) including media data 102A. In some examples, packet 100B is the first packet carrying a larger media data (e.g., a video frame) including media data 102B. In some examples, packet 100A is any packet carrying a larger media data (e.g., a video frame) including media data 102A. In some examples, packet 100B is any packet carrying a larger media data (e.g., a video frame) including media data 102B. The delta time (ΔT) may be in Coordinated Universal Time (UTC), a truncated version of UTC, or in units of media data sampling periods. In some examples, a frame rate in frames per second may be substituted for the delta time (ΔT), and the frame rate may be equal to the inverse of ΔT.
[0064] In this example, ΔT value 104A illustrates an example of a value representing the time between packet 100A and packet 100B. For example, ΔT value 104A may be the current frame rate of media data 102A. In some examples, ΔT value 104A may represent a frame rate value for a specific number of frames, including the frame of media data 102A. That is, the signaled frame rate value may represent the frame rate of each of a plurality of frames in a sequence of frames, starting with the frame of media data 102A. In this case, ΔT value 104B need not be signaled. Alternatively, ΔT value 104A may represent an expected reception time of packet 100B. For example, ΔT value 104A may represent an expected time between packet 100A and packet 100B. In some examples, ΔT value 104A may represent an expected time between each packet in the sequence of packets, including packets 100A and 100B. In some examples, ΔT value 104A may be signaled in a payload portion of packet 100A accessible to a network device (eg, an unencrypted portion of the payload).
[0065] In some examples, ΔT value 104A may be signaled in a header of packet 100A (such as in an RTP header, RTSP header, SRTP header, or RTCP header). In some examples, packet 100A may be encapsulated with a tunnel header, and ΔT value 104A may be signaled in the tunnel header. The tunnel header may be, for example, a header according to the General Packet Radio Service (GPRS) Tunneling Protocol User Data Tunneling (GTP-U). In some examples, when the RTP packet is encapsulated to form a GTP-U packet, ΔT value 104A may be copied to the GTP-U packet header. Thus, ΔT value 104A may provide benefits to routing devices in 5G networks or other networks, where the routing device may determine when to expect packet 100B and may disable one or more hardware components or perform other operations (such as load balancing). In some examples, ΔT value 104A may be signaled in an IP packet options field of packet 100A.
[0066] Figure 1 Content preparation device 20 or server device 60 may be configured to calculate ΔT value 104A based on one or more factors, such as the complexity of the scene, whether a scene change is occurring, available computing power, past inter-frame delays, etc.
[0067] In some examples, ΔT value 104A can additionally or alternatively be signaled for a protocol data unit (PDU) set and / or a PDU set burst. A PDU set represents a set of IP packets that carry an information unit (such as a slice of a video frame) at the application layer. A PDU set burst is a set of IP packets or PDUs that should be delivered to a UE (such as client device 40) with the same deadline, such as all slices of a video frame or all data of an access unit.
[0068] Figure 3 is a block diagram illustrating elements of an example video file 150. As described above, video files according to the ISO base media file format and its extensions store data in a series of objects called "boxes". Figure 3 In the example of FIG, a video file 150 includes a file type (FTYP) box 152, a movie (MOOV) box 154, a fragment index (sidx) box 162, a movie fragment (MOOF) box 164, and a movie fragment random access (MFRA) box 166. Figure 3An example of a video file is shown, but it should be understood that other media files may also include other types of media data (eg, audio data, timed text data, etc.) structured similarly to the data of the video file 150 according to the ISO base media file format and its extensions.
[0069] A file type (FTYP) box 152 generally describes the file type of the video file 150. The file type box 152 may include data identifying specifications that describe the optimal use of the video file 150. The file type box 152 may alternatively be placed before the MOOV box 154, the movie fragment box 164, and / or the MFRA box 166.
[0070] exist Figure 3 In the example of , the MOOV box 154 includes a movie header (MVHD) box 156, a track (TRAK) box 158, and one or more movie extension (MVEX) boxes 160. Generally speaking, the MVHD box 156 can describe the overall characteristics of the video file 150. For example, the MVHD box 156 can include data describing when the video file 150 was originally created, when the video file 150 was last modified, the time scale of the video file 150, the playback duration of the video file 150, or other data that generally describes the video file 150.
[0071] The TRAK box 158 may include data for a track of the video file 150. The TRAK box 158 may include a track header (TKHD) box that describes characteristics of the track corresponding to the TRAK box 158. In some examples, the TRAK box 158 may include decoded video pictures, while in other examples, the decoded video pictures of the track may be included in a movie fragment 164, which may be referenced by data in the TRAK box 158 and / or the sidx box 162.
[0072] In some examples, the video file 150 may include more than one track. Accordingly, the MOOV box 154 may include a number of TRAK boxes equal to the number of tracks in the video file 150. The TRAK box 158 may describe the characteristics of the corresponding track of the video file 150. For example, the TRAK box 158 may describe the time information and / or spatial information of the corresponding track. Figure 1 ) When a parameter set track is included in a video file, such as video file 150, a TRAK box similar to TRAK box 158 of MOOV box 154 may describe characteristics of the parameter set track. Encapsulation unit 30 may signal the presence of sequence-level SEI messages in the parameter set track within the TRAK box that describes the parameter set track.
[0073] The MVEX box 160 may describe characteristics of the corresponding movie fragments 164, e.g., to signal that the video file 150 includes the movie fragments 164 in addition to the video data (if any) included in the MOOV box 154. In the case of streaming video data, decoded video pictures may be included in the movie fragments 164 instead of the MOOV box 154. Accordingly, all decoded video samples may be included in the movie fragments 164 instead of the MOOV box 154.
[0074] The MOOV box 154 may include a plurality of MVEX boxes 160 equal to the number of movie fragments 164 in the video file 150. Each MVEX box 160 may describe characteristics of a corresponding one of the movie fragments 164. For example, each MVEX box may include a Movie Extension Header Box (MEHD) box that describes the temporal duration of the corresponding one of the movie fragments 164.
[0075] As described above, encapsulation unit 30 may store sequence data sets in video samples that do not include actual decoded video data. A video sample may generally correspond to an access unit, which is a representation of a decoded picture at a specific instance in time. In the context of AVC, a decoded picture comprises one or more VCL NAL units, which contain information about all pixels and other associated non-VCL NAL units (such as SEI messages) used to construct the access unit. Accordingly, encapsulation unit 30 may include a sequence data set in one of the movie fragments 164, and the sequence data set may include a sequence-level SEI message. Encapsulation unit 30 may also signal the presence of a sequence data set and / or a sequence-level SEI message in one of the movie fragments 164 within one of the MVEX boxes 160 corresponding to one of the movie fragments 164.
[0076] SIDX box 162 is an optional element of video file 150. That is, a video file conforming to the 3GPP file format or other such file formats does not necessarily include SIDX box 162. According to the example of the 3GPP file format, a SIDX box can be used to identify sub-segments of a fragment (e.g., a fragment contained within video file 150). The 3GPP file format defines a sub-segment as "a self-contained set of one or more consecutive Movie Fragment boxes with corresponding Media Data boxes, where the Media Data box containing data referenced by the Movie Fragment box must follow the Movie Fragment box and precede the next Movie Fragment box containing information about the same track." The 3GPP file format also states that a SIDX box "contains a sequence of references to sub-segments of the (sub)fragment recorded by the box. The referenced sub-segments are contiguous in presentation time. Similarly, the bytes referenced by the Fragment Index Box are always contiguous within the fragment. The referenced size gives a count of the number of bytes in the referenced material."
[0077] SIDX box 162 generally provides information representing one or more sub-segments of a segment included in video file 150. For example, such information may include the playback time at which the sub-segment starts and / or ends, the byte offset of the sub-segment, whether the sub-segment includes a stream access point (SAP) (e.g., starts with a SAP), the type of SAP (e.g., whether the SAP is an instantaneous decoder refresh (IDR) picture, a clean random access (CRA) picture, a broken link access (BLA) picture, etc.), the location of the SAP in the sub-segment (in terms of playback time and / or byte offset), etc.
[0078] A movie fragment 164 may include one or more decoded video pictures. In some examples, a movie fragment 164 may include one or more groups of pictures (GOPs), each of which may include multiple decoded video pictures, such as frames or pictures. Furthermore, as described above, a movie fragment 164 may include a sequence data set in some examples. Each movie fragment 164 may include a movie fragment header box (MFHD, Figure 3 MFHD box may describe characteristics of the corresponding movie fragment, such as a sequence number of the movie fragment. The movie fragments 164 may be included in the order of the sequence numbers in the video file 150.
[0079] MFRA box 166 can describe random access points within movie fragments 164 of video file 150. This can facilitate the execution of trick modes, such as seeking to a specific temporal position (i.e., playback time) within the fragments encapsulated by video file 150. In some examples, MFRA box 166 is generally optional and need not be included in the video file. Likewise, a client device (such as client device 40) does not necessarily need to reference MFRA box 166 to correctly decode and display the video data of video file 150. MFRA box 166 can include a number of track fragment random access (TFRA) boxes (not shown) equal to the number of tracks in video file 150, or, in some examples, a number of TFRA boxes equal to the number of media tracks (e.g., non-hint tracks) in video file 150.
[0080] In some examples, a movie fragment 164 may include one or more stream access points (SAPs), such as IDR pictures. Similarly, an MFRA box 166 may provide an indication of the location of the SAP within the video file 150. Accordingly, a temporal subsequence of the video file 150 may be formed from the SAPs of the video file 150. The temporal subsequence may also include other pictures, such as P-frames and / or B-frames, that are subordinate to the SAPs. Frames and / or slices of the temporal subsequence may be arranged within the fragment so that frames / slices of the temporal subsequence that are dependent on other frames / slices of the subsequence can be correctly decoded. For example, in a hierarchical arrangement of data, data used to predict other data may also be included in the temporal subsequence.
[0081] The frame rate of the media data of video file 150 can be signaled, for example, in MVHD 156. Furthermore, when content preparation device 20 packets frames of movie fragments 164, content preparation device 20 can encapsulate all or a portion of one of movie fragments 164 and add data indicating the frame rate to the header of the packet. Thus, the data indicating the frame rate can be accessible to network devices downstream from content preparation device 20. Furthermore, client device 40 can determine an expected time of receipt of subsequent packets. Accordingly, client device 40 can disable one or more hardware components associated with receiving the subsequent packets until the expected time to reduce power consumption and processing operations until the expected time.
[0082] Figure 4 is a flow chart illustrating an example method of exchanging media data via a network, according to the techniques of this disclosure. Figure 4 The method can be Figure 1 The content preparation device 20, server device 60 or client device 40, base station (such as gNB), user equipment (UE), router or other various network devices are executed. For the purpose of explanation, the media communication device is explained Figure 4 The method, the media communication device may correspond to Figure 1 any one of the content preparation device 20, server device 60 or client device 40, a base station (such as a gNB), a UE, a router or other such devices.
[0083] Initially, a media communication device may receive data of a media data frame (200). For example, server device 60, a base station, or client device 40 may receive (e.g., from content preparation device 20 or another client device / UE) one or more packets, protocol data unit (PDU) sets, or bursts of PDU sets that include data of a media data frame (such as one or more slices of a frame). As another example, content preparation device 20 or client device 40 may obtain (e.g., receive, capture, or generate) a media data frame to be sent to a client device / UE.
[0084] The media communication device may then determine an expected time to the next frame (202). For example, when the media communication device is generating / encoding / sending media content, such as when the content preparation device 20, the client device 40, or another UE is generating and sending media content, the media communication device may estimate the time to the next frame based on, for example, the complexity of the current scene, whether a scene change is occurring, computing power, past inter-frame delays, etc. In addition, such a device may form a packet including data indicating the expected delay to the next frame, such as a frame rate or a delay value.
[0085] In some examples, a media communication device acting as a source of media content may signal one or more supported frame rates to, for example, client device 40, and receive a selection of one of the supported frame rates from client device 40. The media communication device may then determine an expected time to the next frame based on the selected frame rate.
[0086] Alternatively, when the media communication device is receiving media data, such as when the media communication device is one of the server device 60, the intermediate network routing device, or the client device 40 and is receiving media data, the media communication device may receive the media data based on the data signaled in the packet including the data of the current frame (such as Figure 2 As described above, data may be signaled in a packet in a portion of a header or payload of the packet accessible to a network device.
[0087] The media communication device may then wait to process the next frame until the expected time (204). For example, when the media communication device is a frame source (e.g., a content preparation device 20 or a client device 40 that sends media data), the media communication device may buffer the data of the next frame until the expected time, and then send the buffered data of the next frame at or after the expected time. Similarly, when the media communication device is an intermediate device (such as a router, a base station, or a server device 60) between two endpoint devices (e.g., a content preparation device 20 and a client device 40 or two client devices), the media communication device may buffer any received data of the next frame until the expected time, and then send the data of the next frame at the expected time. Alternatively, when the media communication device is an endpoint (such as a client device 40 or a UE) that receives media data, the media communication device may disable one or more hardware components associated with receiving the media data until the expected time.
[0088] Finally, the media communication device may process the next frame at the expected time (206). For example, when acting as a source or intermediate device, the media communication device may send data for the next frame at the expected time, or when acting as a destination device, the media communication device may activate a hardware component to receive the next frame.
[0089] In this way, Figure 4 The method represents an example of a method including: receiving, by a network device, data representing an expected time between a first frame of media data and a second frame of media data from a media application; receiving, by the network device, the first frame of media data at a first time; waiting, by the network device, to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and processing, by the network device, the second frame of media data at the second time.
[0090] Figure 5 is a flow chart illustrating another example method of exchanging media data via a network, in accordance with the techniques of this disclosure. Figure 5 The method is performed by devices labeled "source device" and "destination device." For example, the source device may be the content preparation device 20, and the destination device may be Figure 1 Alternatively, the source device may be a first client device / UE and the destination device may be a second client device / UE.
[0091] Initially, a source device may determine a frame rate for a media communication session with a destination device (220). For example, the source device may determine the frame rate and determine a delay between frames based on the frame rate. In some examples, the source device may support various frame rates, and the destination device may request one of the supported frame rates (e.g., by sending an index to a list of various frame rates). The frame rate may vary during the communication session (e.g., based on a request from the destination device and / or based on available bandwidth determined by the source device).
[0092] The source device may obtain data for a first frame of the media communication session (222). For example, the source device may capture or generate data for a first frame, which may be a video frame, an XR / AR / MR / VR frame, etc. The "first" frame need not be the sequential first frame of the media communication session, but may correspond to any frame of the media communication session. The source device may also determine (224) an expected delay to the next frame (e.g., based on a frame rate of the media communication session). Additionally or alternatively, the delay may be based on potential processing delays due to complexity of the virtual scene of the frame, computational power, etc.
[0093] The source device may then form a packet including the delay value and the data of the first frame (226). For example, the source device may form a packet such as Figure 2 The packet shown is formed as packet 100A to include a payload containing media data and an indication of a delay value, which may be in the payload or in a packet header such as an RTP header. The source device may also send the packet to the destination device (228).
[0094] After sending the packet to the destination device, the source device may obtain data for the next frame (230). Similarly, the source device may capture or generate data for the next frame. If the expected time to the next frame has not yet passed, the source device may buffer the data for the next frame (232). After waiting for a delay period corresponding to the expected time to the next frame, the source device may send a packet including the data for the next frame to the destination device (234).
[0095] Meanwhile, the destination device may initially receive a packet including data for a first frame (240). The destination device may extract the signaled delay data to determine a delay to a next frame (242). Similarly, the destination device may extract and present the media data for the first frame (244). The destination device may then wait for a delay period based on the signaled data indicating the delay to the next frame to process the next frame (246). For example, the destination device may disable one or more hardware components associated with monitoring a data communication channel (such as a PDCCH) until the expected time of the next frame.
[0096] After waiting for the delay period to end, the destination device can reactivate the hardware components to receive packets for the next frame (248).The destination device can then extract and present the media data for the next frame (250).
[0097] In this way, Figure 5 A source device and a destination device may both perform a method comprising: receiving, by a network device, data representing an expected time between a first frame of media data and a second frame of media data from a media application; receiving, by the network device, the first frame of media data at a first time; waiting, by the network device, to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and processing, by the network device, the second frame of media data at the second time.
[0098] The following clauses summarize various examples of the techniques of this disclosure:
[0099] Clause 1: A method for exchanging media data via a network, the method comprising: receiving, by a network device, data representing an expected time between a first frame of media data and a second frame of media data from a media application; receiving, by the network device, the first frame of media data at a first time; waiting, by the network device, to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and processing, by the network device, the second frame of media data at the second time.
[0100] Clause 2: The method of clause 1, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving a packet comprising data of the first frame of media data, the packet further comprising data representing the expected time.
[0101] Clause 3: The method of clause 2, wherein the packet comprises one of a Real-time Transport Protocol (RTP) packet, a Real-time Streaming Protocol (RTSP) packet, a Secure Real-time Transport Protocol (SRTP) packet, or an RTP Control Protocol (RTCP) packet.
[0102] Clause 4: The method of clause 2, further comprising: encapsulating the packet with a packet tunnel header; and adding data representing an expected time between the first frame and the second frame to the packet tunnel header.
[0103] Clause 5: The method of clause 4, wherein the packet tunnel header comprises a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
[0104] Clause 6: The method of clause 1, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving data representing an expected time between each of a plurality of frames in a sequence of frames starting with the first frame.
[0105] Clause 7: The method of clause 1, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving a packet encapsulated with a packet tunnel header, the packet tunnel header including the data representing the expected time between the first frame and the second frame.
[0106] Clause 8: The method of clause 7, wherein the packet tunnel header comprises a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
[0107] Clause 9: The method of clause 1, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving a packet including the data representing the expected time between the first frame and the second frame in an Internet Protocol (IP) packet options field.
[0108] Clause 10: The method of clause 1, wherein waiting to process the second frame comprises skipping monitoring of the data transmission channel until a second time, and wherein processing the second frame comprises: monitoring the data transmission channel starting at the second time; and receiving the second frame via the data transmission channel.
[0109] Clause 11: The method of clause 10, wherein the data transmission channel comprises a Physical Downlink Control Channel (PDCCH).
[0110] Clause 12: The method of clause 1, wherein waiting to process the second frame comprises: receiving the second frame before a second time; and sending the second frame at the second time.
[0111] Clause 13: The method of clause 1, further comprising sending data representing an expected time between the first frame and the second frame to a downstream network device.
[0112] Clause 14: The method of clause 1, wherein the data representative of the expected time between the first frame and the second frame comprises a frame rate value for at least a portion of the media stream of the media data.
[0113] Clause 15: The method of clause 14, wherein receiving the frame rate value comprises receiving the frame rate value from a Session Description Protocol (SDP) message.
[0114] Clause 16: The method of clause 14, wherein receiving the frame rate value comprises receiving the frame rate value from a payload or header of a Real-time Transport Protocol (RTP) packet or an RTP Control Protocol (RTCP) packet, or from a header of a Secure Real-time Transport Protocol (SRTP) packet.
[0115] Clause 17: The method of clause 14, further comprising sending data representing one or more frame rates supported by the network device to the media application, wherein receiving the data representing the frame rate comprises receiving data representing a selection of a frame rate from the one or more frame rates supported by the network device from the media application.
[0116] Clause 18: The method of clause 17, further comprising configuring network transmission of the media session of the media data based on the selection of the frame rate.
[0117] Item 19: A device for exchanging media data via a network, the device comprising: a memory configured to store the media data; and one or more processors implemented in circuitry and configured to: retrieve data representing an expected time between a first frame of media data and a second frame of media data from a media application; receive the first frame of media data at a first time; wait to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and process the second frame of media data at a second time.
[0118] Clause 20: The apparatus of clause 19, wherein to receive data representing an expected time between the first frame and the second frame, the one or more processors are configured to receive a packet comprising data of the first frame of media data, the packet further comprising data representing the expected time.
[0119] Clause 21: The apparatus of clause 20, wherein the packet comprises one of a Real-time Transport Protocol (RTP) packet, a Real-time Streaming Protocol (RTSP) packet, a Secure Real-time Transport Protocol (SRTP) packet, or an RTP Control Protocol (RTCP) packet.
[0120] Clause 22: The apparatus of clause 20, wherein the one or more processors are further configured to: encapsulate the packet with a packet tunnel header; and add data representing an expected time between the first frame and the second frame to the packet tunnel header.
[0121] Clause 23: The apparatus of clause 22, wherein the packet tunnel header comprises a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
[0122] Clause 24: The apparatus of clause 19, wherein, to receive data representing an expected time between a first frame and a second frame, the one or more processors are configured to receive data representing an expected time between each of a plurality of frames in a sequence of frames starting with the first frame.
[0123] Clause 25: The apparatus of clause 19, wherein, to receive data representing an expected time between the first frame and the second frame, the one or more processors are configured to receive a packet encapsulated with a packet tunnel header, the packet tunnel header including the data representing the expected time between the first frame and the second frame.
[0124] Clause 26: The apparatus of clause 25, wherein the packet tunnel header comprises a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
[0125] Clause 27: The apparatus of clause 19, wherein, to receive the data representing the expected time between the first frame and the second frame, the one or more processors are configured to receive a packet including the data representing the expected time between the first frame and the second frame in an Internet Protocol (IP) packet option field.
[0126] Clause 28: An apparatus according to clause 19, wherein, to wait for processing of the second frame, the one or more processors are configured to skip monitoring of the data transmission channel until a second time, and wherein, to process the second frame, the one or more processors are configured to: monitor the data transmission channel starting at the second time; and receive the second frame via the data transmission channel.
[0127] Clause 29: The apparatus of clause 28, wherein the data transmission channel comprises a physical downlink control channel (PDCCH).
[0128] Clause 30: The apparatus of clause 19, wherein, to await processing of the second frame, the one or more processors are configured to: receive the second frame before a second time; and send the second frame at the second time.
[0129] Clause 31: The device of clause 19, wherein the one or more processors are further configured to send data representing an expected time between the first frame and the second frame to a downstream network device.
[0130] Clause 32: The apparatus of clause 19, wherein the data representative of the expected time between the first frame and the second frame comprises a frame rate value of at least a portion of the media stream of the media data.
[0131] Clause 33: The apparatus of clause 32, wherein to receive the frame rate value, the one or more processors are configured to receive the frame rate value from a Session Description Protocol (SDP) message.
[0132] Clause 34: The apparatus of clause 32, wherein to receive the frame rate value, the one or more processors are configured to receive the frame rate value from a payload or header of a Real-time Transport Protocol (RTP) packet or an RTP Control Protocol (RTCP) packet or from a header of a Secure Real-time Transport Protocol (SRTP) packet.
[0133] Clause 35: A device according to clause 32, wherein the one or more processors are further configured to send data representing one or more frame rates supported by the network device to the media application, wherein, in order to receive the data representing the frame rate, the one or more processors are configured to receive data representing a selection of a frame rate from the one or more frame rates supported by the network device from the media application.
[0134] Clause 36: The apparatus of clause 35, wherein the one or more processors are further configured to configure network transmission of the media session of the media data based on the selection of the frame rate.
[0135] Clause 37: The apparatus of clause 19, wherein the apparatus comprises at least one of: an integrated circuit; a microprocessor; and a wireless communication device.
[0136] Clause 38: A computer-readable storage medium having instructions stored thereon that, when executed, cause a processor of a network device to: receive data from a media application representing an expected time between a first frame of media data and a second frame of media data; receive the first frame of media data at a first time; wait to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and process the second frame of media data at the second time.
[0137] Clause 39: A device for exchanging media data via a network, the device comprising: a component for receiving data representing an expected time between a first frame of media data and a second frame of media data from a media application; a component for receiving the first frame of media data at a first time; a component for waiting to process the second frame of media data until a second time equal to or greater than the first time plus the expected time; and a component for processing the second frame of media data at the second time.
[0138] Clause 40: A method for exchanging media data via a network, the method comprising: receiving, by a network device, data representing an expected time between a first frame of media data and a second frame of media data from a media application; receiving, by the network device, the first frame of media data at a first time; waiting, by the network device, to process the second frame of media data until a second time that is equal to or greater than the first time plus the expected time; and processing, by the network device, the second frame of media data at the second time.
[0139] Clause 41: The method of clause 40, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving a packet comprising data of the first frame of media data, the packet further comprising data representing the expected time.
[0140] Clause 42: The method of clause 41, wherein the packet comprises one of a Real-time Transport Protocol (RTP) packet, a Real-time Streaming Protocol (RTSP) packet, a Secure Real-time Transport Protocol (SRTP) packet, or an RTP Control Protocol (RTCP) packet.
[0141] Clause 43: The method of any of clauses 41 and 42, further comprising: encapsulating the packet with a packet tunnel header; and adding data representing an expected time between the first frame and the second frame to the packet tunnel header.
[0142] Clause 44: The method of clause 43, wherein the packet tunnel header comprises a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
[0143] Clause 45: The method of any of clauses 40-42, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving data representing an expected time between each of a plurality of frames in a sequence of frames starting with the first frame.
[0144] Clause 46: A method according to any of clauses 40-42 or 45, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving a packet encapsulated with a packet tunnel header, the packet tunnel header comprising data representing the expected time between the first frame and the second frame.
[0145] Clause 47: The method of clause 46, wherein the packet tunnel header comprises a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
[0146] Clause 48: The method of any of clauses 40-42 or 45-47, wherein receiving data representing an expected time between the first frame and the second frame comprises receiving a packet including data representing an expected time between the first frame and the second frame in an Internet Protocol (IP) packet options field.
[0147] Clause 49: A method according to any of clauses 40-42 or 45-48, wherein waiting to process the second frame includes skipping monitoring of the data transmission channel until a second time, and wherein processing the second frame includes: monitoring the data transmission channel starting at the second time; and receiving the second frame via the data transmission channel.
[0148] Clause 50: The method of clause 49, wherein the data transmission channel comprises a physical downlink control channel (PDCCH).
[0149] Clause 51: The method of any of clauses 40-44, wherein waiting to process the second frame comprises: receiving the second frame before a second time; and sending the second frame at the second time.
[0150] Clause 52: The method of any of clauses 40-44 or 51, further comprising sending data representing an expected time between the first frame and the second frame to a downstream network device.
[0151] Clause 53: A method as recited in any of clauses 40-44, 51 or 52, wherein the data representative of the expected time between the first frame and the second frame comprises a frame rate value for at least a portion of the media stream of media data.
[0152] Clause 54: The method of clause 53, wherein receiving the frame rate value comprises receiving the frame rate value from a Session Description Protocol (SDP) message.
[0153] Clause 55: The method of clause 53, wherein receiving the frame rate value comprises receiving the frame rate value from a payload or header of a Real-time Transport Protocol (RTP) packet or an RTP Control Protocol (RTCP) packet, or from a header of a Secure Real-time Transport Protocol (SRTP) packet.
[0154] Clause 56: The method of clause 53 further comprising sending data representing one or more frame rates supported by the network device to the media application, wherein receiving the data representing the frame rate comprises receiving data representing a selection of a frame rate from the one or more frame rates supported by the network device from the media application.
[0155] Clause 57: The method of clause 56, further comprising configuring network transmission of the media session of the media data based on the selection of the frame rate.
[0156] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media (which corresponds to tangible media such as data storage media) or communication media (including any medium that facilitates (e.g., according to a communication protocol) the transfer of a computer program from one place to another). In this manner, computer-readable media may generally correspond to (1) non-transitory, tangible computer-readable storage media or (2) communication media (such as a signal or carrier wave). Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include computer-readable media.
[0157] By way of example and not limitation, these computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing the desired program code in the form of instructions or data structures and capable of being accessed by a computer. In addition, any connection is properly referred to as a computer-readable medium. For example, if instructions are sent from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) are all included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but rather refer to non-temporary tangible storage media. Disks and optical disks, as used herein, include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0158] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the term "processor" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Furthermore, these techniques may be implemented entirely in one or more circuits or logic elements.
[0159] The techniques of this disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but they do not necessarily require implementation by distinct hardware units. Instead, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units (including one or more processors described above) in conjunction with appropriate software and / or firmware.
[0160] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for exchanging media data via a network, the method comprising: receiving, by a network device from a media application, data representing an expected time between a first frame of media data and a second frame of the media data; Receiving, by the network device, a first frame of the media data at a first time; Waiting, by the network device, to process a second frame of the media data until a second time equal to or greater than the first time plus the expected time; as well as A second frame of the media data is processed by the network device at the second time.
2. The method according to claim 1, wherein Receiving data representing an expected time between the first frame and the second frame includes receiving a packet including data of the first frame of the media data, the packet also including data representing the expected time.
3. The method according to claim 2, wherein: The packet includes one of a Real-time Transport Protocol (RTP) packet, a Real-time Streaming Protocol (RTSP) packet, a Secure Real-time Transport Protocol (SRTP) packet, or an RTP Control Protocol (RTCP) packet.
4. The method according to claim 2, further comprising: encapsulating the packet with a packet tunnel header; as well as Data representing an expected time between the first frame and the second frame is added to the packet tunnel header.
5. The method according to claim 4, wherein The packet tunnel header includes a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
6. The method according to claim 1, wherein Receiving data representing an expected time between the first frame and the second frame includes receiving data representing an expected time between each of a plurality of frames in a sequence of frames starting with the first frame.
7. The method according to claim 1, wherein Receiving data representing an expected time between the first frame and the second frame includes receiving a packet encapsulated with a packet tunnel header, the packet tunnel header including data representing an expected time between the first frame and the second frame.
8. The method according to claim 7, wherein: The packet tunnel header includes a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
9. The method according to claim 1, wherein Receiving data representing an expected time between the first frame and the second frame includes receiving a packet including data representing an expected time between the first frame and the second frame in an Internet Protocol (IP) packet options field.
10. The method according to claim 1, in, Waiting to process the second frame includes skipping monitoring of the data transmission channel until the second time, and The processing of the second frame includes: monitoring the data transmission channel starting at the second time; and The second frame is received via the data transmission channel.
11. The method according to claim 10, wherein: The data transmission channel includes a physical downlink control channel (PDCCH).
12. The method according to claim 1, wherein Waiting for processing the second frame includes: receiving the second frame before the second time; and The second frame is sent at the second time.
13. The method of claim 1, further comprising sending data representing an expected time between the first frame and the second frame to a downstream network device.
14. The method according to claim 1, wherein The data representing the expected time between the first frame and the second frame comprises a frame rate value for at least a portion of a media stream of the media data.
15. The method according to claim 14, wherein Receiving the frame rate value includes receiving the frame rate value from a Session Description Protocol (SDP) message.
16. The method according to claim 14, wherein Receiving the frame rate value includes receiving the frame rate value from a payload or header of a Real-time Transport Protocol (RTP) packet or an RTP Control Protocol (RTCP) packet, or from a header of a Secure Real-time Transport Protocol (SRTP) packet.
17. The method of claim 14, further comprising sending data representing one or more frame rates supported by the network device to the media application, wherein Receiving data representing the frame rate includes receiving data representing a selection of a frame rate from one or more frame rates supported by the network device from the media application.
18. The method of claim 17, further comprising configuring network transmission of a media session of the media data based on the selection of the frame rate.
19. A device for exchanging media data via a network, the device comprising: a memory configured to store media data; as well as One or more processors embodied in circuitry and configured to: retrieving data from a media application representing an expected time between a first frame of media data and a second frame of the media data; receiving a first frame of the media data at a first time; Waiting to process a second frame of the media data until a second time equal to or greater than the first time plus the expected time; as well as A second frame of the media data is processed at the second time.
20. The apparatus according to claim 19, wherein To receive data representing an expected time between the first frame and the second frame, the one or more processors are configured to receive a packet comprising data of the first frame of the media data, the packet further comprising data representing the expected time.
21. The apparatus according to claim 20, wherein The packet includes one of a Real-time Transport Protocol (RTP) packet, a Real-time Streaming Protocol (RTSP) packet, a Secure Real-time Transport Protocol (SRTP) packet, or an RTP Control Protocol (RTCP) packet.
22. The apparatus according to claim 20, wherein The one or more processors are further configured to: encapsulating the packet with a packet tunnel header; and Data representing an expected time between the first frame and the second frame is added to the packet tunnel header.
23. The apparatus of claim 22, wherein: The packet tunnel header includes a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
24. The apparatus of claim 19, wherein To receive data representing an expected time between the first frame and the second frame, the one or more processors are configured to receive data representing an expected time between each of a plurality of frames in a sequence of frames starting with the first frame.
25. The apparatus of claim 19, wherein To receive data representing an expected time between the first frame and the second frame, the one or more processors are configured to receive a packet encapsulated with a packet tunnel header including data representing an expected time between the first frame and the second frame.
26. The apparatus of claim 25, wherein: The packet tunnel header includes a General Packet Radio Service (GPRS) Tunneling Protocol-User (GTP-U) packet header.
27. The apparatus of claim 19, wherein: To receive data representing an expected time between the first frame and the second frame, the one or more processors are configured to receive a packet including data representing an expected time between the first frame and the second frame in an Internet Protocol (IP) packet option field.
28. The apparatus according to claim 19, in, To await processing of the second frame, the one or more processors are configured to skip monitoring of the data transmission channel until the second time, and Wherein, in order to process the second frame, the one or more processors are configured to: monitoring the data transmission channel starting at the second time; and The second frame is received via the data transmission channel.
29. The apparatus of claim 28, wherein The data transmission channel includes a physical downlink control channel (PDCCH).
30. The apparatus of claim 19, wherein To wait for processing of the second frame, the one or more processors are configured to: receiving the second frame before the second time; and The second frame is sent at the second time.
31. The apparatus of claim 19, wherein The one or more processors are further configured to send data representing an expected time between the first frame and the second frame to a downstream network device.
32. The apparatus of claim 19, wherein: The data representing the expected time between the first frame and the second frame comprises a frame rate value for at least a portion of a media stream of the media data.
33. The apparatus of claim 32, wherein: To receive the frame rate value, the one or more processors are configured to receive the frame rate value from a Session Description Protocol (SDP) message.
34. The apparatus of claim 32, wherein: To receive the frame rate value, the one or more processors are configured to receive the frame rate value from a payload or header of a Real-time Transport Protocol (RTP) packet or an RTP Control Protocol (RTCP) packet or from a header of a Secure Real-time Transport Protocol (SRTP) packet.
35. The apparatus of claim 32, wherein: The one or more processors are further configured to send data representing one or more frame rates supported by the network device to the media application, wherein, to receive the data representing the frame rate, the one or more processors are configured to receive data representing a selection of a frame rate from the one or more frame rates supported by the network device from the media application.
36. The apparatus of claim 35, wherein The one or more processors are further configured to configure network transmission of the media session of the media data based on the selection of the frame rate.
37. The apparatus of claim 19, wherein: The apparatus comprises at least one of the following: integrated circuit; microprocessor; and Wireless communication equipment.
38. A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor of a network device to: receiving, from a media application, data representing an expected time between a first frame of media data and a second frame of the media data; receiving a first frame of the media data at a first time; Waiting to process a second frame of the media data until a second time equal to or greater than the first time plus the expected time; as well as A second frame of the media data is processed at the second time.
39. An apparatus for exchanging media data via a network, the apparatus comprising: means for receiving, from a media application, data representing an expected time between a first frame of media data and a second frame of said media data; means for receiving a first frame of said media data at a first time; means for waiting to process a second frame of said media data until a second time equal to or greater than said first time plus said expected time; as well as Means for processing a second frame of the media data at the second time.