Data redundancy request

By employing RTP and RTCP packets to manage redundancy requests for PI data, the solution addresses the lack of redundancy techniques in IVAS implementations, enhancing the reliability and quality of XR applications.

WO2025212616A1PCT designated stage Publication Date: 2025-10-09QUALCOMM INC

Patent Information

Application Number
PCT/US2025/022498
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2025-04-01
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing codecs, such as IVAS implementations, lack techniques for requesting redundancy for processing information (PI) data, which is crucial for rendering decoded audio data.

Method used

Implementing techniques for redundancy requests involving PI data using Real-Time Transport Protocol (RTP) packets, Real-Time Transport Control Protocol (RTCP) packets, and/or QUIC packets, allowing devices to send and receive redundancy requests for audio data, metadata, and PI data.

Benefits of technology

Enhances the reliability and robustness of audio data rendering by ensuring redundancy requests for PI data are handled, improving the overall quality of extended reality (XR) applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025022498_09102025_PF_FP_ABST
    Figure US2025022498_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Example devices, methods, and computer-readable media are described. An example device includes one or more processors. The one or more processors are configured to determine to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data. The one or more processors are configured to generate the redundancy request. The one or more processors are configured to send the redundancy request to a second computing device.
Need to check novelty before this filing date? Find Prior Art

Description

DATA REDUNDANCY REQUEST

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 573,172, filed 2 April 2024, U.S. Provisional Patent Application No. 63 / 680,548, filed 7 August 2024, and U.S. Provisional Patent Application No. 63 / 778,771, filed 27 March 2025, the entire content of each is incorporated herein by reference.TECHNICAL FIELD

[0002] This disclosure relates to transport of data, such audio data, metadata, and / or processing information.BACKGROUND

[0003] Applications, such as extended reality (XR) applications, may be accessed by a device over one or more networks from another device. Such applications may be used with / by encoder / decoders (codecs), such as Immersive Voice and Audio Services (IVAS) codecs and / or other standards or air interfaces, as examples.SUMMARY

[0004] Some codecs, such as IVAS implementations and / or other standards or air interfaces, as examples, may use 1) metadata: information for rendering and maybe sent along with the audio data to the decoder; 2) processing information (PI): other information for rendering decoded audio data.

[0005] While IVAS implementations and / or other standards or air interfaces, as examples, may include techniques for requesting redundancy for audio data and the associated metadata, IVAS implementations and / or other standards or air interfaces, as examples, may not have techniques for requesting redundancy for PI. In general, this disclosure describes techniques for redundancy requests involving PI, for example, for IVAS and / or other codec implementations.

[0006] Implementations of the techniques of this disclosure may include the use of Real- Time Transport Protocol (RTP) packets, Real-Time Transport Control Protocol (RTCP) packets (which may include RTCP-APP packets), and / or QUIC packets.

[0007] In one example, a device includes one or more processors configured to: determine to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generate the redundancy request; and send the redundancy request to a second computing device.

[0008] In one example, a method includes: determining, by a first computing device, to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generating, by the first computing device, the redundancy request; and sending, by the first computing device and to the second computing device, the redundancy request.

[0009] In another example, a device includes one or more processors configured to: receive a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generate, in response to the redundancy request, one or more packets comprising redundant data; and send, to a first computing device, the one or more packets.

[0010] In another example, a method includes: receiving, by a second computing device and from a first computing device, a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generating, by the second computing device and in response to the redundancy request, one or more packets comprising redundant data; and sending, by the second computing device and to the first computing device, the one or more packets.

[0011] In yet another example, computer readable media stores instructions, which, when executed, cause one or more processors to perform any of the techniques of this disclosure.

[0012] In yet another example, a computing device includes one or more means for performing any of the techniques of this disclosure.

[0013] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF DRAWINGS

[0014] FIG. 1A is a block diagram illustrating an example system that implements techniques for streaming media data over a network.

[0015] FIG. IB is a block diagram illustrating another example system that implements techniques for streaming media data over a network.

[0016] FIG. 2 is a conceptual diagram illustrating examples of a new voice packet, a new voice packet, and a legacy voice packet according to one or more aspects of this disclosure.

[0017] FIG. 3 is a conceptual diagram illustrating an example redundancy request according to one or more aspects of this disclosure.

[0018] FIG. 4 is a conceptual diagram illustrating an example message 1001 according to one or more aspects of this disclosure.

[0019] FIG. 5 is a conceptual diagram illustrating an example 1010 message according to one or more aspects of this disclosure.

[0020] FIG. 6 is a conceptual diagram illustrating an example 1011 message according to one or more aspects of this disclosure.

[0021] FIG. 7 is a conceptual diagram illustrating an example message 1001 according to one or more aspects of this disclosure.

[0022] FIG. 8 is a conceptual diagram illustrating a flow of parameter sets for encoded frames.

[0023] FIG. 9 is a conceptual diagram illustrating packetization for RTCP APP REQ RED AUDIO request with 200% redundancy without frame aggregation.

[0024] FIG. 10 is a conceptual diagram illustrating example packetization for RTCP APP REQ RED METADATA request with 300% redundancy without frame aggregation.

[0025] FIG. 11 is a conceptual diagram illustrating example packetization for RTCP APP REQ RED AUDIO METADATA request with 200% redundancy without frame aggregation.

[0026] FIG. 12 is a table illustrating example conditions for redundancy requests according to one or more aspects of this disclosure.

[0027] FIG. 13 is a table illustrating further example conditions for redundancy requests according to one or more aspects of this disclosure.

[0028] FIG. 14 is a table illustrating further example conditions for redundancy requests according to one or more aspects of this disclosure.

[0029] FIG. 15 is a table illustrating further example conditions for redundancy requests according to one or more aspects of this disclosure.

[0030] FIG. 16 is a conceptual diagram illustrating an example packet format at various levels of detail.

[0031] FIG. 17 is a conceptual diagram illustrating an example Real-time Transport Protocol (RTP) packet with PI data.

[0032] FIG. 18 is a conceptual diagram of an example RTCP APP REQ RED request.

[0033] FIG. 19 is a conceptual diagram illustrating an example subsequent RTP packet including redundant PI data.

[0034] FIG. 20 is a conceptual diagram of an example redundancy request for PI data according to one or more aspects of this disclosure.

[0035] FIG. 21 is a conceptual diagram illustrating an example QUIC packet according to one or more aspects of this disclosure.

[0036] FIG. 22 is a flow diagram illustrating example PI redundancy request techniques according to one or more aspects of this disclosure.

[0037] FIG. 23 is a flow diagram illustrating example response to PI redundancy request techniques according to one or more aspects of this disclosure.DETAILED DESCRIPTION

[0038] Some codecs, such as IVAS implementations, and / or other standards or air interfaces, as examples, may use 1) metadata: information for rendering and maybe sent along with the audio data to the decoder; 2) processing information (PI): other information for rendering decoded audio data. Metadata may include information used for rendering of the audio data, such as azimuth, elevation, radius, pitch, and yaw. Metadata may be packaged together with audio data, for example, in a same data frame, and sent to an audio decoder of a receiving device for decoding. For example, a Real-time Transport Protocol (RTP) sender may not distinguish between audio data and metadata. In some examples, such metadata may be used by the audio decoder when decoding the audio data and / or, in some examples, may be used by a Tenderer of the receiving device when rendering the decoded audio data. A frame of such data may be referred to herein as an audio frame and it should be understood that an audio frame may include such metadata.

[0039] While the discussion herein is primarily directed to IVAS implementations, it should be understood that the scope of this disclosure is not limited to IVAS implementation, and may include other standard implementations and / or air interfaces.

[0040] PI data may include additional information used by a Tenderer of a receiving device for rendering. Examples of PI data include scene orientation data, device orientation (compensated) data, device orientation (uncompensated) data, acoustic environment data, no PI data, etc. These different examples of PI data are set forth in 3GPP TS26.253 as different PI data types. The definition of PI data in 3GPP TS26.114 vl 8.6.0 is “Processing information data: Provides additional information that may be used for the IVAS media receiver for rendering the audio signal.” In some examples, PI data is not used by the audio decoder of the receiving device, but is used by the Tenderer of the receiving device. An audio decoder may decode audio data, while a Tenderer may render already decoded audio data. While PI data is described herein primarily with respect to IVAS, PI data may be defined for other coder / decoders (“codecs”) such as an Enhanced Voice Services (EVS) codec.

[0041] Currently, IVAS implementations may include techniques for requesting redundancy for speech (e.g., voice), and / or audio data (including the metadata to be sent to the audio decoder). However, IVAS implementations may not include techniques for requesting redundancy for PI data. In general, this disclosure describes techniques for redundancy requests involving PI data, for example, for IVAS implementations and / or other codecs.

[0042] FIG. 1A is a block diagram illustrating an example system 10 that implements techniques for streaming media data over a network. In this example, system 10 includes content preparation device 20, server device 60, and client device 40. Server device 60 may be an XR application server. Client device 40 and server device 60 are communicatively coupled by network 74, which may comprise a wireless wide area network, a wireless local area network, the Internet, and / or the like. In some examples, content preparation device 20 and server device 60 may also be coupled by network 74 or another network, or may be directly communicatively coupled. In some examples, content preparation device 20 and server device 60 may comprise the same device. It should be noted that an IVAS codec may not process video data and, as such, in IVAS only implementations, video components of FIG. 1 A may be deleted. Other codecs may process video data and include the video components of FIG. 1A.

[0043] Content preparation device 20, in the example of FIG. 1 A, may include audio source 22 and video source 24. Audio source 22 may comprise, for example, a microphone that produces electrical signals representative of captured audio data to be encoded by audio encoder 26. In some examples, audio encoder 26 represents an IVAS encoder. Alternatively, audio source 22 may comprise storage media storing previously recorded audio data, an audio data generator such as a computerized synthesizer, or any other source of audio data. Video source 24 may comprise a video camera that produces video data to be encoded by video encoder 28, storage media encoded with previously recorded video data, a video data generation unit such as a computer graphics source, or any other source of video data. Content preparation device 20 is not necessarily communicatively coupled to server device 60 in all examples, but may store multimedia content to separate media that is read by server device 60.

[0044] Raw audio and video data may comprise analog or digital data. Analog data may be digitized before being encoded by audio encoder 26 and / or video encoder 28. Audio source 22 may obtain audio data from a speaking participant while the speaking participant is speaking, and video source 24 may simultaneously obtain video data of the speaking participant. In other examples, audio source 22 may comprise computer- readable storage media comprising stored audio data, and video source 24 may comprise computer-readable storage media comprising stored video data. In this manner, the techniques described in this disclosure may be applied to live, streaming, real-time audio and video data or to archived, pre-recorded audio and video data.

[0045] Audio frames that correspond to video frames are generally audio frames containing audio data that was captured (or generated) by audio source 22 contemporaneously with video data captured (or generated) by video source 24 that is contained within the video frames. For example, while a speaking participant generally produces audio data by speaking, audio source 22 captures the audio data, and video source 24 captures video data of the speaking participant at the same time, that is, while audio source 22 is capturing the audio data. Hence, an audio frame may temporally correspond to one or more particular video frames. Accordingly, an audio frame corresponding to a video frame generally corresponds to a situation in which audio data and video data were captured at the same time and for which an audio frame and a video frame comprise, respectively, the audio data and the video data that was captured at the same time.

[0046] In some examples, audio encoder 26 may encode a timestamp in each encoded audio frame that represents a time at which the audio data for the encoded audio frame was recorded, and similarly, video encoder 28 may encode a timestamp in each encoded video frame that represents a time at which the video data for an encoded video frame was recorded. In such examples, an audio frame corresponding to a video frame may comprise an audio frame comprising a timestamp and a video frame comprising the same timestamp. Content preparation device 20 may include an internal clock from which audio encoder 26 and / or video encoder 28 may generate the timestamps, or that audio source 22 and video source 24 may use to associate audio and video data, respectively, with a timestamp.

[0047] In some examples, audio source 22 may send data to audio encoder 26 corresponding to a time at which audio data was recorded, and video source 24 may send data to video encoder 28 corresponding to a time at which video data was recorded. In some examples, audio encoder 26 may encode a sequence identifier in encoded audio data to indicate a relative temporal ordering of encoded audio data, but without necessarily indicating an absolute time at which the audio data was recorded, and similarly, video encoder 28 may also use sequence identifiers to indicate a relative temporal ordering of encoded video data. Similarly, in some examples, a sequence identifier may be mapped or otherwise correlated with a timestamp.

[0048] Audio encoder 26 generally produces a stream of encoded audio data, while video encoder 28 produces a stream of encoded video data. Each individual stream of data (whether audio or video) may be referred to as an elementary stream. An elementary stream is a single, digitally coded (possibly compressed) component of a representation. For example, the coded video or audio part of the representation can be an elementary stream. An elementary stream may be converted into a packetized elementary stream (PES) before being encapsulated within a video file. Within the same representation, a stream ID may be used to distinguish the PES-packets belonging to one elementary stream from the other. The basic unit of data of an elementary stream is a packetized elementary stream (PES) packet. Thus, coded video data generally corresponds to elementary video streams. Similarly, audio data corresponds to one or more respective elementary streams.

[0049] Many video coding standards, such as ITU-T H.264 / AVC and the High Efficiency Video Coding (HEVC) standard, define the syntax, semantics, and decoding process for error-free bitstreams, any of which conform to a certain profile or level. Video codingstandards typically do not specify the encoder, but the encoder is tasked with guaranteeing that the generated bitstreams are standard-compliant for a decoder. In the context of video coding standards, a “profile” corresponds to a subset of algorithms, features, or tools and constraints that apply to them. As defined by the H.264 standard, for example, a “profile” is a subset of the entire bitstream syntax that is specified by the H.264 standard. A “level” corresponds to the limitations of the decoder resource consumption, such as, for example, decoder memory and computation, which are related to the resolution of the pictures, bit rate, and block processing rate. A profile may be signaled with a profile idc (profile indicator) value, while a level may be signaled with a level idc (level indicator) value.

[0050] The H.264 standard, for example, recognizes that, within the bounds imposed by the syntax of a given profile, it is still possible to require a large variation in the performance of encoders and decoders depending upon the values taken by syntax elements in the bitstream such as the specified size of the decoded pictures. The H.264 standard further recognizes that, in many applications, it is neither practical nor economical to implement a decoder capable of dealing with all hypothetical uses of the syntax within a particular profile. Accordingly, the H.264 standard defines a “level” as a specified set of constraints imposed on values of the syntax elements in the bitstream. These constraints may be simple limits on values. Alternatively, these constraints may take the form of constraints on arithmetic combinations of values (e.g., picture width multiplied by picture height multiplied by number of pictures decoded per second). The H.264 standard further provides that individual implementations may support a different level for each supported profile.

[0051] A decoder conforming to a profile ordinarily supports all the features defined in the profile. For example, as a coding feature, B-picture coding is not supported in the baseline profile of H.264 / AVC but is supported in other profiles of H.264 / AVC. A decoder conforming to a level should be capable of decoding any bitstream that does not require resources beyond the limitations defined in the level. Definitions of profiles and levels may be helpful for interpretability. For example, during video transmission, a pair of profile and level definitions may be negotiated and agreed for a whole transmission session. More specifically, in H.264 / AVC, a level may define limitations on the number of macroblocks that need to be processed, decoded picture buffer (DPB) size, coded picture buffer (CPB) size, vertical motion vector range, maximum number of motion vectors per two consecutive macroblocks (MBs), and whether a B-block can have sub-macroblock partitions less than 8x8 pixels. In this manner, a decoder may determine whether the decoder is capable of properly decoding the bitstream.

[0052] In the example of FIG. 1 A, encapsulation unit 30 of content preparation device 20 receives elementary streams comprising coded video data from video encoder 28 and elementary streams comprising coded audio data from audio encoder 26. In some examples, video encoder 28 and audio encoder 26 may each include packetizers for forming PES packets from encoded data. In other examples, video encoder 28 and audio encoder 26 may each interface with respective packetizers for forming PES packets from encoded data. In still other examples, encapsulation unit 30 may include packetizers for forming PES packets from encoded audio and video data.

[0053] Video encoder 28 may encode video data of multimedia content in a variety of ways, to produce different representations of the multimedia content at various bitrates and with various characteristics, such as pixel resolutions, frame rates, conformance to various coding standards, conformance to various profiles and / or levels of profiles for various coding standards, representations having one or multiple views (e.g., for two- dimensional or three-dimensional playback), or other such characteristics. A representation, as used in this disclosure, may comprise one of audio data, video data, text data (e.g., for closed captions), or other such data. The representation may include an elementary stream, such as an audio elementary stream or a video elementary stream. Each PES packet may include a stream id that identifies the elementary stream to which the PES packet belongs. Encapsulation unit 30 is responsible for assembling elementary streams into video files (e.g., segments) of various representations.

[0054] Encapsulation unit 30 receives PES packets for elementary streams of a representation from audio encoder 26 and video encoder 28 and forms corresponding network abstraction layer (NAL) units from the PES packets. Coded video segments may be organized into NAL units, which provide a “network-friendly” video representation addressing applications such as video telephony, storage, broadcast, or streaming. NAL units can be categorized to Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL units may contain the core compression engine and may include block, macroblock, and / or slice level data. Other NAL units may be non-VCL NAL units. In some examples, a coded picture in one time instance, normally presented as a primary coded picture, may be contained in an access unit, which may include one or more NAL units.

[0055] Non-VCL NAL units may include parameter set NAL units and SEI NAL units, among others. Parameter sets may contain sequence-level header information (in sequence parameter sets (SPS)) and the infrequently changing picture-level header information (in picture parameter sets (PPS)). With parameter sets (e.g., PPS and SPS), infrequently changing information need not to be repeated for each sequence or picture; hence, coding efficiency may be improved. Furthermore, the use of parameter sets may enable out-of-band transmission of the important header information, avoiding the need for redundant transmissions for error resilience. In out-of-band transmission examples, parameter set NAL units may be transmitted on a different channel than other NAL units, such as SEI NAL units.

[0056] Supplemental Enhancement Information (SEI) may contain information that is not necessary for decoding the coded pictures samples from VCL NAL units, but may assist in processes related to decoding, display, error resilience, and other purposes. SEI messages may be contained in non-VCL NAL units. SEI messages are the normative part of some standard specifications, and thus are not always mandatory for standard compliant decoder implementation. SEI messages may be sequence level SEI messages or picture level SEI messages. Some sequence level information may be contained in SEI messages, such as scalability information SEI messages in the example of Scalable Video Coding (SVC) and view scalability information SEI messages in Multiview Video Coding (MVC). These example SEI messages may convey information on, e.g., extraction of operation points and characteristics of the operation points. In addition, encapsulation unit 30 may form a manifest file, such as a media presentation descriptor (MPD) that describes characteristics of the representations. Encapsulation unit 30 may format the MPD according to extensible markup language (XML).

[0057] Encapsulation unit 30 may provide data for one or more representations of multimedia content, along with the manifest file (e.g., the MPD) to output interface 32. Output interface 32 may comprise a network interface or an interface for writing to storage media, such as a universal serial bus (USB) interface, a CD or DVD writer or burner, an interface to magnetic or flash storage media, or other interfaces for storing or transmitting media data. Encapsulation unit 30 may provide data of each of the representations of multimedia content to output interface 32, which may send the data to server device 60 via network transmission or storage media. In the example of FIG. 1A, server device 60 includes storage media 62 that stores various multimedia content 64,each including a respective manifest file 66 and one or more representations 68A-68N (representations 68). In some examples, output interface 32 may also send data directly to network 74.

[0058] In some examples, representations 68 may be separated into adaptation sets. That is, various subsets of representations 68 may include respective common sets of characteristics, such as codec, profile and level, resolution, number of views, file format for segments, text type information that may identify a language or other characteristics of text to be displayed with the representation and / or audio data to be decoded and presented, e.g., by speakers, camera angle information that may describe a camera angle or real -world camera perspective of a scene for representations in the adaptation set, rating information that describes content suitability for particular audiences, or the like.

[0059] Manifest file 66 may include data indicative of the subsets of representations 68 corresponding to particular adaptation sets, as well as common characteristics for the adaptation sets. Manifest file 66 may also include data representative of individual characteristics, such as bitrates, for individual representations of adaptation sets. In this manner, an adaptation set may provide for simplified network bandwidth adaptation. Representations in an adaptation set may be indicated using child elements of an adaptation set element of manifest file 66.

[0060] Server device 60 includes request processing unit 70 and network interface 72. In some examples, server device 60 may include a plurality of network interfaces. Furthermore, any or all of the features of server device 60 may be implemented on other devices of a content delivery network, such as routers, bridges, proxy devices, switches, or other devices. In some examples, intermediate devices of a content delivery network may cache data of multimedia content 64, and include components that conform substantially to those of server device 60. In general, network interface 72 is configured to send and receive data via network 74.

[0061] Request processing unit 70 is configured to receive network requests from client devices, such as client device 40, for data of storage media 62. In some examples, request processing unit 70 may receive network requests from client device 40 in the form of RTP / SRTP packets and may deliver content, such as XR application content, to client device 40 in the form of RTP / SRTP packets.

[0062] Additionally, or alternatively, request processing unit 70 may implement hypertext transfer protocol (HTTP) version 1.1, as described in Request for Comments(RFC) 2616, “Hypertext Transfer Protocol - HTTP / 1.1,” by R. Fielding et al, Network Working Group, Internet Engineering Task Force (IETF), June 1999. That is, request processing unit 70 may be configured to receive HTTP GET or partial GET requests and provide data of multimedia content 64 in response to the requests. The requests may specify a segment of one of representations 68, e.g., using a Uniform Resource Locator (URL) of the segment. In some examples, the requests may also specify one or more byte ranges of the segment, thus comprising partial GET requests. Request processing unit 70 may further be configured to service HTTP HEAD requests to provide header data of a segment of one of representations 68. In any case, request processing unit 70 may be configured to process the requests to provide requested data to a requesting device, such as client device 40.

[0063] Additionally, or alternatively, request processing unit 70 may be configured to deliver media data via a broadcast or multicast protocol, such as eMBMS. Content preparation device 20 may create DASH segments and / or sub-segments in substantially the same way as described, but server device 60 may deliver these segments or subsegments using eMBMS or another broadcast or multicast network transport protocol. For example, request processing unit 70 may be configured to receive a multicast group join request from client device 40. That is, server device 60 may advertise an Internet protocol (IP) address associated with a multicast group to client devices, including client device 40, associated with particular media content (e.g., a broadcast of a live event). Client device 40, in turn, may submit a request to join the multicast group. This request may be propagated throughout network 74, e.g., routers making up network 74, such that the routers are caused to direct traffic destined for the IP address associated with the multicast group to subscribing client devices, such as client device 40.

[0064] As illustrated in the example of FIG. 1 A, multimedia content 64 includes manifest file 66, which may correspond to a media presentation description (MPD). Manifest file 66 may contain descriptions of different alternative representations 68 (e.g., video services with different qualities) and the description may include, e.g., codec information, a profile value, a level value, a bit rate, and other descriptive characteristics of representations 68. Client device 40 may retrieve the MPD of a media presentation to determine how to access segments of representations 68.

[0065] In particular, retrieval unit 52 may retrieve configuration data (not shown) of client device 40 to determine decoding capabilities of video decoder 48 and renderingcapabilities of video output 44. The configuration data may also include any or all of a language preference selected by a user of client device 40, one or more camera perspectives corresponding to depth preferences set by the user of client device 40, and / or a rating preference selected by the user of client device 40. Retrieval unit 52 may comprise, for example, a web browser or a media client configured to submit HTTP GET and partial GET requests. Retrieval unit 52 may correspond to software instructions executed by one or more processors or processing units (not shown) of client device 40. In some examples, all or portions of the functionality described with respect to retrieval unit 52 may be implemented in hardware, or a combination of hardware, software, and / or firmware, where requisite hardware may be provided to execute instructions for software or firmware.

[0066] Retrieval unit 52 may compare the decoding and rendering capabilities of client device 40 to characteristics of representations 68 indicated by information of manifest file 66. Retrieval unit 52 may initially retrieve at least a portion of manifest file 66 to determine characteristics of representations 68. For example, retrieval unit 52 may request a portion of manifest file 66 that describes characteristics of one or more adaptation sets. Retrieval unit 52 may select a subset of representations 68 (e.g., an adaptation set) having characteristics that can be satisfied by the coding and rendering capabilities of client device 40. Retrieval unit 52 may then determine bitrates for representations in the adaptation set, determine a currently available amount of network bandwidth, and retrieve segments from one of the representations having a bitrate that can be satisfied by the network bandwidth.

[0067] In general, higher bitrate representations may yield higher quality video playback, while lower bitrate representations may provide sufficient quality video playback when available network bandwidth decreases. Accordingly, when available network bandwidth is relatively high, retrieval unit 52 may retrieve data from relatively high bitrate representations, whereas when available network bandwidth is low, retrieval unit 52 may retrieve data from relatively low bitrate representations. In this manner, client device 40 may stream multimedia data over network 74 while also adapting to changing network bandwidth availability of network 74.

[0068] Additionally, or alternatively, retrieval unit 52 may be configured to receive data in accordance with a broadcast or multicast network protocol, such as eMBMS or IP multicast. In such examples, retrieval unit 52 may submit a request to join a multicastnetwork group associated with particular media content. After joining the multicast group, retrieval unit 52 may receive data of the multicast group without further requests issued to server device 60 or content preparation device 20. Retrieval unit 52 may submit a request to leave the multicast group when data of the multicast group is no longer needed, e.g., to stop playback or to change channels to a different multicast group.

[0069] Network interface 54 may receive and provide data of segments of a selected representation to retrieval unit 52, which may in turn provide the segments to decapsulation unit 50. Decapsulation unit 50 may decapsulate elements of a video file into constituent PES streams, depacketize the PES streams to retrieve encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending on whether the encoded data is part of an audio or video stream, e.g., as indicated by PES packet headers of the stream. Audio decoder 46 decodes encoded audio data and sends the decoded audio data to Tenderer 56 for rendering audio output 42, while video decoder 48 decodes encoded video data and sends the decoded video data, which may include a plurality of views of a stream, to video output 44. In some examples, audio decoder 46 represents an IVAS decoder. In some examples, audio decoder 46 represents a decoder that is not an IVAS decoder. For example, audio decoder 46 may decode frames of audio data. These frames of audio data may include metadata, such as azimuth, elevation, radius, pitch, and yaw information. Renderer 56 may receive PI data frames which may include PI data such as scene orientation data, device orientation (compensated) data, device orientation (uncompensated) data, acoustic environment data, or the like, which renderer 56 may use when rendering audio output 42.

[0070] Video encoder 28, video decoder 48, audio encoder 26, audio decoder 46, renderer 56, encapsulation unit 30, retrieval unit 52, and decapsulation unit 50 each may be implemented as any of a variety of suitable processing circuitry, as applicable, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. Each of video encoder 28 and video decoder 48 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined video encoder / decoder (CODEC). Likewise, each of audio encoder 26 and audio decoder 46 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined CODEC. An apparatus including video encoder 28, video decoder 48, audio encoder 26, audiodecoder 46, Tenderer 56, encapsulation unit 30, retrieval unit 52, and / or decapsulation unit 50 may comprise an integrated circuit, a microprocessor, and / or a wireless communication device, such as a cellular telephone.

[0071] Client device 40, server device 60, and / or content preparation device 20 may be configured to operate in accordance with the techniques of this disclosure. For purposes of example, this disclosure describes these techniques with respect to client device 40 and server device 60. However, it should be understood that content preparation device 20 may be configured to perform these techniques, instead of (or in addition to) server device 60.

[0072] Encapsulation unit 30 may form NAL units comprising a header that identifies a program to which the NAL unit belongs, as well as a payload, e.g., audio data, video data, or data that describes the transport or program stream to which the NAL unit corresponds. For example, in H.264 / AVC, a NAL unit includes a 1-byte header and a payload of varying size. A NAL unit including video data in its payload may comprise various granularity levels of video data. For example, a NAL unit may comprise a block of video data, a plurality of blocks, a slice of video data, or an entire picture of video data. Encapsulation unit 30 may receive encoded video data from video encoder 28 in the form of PES packets of elementary streams. Encapsulation unit 30 may associate each elementary stream with a corresponding program.

[0073] Encapsulation unit 30 may also assemble access units from a plurality of NAL units. In general, an access unit may comprise one or more NAL units for representing a frame of video data, as well as audio data corresponding to the frame when such audio data is available. An access unit generally includes all NAL units for one output time instance, e.g., all audio and video data for one time instance. For example, if each view has a frame rate of 20 frames per second (fps), then each time instance may correspond to a time interval of 0.05 seconds. During this time interval, the specific frames for all views of the same access unit (the same time instance) may be rendered simultaneously. In one example, an access unit may comprise a coded picture in one time instance, which may be presented as a primary coded picture.

[0074] Accordingly, an access unit may comprise all audio and video frames of a common temporal instance, e.g., all views corresponding to time X. This disclosure also refers to an encoded picture of a particular view as a “view component.” That is, a view component may comprise an encoded picture (or frame) for a particular view at aparticular time. Accordingly, an access unit may be defined as comprising all view components of a common temporal instance. The decoding order of access units need not necessarily be the same as the output or display order.

[0075] A media presentation may include a media presentation description (MPD), which may contain descriptions of different alternative representations (e.g., video services with different qualities) and the description may include, e.g., codec information, a profile value, and a level value. An MPD is one example of a manifest file, such as manifest file 66. Client device 40 may retrieve the MPD of a media presentation to determine how to access movie fragments of various presentations. Movie fragments may be located in movie fragment boxes (moof boxes) of video files.

[0076] Manifest file 66 (which may comprise, for example, an MPD) may advertise availability of segments of representations 68. That is, the MPD may include information indicating the wall-clock time at which a first segment of one of representations 68 becomes available, as well as information indicating the durations of segments within representations 68. In this manner, retrieval unit 52 of client device 40 may determine when each segment is available, based on the starting time as well as the durations of the segments preceding a particular segment.

[0077] After encapsulation unit 30 has assembled NAL units and / or access units into a video file based on received data, encapsulation unit 30 passes the video file to output interface 32 for output. In some examples, encapsulation unit 30 may store the video file locally or send the video file to a remote server via output interface 32, rather than sending the video file directly to client device 40. Output interface 32 may comprise, for example, a transmitter, a transceiver, a device for writing data to computer-readable storage media such as, for example, an optical drive, a magnetic media drive (e.g., floppy drive), a universal serial bus (USB) port, a network interface, and / or other output interface. Output interface 32 outputs the video file to computer-readable media, such as, for example, a transmission signal, a magnetic medium, an optical medium, a memory, a flash drive, and / or other computer-readable media.

[0078] Network interface 54 may receive a NAL unit or access unit via network 74 and provide the NAL unit or access unit to decapsulation unit 50, via retrieval unit 52. Decapsulation unit 50 may decapsulate a elements of a video file into constituent PES streams, depacketize the PES streams to retrieve encoded data, and send the encoded data to either audio decoder 46 or video decoder 48, depending on whether the encoded datais part of an audio or video stream, e.g., as indicated by PES packet headers of the stream. Audio decoder 46 decodes encoded audio data and sends the decoded audio data to audio output 42, while video decoder 48 decodes encoded video data and sends the decoded video data, which may include a plurality of views of a stream, to video output 44.

[0079] The example of FIG. 1A describes the use of RTP, DASH, and HTTP -based streaming for purposes of example. However, it should be understood that other types of protocols may be used to transport media data. For example, request processing unit 70 and retrieval unit 52 may be configured to operate according to Real-time Streaming Protocol (RTSP), or the like, and use supporting protocols such as Session Description Protocol (SDP) or Session Initiation Protocol (SIP).

[0080] FIG. IB is a block diagram illustrating another example system that implements techniques for streaming media data over a network. FIG. IB is similar to the example of FIG. 1 A, but FIG. IB includes two end devices, rather than a client device and a server device. For example, each end device 80 A and 80B of system 10B may be configured to both consume content from and provide content to the other of end device 80 A and 80B. The system of FIG. IB may implement the techniques disclosed herein.

[0081] FIG. 2 is a conceptual diagram illustrating examples of a new voice packet , a new voice packet, and a legacy voice packet according to one or more aspects of this disclosure.

[0082] IVAS voice codec support is set forth in TS26.114. TS 26.114 VI 8.6.0 (Feb 2024) - S4-240128 - specifies that for IVAS , meta data (data that helps render the speech frames, e.g., azimuth, elevation, radius, pitch, and yaw, energy ratio, diffuseness, spread coherence, surround coherence as described in TS 26.253) can be (1) multiplexed with audio data (e.g., with one or more audio frames) in an RTP packet, (2) sent alone, or (3) not sent. As such, IVAS voice codec support of TS26.114 results in 3 types of RTP packets. FIG. 2 is a conceptual diagram illustrating examples of a new voice packet, another new voice packet, and a legacy voice packet according to one or more aspects of this disclosure. For example, RTP packet 200 includes RTP header 202, meta data 204, and audio data 206. RTP packet 210 includes RTP header 212 and meta data 214, but does not include audio data. Legacy RTP packet 220 includes RTP header 222 and audio data 224.

[0083] The meta data is directly associated with the audio data if the meta data and audio data are encapsulated in the same RTP packet (e.g., as in RTP packet 200), or is indirectlyassociated with audio data through the RTP timestamps (of the same value) otherwise. For example, meta data may be said to be directly associated with the audio data if the meta data is encapsulated in the same RTP packet as the audio data, or is indirectly associated with audio data through the use of RTP timestamps having a same value.

[0084] To better fit the characteristics of the current transport network, the receiver (e.g., a receiver of an RTP packet, such as a speech and / or audio decoder) may send an RTCP- APP packet to request to adapt the speech encoder. In particular, the receiver may send a redundancy request via the RTCP APP REQ RED message in the payload of the RTCP- APP packet.

[0085] A redundancy request is now discussed. FIG. 3 is a conceptual diagram illustrating an example redundancy request according to one or more aspects of this disclosure. A redundancy request is a type of message carried in an RTCP-APP message. In some examples, a payload chunk is defined as a speech frame or aggregated multiple speech frames (to be carried in a single RTP packet). This definition only allows for repetition of speech frames.

[0086] However, a redundancy request for meta data and / or audio data may be desirable. Based on the transport network characteristics and the relative importance of meta data compared to audio data, the receiver may make one of the following redundancy requests to the audio sender to repeat or perform application layer forward error correction (FEC) for audio data, audio data and the associated meta data, or meta data. As used herein, the receiver may be a computing device, such as content preparation device 20, client device 40, server device 60, end device 80A, end device 80B, and / or the like.

[0087] Audio data may be a payload chunk including one speech frame to be sent in a single RTP packet or aggregated multiple speech frames to be sent in a single RTP packet. Meta data may include data that helps the receiver process the audio data.

[0088] The redundancy request may be implemented by a RTCP APP REQ RED request in a RTCP-APP packet. In one example, the type of request may be indicated by an identifier. For example, the identifier may be a 2 -bit flag. For example, 00 may indicate audio data, 01 may indicate metadata, and 10 may indicate audio data and the associated meta data. In another example, a different message may exist for each type of payload chunk, with each message having a different message ID. For example, 1001 may indicate audio data, 1010 may indicate audio data and the associated meta data, and 1011 may indicate meta data.

[0089] FIG. 4 is a conceptual diagram illustrating an example message 1001 according to one or more aspects of this disclosure. RTP packet 400 may include a message ID of 1001 and bit field 402. RTP packet 400 may carry audio data in bit field 402. In the example of FIG. 4, bit field 402 may be a 12-bit bitmask that signals a request on how the audio data portions of non-redundant payloads chunks are to be repeated in subsequent packets. The position of the bit set in bit field 402 may indicate which earlier non- redundant payload chunks are requested to be added as redundant payload chunks to the current packet.

[0090] For example, if the least significant bit (e.g., rightmost bit) is set equal to 1, that indicates that the audio data portion of the last previous payload chunk is requested to be repeated as a redundant payload in the current packet. If the most significant bit (e.g., leftmost bit) is set equal to 1, that indicates that the audio data portion of the payload chunk that was transmitted 12 packets ago is requested to be repeated as a redundant payload chunk in the current packet. Note that it is not guaranteed that the sender has access to such old payload chunks. In some examples, the maximum amount of redundancy is 300 %, e.g., at maximum three bits can be set in the bit field.

[0091] FIG. 5 is a conceptual diagram illustrating an example 1010 message according to one or more aspects of this disclosure. RTP packet 500 may include a message ID of 1010 and bit field 502. RTP packet 500 may carry meta data in bit field 502. In the example of FIG. 5, bit field 502 may be a 12-bit bitmask that signals a request on how the metadata portions of non-redundant payloads chunks are to be repeated in subsequent packets. The position of the bit set in bit field 502 may indicate which earlier non- redundant payload chunks are requested to be added as redundant payload chunks to the current packet.

[0092] For example, if the least significant bit (e.g., rightmost bit) is set equal to 1, that indicates that the metadata portion of the last previous payload chunk is requested to be repeated as a redundant payload in the current packet. If the most significant bit (e.g., leftmost bit) is set equal to 1, that indicates that the metadata portion of the payload chunk that was transmitted 12 packets ago is requested to be repeated as redundant payload chunk in the current packet. Note that it is not guaranteed that the sender has access to such old payload chunks. In some examples, the maximum amount of redundancy is 500 %, e.g., at maximum five bits can be set in the bit field.

[0093] Since not all packets carry metadata, it is possible that the receiver’s request to repeat the metadata for a particular non-redundant payload chunk cannot be fulfilled by the sender and this can cause ambiguity for the receiver in determining from which packets the redundant metadata chunks have been repeated. Therefore, when replying to this redundancy request for metadata, the sender may (in some examples, shall) indicate packets from which the non-redundant metadata chunks are repeated by setting the corresponding bit fields in the bit mask.

[0094] FIG. 6 is a conceptual diagram illustrating an example 1011 message according to one or more aspects of this disclosure. RTP packet 600 may include a message ID of 1011 and bit field 602. RTP packet 600 may carry both audio data and meta data in bit field 602. In the example of FIG. 6, bit field 602 may include a 12-bit bitmask that signals a request on how the combinations of audio data and metadata of non-redundant payloads chunks are to be repeated in subsequent packets. For example, if the packet that has been requested to be repeated does not have both audio data and metadata, the sender may repeat whichever of the type of data (audio or metadata) is available. The position of the bit set may indicate which earlier non-redundant payload chunks is requested to be added as redundant payload chunks to the current packet.

[0095] If the least significant bit (e.g., rightmost bit) is set equal to 1, that indicates that the combinations of audio data and metadata of the last previous payload chunk is requested to be repeated as redundant payload in the current packet. If the most significant bit (e.g., leftmost bit) is set equal to 1, that indicates that the combinations of audio data and metadata of the payload chunk that was transmitted 12 packets ago is requested to be repeated as a redundant payload chunk in the current packet. Note that it is not guaranteed that the sender has access to such old payload chunks. In some examples, the maximum amount of redundancy is 300 %, e.g., at maximum three bits can be set in the bit field.

[0096] In some examples, the requests of FIGS. 4-6 may (in some examples, shall) be used for IVAS for codecs.

[0097] In another example, a message may include a bit field to indicate the requested type of payload chunks. For example, the message ID may be 1001. FIG. 7 is a conceptual diagram illustrating an example message 1001 according to one or more aspects of this disclosure. RTP packet 700 may include message ID 1001 and bit field 702. Bit field 702 may include of pairs of bits, with each pair indicating one of the 3 types of payload chunks. For example, 00 may indicate audio data, 01 may indicate metadata,10 may indicate audio data and the associated meta data, and 11 may be a reserved value. In this example, the position of the bit pair within bit field 702 indicates which earlier non-redundant payload chunk is requested to be added as redundant chunks to the current packet.

[0098] If the two least significant bits (e.g., two rightmost bits) are set equal to 00 (or 01, or 10), that indicates audio data (or metadata, or audio data and the associated meta data) of the last non-redundant payload chunk is to be repeated in the current packet. If the two most significant bits (e.g., two leftmost bits) are set equal to 00 (or 01, or 10), that indicates audio data (or metadata, or audio data and the associated meta data) of the payload chunk that was transmitted 10 packets ago is to be repeated in the current packet.

[0099] In some examples, the sender indicates which payload chunk is included in an RTP packet for a type of the payload chunk. This is because the ‘audio+metadata’ type and the ‘metadata’ type may not be sent in each (e.g., every) frame period. In one example, the indication may be in the form of a bit map in the payload (e.g., bit field 702). In one example, the indication may be in the form of an RTP header extension, and may, in some examples, also include the lengths of each of the included data units (audio frame, metadata, or a combination of audio frame and metadata).

[0100] The sender and receiver may negotiate, e.g., via SDP, the use of the above finer granularity redundancy request, and the negotiation may include: the support for a message, whether repetition or application-layer FEC is used, and if FEC is used, which FEC scheme (e.g., flex FEC or RaptorQ) is used, and what is the configuration for the FEC scheme (e.g., redundancy ratio, coding rate).

[0101] For example, more details on the support for a message may include, where the portions between **< and >** indicate new material relative to current standards: RTCP-APP request messages that can be used are negotiated with SDP using the ‘3gpp_mtsi_app_adapt’ attribute. The syntax for the 3GPP MTSI RTCP-APP adaptation attribute is: a=3 gpp mtsi app adapt : <reqN ames> where:<reqNames> is a comma-separated list identifying the different request messages (see below).The ABNF for the RTCP-APP adaptation messages negotiation attribute is the following:adaptation attribute = "a" "=" "3gpp_mtsi_app_adapt" reqNamereqName) reqName = "RedReq" / "FrameAggReq" / "AmrCmr" / "EvsRateReq" / "EvsBandwidthReq" / "EvsParRedReq" / "EvsIoModeReq" / "EvsPrimaryModeReq'7 **< "IvasRedReqAudio" " / "IvasRedReqMetadata" / "IvasRedReqAudioMetadata« >**The name denotes the RTCP APP packet types the SDP sender supports to receive. The meaning of the values is as follows:RedReq: Redundancy Request, clause 10.2.1.3FrameAggReq: Frame Aggregation Request, clause 10.2.1.4AmrCmr: Codec Mode Request for AMR and AMR-WB, clause 10.2.1.5EvsRateReq: EVS Primary Rate Request, clause 10.2.1.7EvsBandwidthReq: EVS Bandwidth Request, clause 10.2.1.8EvsParRedReq: EVS Partial Redundancy Request, clause 10.2.1.9 EvsIoModeReq: EVS Primary mode to EVS AMR-WB IO mode SwitchingRequest, clause 10.2.1.10EvsPrimaryModeReq: EVS AMR-WB IO mode to EVS Primary mode Switching Request, clause 10.2.1.11IvasRedReqAudio: Redundancy Request for audio data for IVAS IvasRedReqMetadata: Redundancy Request for metadata for IVAS IvasRedReqAudioMetadata: Redundancy Request for a combination of audio data and metadata for IVASAn MT SI client supporting only AMR, AMR-WB and, EVS and IVAS may for instance include the following in the SDP offer: a=3 gpp mtsi app adapt :RedReq, FrameAggReq, AmrCmr, EvsRateReq, EvsBandwidthReq, EvsParRedReq, Evslo ModeReq, EvsPrimaryModeReq, **<IvasRedReqAudio, IvasRedReqMetadata, IvasRedR eqAudioMetadata>* *

[0102] FIG. 8 is a conceptual diagram illustrating a flow of parameter sets for encoded frames. The flow of FIG. 8 may include audio frames 800, metadata 802, and metadata 804. For example, speech codec may generate frames numbered N-12. . .N in a continuous flow. In such a case, the sender may also generate five metadata units numbered N-8, N- 6, N-4, N-2 and N, among which metadata unit N-6 and speech frame N-6 were sent in asingle packet and metadata unit N-8 and speech frame N-8 were sent in two separate packets and so on. Each increment in FIG. 8 may correspond to a time difference of 20 ms and metadata units.

[0103] In one example, an RTCP APP REQ RED AUDIO request with bit field 000000000011 (200% redundancy) and an RTCP APP REQ AGG request with value = 0 (no frame aggregation 2) yield packets as shown in FIG. 9. FIG. 9 is a conceptual diagram illustrating packetization for RTCP APP REQ RED AUDIO request with 200% redundancy without frame aggregation. P-1 . . P may denote the sequence numbers of the packets.

[0104] FIG. 10 is a conceptual diagram illustrating example packetization for RTCP APP REQ RED METADATA request with 300% redundancy without frame aggregation. An RTCP APP REQ RED METADATA request with bit field 000000000111 (300% redundancy) and an RTCP APP REQ AGG request with value = 0 (no frame aggregation 2) may yield packets as shown in FIG. 10. Packet P may carry a bit field 000000000010 that indicates the time relation of the metadata unit (numbered N- 2 but unknown to the receiver) with respect to the speech frame N, for example, the metadata unit is two frames earlier than speech frame N. The receiver may have requested three metadata units numbered N-l, N-2, and N-3, but because only N-2 had metadata, the receiver received only one metadata unit. Since this metadata chunk is without a sequence number or timestamp (which were carried in the earlier non-redundant RTP packet but stripped off for the current packet P), the receiver does not know which packet sequence number or timestamp the received metadata unit corresponds to. The inserted bit field may be used to resolve the ambiguity.

[0105] An RTCP APP REQ RED AUDIO METADATA request with bit field 000000000011 (200% redundancy) and an RTCP APP REQ AGG request with value = 0 (no frame aggregation 2) may yield packets as shown in FIG. 11. FIG. 11 is a conceptual diagram illustrating example packetization forRTCP APP REQ RED AUDIO METADATA request with 200% redundancy without frame aggregation.

[0106] The following description is a first example of how the redundancy REQ of this disclosure might be used. This example may be based on a packet error rate estimated at the receiver, which may be referred to as PER. PERmetadata may be a maximum PER that the metadata can tolerate before causing unacceptable quality of experience (QoE).PERaudio may be a maximum PER that the audio can tolerate before causing unacceptable QoE. PERRx may be a PER measured at the receiver side of the link.

[0107] FIG. 12 is a table illustrating example conditions for redundancy requests according to one or more aspects of this disclosure. For some values of PERRX, the REQ redundancy for metadata may be different than for the audio data, in which case the receiver can toggle between using 10, 01, or even 00 (audio data only) to achieve different levels of redundancy. With respect to RTCP APP REQ RED including 10, an amount of redundancy may be indicated which may include 100%, 200%, 300%, etc. Parity symbols of forward error correction (FEC) requested increases as the PERRXincreases to achieve effective PER below PERMETADATA. With respect to RTCP APP REQ RED including 01, an amount of redundancy / parity symbols of FEC requested for meta data increases as the PERRXincreases to achieve effective PER below PERMETADATA. Additionally, an amount of redundancy / parity symbols of FEC requested for audio data increases as the PERRXincreases to achieve effective PER below PERAUDIO.

[0108] The following description is a second example of how the redundancy REQ of this disclosure might be used. This example may be based on a receiver’s prediction of PER or link connectivity. For example, a receiver may predict that the link will be disrupted due to a handoff or beam steering. In such a case, the receiver may request more redundancy for the metadata than audio, and may request that the repetition be spread out in time (e.g., time diversity) over the expected handoff period. For example, the receiver may predict that the PER will change using link perception or prediction techniques. The receiver may predict the future PERRXand then use a similar algorithm.

[0109] For example, the receiver may predict a link disruption or changes in PER based on (1) observed RTP packet loss rates in the past, (2) signaling of network conditions (e.g., an ECN (explicit congestion control) indication or L4S ECN (RFC 9331) indication on network congestion from the routers, indication of upcoming network congestion or poor channel conditions from a cellular base station), (3) observed radio channel conditions, (4) user mobility (high speed means higher Doppler which affects channel conditions) and trajectory in a radio propagation map, and / or the like. In some examples, the receive may predict the link disruption based on a predicted packet loss pattern. For example, when packet losses are bursty (e.g., unevenly spaced), the receiver may request repetitions to be spread out, which may be more robust, but may add more delay. When packet losses are more evenly spaced, the receiver may request repetitions to be moreconcentrated, which may reduce delay compared to examples where repetitions are more spread out.

[0110] FIG. 13 is a table illustrating further example conditions for redundancy requests according to one or more aspects of this disclosure. For some values of PERRX, the REQ redundancy for metadata may be different than for the audio data, in which case the receiver can toggle between using 10, 01, or even 00 (audio data only) to achieve different levels of redundancy. With respect to RTCP APP REQ RED including 10, an amount of redundancy may be indicated which may include 100%, 200%, 300%, etc. Parity symbols of forward error correction (FEC) requested increases as the PERRXincreases to achieve effective PER below PERMETADATA. With respect toRTCP APP REQ RED including 01, an amount of redundancy / parity symbols of FEC requested for meta data increases as the PERRXincreases to achieve effective PER below PERMETADATA. Additionally, an amount of redundancy / parity symbols of FEC requested for audio data increases as the PERRXincreases to achieve effective PER belowPERAUDIO.[OHl] The following description is a third example of how the redundancy REQ of this disclosure might be used. This example may be based on the how quickly the user’s pose changes.

[0112] When the user pose changes slowly, error concealment using past pose may work well at the receiver. However, when the user pose changes quickly, error concealment may not work well.

[0113] It may be difficult to compare the changes in the audio frames and the changes in the pose. However, the receive may look at the changes in the pose relative to some thresholds.

[0114] For example, if the pose changes quickly, it may be more important to receive the metadata accurately and the receiver may request more redundancy of the metadata. If the pose changes slowly, it may be less important to receive metadata accurately, and the receiver may request less redundancy of the metadata.

[0115] A receiver may determine the rate of change of a user's pose. For example, the receiver may determine the rate of change of a user’s pose based on (1) measurements from sensors, such as motion sensors, inertial sensors, accelerometers on a headmounted display or AR glasses, (2) the rate of change of a user's view (e.g., when the user uses AR glasses to look at the surroundings), and / or the like.

[0116] FIG. 14 is a table illustrating further example conditions for redundancy requests according to one or more aspects of this disclosure. For example, T may be an audio frame period. AR may be the change of roll in T. AP may be the change of pitch in T. AY may be the change of yaw in T. C (change) may be equal to | AR|+| AP|+| AY|. A may be a first threshold, e.g., 2 degrees per T. B (B>A) may be a second threshold, e.g., 4 degrees per T. It should be noted that the required redundancy for audio data may be requested separately (e.g., independently).

[0117] The following description is a fourth example of how the redundancy REQ might be used. This example may be based on the network delay (e.g., round trip time (RTT)).

[0118] When the network delay is small (e.g., if the RTT is small relative to the audio frame period), it may be acceptable for the receiver to request a retransmission of an RTP packet if the RTP packet gets lost in the communication network, and in this case the redundancy REQ may not be needed for meta data, audio data, or both.

[0119] In some examples, a redundancy REQ configuration is based on the network delay relative to the delay that can be tolerated by the user without significant degradation in the user experience (e.g., QoE).

[0120] FIG. 15 is a table illustrating further example conditions for redundancy requests according to one or more aspects of this disclosure. For example, T may be an audio frame period, e.g., 20ms. The parameter a may be a first threshold, e.g., 0.4. b (b>a) may be a second threshold, e.g., 0.8.

[0121] FIG. 16 is a conceptual diagram illustrating an example packet format at various levels of detail. For example, packet 1600 may be an RTP packet whose format is defined in TS26.253. Packet 1600 may include RTP header 1602 which may, in some examples, include a header extension (HDREXT), payload header 1604, frame data 1606, and PI data 1608. Payload header 1604, frame data 1606, and PI data 1608 may together make up payload 1610 of packet 1600. In some examples, payload 1610 may be an IVAS payload. In some examples, rather than making up an IVAS payload, payload header 1604, frame data 1606, and PI data 1608 may together make up a different type of payload, such as that of an EVS frame, another codec frame, or a NO D ATA frame.

[0122] Payload header 1604 may include a table of contents (ToC) byte and / or extra (E) byte. The ToC byte may define or identify the content in frame data 1606. In some examples, there is a ToC byte for each IVAS, EVS frame, another codec frame, and / orNO DATA frame. The E byte may contain extra information and may precede the ToC byte of the coded frame.

[0123] Frame data 1606 may include audio data and metadata. For example, frame data 1606 may include one or more frames and those frames may include audio data and / or metadata intended to be available to audio decoder 46 (e.g., not PI data).

[0124] PI data 1608 may include PI data, including a PI header data 1612 and PI frame data 1614. PI header data 1612 may include one or more PI data headers, such as PI data header 1620, each of the PI data headers identifying the type and size for a corresponding PI data frame in PI frame data 1614, and identifying with which audio frame(s) the corresponding PI data frame is associated. PI frame data 1614 may include one or more PI data frames, each of the PI data frames being of a specific PI data type. For example, a sending computing device, such as server device 60 may generate packet 1600 and include different types of PI data into different PI data frames. The PI data frames may include PI data intended to be available to Tenderer 56.

[0125] For example, PI data header 1620 may include a PF which may be a bit that is indicative of whether another PI data header follows this particular PI data header in the packet. PI data header 1620 may include a PM which may be two bits indicative of which audio frame is associated with the PI date frame for which PI data header 1620 corresponds. PI data header 1620 may include a PI type which may be five bits indicative of the type of PI data included in a corresponding PI data frame. PI data header 1620 may include a PI size that indicates a size of the corresponding PI data frame in bytes.

[0126] FIG. 17 is a conceptual diagram illustrating an example RTP packet with PI data. RTP packet 1700 may include RTP header 1702 which may be an example of RTP header 1602. Payload 1730 may include payload header 1704 which may be an example of payload header 1604. Payload 1610 may also include 1stframe 1706 and 2ndframe 1708. 1stframe 1706 and 2ndframe 1708 may be audio and metadata frames and together be referred to as frame data 1760. Frame data 1760 may be an example of frame data 1606.

[0127] Payload 1730 also includes PI data header for all frames (PIDHAF) 1710, 1stPI data header for frame 1 (1stPIDHF1) 1712, 2ndPI data header for frame 1 (2ndPIDHF1) 1714, and PI data header for frame 2 (PIDHF2) 1716, which together may be referred to as PI header section 1750. PI header section 1750 may be an example of PI header data 212. PI data header for all frames 1710 may include an instance of PI data header 1620 with the PM bits set to 11. 1stPI data header for frame 1 1712 may be a PI data headerfor PI data frame 1 1720. 2ndPI data header for frame 1 may be a PI data header for PI data frame 2 1722. PI data header for frame 2 1716 may be a PI data header for PI data frame 3. Thus, in the example of FIG. 17 there are two PI data frames (PI data frame 1 1720 and PI data frame 2 1722) that are associated with or provide PI data for use with 1stframe 1706 (of audio and metadata). Additionally, there is one PI data frame (PI data frame 3 1724) that is associated with or provides PI data for use with 2ndframe 1708 (of audio and metadata). Associations between the different PI data frames, PI data headers, and the frames of audio and metadata are depicted using different dashed or dotted lines.

[0128] Payload 1730 also includes PI frame data for all frames (PIDFAF) 1718 which may include PI frame data that is associated with all the audio frames of the packet, PI data frame 1 (PIDF1) 1720, PI data frame 2 (PIDF2) 1722, and PI data frame 3 (PIDF3) 1724, which together may be referred to as PI frame data 1770, which may be an example of PI frame data 1614.

[0129] In some examples, the contents of RTP packet 1700 and / or the order of the contents of RTP packet 1700 may be different than depicted.

[0130] In the example of FIG. 17, 1stframe 1706 of frame data 1760 includes audio data and metadata and has two associated PI data frames, namely PI data frame 1 1720 and PI data frame 2 1722. For example, the PI data in PI data frame 1 1720 and PI data frame 2 1722 may include different types of PI data (e.g., scene orientation data and acoustic environment data) that are both applicable to the audio data in 1stframe 1706. Renderer 56 of client device 40 may use the PI data of PI data frame 1 1720 and PI data frame 2 1722 when rendering the audio data decoded from 1stframe 1706. 2ndframe 1708 of frame data 1760 has one associated PI data frame, namely PI data frame 3 1724. For example, the PI data in PI data frame 3 1724 may include a type of PI data (e.g., scene orientation data) that is applicable to the audio data in 2ndframe 1708. Renderer 56 of client device 40 may use the PI data of PI data frame 3 1724 when rendering the audio data decoded from 2ndframe 1708.

[0131] A receiving computing device, such as client device 40, may use redundancy requests to request redundancy in PI data sent by a sending computing device, such as server device 60. In some examples, such redundancy requests may be sent for general PI data redundancy purposes, as opposed to requesting a retransmission of a dropped or corrupted packet or portion of a packet. As such, a redundancy request may be sent on an infrequent basis. For the following discussion, the receiving computing device will bereferred to as client device 40 and the sending computing device will be referred to as server device 60. It should be understood that this usage is for explanatory purposes and that any device capable of performing the techniques set forth herein may perform such techniques.

[0132] As it may be desirable for client device 40 to be able to request redundancy for PI data, client device 40 may generate and send a redundancy request for redundancy for PI data to server device 60. Server device 60 may, in response to receiving the redundancy request, send redundant PI data to client device 40 in one or more packets, subsequent to a packet containing non-redundant (e.g., original or first sent) data.

[0133] In some examples, client device 40 may use a joint redundancy request for PI data and frame data to request redundancy from server device 60. For example, client device 40 may generate a joint redundancy request and send the joint redundancy request to server device 60. The joint redundancy request may be a single redundancy request that indicates a redundancy request for both non-redundant PI data and non-redundant frame data. For example, client device 40 may send a joint redundancy request that is a request for both non-redundant PI data and non-redundant frame data to server device 60. In some examples, non-redundant PI data is PI data before any repetition. In some examples, non- redundant frame data is frame data before any repetition. In other words, non-redundant PI data is original PI data that is sent, but not sent in response to a redundancy request and non-redundant frame data is original frame data that is sent, but not sent in response to a redundancy request.

[0134] In some examples, the joint redundancy request may include an indication of redundancy level. A redundancy level may be indicative of how many times the sending device (e.g., server device 60), should repeat both non-redundant PI data and non- redundant frame data. For example, a redundancy level of 100% may indicate that both non-redundant PI data and non-redundant frame data should be repeated once and a redundancy level of 200% may indicate that both non-redundant PI data and non- redundant frame data should be repeated twice. While redundancy level is described herein in terms of a percentage, redundancy level may be indicated in any other manner, such as a 1, 2, etc.

[0135] In some examples, upon receiving the joint redundancy request, server device 60 may send redundant PI data and redundant frame data in the packet(s) immediately following the packet including the non-redundant PI data and non-redundant frame datauntil the requested redundancy level is achieved. However, such an arrangement may be susceptible to bursty packet loss, which may cause the loss of a number of contiguous packets.

[0136] In order to better protect against potential bursty packet loss, in some examples, the joint redundancy request may further indicate a repetition pattern, e.g., using a bit map or bitmask to indicate in which data packet (e.g., RTP packet or QUIC packet) the non- redundant PI data and the non-redundant frame data are to be repeated. The joint redundancy request may further indicate which type of PI data frames are to be repeated, such as when the frame data is associated with multiple types of PI data frames. For example, the joint redundancy request may indicate that only a specific type (or types) of PI data frames are to be repeated. The use of a data map to indicate a repetition pattern may cause a sending device (e.g., server device 60) send the redundant PI data and the redundant frame data according to the pattern which may include non-contiguous packets, thereby providing improved time diversity to the redundancy transmissions. The improved time diversity may be helpful in combating bursty packet losses.

[0137] FIG. 14 is a conceptual diagram of an example RTCP APP REQ RED request. In some examples, the RTCP APP REQ RED request 1800 (including an ID 0001 and a 12-bit bit map, as shown in FIG. 18) in an RTCP-APP message may be reused with the following re-interpretation. In some examples, for IVAS, the “payload chunk” is interpreted as a combination of payload header 1704, frame data 1760, and PI data 1740. For example, payload 1730 may be referred to a payload chunk when payload 1730 represents an IVUS payload. In some examples, request 1800 may be used as a PI redundancy request. In some examples, bit field 1802 may include the bit map indicating the repetition pattern, the type of PI data frames to be repeated, and / or a redundancy level. In some examples, in a data packet containing the requested redundancy, the redundant payload header 1704, frame data 1760, and PI data 1740 may be placed next to (e.g., in front of) the respective payload header 1704, frame data 1760, and PI data 1740 of a non- redundant payload chunk. In some examples, in a data packet containing the requested redundancy, the redundant payload header 1704, frame data 1760, and PI data 1740 may maintain their order and be placed as a whole next to (e.g., in front of) the combination of the payload header 1704, frame data 1760, and PI data 1740 of a non-redundant payload chunk.

[0138] For the bit map, in one example, if the / / th left bit of the bit map is 1 (or alternatively, a 0), then the combination may be repeated in RTP packet m+n-1, where m is the current RTP packet. In some examples, the redundancy level is limited by how many l’s (or alternatively, Os) the bit map is allowed to contain.

[0139] FIG. 19 is a conceptual diagram illustrating an example subsequent RTP packet including redundant PI data. In some examples, in response to receiving the joint redundancy request, server device 60 sends one or more packets, like packet 500, including redundancy for the PI data and the frame data as subsequent packet(s) in accordance with the joint redundancy request. For example, packet 1900 may include RTP header 1902 which may be an example of RTP header 1602. Payload 1930 may include payload header 1904 which may be an example of payload header 1604. Payload 1930 may also include 3rdframe 1906 and 2ndframe 1708. For example, 2ndframe 1708 may be a redundant frame including redundant frame data in response to the joint redundancy request, while 3rdframe 1906 may include original or non-redundant frame data.

[0140] Payload 1930 also includes PI data header for all frames (PIDHAF) 1910, PI data header for frame 3 (PIDHF3) 1912, and PI data header for frame 2 1716 (which may be a redundant PI data header), which together may be referred to as PI header section 1950. PI header section 1950 may be an example of PI header data 1612. PI data header for all frames 1910 may be similar to PI data header for all frames 1710. PI data header for frame 3 1912 may be a PI data header for PI data frame 4 1920. PI data header for frame 2 1716 may be a PI data header for PI data frame 3 1724. Thus, in the example of FIG. 19 there is one redundant data frame (2ndframe 1708) and associated PI data header (PI data header frame 2 1716) and PI data frame (PI data frame 3 1724) and one original or non-redundant data frame (3rdframe 1906) and associated PI data header (PI data header frame 3 1912) and PI data frame (PI data frame 4 1920). PI frame data 1970 includes PI frame data for all frames 1918 which may be similar to PI frame data 1718, PI data frame 3 1724, and PI data frame 4 1920. While shown PI data frame 3 1724, in some examples, PI data frame for all frames 1918 may appear after PI data frame 3 1724 or may not be included. In some examples, the contents of RTP packet 1900 and / or the order of the contents of RTP packet 1900 may be different than depicted. If an audio frame for which redundant PI data is requested does not have associated PI data, the associated PI data header framemay indicate so, e.g., with NO PI DATA as the PI Type, and the corresponding PI data frame is empty.

[0141] When there is a bandwidth limitation, for example, between client device 40 and server device 60, it may be more efficient to request redundancy for some PI data frames rather than send a joint redundancy request. For example, a PI data frame may be associated with a plurality of audio frames, or an audio frame may be associated with a plurality of PI data frames. Additionally, a relative importance between the audio frame and the PI data frames may be different, and the importance of PI data frames amongst the PI data frames may be different. As such, it may be desirable, in the interest of saving bandwidth, to use a separate redundancy request for PI data frames, which may result in a sending device (e.g., server device 60) not sending unnecessary redundant frame data. In such examples, the binding or association between any redundant PI data frame and particular frame data (e.g., which audio and metadata frame(s) the PI data frames to be repeated are associated with) should be indicated in any subsequent frame that includes a redundant PI data frame. Thus, server device 60 may include an indication of the binding or association in a subsequent packet that includes redundant PI data. In the example, where PI data frames apply to a plurality of audio frames and may be valid for a period of time, indication of binding may not be needed.

[0142] In some examples, client device 40 may send a redundancy request to indicate a redundancy request for PI data as a whole. The request may indicate that the sender use repetitions contiguously in subsequent data packets or indicate a redundancy pattern (e.g., by a bit map, as in the example described above) of selected data packets in the subsequent data packets in which to include redundant PI data. For example, the data packet may be an RTP packet or QUIC packet. In some examples, request 1800 may represent the redundancy request for PI data as a whole. In some examples, client device 40 may use bit field 1802 to indicate a redundancy pattern, for example, using a bit map.

[0143] In some examples, in response to receiving the redundancy request for PI data as a whole, server device 60 sends one or more packets including redundancy for the PI data as subsequent packet(s) in accordance with the joint redundancy request. Such a subsequent packet may be similar to packet 1900, but not necessarily include 2ndframe 1708, as the redundancy request, in this example, is not a joint redundancy request. In some examples, the subsequent packet may also include 1stPI data header for frame 11712, 2ndPI data header for frame 2 1714, PI data frame 1 1720, and PI data frame 2 1722 (not shown in FIG. 19).

[0144] In some examples, the subsequent packet including the redundant PI data may include an indication of the frame data or audio frame(s) with which the redundant PI data is associated. The indication may include a packet identifier (ID) of the data packet that carried the original, non-redundant frame data. In some examples, the indication may be in a payload header, e.g., payload header 1904, as part of a ToC. The data packet may carry the requested redundant PI data frame(s) and the associated PI data header(s).

[0145] In some examples, the packet ID may be an RTP sequence number for an RTP packet, or packet number for a QUIC packet. For example, the indication may include the packet ID of the data packet that carried the non-redundant frame data, or may include a difference of the packet ID of the current data packet (which may carry redundant PI data) from the packet ID of the data packet that carried the non-redundant frame data (e.g., the original packet). The use of a difference rather than a packet ID itself, may reduce a number of bits used for the indication. In some examples, the indication may further indicate the position of the audio frame in the non-redundant frame data to which the redundant PI data is associated, e.g., using an offset (e.g., 0, 1, 2, ...). For example, an offset of zero may indicate that the audio frame of the non-redundant frame data to which the redundant PI data is associated is the first audio frame in the non-redundant frame data, an offset of one may indicate that the audio frame of the non-redundant frame data to which the redundant PI data is associated is the second audio frame in the non- redundant frame data, and so on.

[0146] In some situations, it may be desirable to only use redundancy for one or more particular types of PI data. In such examples, client device 40 may send a redundancy request to indicate a redundancy request one or more particular types of PI data (e.g., scene orientation, device orientation compensated, device orientation uncompensated, acoustic environment, etc.). In some examples, the redundancy request may indicate that the sender use repetitions contiguously in subsequent data packets or in a redundancy pattern (e.g., as set forth in a bit map as described above) of requested type(s) of PI data in the subsequent packets. The requested type(s) of PI data may be indicated by an index, for example, in the redundancy request. For example, one or more bits in bit field 1802 may be used for such an index.

[0147] Server device 60 may send a data packet containing the requested PI data and an indication of the associated frame data or audio frames. The indication may indicate the packet ID of the packet that carried the non-redundant frame data. The PI data may include the PI data headers, and the associated PI data frames requested by client device 40. In some examples, the PI data headers and the PI data frames included in the packet providing the redundancy may be different than the PI data headers and the PI data frames in the data packet that carried the non-redundant data. For example, the original packet including the non-redundant data may include one or more other PI data headers and PI data frames not included in the packet providing the redundant data.

[0148] FIG. 20 is a conceptual diagram of an example redundancy request for PI data according to one or more aspects of this disclosure. Redundancy request 2000 may represent an example dedicated redundancy request for PI data. In some examples, redundancy request 2000 may be used rather than repurposing RTCP APP REQ RED request 1800 for use with requesting redundancy for PI data. Redundancy request 2000 may include a first four bits of 1011, as shown, which may be indicative of redundancy request 2000 being a redundancy request for PI data.

[0149] In some examples, redundancy request 2000 may be used by a receiving device to request redundancy for IVAS, EVS, and / or other PI data. Bit field 2002 may include a 12-bit bitmask or bit map that signals a request on how the PI data portions of non- redundant payloads chunks are to be repeated in subsequent packets. For example, the position of the bit set indicates which earlier non-redundant payload chunks are requested to be added as redundant payload chunks to the current packet. In some examples, if the least significant bit (LSB) (which may be a rightmost bit) is set equal to 1, this indicates that the PI data portion of the last previous payload chunk is requested to be repeated as redundant payload in the current packet. If the most significant bit (MSB) (which may be a leftmost bit) is set equal to 1, this indicates that the PI data portion of the payload chunk that was transmitted 12 packets ago is requested to be repeated as redundant payload chunk in the current packet. Note that it is not guaranteed that the sender has access to such old payload chunks. For example, the sender may not have access to a payload chunk sent 12 packets ago and may therefore not be able to provide the requested redundancy. In some examples, a maximum amount of redundancy that may be requested is 500%, e.g., at maximum five bits can be set in the bit field 2002.

[0150] In some examples, in FIG. 20 additional bits may be added to the bit field (e.g., bit field 2002) to indicate which types of PI data is requested for redundancy. The additional bits may include indices of the types (e.g., the 5-bit PI Type) and / or a count of the types.

[0151] FIG. 21 is a conceptual diagram illustrating an example QUIC packet according to one or more aspects of this disclosure. QUIC packet 2100 may include header 2102 which may be a QUIC header. QUIC packet 2100 may also include QUIC frame 2104 and QUIC frame 2106, which together may form a QUIC payload of QUIC packet 2100. QUIC frame 2104 may include PI and QUIC frame 2106 may include speech data.

[0152] In some examples, a receiver may send the redundancy request in a QUIC packet. The QUIC packet may include an initial packet, a handshake packet, a 1-RTT packet, or another QUIC packet, such as those specified in IETF RFC 9000. The QUIC packet may be exchanged between the receiver and sender during the handshake phase of a QUIC connection setup. Alternatively, the QUIC packet for redundancy request may be sent after the handshake phase of a QUIC connection setup, and may include a QUIC frame for redundancy request, where the QUIC frame includes a frame type field and typedependent fields which indicate the requested redundancy pattern. Such a frame type may be PI RED REQ.

[0153] In some examples, for a QUIC packet that carries the PI (data), the PI is carried in a QUIC frame of a QUIC packet that is dedicated to PI (data). The QUIC frame may include a frame type field and type-dependent fields which include the PI (data). The frame type may be PI VOICE CODEC. For example, QUIC packet 2100 may include a dedicated frame, QUIC frame 2104, for carrying the PI data.

[0154] FIG. 22 is a flow diagram illustrating example redundancy request techniques according to one or more aspects of this disclosure. Client device 40 may determine to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data (2200). For example, client device 40 may determine that audio data, metadata, audio data and associated metadata, or PI data from server device 60 should have redundancy, such as a particular level of redundancy. The audio data, metadata, audio data and associated metadata, or PI data may be data destined for Tenderer 56 of client device 40 or, in some examples, a separate Tenderer (not shown in FIG. 1A).

[0155] Client device 40 may generate the redundancy request (2202). For example, client device 40 may generate a redundancy request like request 1800 to request the audio data, metadata, audio data and associated metadata, or PI data redundancy. Client device 40 may send, to the second computing device, the redundancy request (2204). For example, client device 40 may send request 1800 to server device 60 to request the redundancy.

[0156] In some examples, the redundancy request includes a 2-bit identifier identifying whether the redundancy request is for the audio data, the metadata, or the audio data and the associated metadata. In some examples, the redundancy request includes a joint redundancy request for the PI data and for frame data. In some examples, redundancy request includes a dedicated redundancy request for the PI data. In some examples, the dedicated redundancy request for the PI data includes a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data. In some examples, the one or more specified types of the PI data include one or more of scene orientation data, compensated device orientation data, uncompensated device orientation data, acoustic environment data, or no PI data.

[0157] In some examples, the redundancy request includes a message for a type of payload chunk. In some examples, the message includes a 4-bit identifier identifying the type of payload chunk, and wherein the type of payload chunk comprises an audio data type, a metadata type, or an audio data and associated metadata type. In some examples, the message includes a bit field, wherein a position of a bit set within the bit field is indicative of which earlier non-redundant payload chunk is requested to be added as redundant payload chunks to a response to the redundancy request.

[0158] In some examples, the bit field comprises a 12-bit bitmask. In some examples, the redundancy request includes a message having a bit field for indicating a requested type of payload chunk. In some examples, the bit field includes two pairs of bits, each pair of bits indicating a respective one of a plurality of types of payload chunks. In some examples, the plurality of types of payload chunks includes: an audio data type; a metadata type; and an audio data and associated metadata type.

[0159] In some examples, client device 40 is configured to negotiate with the second computing device (e.g., server device 60) for use of the redundancy request. In some examples, negotiating for the use of the redundancy request includes negotiating for at least one of support for a message; whether repetition or application-layer forward errorcorrection (FEC) is used; or which FEC scheme is used and which configuration is used for the FEC scheme.

[0160] In some examples, client device 40 is configured to determine to send the redundancy request based on a measured packet error rate (PER) of the device meeting a predetermined maximum metadata PER threshold and not meeting a predetermined maximum audio data PER threshold or a PER of the device meeting the predetermined maximum audio data PER threshold. In some examples, client device 40 is configured to determine to send the redundancy request based on predicting that a link will be disrupted or that a packet error rate (PER) will change. In some examples, client device 40 is configured to determine to send the redundancy request based on a rate of change of a pose of a user or a round-trip network delay.

[0161] In some examples, client device 40 is configured to receive, in response to sending the redundancy request, one or more subsequent packets comprising redundant data from the second computing device. In some examples, at least one of the one or more subsequent packets includes an indication associating redundant PI data of the at least one of the one or more subsequent packets to an audio frame of a packet previously sent by the second computing device, the redundant PI data is a dummy PI data if the audio frame for which the redundant PI data is requested does not have associated PI data, or the dummy PI data is indicated by a PI data type NO PI DATA and a null PI frame data.

[0162] In some examples, the redundancy request includes an indication of a redundancy level, the redundancy level including an amount of redundancy requested. In some examples, the redundancy request includes an indication of a repetition pattern indicative of which subsequent packets are to include redundant PI data. In some examples, the indication of the repetition pattern includes a bit map.

[0163] In some examples, client device 40 may receive, in response to sending the redundancy request and from the second computing device (e.g., server device 60), one or more subsequent packets including redundant PI data. In some examples, one of the one or more subsequent packets includes an indication associating redundant PI data of the at least one of the one or more subsequent packets to an audio frame of a packet previously sent by the second computing device.

[0164] In some examples, the redundant PI data is a dummy PI data if the audio frame for which the redundant PI data is requested does not have associated PI data. In someexamples, the dummy PI data is indicated by a PI data type NO PI DATA and a null PI frame data.

[0165] In some examples, the redundancy request is carried in an RTCP-APP packet, and wherein subsequent packets carrying redundant PI data are RTP packets. In some examples, the redundancy request is carried in a QUIC frame during a QUIC handshake phase, and wherein subsequent packets carrying redundant PI data are QUIC packets.

[0166] FIG. 23 is a flow diagram illustrating example response to redundancy request techniques according to one or more aspects of this disclosure. Server device 60 may receive, from a first computing device, a redundancy request for audio data, metadata, audio data and associated metadata, or PI data for a Tenderer (2300). For example, server device 60 may receive a redundancy request for audio data, metadata, audio data and associated metadata, or PI data from client device 40. The PI data may be intended for Tenderer 56 or an external Tenderer.

[0167] Server device 60 may generate, in response to the redundancy request, one or more packets comprising redundant data (2302). For example, server device may generate packets including the requested redundant data.

[0168] Server device 60 may send, to the first computing device, the one or more packets (2304). For example, server device 60 may send the one or more packets including the redundant data to client device 40.

[0169] In some examples, the redundancy request includes a 2 -bit identifier identifying whether the redundancy request is for the audio data, the metadata, or the audio data and the associated metadata. In some examples, the redundancy request includes a joint redundancy request for the PI data and for frame data. In some examples, redundancy request includes a dedicated redundancy request for the PI data. In some examples, the dedicated redundancy request for the PI data includes a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data. In some examples, the one or more specified types of the PI data include one or more of scene orientation data, compensated device orientation data, uncompensated device orientation data, acoustic environment data, or no PI data.

[0170] In some examples, the redundancy request includes a message for a type of payload chunk. In some examples, the message includes a 4-bit identifier identifying the type of payload chunk, and wherein the type of payload chunk comprises audio data, metadata, or audio data and associated metadata. In some examples, the message includesa bit field, wherein a position of a bit set within the bit field is indicative of which earlier non-redundant payload chunk is requested to be added as redundant payload chunks to a response to the redundancy request.

[0171] In some examples, the bit field comprises a 12-bit bitmask. In some examples, the redundancy request includes a message having a bit field for indicating a requested type of payload chunk. In some examples, the bit field includes two pairs of bits, each pair of bits indicating a respective one of a plurality of types of payload chunks

[0172] In some examples, the redundancy request includes an indication of a redundancy level, the redundancy level including an amount of redundancy requested. In some examples, the redundancy request includes an indication of a repetition pattern indicative of which subsequent packets are to include redundant PI data. In some examples, the indication of the repetition pattern includes a bit map.

[0173] In some examples, one of the one or more subsequent packets includes an indication associating redundant PI data of the at least one of the one or more subsequent packets to an audio frame of a packet previously sent by the second computing device.

[0174] In some examples, the redundant PI data is a dummy PI data if the audio frame for which the redundant PI data is requested does not have associated PI data. In some examples, the dummy PI data is indicated by a PI data type NO PI DATA and a null PI frame data.

[0175] In some examples, the redundancy request is carried in an RTCP-APP packet, and wherein subsequent packets carrying redundant PI data are RTP packets. In some examples, the redundancy request is carried in a QUIC frame during a QUIC handshake phase, and wherein subsequent packets carrying redundant PI data are QUIC packets.

[0176] In some examples, the redundancy request includes a joint redundancy request for the PI data and for frame data. In some examples, the redundancy request includes a dedicated redundancy request for the PI data. In some examples, the dedicated redundancy request for the PI data includes a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data. In some examples, the one or more specified types of the PI data include one or more of scene orientation data, compensated device orientation data, uncompensated device orientation data, acoustic environment data, or no PI data.

[0177] In some examples, the redundancy request includes an indication of a redundancy level, the redundancy level including an amount of redundancy requested. In someexamples, the redundancy request includes an indication of a repetition pattern indicative of which subsequent packets are to include redundant PI data. In some examples, the indication of the repetition pattern comprises a bit map. In some examples, as part of generating the one or more packets including the redundant PI data, server device 60 may generate the one or more packets including the redundant PI data in accordance with the repetition pattern. In some examples, at least one of the one or more packets includes an indication associating the redundant PI data of the at least one of the one or more packets to an audio frame of a packet previously sent by server device 60 to the first computing device (e.g., client device 40).

[0178] Various examples of the techniques of this disclosure are summarized in the following clauses:

[0179] Clause 1 A: A method comprising: determining, by a first computing device, to send a redundancy request for audio data, metadata, or audio data and associated metadata; and sending, by the first computing device and to a second computing device, the redundancy request.

[0180] Clause 2 A. The method of clause 1 A, wherein the redundancy request comprises a 2-bit identifier, the 2-bit identifier identifying whether the request is for the audio data, the metadata, or the audio data and the associated metadata.

[0181] Clause 3A. The method of clause 1A, wherein the redundancy request comprises a message for a type of payload chunk.

[0182] Clause 4A. The method of clause 3 A, wherein the message comprises a 4-bit identifier identifying the type of payload chunk, and wherein the type of payload chunk comprises audio data, metadata, or audio data and associated metadata.

[0183] Clause 5A. The method of any of clauses 3 A-4A, wherein the message further comprises a bit field, wherein a position of a bit set within the bit field is indicative of which earlier non-redundant payload chunk is requested to be added as redundant payload chunks to a response to the redundancy request.

[0184] Clause 6A. The method of any of clauses 3A-5A, wherein the bit field comprises a 12-bit bitmask.

[0185] Clause 7 A. The method of clause 1A, wherein the redundancy request comprises a message having a bit field for indicating a requested type of payload chunk.

[0186] Clause 8A. The method of clause 7A, wherein the bit field comprises two pairs of bits, each pair of bits indicating a respective one of a plurality of types of payload chunks.

[0187] Clause 9A. The method of clause 8A, wherein the plurality of types of payload chunks comprises audio data, metadata, and audio data and associated metadata.

[0188] Clause 10 A. The method of any of clauses 7A-9A, wherein a position of a bit pair within the bit field is indicative of which earlier non-redundant payload chunk is requested to be added as redundant payload chunks to a response to the redundancy request.

[0189] Clause 11 A. The method of any of clauses 1A-10A, further comprising negotiating, by the first computing device with the second computing device, for use of the redundancy request.

[0190] Clause 12A. The method of clause 11 A, wherein negotiating for the use of the redundancy request comprises negotiating for at least one of a) support for a message, b) whether repetition or application-layer forward error correction (FEC) is used, or c) which FEC scheme is used and which configuration is used for the FEC scheme.

[0191] Clause 13A. The method of any of clauses 1A-12A, wherein determining to send a redundancy request comprises determining that a measured packet error rate (PER) of the first computing device meets a predetermined maximum metadata PER threshold and does not meet a predetermined maximum audio data PER threshold.

[0192] Clause 14 A. The method of any of clauses 1A-12A, wherein determining to send a redundancy request comprises determining that a measured packet error rate (PER) of the first computing device meets a predetermined maximum audio data PER threshold.

[0193] Clause 15 A. The method of any of clauses 1A-12A, wherein determining to send a redundancy request comprises predicting that a link will be disrupted or that a packet error rate (PER) will change.

[0194] Clause 16 A. The method of any of clauses 1A-12A, wherein determining to send a redundancy request is based on a rate of change of a pose of a user.

[0195] Clause 17 A. The method of any of clauses 1A-12A, wherein determining to send a redundancy request is based on a round trip network delay.

[0196] Clause 18 A. A computing device, comprising: one or more memories configured to store a redundancy request; and one or more processors coupled to thememory, the one or more processors being configured to perform any of the methods of clauses 1A-17A.

[0197] Clause 19 A. The computing device of clause 18 A, wherein the computing device comprises a mobile device or an application server.

[0198] Clause 20A. A computing device comprising at least one means for performing any of the methods of clauses 1A-17A.

[0199] Clause 21A. Computer-readable storage media storing instructions, which, when executed, cause one or more processors to perform any of the methods of clauses 1A-17A.

[0200] Clause IB: A method comprising: determining, by a first computing device, to send a redundancy request for processing information (PI) data for a Tenderer to a second computing device; generating, by the first computing device, the redundancy request; and sending, by the first computing device and to the second computing device, the redundancy request.

[0201] Clause 2B. The method of clause IB, wherein the redundancy request comprises a joint redundancy request for the PI data and for frame data.

[0202] Clause 3B. The method of clause IB, wherein the redundancy request comprises a dedicated redundancy request for the PI data.

[0203] Clause 4B. The method of clause 3B, wherein the dedicated redundancy request for the PI data comprises a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data.

[0204] Clause 5B. The method of clause 4B, wherein the one or more specified types of the PI data comprise one or more of scene orientation data, compensated device orientation data, uncompensated device orientation data, acoustic environment data, or no PI data.

[0205] Clause 6B. The method of any of clauses 1B-5B, wherein the redundancy request comprises an indication of a redundancy level, the redundancy level comprising an amount of redundancy requested.

[0206] Clause 7B. The method of any of clauses 1B-6B, wherein the redundancy request comprises an indication of a repetition pattern indicative of which subsequent packets are to include redundant PI data.

[0207] Clause 8B. The method of clause 7B, wherein the indication of the repetition pattern comprises a bit map.

[0208] Clause 9B. The method of any of clauses 1B-8B, further comprising receiving, in response to sending the redundancy request and from the second computing device, one or more subsequent packets comprising redundant PI data.

[0209] Clause 10B. The method of clause 9B, wherein at least one of the one or more subsequent packets comprises an indication associating redundant PI data of the at least one of the one or more subsequent packets to an audio frame of a packet previously sent by the second computing device.

[0210] Clause 11B. A device for decoding audio data, the device comprising: one or more memories configured to store the audio data; and one or more processors implemented in circuitry and communicatively coupled to the one or more memories, the one or more processors configured to: perform the method of any of clauses 1B-10B.

[0211] Clause 12B. Non-transitory computer-readable storage media storing instructions, which when executed, cause one or more processors to perform the method of any of clauses 1B-10B.

[0212] Clause 13B. A device for decoding audio data, the device comprising at least one means for performing the method of any of clauses 1B-10B.

[0213] Clause 14B. A method comprising: receiving, by a second computing device and from a first computing device, a redundancy request for processing information (PI) data for a Tenderer; generating, by the second computing device and in response to the redundancy request, one or more packets comprising redundant PI data; and sending, by the second computing device and to the first computing device, the one or more packets.

[0214] Clause 15B. The method of clause 14B, wherein the redundancy request comprises a joint redundancy request for the PI data and for frame data.

[0215] Clause 16B. The method of clause 14B, wherein the redundancy request comprises a dedicated redundancy request for the PI data.

[0216] Clause 17B. The method of clause 16B, wherein the dedicated redundancy request for the PI data comprises a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data.

[0217] Clause 18B. The method of clause 17B, wherein the one or more specified types of the PI data comprise one or more of scene orientation data, compensated device orientation data, uncompensated device orientation data, acoustic environment data, or no PI data.

[0218] Clause 19B. The method of any of clauses 14B-18B, wherein the redundancy request comprises an indication of a redundancy level, the redundancy level comprising an amount of redundancy requested.

[0219] Clause 20B. The method of any of clauses 14B-19B, wherein the redundancy request comprises an indication of a repetition pattern indicative of which subsequent packets are to include redundant PI data.

[0220] Clause 21B. The method of clause 20B, wherein the indication of the repetition pattern comprises a bit map.

[0221] Clause 22B. The method of clause 20B or clause 21B, wherein generating the one or more packets comprising the redundant PI data comprises generating the one or more packets comprising the redundant PI data in accordance with the repetition pattern.

[0222] Clause 23B. The method of any of clauses 14B-22B, wherein at least one of the one or more packets comprises an indication associating the redundant PI data of the at least one of the one or more packets to an audio frame of a packet previously sent by the second computing device to the first computing device.

[0223] Clause 24B. A device for encoding audio data, the device comprising: one or more memories configured to store the audio data; and one or more processors implemented in circuitry and communicatively coupled to the one or more memories, the one or more processors configured to: perform the method of any of clauses 14B-23B.

[0224] Clause 25B. Non-transitory computer-readable storage media storing instructions, which when executed, cause one or more processors to perform the method of any of clauses 14B-23B.

[0225] Clause 26B. A device for encoding audio data, the device comprising at least one means for performing the method of any of clauses 14B-23B.

[0226] Clause 1C. A device comprising one or more processors configured to: determine to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generate the redundancy request; and send the redundancy request to a second computing device.

[0227] Clause 2C. The device of clause 1C, wherein the redundancy request comprises an identifier identifying whether the request is for the audio data, the metadata, or the audio data and the associated metadata.

[0228] Clause 3C. The device of clause 1C, wherein the redundancy request comprises a joint redundancy request for the PI data and for frame data.

[0229] Clause 4C. The device of clause 1C, wherein the redundancy request comprises a dedicated redundancy request for the PI data.

[0230] Clause 5C. The device of clause 4C, wherein the dedicated redundancy request for the PI data comprises a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data.

[0231] Clause 6C. The device of clause 5., wherein the one or more specified types of the PI data comprise one or more of: scene orientation data; compensated device orientation data; uncompensated device orientation data; acoustic environment data; or no PI data.

[0232] Clause 7C. The device of any of clauses 1C-6C, wherein the redundancy request comprises a message for a type of payload chunk.

[0233] Clause 8C. The device of clause 7C, wherein the message comprises a 4-bit identifier identifying the type of payload chunk, and wherein the type of payload chunk comprises an audio data type, a metadata type, or an audio data and associated metadata type.

[0234] Clause 9C. The device of clause 7C, wherein the message further comprises a bit field, wherein a position of a bit set within the bit field is indicative of which earlier non-redundant payload chunk is requested to be added as redundant payload chunks to a response to the redundancy request.

[0235] Clause 10C. The device of clause 9C, wherein the bit field comprises a 12-bit bitmask.

[0236] Clause 11C. The device of clause 1C, wherein the redundancy request comprises a message having a bit field for indicating a requested type of payload chunk.

[0237] Clause 12C. The device of clause 11C, wherein the bit field comprises two pairs of bits, each pair of bits indicating a respective one of a plurality of types of payload chunks.

[0238] Clause 13C. The device of clause 12C, wherein the plurality of types of payload chunks comprises: an audio data type; a metadata type; and an audio data and associated metadata type.

[0239] Clause 14C. The device of any of clauses 1C-13C, wherein the one or more processors are further configured to negotiate with the second computing device for use of the redundancy request.

[0240] Clause 15C. The device of clause 14C, wherein negotiating for the use of the redundancy request comprises negotiating for at least one of: support for a message; whether repetition or application-layer forward error correction (FEC) is used; or which FEC scheme is used and which configuration is used for the FEC scheme.

[0241] Clause 16C. The device of any of clauses 1C-15C, wherein the one or more processors are further configured to determine to send the redundancy request based on a measured packet error rate (PER) of the device meeting a predetermined maximum metadata PER threshold and not meeting a predetermined maximum audio data PER threshold or a measured PER of the device meeting the predetermined maximum audio data PER threshold.

[0242] Clause 17C. The device of any of clauses 1C-15C, wherein the one or more processors are further configured to determine to send the redundancy request based on predicting that a link will be disrupted or that a packet error rate (PER) will change.

[0243] Clause 18C. The device of any of clauses 1C-15C, wherein the one or more processors are further configured to determine to send the redundancy request based on a rate of change of a pose of a user or a round trip network delay.

[0244] Clause 19C. The device of any of clauses 1C-18C, wherein the one or more processors are further configured to receive, in response to sending the redundancy request, one or more subsequent packets comprising redundant data from the second computing device.

[0245] Clause 20C. The device of clause 19C, wherein: at least one of the one or more subsequent packets comprises an indication associating redundant PI data of the at least one of the one or more subsequent packets to an audio frame of a packet previously sent by the second computing device; the redundant PI data is a dummy PI data if the audio frame for which the redundant PI data is requested does not have associated PI data; or the dummy PI data is indicated by a PI data type NO PI DATA and a null PI frame data.

[0246] Clause 21C. The device of any of clauses 1C-20C, wherein the redundancy request is carried in a Real-Time Transport Control Protocol (RTCP)-APP packet, and wherein subsequent packets carrying redundant data are Real-Time Transport Protocol (RTP) packets.

[0247] Clause 22C. The device of any of clauses 1C-20C, wherein the redundancy request is carried in a QUIC frame during a QUIC handshake phase, and wherein subsequent packets carrying redundant data are QUIC packets.

[0248] Clause 23C. A method of comprising: determining, by a first computing device, to send a redundancy request for processing information (PI) data; generating, by the first computing device, the redundancy request; and sending, by the first computing device and to a second computing device, the redundancy request.

[0249] Clause 24C. A device comprising: means for determining to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; means for generating the redundancy request; and means for sending the redundancy request to the second computing device.

[0250] Clause 25 C. Non-transitory computer-readable storage media storing instructions, which when executed, cause one or more processors to determine to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generate the redundancy request; and send the redundancy request to a second computing device.

[0251] Clause 26C. The device of clause 12C, wherein the plurality of types of payload chunks comprises: an audio data type; a metadata type; and an audio data and associated metadata type, and wherein the redundancy request is carried in a Real-Time Transport Control Protocol (RTCP)-APP packet and subsequent packets comprising redundant data are Real-Time Transport Protocol (RTP) packets, or the redundancy request is carried in a QUIC frame during a QUIC handshake phase and the subsequent packets comprising the redundant data are QUIC packets.

[0252] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on one or more computer-readable media and executed by one or more hardware-based processing units. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer- readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions,code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include one or more computer-readable media.

[0253] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM and / or other optical disk storage, magnetic disk storage, and / or other magnetic storage devices, flash memory, and / or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but are instead directed to non-transitory, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0254] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0255] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure toemphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0256] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A device comprising one or more processors configured to: determine to send a redundancy request for audio data, metadata, audio data and associated metadata, or processing information (PI) data; generate the redundancy request; and send the redundancy request to a second computing device.

2. The device of claim 1, wherein the redundancy request comprises an identifier identifying whether the redundancy request is for the audio data, the metadata, or the audio data and the associated metadata.

3. The device of claim 1, wherein the redundancy request comprises a joint redundancy request for the PI data and for frame data.

4. The device of claim 1, wherein the redundancy request comprises a dedicated redundancy request for the PI data.

5. The device of claim 4, wherein the dedicated redundancy request for the PI data comprises a dedicated redundancy request for all PI data or a dedicated redundancy request for one or more specified types of the PI data.

6. The device of claim 5, wherein the one or more specified types of the PI data comprise one or more of scene orientation data; compensated device orientation data; uncompensated device orientation data; acoustic environment data; or no PI data.

7. The device of claim 6, wherein the redundancy request comprises a message for a type of payload chunk.

8. The device of claim 7, wherein the message comprises a 4-bit identifier identifying the type of payload chunk, and wherein the type of payload chunk comprises an audio data type, a metadata type, or an audio data and associated metadata type.

9. The device of claim 7, wherein the message further comprises a bit field, wherein a position of a bit set within the bit field is indicative of which earlier non- redundant payload chunk is requested to be added as redundant payload chunks to a response to the redundancy request.

10. The device of claim 9, wherein the bit field comprises a 12-bit bitmask.

11. The device of claim 1, wherein the redundancy request comprises a message having a bit field for indicating a requested type of payload chunk.

12. The device of claim 11, wherein the bit field comprises two pairs of bits, each pair of bits indicating a respective one of a plurality of types of payload chunks.

13. The device of claim 12, wherein the plurality of types of payload chunks comprises: an audio data type; a metadata type; and an audio data and associated metadata type, and wherein the redundancy request is carried in a Real-Time Transport Control Protocol (RTCP)-APP packet and subsequent packets comprising redundant data are Real-Time Transport Protocol (RTP) packets, or the redundancy request is carried in a QUIC frame during a QUIC handshake phase and the subsequent packets comprising the redundant data are QUIC packets.

14. The device of claim 1, wherein the one or more processors are further configured to negotiate with the second computing device for use of the redundancy request.

15. The device of claim 14, wherein negotiating for the use of the redundancy request comprises negotiating for at least one of: support for a message; whether repetition or application-layer forward error correction (FEC) is used; orwhich FEC scheme is used and which configuration is used for the FEC scheme.

16. The device of claim 1, wherein the one or more processors are further configured to determine to send the redundancy request based on a measured packet error rate (PER) of the device meeting a predetermined maximum metadata PER threshold and not meeting a predetermined maximum audio data PER threshold or a PER of the device meeting the predetermined maximum audio data PER threshold.

17. The device of claim 1, wherein the one or more processors are further configured to determine to send the redundancy request based on predicting that a link will be disrupted or that a packet error rate (PER) will change.

18. The device of claim 1, wherein the one or more processors are further configured to determine to send the redundancy request based on a rate of change of a pose of a user or a round-trip network delay.

19. The device of claim 1, wherein the one or more processors are further configured to receive, in response to sending the redundancy request, one or more subsequent packets comprising redundant data from the second computing device.

20. The device of claim 19, wherein: at least one of the one or more subsequent packets comprises an indication associating redundant PI data of the at least one of the one or more subsequent packets to an audio frame of a packet previously sent by the second computing device; the redundant PI data is a dummy PI data if the audio frame for which the redundant PI data is requested does not have associated PI data; or the dummy PI data is indicated by a PI data type NO PI DATA and a null PI frame data.

Citation Information

Patent Citations

  • Signalling of a request to adapt a voice-over-IP communication session

    US20200236154A1

  • US202463573172P

  • US202463680548P

  • US202563778771P

Cited By

  • Representing descriptive metadata in spatial audio

    WO2026162311A1