Audio data transport format for an immersive audio codec during a voice over IP communication session
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-08-13
Smart Images

Figure EP2026052607_13082026_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] Title: Audio data transport format for an immersive audio codec during a Voice over IP communication session
[0003] Technical Field
[0004] The present invention relates to the field of telecommunications and more particularly to packet-switched communication networks. In this type of network, it is possible to route data streams associated with real-time services.
[0005] The Internet Protocol, called IP for "Internet Protocol", developed by the IETF, for "Internet Engineering Task Force", is implemented on packet communication networks to support both non-real-time services such as data transfer services, web page viewing, email, and real-time or conversational services, such as IP telephony, IP video telephony or IP video streaming.
[0006] The invention relates more particularly to the generation and reception of payloads (hereafter referred to as transport formats) for transporting data encoded according to the Real-Time Transport Protocol (RTP) during real-time communication between two communication devices, for example, between two communication terminals or between a communication terminal and network equipment, for processing real-time signals such as voice or video signals. The invention is particularly applicable to the data transport format of an immersive audio codec.
[0007] Previous technique
[0008] Reference is made to IETF RFC 3550 and RFC 4566 specifications for the known state-of-the-art basics of the RTP protocol as well as the SDP (Session Description Protocol).
[0009] The payload format for transporting data via the RTP protocol, "RTP payload format" in English, also referred to as the transport format in the rest of the document, is defined for the EVS codec (for "Enhanced Voice services" in English) in the 3GPP TS 26.445 Annex A specification.
[0010] For the sake of brevity, we will not repeat all the details of this specification here. We will not reiterate the complete list of EVS codec SDP parameters (see 3GPP TS 26.445 Appendix A), but we will recall that there are global parameters, parameters specific to EVS Primary mode (e.g., "br" or "bw"), and parameters specific to EVS AMR-WB 10 (e.g., "mode-set"). The IVAS codec (for "Immersive Voice and Audio Services") was specified in 3GPP Release 18 in 2024. This codec is currently defined in a floating-point source code version (3GPP TS 26.258), but a fixed-point source code version is under development and is to be standardized in Release 19 (in 3GPP TS 26.251). A high-level description of the IVAS codec is given in 3GPP TS 26.250.
[0011] The IVAS codec is an extension of the EVS codec for representing stereo and immersive formats. It is identical to EVS for mono encoding. For stereo and immersive formats, the IVAS codec bitrate ranges from 13.2 to 512 kbit / s (with restrictions depending on the exact format). The different input formats of the IVAS encoder are defined in Table 1 below.Some formats contain metadata: the MASA format is a parametric format with one or two audio channels with spatial metadata that describes the spatial characteristics of an audio scene such as spatial direction information, ratios between directional and non-directional energy, or spatial coherence information; objects (called ISM in IVAS) are mono sources with metadata that represent information describing the audio stream and the artistic intention (e.g. azimuth and elevation direction, distance, source orientation, scaling factor to apply to the audio source) to be used to convert this audio stream in a playback system.
[0012] Table 1
[0013]
[0014] " "
[0015]
[0016] The IVAS codec also includes encoding modes called "split rendering," illustrated in Figure 1, for applications where immersive rendering is split, distributed, or
[0017] "Split" into two parts. This distributed rendering method will subsequently be referred to by its English terminology "split rendering" or its abbreviation SR. Figure 1 assumes IVAS communication between a terminal A (for example, a smartphone) and a terminal C (for example, augmented reality glasses), the latter being limited in computing power. To reduce the computing load on terminal C, an intermediate terminal B performs a binaural pre-rendering (either directly or via a cutoff). The signal captured (001) and encoded (002) by A is decoded (010) and re-encoded in an SR format (011) in B; the SR format corresponds to a binaural pre-rendering which is then sent to C. A "light" decoding of the binaural pre-rendering and a final rendering (020) are then performed at low cost in terminal C before playback (021); terminal C also includes a sensor (023) measuring the head movement (pos., called "pose" in English) of the listener; the "pos.The head movement signal is encoded (022) by C and multiplexed with the captured audio signal (025) encoded (024) by C. The "pos." information is decoded (012) by B (after extraction by 014) to adapt the pre-rendering in 010, knowing that the initially captured "pos." information is also used to make any final corrections in 020. The stream from terminal C to terminal A is processed in terminal B by IVAS decoding and re-encoding in 014 and 013 respectively, before being decoded (003) and rendered in 004. In a variant, it would be possible to combine blocks 013 and 014 into a single block that would only extract the pos information. If the encoding block 024 is an IVAS encoder that does not perform split rendering.Figure 1 does not illustrate all the variants of split rendering implementation, in particular terminal B could be replaced by a network entity in a cloud or a radio cell edge, and different options are possible for multiplexing pos. information and audio data in 024.
[0018] Details of split rendering (IVAS) are provided in the 3GPP specifications TS 26.253 (detailed description of IVAS) and TS 26.258 (floating-point source code), as well as TS 26.249 (detailed description of split rendering). Study reports on split rendering are also published in TS 26.865 (requirements) and TS 26.996 (characterization).
[0019] The bitrate of the IVAS codec actually depends on the input format (in bits / s):
[0020] - Mono: identical to EVS
[0021] - Stereo: 13200, 16400, 24400, 32000, 48000, 64000, 80000, 96000, 128000
[0022] 160,000, 192,000, 256,000
[0023] - ISM:
[0024] o ISM1: 13200, 16400, 24400, 32000, 48000, 64000, 80000, 96000, 128000 o ISM2: 16400, 24400, 32000, 48000, 64000, 80000, 96000, 128000, 160000, 192000, 256000
[0025] o ISM3: 24400, 32000, 48000, 64000, 80000, 96000, 128000, 160000, 192000, 256000, 384000
[0026] o ISM4: 24400, 32000, 48000, 64000, 80000, 96000, 128000, 160000, 192000, 256000, 384000
[0027] - SBA, MASA, MC, OSBA, OMASA: 13200, 16400, 24400, 32000, 48000, 64000,
[0028] 80000, 96000, 128000, 160000, 192000, 256000, 384000, 512000
[0029] - Split rendering: 256000, 384000 and 512000
[0030] The different output formats of the IVAS decoder / renderer are defined in Table 2 below:
[0031] Table 2
[0032]
[0033] "
[0034] "
[0035]
[0036] It is important to note that the algorithmic delay of the IVAS codec depends on the encoding / decoding format; it is 32 ms for mono and stereo formats and is between 32 and 38 ms for other combinations of input / output formats.
[0037] The payload format for transporting data via the RTP protocol, also called the transport format (“payload format”) for the IVAS codec, was defined in Release 18 in Annex A of 3GPP TS 26.253.
[0038] Figures 2a to 2g illustrate different relevant aspects of the IVAS codec's RTP transport format.
[0039] The structure of an RTP packet for the IVAS codec is illustrated in Figure 2a.
[0040] Unlike EVS, the IVAS codec transport format does not use any compact mode; the payload is entirely header-full, with a header always present, followed by the encoded data. In IVAS, an optional additional section called "PI data" (PI stands for "Processing Information") can be added (if negotiated at the SDP level) after the encoded data. This feature, unlike EVS, is used to carry additional metadata (besides the encoded data), such as the orientation of the transmitting or receiving terminal, descriptors of the encoded audio source (depending on the format), and rendering parameters. The "PI" section is similar to the SEI (Supplemental Enhancement Information) messages of some video codecs.This data provides information in particular for the processing of audio data to be applied during playback (rendering) by the playback device (renderer).
[0041] Regarding the header of the IVAS transport format, which may subsequently be called the IVAS payload, the concept inherited from EVS of the CMR byte ("CMR byte") specifying codec mode change requests (for "Codec Mode Request") and the ToC byte to describe multiplexed frames and allow their extraction (for "Table of Content"), distinguished by their first bit (H), is retained. However, the CMR byte is extended and the ToC byte is adapted for IVAS. The CMR byte (8 bits) is thus renamed "E byte" ("E" for "Extra") with an initial bit H always set to 1 (specifying an "E-byte"). This type of byte may subsequently be called either "Extra Byte" or the English terminology "E-byte". The bits following the first bit H depend on the nature of the "E-byte":
[0042] - If it is the first E-byte (first extra byte) (see figure 2b), then it is a CMR where the reserved values ('Reserved') of the CMR defined in the EVS payload are replaced by IVAS rate indications, while keeping a value indicated at 'NO_REQ' for an empty CMR request.
[0043] - If other E-bytes or additional bytes (header bytes with first bit set to 1) are present in the IVAS payload header, their meaning is defined in Figure 2c, we currently distinguish 3 types of other E-bytes (additional bytes outside CMR which we can subsequently call secondary additional bytes): coded audio tape request (Bw Req.), coded format request (Format Req.), and indication of the presence of a "PI" section (PI indic.).
[0044] Furthermore, the ToC byte (8 bits) in IVAS extends the ToC byte of EVS; it is divided into 3 parts represented in figure 2d:
[0045] • H (1 bit): always 0 (to distinguish the byte as a ToC field)
[0046] • F (1 bit): If F=1, another frame follows the current frame, if F=0, it is the last frame of the packet.
[0047] • FT (6 bits): bits indicating the frame type, either EVS Primary or EVS AMR-WB 10 (including SID mode), or IVAS – in the latter case, the FT field includes a subfield of the IVAS bitrate (BR), and the first bits of the FT field inherited from EVS are set to "01". The ToC byte format is therefore simply modified to indicate the bitrates of the IVAS codec. Thus, the bitrates associated with EVS are also listed, which implies that the mono format (EVS) is also carried in the IVAS payload.
[0048] The method for decoding the IVAS payload header can be found in Annex A of TS 26.253, Figure A.3.3.3.3.1-2.
[0049] Figure 2e illustrates the structure of the 2-byte header associated with each part of the PI (for "Processing Information") section, the PF bit indicates whether it is the last header of the PI section, the 2-bit PM field gives a PI marker, the type of PI data associated with the header is coded on 5 bits, and the length (in bits) of this data is coded on the second byte of the header (PF size).
[0050] PI data types are classified into PI data for the direction of transmission or "forward" (Figure 2f) and PI data for the direction of reception or "reverse" (Figure 2g). We will not detail the different types of PI data here; however, it should be noted that no type is currently defined in Annex A of TS 26.253 for split rendering, even though orientation data—more specifically the PI type "HEAD-ORIENTATION"—can be used to signal the movement of a terminal towards the entity performing the binaural pre-rendering.
[0051] Currently, the transport format of the IVAS codec defined in Annex A of TS 26.253 is incomplete. Several features of the IVAS codec are not yet officially supported:
[0052] - Split rendering (SR) is not yet explicitly included;
[0053] - Certain codec characteristics, such as the detailed immersive format or "sub-coded format" or secondary format (e.g., FOA, HOA2, HOA3) of a "coded format" or primary format (e.g., SBA), are not supported at the SDP negotiation level or in a coded format change request. Thus, the current IVAS codec payload only allows for the reporting and negotiation of "high-level" coded formats (hereafter referred to as primary formats) as listed in Table 3 below; it therefore does not allow for specifying the details of sub-formats (hereafter referred to as secondary formats) such as binaural for stereo, FOA (for First-Order Ambisonics) corresponding to first order - planar or 3D - in SBA, the number of objects (or channels) in ISM, etc.
[0054] - The retro-compatible stereo-to-mono reduction using EVS, a special mode of the IVAS codec that takes a stereo signal as input and outputs a mono signal encoded by EVS as its bitstream, is not defined at the SDP level. - The type of audio source (diegetic / non-diegetic) is not specified in the payload or in the negotiation of an IVAS session; it can be assumed that the current transport format only includes "diegetic" sources. The definition is reiterated here:
[0055] o Audio diegetic (or “head-tracked”): Audio intended to be presented to the listener in such a way that it is perceived as fixed relative to the listening environment.
[0056] o Non-diegetic (or "head-locked") audio: intended to be presented to the listener in such a way that it is perceived as fixed relative to the listener's head. - The current IVAS payload does not distinguish between the encoding of a stereo input and that of a binaural input. Table 3
[0057]
[0058] Regarding split rendering, it should be noted that proposals have already been presented to 3GPP, including the contribution 3GPP S4-241961, CR26253 - Further corrections to Annex A (source: Dolby Sweden AB, Ericsson LM, Nokia), November 18-22, 2024.
[0059] This document proposes to introduce a new coded audio format (noted "SR") in the IVAS payload format, new media attributes or SDP parameters (noted "sr-dof", "sr-tc"), the latter are not reviewed in detail here.
[0060] We are particularly interested here in the proposals concerning the IVAS payload header and the IVAS PI section:
[0061] - It is proposed in S4-241961 to modify the definition of the "ToC byte" to replace the reserved value "Reserved" (Res.) in Figure 2d, with a special code for BR=1110 as in Figure 3a indicating that a second ToC byte specific to split rendering follows this specific ToC as in Figure 3b.
[0062] - It is proposed in S4-241961 to modify the definition of the "E-byte" in figure 2c (excluding the first E-byte) as shown in figure 3c, where the reserved code ET = 11 is used to indicate a split rendering request (SR Req.) for the D, Y, P, R information:
[0063] o D (1 bit): Bit D controls whether the flow should be linked to head movement or not
[0064] o Y (1 bit): Enables or disables position correction (pose) along the z-axis ("yaw").
[0065] o P (1 bit): Enables or disables position correction (pose) along the y-axis ("pitch").
[0066] o R (1 bit): Enables or disables position correction (pose) along the x-axis ("roll").
[0067] This solution proposed in 3GPP S4-241961 has several drawbacks: the use of an escape code (BR=1110) for the ToC byte (byte in Figure 3a) implies
[0068] "Wasting" an additional ToC byte (byte in Figure 3b) to then signal the split rendering rate with a list of split rendering rates that is a subset of the rates already defined in the "normal" ToC. If several split rendering frames are multiplexed in the same RTP packet, the ToC byte with escape code (byte in Figure 3a) must be repeated for each of the split rendering encoded (SR) frames in the current packet, whereas a single encoded format (SR) indication would suffice, since switching on the fly in an IVAS RTP session from one SR format to another immersive IVAS format or vice versa (without SDP renegotiation) is not intended.
[0069] Furthermore, the function of the ToC byte is to signal the structure of successive frames of multiplexed encoded audio data within the current packet—to allow for the demultiplexing of the encoded frames and the derivation of their respective temporal "position"—and not to specify rendering modes or formats. Therefore, it is inconsistent to use this type of field to provide information specific to split rendering.
[0070] It should also be noted that document S4-241961 does not address the other incomplete parts listed above for the IVAS format. It uses the additional E-byte with code ET = 11 for split rendering-related requests, whereas other existing mechanisms of the IVAS payload could be used more advantageously for these specific requests.
[0071] Therefore, there is a need to extend the IVAS payload in order to define different encoding modes (such as secondary formats or split rendering) or different audio data rendering formats (diegetic or non-diegetic, binaural, etc.) for an immersive codec of the IVAS type.
[0072] Description of the invention
[0073] The invention improves upon the existing state of the art.
[0074] To this end, the invention aims at a method for generating an audio data transport format for an immersive audio codec (IVAS) during a communication session between communication devices, in which secondary format information associated with a primary format, for rendering audio data, is associated with an additional secondary byte (E-byte) of the transport format header.
[0075] This information can therefore be inserted into an additional secondary byte of the codec transport format header without it being necessary to modify the ToC byte which is not intended to contain this type of information.
[0076] In a first embodiment, secondary format information is inserted into the transport format header by a dedicated request type in the secondary supplementary byte.
[0077] This cleverly allows for the definition of all possible secondary formats for each of the primary formats already defined in the additional secondary byte of the transport format header. Furthermore, this solution maintains backward compatibility with the current format defined for the IVAS immersive audio codec.
[0078] In a second embodiment, a value is allocated to a data field of the data format indicating the primary format to inform of the existence or not of additional secondary format information associated with the audio data.
[0079] This solution allows a field indicating a type of request for the secondary additional byte to be left available for another use.
[0080] In a particular mode of this second embodiment, when the value indicates the existence of additional secondary format information, this additional secondary format information is given in an additional byte following the byte indicating the primary format.
[0081] Thus, it is possible to define a large number of possible secondary formats (or sub-formats) in this additional byte.
[0082] In a third embodiment, a data field whose value indicates the existence or not of additional secondary format information associated with the audio data is included in a data field of the secondary supplementary byte in addition to the information indicating the type of request of the supplementary byte.
[0083] This solution allows all possible secondary formats to be defined for each of the primary formats on a single byte of the additional secondary byte.
[0084] In a fourth embodiment, secondary format information is inserted into the transport format header by repeating the corresponding primary format query.
[0085] This also allows us to define all possible secondary formats for each of the primary formats already defined in the first request to change the format of the additional secondary byte in the transport format header. The second request is therefore linked to the first and thus to the corresponding primary format.
[0086] This embodiment is backward compatible with an implementation of the current transport format of the IVAS codec for at least the first request.
[0087] In a particular embodiment, information indicating a mode of rendering audio data is further associated with information indicating the existence of metadata associated with the rendering of audio data from the transport format header.
[0088] Since the PI section was defined in the IVAS payload to carry information related to the rendering of audio data ("rendering") and to send requests or feedback, it is therefore appropriate to use "PI" (Processing Information) data to indicate a request for the audio data rendering mode. According to one embodiment, information indicating an audio data rendering mode is further inserted into the transport format header by a dedicated request type in the secondary supplementary byte.
[0089] This allows several useful pieces of information to be defined to adapt the rendering of audio data, not only a specific mode of distributed rendering but also to indicate, for example, whether the stream is diegetic or not, or whether it is in binaural format or not in the case of a stereo stream.
[0090] The invention also relates to a communication terminal comprising a processing circuit capable of implementing a method for generating an audio data transport format for an immersive audio codec (IVAS) during a real-time communication session with a communication device, in which secondary format information associated with a primary format, for rendering audio data, is associated with an additional secondary byte (E-byte) of the transport format header.
[0091] It relates to an audio data transport format for an immersive audio codec (IVAS), in which secondary format information associated with a primary format, for rendering audio data, is associated with an additional secondary byte (E-byte) of the transport format header.
[0092] Such a transport format conforms to any one of the embodiments described above.
[0093] The invention also relates to a computer program comprising instructions for implementing the generation process according to the invention, according to any one of the particular embodiments described above, when said program is executed by a processor.
[0094] Such instructions can be stored permanently in a non-transient memory medium of the communication terminal implementing the transport method according to the invention.
[0095] This program can use any programming language, and be in the form of source code, object code, or code somewhere between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0096] The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions for a computer program as mentioned above.
[0097] The recording medium can be any entity or device capable of storing the program. For example, the medium may include a storage means, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a removable storage medium, a hard drive, or an SSD. Alternatively, the recording medium may be a transmissible medium, such as an electrical or optical signal, which can be transmitted via an electrical or optical cable, by radio, or by other means, so that the computer program it contains is executable remotely. The program according to the invention can, in particular, be uploaded to a network, for example, an Internet-type network.
[0098] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the aforementioned transport process.
[0099] According to an example implementation, the present technique is implemented using software and / or hardware components.
[0100] Brief description of the drawings
[0101] Other features and advantages of the invention will become clearer upon reading the following description of particular embodiments, given by way of simple illustrative and non-limiting examples, and the accompanying drawings, among which:
[0102] [Fig. 1] illustrates the operating principle of the distributed rendering mode (“split rendering”) in the IVAS codec;
[0103] [Fig. 2a], [Fig. 2b], [Fig. 2c], [Fig. 2d], [Fig. 2e], [Fig. 2f] and [Fig. 2g] illustrate the current transport format defined for the IVAS codec, described previously;
[0104] [Fig. 3a], [Fig. 3b], [Fig. 3c] illustrate the proposed modification of the IVAS transport format according to the aforementioned document S4-241961;
[0105] [Fig. 4] illustrates an embodiment of a voice over IP communication system and a method for generating a data transport format according to an embodiment of the invention as well as a method for receiving such a data transport format;
[0106] [Fig. 5a] illustrates a first embodiment according to the invention of a transport format adapted to the IVAS codec for the explicit support of coded "sub-formats" (secondary formats);
[0107] [Fig. 5b] illustrates a second embodiment according to the invention of a transport format adapted to the IVAS codec for the explicit support of coded "sub-formats" (secondary formats);
[0108] [Fig. 5c] illustrates a third embodiment according to the invention of a transport format adapted to the IVAS codec for the explicit support of coded "sub-formats" (secondary formats); [Fig. 5d] illustrates a fourth embodiment according to the invention of a transport format adapted to the IVAS codec for the explicit support of coded "sub-formats" (secondary formats);
[0109] [Fig. 6a], [Fig. 6b], [Fig. 6c] and [Fig. 6d] illustrate reading methods for a transport format adapted to the IVAS codec according to the respective embodiments of [Fig. 5a], [Fig.
[0110] 5b], [Fig. 5c] and [Fig. 5d];
[0111] [Fig. 7a] illustrates an embodiment according to the invention of a transport format adapted to the IVAS codec for the explicit support of a distributed rendering mode (“split rendering”); [Fig. 7b] illustrates an embodiment according to the invention of a transport format adapted to the IVAS codec for the support of a rendering mode including the distributed rendering mode (“split rendering”);
[0112] [Fig. 7c] illustrates a description of a data type related to a split rendering request according to an embodiment of the invention; and [Fig. 8] illustrates a hardware embodiment of a terminal implementing the data transport format generation method according to an embodiment as well as a method for receiving such a data transport format.
[0113] Description of the implementation methods
[0114] The invention will now be described below with reference to Figures 4 to 8.
[0115] Figure 4 illustrates an example of a two-way communication system between two communication devices A and B, for example two communication terminals A and B or a communication terminal A and a network device B (or vice versa).
[0116] The communication devices implement a method for generating an audio data transport format according to an embodiment of the invention. The ambient acoustic signal is, for example, captured by one or more microphones (101 and 151) on each side of the communication.
[0117] The remote signal is reproduced on one or more loudspeakers (102 and 152), which also includes headphone listening. The captured and reproduced audio signals generally undergo various acoustic treatments during transmission (103 and 153) and reception (104 and 154), such as:
[0118] • Analog-to-digital conversion, gain control, noise reduction, echo cancellation, etc., during transmission
[0119] • digital / analog conversion, gain control, etc., in reception. The figure does not represent here the acoustic treatments nor the network (including treatments and equipment such as antennas, modem, etc.) linking the two terminals, for the sake of conciseness and clarity.
[0120] In one embodiment, one of the devices A or B can be replaced by a network device, which could be a media gateway (e.g., IMS-AGW), a conference bridge (e.g., MRFP), a voicemail server, etc. To avoid making the text cumbersome, we subsequently describe an example of communication between two terminals A and B, with the understanding that one of the terminals could be a network termination (generally without audio capture / acoustic rendering).
[0121] In the case where one of the terminals is a network device, the interface providing the input signals to the encoder or retrieving the output signals from the decoder / renderer is directly electrical.
[0122] Terminal A comprises a transmitter 103 and a receiver 104, while terminal B comprises a transmitter 153 and a receiver 154. Both terminals A and B are typically configured for an exchange of offer(s) and response(s) according to the Session Description Protocol (SDP), via blocks 100 and 150, respectively. Typically, this SDP configuration is used at least during the instantiation (initialization or re-initialization) of the transmitter and receiver; in variations of the invention, this configuration may be considered internal to the terminal and allows for the configuration of the terminal's SDP offer and response. In variations of the invention, other equivalent or derivative media capacity / configuration signaling protocols (session descriptions) derived from SDP may be used in an equivalent manner.
[0123] Transmitters have functions for encoding and generating RTP packets by adding protocol headers to create a payload format corresponding to the encoding format. A transmitter also has a function for adapting the transport (aggregation, redundancy, or other) as well as a function for transmitting packets.
[0124] Thus, a transmitter of a communication device as described implements a process for generating an audio data transport format during a real-time communication session with another communication device.
[0125] To this end, as described in more detail later, the data transport format for an immersive audio codec (IVAS) includes secondary mode and / or format information associated with a primary audio data rendering format, along with at least one additional secondary byte (E-byte) in the transport format header. Receivers have functions for receiving RTP packets, extracting fields from the RTP header, decoding the transport format, and decoding frames, which may include jitter buffer management, audio decoding, and correction of lost frames. Thus, a receiver on a communication device is capable of receiving and decoding a transport format as generated by a transmitter on another communication device.
[0126] Terminals A and B include a request decoding block (blocks 107 and 157), in particular for CMR requests which allow, if necessary, changing the encoding rate of the remote terminal.
[0127] Figure 4 also includes blocks 108 and 158 for decoding PI data in the receive direction (reverse PI or rev. PI) and transmit blocks 103 and 153 can also encode PI data in the transmit direction (forward PI or fwd. PI).
[0128] Several embodiments for generating the data transport format during communication between two terminals A and B are now described. This data packet transport format exchanged between the terminals is compatible with immersive audio signal encoding and decoding as implemented by an IVAS codec.
[0129] This data transport format can also be used for data storage and therefore be recorded in memory accessible by communication terminals, or be used for messaging applications.
[0130] Before detailing an example of a transport format for IVAS, we first look at the use cases of the IVAS codec.
[0131] The simplest use case for IVAS involves a call between two devices connected to a network, for example, 5G. These devices can be smartphones equipped with multiple microphones and support the following formats:
[0132] - Mono: fallback case on an EVS-compatible call, given that EVS is included in IVAS
[0133] - Stereo with sound captured from 2 microphones on the phone or with stereo sound produced by processing a microphone antenna
[0134] - Binaural: particularly when a phone is connected to earphones that incorporate microphones, allowing them to capture natural binaural sound.
[0135] - MASA: format with a single or two transport channels, supplemented by spatial information from scene analysis as described in 3GPP Tdoc S4-220443, "MASA format updates," April 2022, Source: Nokia, Orange. This list is not exhaustive; this example simply demonstrates that the encoding format supported and displayed in SDP for IVAS would be, for example, a list:
[0136] Cf=Stereo,SBA,MASA
[0137] The Mono format is not listed because it is necessarily included as a basic format in the IVAS codec since EVS is included in IVAS and the IVAS payload allows the transport of data encoded by EVS.
[0138] If both terminals A and B have the same encoding / decoding / rendering capabilities, the terminal receiving such an offer could either choose a single format or maintain a list (complete or partial), implying a potential format switch during the call. In reality, it seems more realistic and simpler to select only a stereo or immersive format, possibly with the option to reduce the number of channels within the same format type, particularly during transmission, to maintain a single audio capture configured for the negotiated session and decoupled from the encoding.
[0139] In IVAS, there is no internal signaling within the codec's binary stream to differentiate between STEREO and BINAURAL. Therefore, it is important to be able to negotiate distinct formats (Stereo and Binaural), or at least to signal that the input signal is Binaural, so that the rendering can adapt and the receiving device preferentially accepts BINAURAL only when headphones or earphones are used for direct listening. Methods exist to "reverse" the binauralization before listening on speakers, but these methods do not need to be described here.
[0140] Retaining more than one stereo / immersive format also raises an ambiguity related to the order of formats, which normally indicates a preference (from best to worst). It is assumed that only one stereo / immersive format is selected in SDP by the receiver of an IVAS SDP offering. However, this allows for potentially reducing or increasing the number of channels (e.g., from HOA3 to FOA, or 5.1 to stereo and vice versa) within the limit of the maximum negotiated format (e.g., HOA3 or 5.1). The format selection rules are beyond the scope of this invention. If a change in encoding format is necessary when only one encoded format has been chosen, then an SDP renegotiation is required.
[0141] Before describing various embodiments of the invention, examples are first given illustrating how the IVAS codec could be negotiated at the SDP level. These examples are by no means definitive, as the payload format is still partially open in 3GPP Release 19. However, it is desirable to limit modifications to the IVAS payload defined in Release 18 and, as far as possible, to remain backward compatible with any existing implementations of the current IVAS payload, to avoid any interoperability risks. Appendix 1 describes a (simplified) SDP offering for Terminal A.
[0142] This first example deals with an SDP offering that allows initiating an audio call using the IVAS codec, which includes the EVS codec. The UDP port, AVP profile, and payload type (PT) are indicated in the "m" media line. The payload type numbers are PT=96 for IVAS or 97 for EVS.
[0143] The "rtpmap" and "fmtp" parts indicate the media format parameters.
[0144] It is assumed here that IVAS and EVS are offered in the package with the usual "telephone event" in telephony.
[0145] Annex 2 describes a possible example of SDP offering of terminal A being the negotiation of the coded format (cf) according to the invention.
[0146] Details of the coded "sub-formats" or secondary coded formats (for example SBA varies according to the ambisonic order and the planar or non-planar nature of the representation) are given later with reference to figures 5a, 5b, 5c and 5d.
[0147] According to the invention, the SDP parameters of an IVAS-compatible terminal could be extended in all embodiments described below to include:
[0148] - A declarative indication that the stereo-encoded format actually corresponds to a binaural-encoded format – this can be implemented with a parameter simply named "binaural," which can also be declined into two directional variants, "binaural-send" and "binaural-recv," for the send and receive directions. When this parameter is present, the stereo format should be interpreted as binaural; in one variant, a binary value can be given to this parameter, resulting in "binaural=0" (the default value if the parameter is absent) for non-binaural input formats and "binaural=l" for binaural input formats. However, this approach will be advantageously replaced by the secondary-encoded format specification of the type "cf=Stereo[binaural=l]" described later.
[0149] - A declarative indication that the coded audio stream is diegetic or not - this can be implemented with a parameter simply named "diegetic", with a binary value, to have "diegetic= 1" (default value if the parameter is absent) in the case of a diegetic stream and "diegetic=O" in the case of a non-diegetic stream.
[0150] - A declarative indication that a stereo input audio stream must undergo channel reduction (downmixing) before being encoded to mono by EVS – this can be implemented simply with a parameter named, for example, "dmx," which can also be expressed in two directional variants, "dmx-send" and "dmx-recv," for the send and receive directions. When this parameter is present, the stereo format must be reduced to mono and encoded in EVS; alternatively, this parameter can be given a binary value, resulting in "dmx=0" (the default value if the parameter is absent) for stereo encoding without downmixing, and "dmx=l" for EVS encoding after downmixing a stereo input. When this indication is present, the parameters of the EVS codec included in IVAS must be set appropriately and the IVAS codec must be configured to start in mono mode (EVS) with the parameter "ivas-mode-switch" which could be renamed (e.g. to "ims" or "mono").These SDP parameters allow for finer configuration of a bidirectional or unidirectional RTP session according to the invention. Other extensions of the IVAS payload's SDP parameters are described below (in relation to the encoded formats).
[0151] In cases where only the recorded RTP stream is available, without all the SDP parameters associated with the session, it can be useful to also insert information about whether a stereo stream is binaural or not, or whether an audio stream is diegetic or not, into the payload header. Therefore, embodiments where certain information is also defined in the IVAS payload header are described later.
[0152] The case of the stereo downmix compatible with EVS is a special one, because in all cases the stream is encoded by EVS and seen as a mono stream already covered by the existing IVAS payload; however, it is necessary to indicate the use of this feature during negotiation so that a stereo recording is instantiated before encoding.
[0153] In a first embodiment, the RTP payload format of the current IVAS codec is modified to allow signaling and negotiation of secondary encoded formats, and sending a secondary encoded format request for appropriate rendering.
[0154] A list of 55 secondary formats is given below in Table 4, divided into two columns: one for "simple" secondary formats and one for ISM combinations with MASA or SBA. Note that the ISM format does not have a variant with extended metadata (ex=0) when combined with MASA or SBA. The numbering Td.' in Table 4 provides a possible index for each sub-format from 0 to 54 (inclusive). In some variations, it will be possible to change the order of the formats and / or number the empty cells in formats on the left, which would result in adding the value 9 to all Td.' numbers for formats on the right, whose numbers would then range from 32 to 63 (instead of 23 to 54).
[0155] Table 4: Secondary coded format (excluding mono): 55
[0156]
[0157] Table 4bis also details the modifications to the SDP parameter "cf" that could be defined in an IVAS payload extension to signal a specific sub-format—the asymmetric variants cf-send and cf-recv would, of course, be modified in the same way. These modifications to the SDP parameter introduce additions to each of the already defined encoded formats: Stereo, SBA, MASA, ISM, MC, OMASA, and OSBA. The logic here is to add an addition within square brackets using information dependent on the encoded format, as listed in Table 4ter.
[0158] The 'Id. Sub' numbering in Table 4 bis indicates the relative index of each secondary subformat within a given primary format. As can be seen, this 'Id. Sub' index ranges from 0 to a maximum value of 1, 4, 5, 7, or 23, depending on the format, which can always be represented using 5 bits.
[0159] Table 4bis: Modified SDP parameter “cf” according to the invention
[0160]
[0161] Table 4ter: Indication of secondary format according to the invention
[0162]
[0163] Subsequently, we can equivalently interchange the name of the sub-formats according to their technical name (e.g., FOA), their description by SDP parameter (e.g., SBA(order=l,planar=0]) or allow ourselves editorial differences to use parentheses and spaces rather than brackets without spaces (e.g., SBA (order=l, planar=0)), without there being any ambiguity about the corresponding secondary format.
[0164] Currently, the IVAS E-byte only allows for a primary coded format request (see Figure 2c with ET=01). To signal a secondary coded format request, the IVAS E-byte can be modified as shown in Figure 5a by assigning the currently reserved ET=11 case to a new type of request. Since the E-byte already has the first bit H set to 1 and the following two bits are already reserved for the ET field, only 5 bits remain available after these 3 E-byte header bits, whereas Table 4 lists 55 secondary coded formats, which require 6 bits. According to the invention, the existing primary format request ET=01, where the FMT field ranges from 000 (Stereo) to 110 (OSBA), can therefore be used, and a new type of E-byte for ET=11 can be added as shown in Figure 5a. The sbFMT (for SubFormat) field on the last 5 bits is defined according to (in addition to) the FMT field assumed also to be present.In this case, the sbFMT field indicates only the numbering of the secondary encoded format within a given primary format (starting from 0 up to the number of subformats minus 1), as explained in Table 5 below. Given a given FMT field value, only the first 32 possible 5-bit sbFMT values are used; unused values can be considered "Not Used" (CNot Used) or "Reserved" (res. C reserved). According to this embodiment, the secondary encoded format (or subformat) request therefore uses a different type of E-byte (e.g., ET=11), and the sbFMT field is relative to the FMT field of the format request of one E-byte with ET=01. Thus, secondary format information associated with a primary format for rendering audio data is associated with an additional secondary byte (E-byte) in the transport format header.In this scenario, secondary format information associated with a primary format is inserted into the transport format header by a dedicated request type (ET=11;
[0165] SbFormat Req.) in the secondary extra byte.
[0166] This cleverly allows all possible secondary formats to be defined for each of the primary formats already defined in the E-byte (the additional secondary byte) of the transport format header.
[0167] This minimizes changes to the IVAS codec's RTP payload and maintains backward compatibility with the current format, since the ET=11 code was previously reserved (for future uses).
[0168] This embodiment allows a remote terminal (terminal in the broadest sense) to be asked to change the encoded format to a finer level of detail than the main encoded format. For example, the request with ET=11 allows the ambisonic order to be changed for the purpose of testing the acoustic quality of sound recordings based on the order. In some cases, this request is not expected to be used; for example, a call configured in stereo (FMT=000) does not need to switch to binaural if the recording was not already specified as binaural during the call negotiation.
[0169] The subformat request can also be used by a terminal to request to receive a "less complex" format (with fewer input audio channels) in order to reduce the computational load for sound rendering upon reception.
[0170] Table 5
[0171]
[0172]
[0173] For clarity, Figure 5a only shows secondary formats related to the primary formats stereo (FMT=000), OMasa (FMT=101), and OSBA (FMT=110). Table 5 above lists the secondary formats related to the other primary formats.
[0174] Furthermore, the 'Reserved' (res.) code can be changed to 'Not used'.
[0175] In a second embodiment, illustrated in Figure 5b, rather than using a new E-byte type for the secondary encoded format request, the E-byte type is modified with ET=01 (for the primary format request) by taking one of the reserved bits, for example bit 4, denoted S (S for SubFormat), to indicate whether the E-byte is extended to 2 bytes (16 bits) or left as one byte (8 bits). If S=0, the request fits on 8 bits and remains a primary format request with the FMT field. If S=1, the E-byte is extended to 2 bytes with an additional byte containing 3 reserved bits (8, 9, and 10) and 5 bits denoted sbFMT to indicate the secondary format. The sbFMT (for SubFormat) field on the last 5 bits is defined as a complement to the FMT field given in the first byte.In this case, the sbFMT field only indicates the numbering of the secondary encoded format within a given primary format (starting from 0 up to the number of subformats minus 1) as explained in Table 5 above. Given a given FMT field value, only the first values of the 32 possible 5-bit sbFMT values are used; unused values can be considered 'Not Used' or 'Reserved' (noted res.). 7 In one embodiment, the second byte can contain 2 reserved bits (8 and 9), the sbFMT field is then 6 bits long and directly defines a subformat as shown in Figure 5c. The main problem with this embodiment is that the FMT field becomes redundant when bit S is set to 1.
[0176] This embodiment changes the nature of the E-byte, which can go from 8 to 16 bits; it is not backward compatible with an implementation of the current IVAS codec payload format, because the additional byte containing the sbFMT field would be seen as a ToC byte since it starts with reserved bits which would be set to 0 by default.
[0177] This solution, however, allows defining all possible secondary formats associated with a primary format and leaves the ET=11 code of the E-byte available for other uses. In this embodiment, a value is allocated to a data field of the data format indicating the primary format to indicate the existence or absence of additional secondary format information associated with the audio data. As mentioned above, if this value is 1, it indicates that additional information, in an extra byte, is provided to define the secondary format. The existence of this extra byte is therefore conditional on the value of this field, here the S field of the primary format query.
[0178] In a third embodiment, illustrated in Figure 5c, the AND field of a secondary E-byte (excluding the first E-byte indicating a CMR) is modified to have a variable length for indicating the E-byte type. Bit 1 is defined as a new E-bit; if E=1, then the remaining 6 bits are used to directly define a subformat. Note that 9 out of 64 values are unused; they are marked 'res.' for 'Reserved' in Figure 5c, but they could be prohibited and indicated as 'Not Used'. Furthermore, the order of the subformats is indicative; it would be possible, for example, to directly follow the code sbFMT = 010110, leaving the last values sbFMT = 110111 to 111111 as unused.
[0179] For E=0, the existing AND field is taken up and the remaining bits of an E-byte are taken up as in figure 3c except that the previously reserved bit 3 is now used.
[0180] This embodiment has the advantage of keeping the E-byte on a single byte to indicate a sub-format request.
[0181] In this embodiment, a data field whose value indicates the presence or absence of additional secondary format information associated with the audio data is included in a data field of the secondary byte, in addition to the information indicating the request type of the secondary byte. The additional information indicating the secondary format is stored on a byte linked to this value. Here, it is field E, as described above and illustrated in Figure 5c, that defines whether the secondary format byte is to be taken into account or not.
[0182] This variant is not backward compatible with an implementation of the current IVAS codec payload format, unless we add the constraint that for E=1 the sbFMT = 000000 field (on 6 bits) is unused, because this allowed us to keep an indication of the presence of PI data for this precise combination - this would require shifting the numbering of the sub-formats corresponding to the values sbFMT = 00000 to 010110 shown in Figure 5c by one notch (by adding 1).
[0183] In a fourth embodiment illustrated in Figure 5d, we take up the principle of the sbFMT field (on 5 bits) of Figure 5a and we consider that if the ET=01 format change request of the E-byte is repeated a second time in the same current packet then the second E-byte request contains this time the 5-bit sbFMT field.
[0184] Thus, in this embodiment, secondary format information associated with a primary format is inserted into the transport format header by repeating the corresponding primary format query. This also allows for the definition of all possible secondary formats for each of the primary formats already defined in the first format change query of the E-byte (the secondary supplementary byte) in the transport format header. The second query is therefore linked to the first and thus to the corresponding primary format.
[0185] This implementation is backward compatible with an implementation of the current IVAS codec payload format for the first request, and it exploits the fact that the repetition of a certain type of E-byte is not regulated in the current IVAS payload. There is ambiguity regarding the behavior of an existing terminal if the repetition of the ET=01 field is observed, and regarding the processing associated with the second format request. It is assumed here that this ambiguity is not yet handled in the IVAS payload, thus allowing some flexibility in defining the secondary format request according to this method and, at the same time, clarifying how to handle a potential repetition of a certain type of E-byte in IVAS.
[0186] In Figure 5d, the two requests are assumed to follow each other immediately, which explains why the bits are numbered consecutively from 0 to 15. In some variations, the two E-bytes of type ET=01 could be non-contiguous in the payload header; however, it is important that the first E-byte corresponds to a (primary) format request in order to interpret the related secondary format information in the second E-byte. Figure 6a illustrates a method for reading the IVAS codec transport format according to the first embodiment described in Figure 5a.
[0187] For each byte read from the IVAS codec transport format header (E600), the first bit (H) is read (E601) to determine whether the byte is a ToC byte as defined above or an E-byte. We are only concerned here with reading E-bytes, i.e., when 1-1=1. The reading of the ToC byte (E602) remains unchanged from that performed in the IVAS Release 18 specification (Annex A of 3GPP TS 26.253) cited previously. If the F bit of the ToC byte is F=1 (E611), then the reading of the next header byte continues (E600); otherwise, it is the last byte in the payload header, and the header byte reading is complete.
[0188] Similarly, for reading E-byte type bytes, the first byte (O in E603) of this type is the one carrying a CMR type request as defined in Annex A of 3GPP TS 26.253 Release 18, we will not describe its reading (E604) further here.
[0189] For secondary E-bytes (N in 603) read in E605, the ET field is first read in E606; it defines the type of information transported in that byte. In the case of the transport format illustrated in Figure 5a, an additional ET value is read. When the ET value is 00, the coded audio tape request byte (Bw Req.) is read (E607) as specified in Figure 2c; this request remains unchanged. When the ET value is 01, a coded format request (Format Req.) is read (E608) as specified in Figure 2c. The primary format is then specified according to the FMT code read. This primary format will then allow the specification of a secondary format for ET=11. When the ET value is 10, the indication of the presence of a "PI" section (PI indic.) is read (E609) according to the format specified in Figure 2c or according to the format specified in Figure 7a. In the latter case, several indications can be read as shown later with reference to Figure 7a.
[0190] When the ET value is 11, the secondary format (E610) is read by the value given to sbFMT as described with reference to Figure 5a. This value depends on the FMT value read previously.
[0191] Figure 6b illustrates a method for reading the IVAS codec transport format according to the second embodiment described in Figure 5b.
[0192] As in Figure 6a, we are only interested here in reading an additional secondary byte in E605. As in Figure 6a, the AND field read in E606 defines the type of information carried in this byte. When the value ET=00, the coded audio tape request byte (Bw Req.) is read (E607) as specified in Figure 2c; this request remains unchanged. When the value ET=01, the value of the S field, as defined with reference to Figure 5b, determines the reading of the coded format request. Thus, the value of S is read in E612. In the case where S=0, the format request is read in E614; it is 8 bits long. The primary format is then specified according to the FMT code read. If S=1, two bytes are read in E615. The first byte specifies the primary format according to the FMT value read, the second byte specifies the secondary format according to the value read in the sbFMT field as described with reference to Figure 5b.This value therefore depends on the FMT value read in the previous byte.
[0193] When the ET value is 10, the indication of the presence of a "PI" section (PI indic.) is read (E609) according to the format specified in Figure 2c or according to the format specified in Figure 7a. In the latter case, several indications can be read as shown later with reference to Figure 7a.
[0194] Figure 6c illustrates a method for reading the IVAS codec transport format according to the third embodiment described in Figure 5c.
[0195] As with figures 6a and 6b, we are only interested here in reading an additional secondary byte in E605.
[0196] Here, the value of the E field read in E620 determines the existence of a secondary format request associated with a primary format. Thus, if E=1, then a secondary format (E622) is read by the value given to the sbFMT of the byte, as described with reference to Figure 5c. Otherwise, for E=0, the ET field is read (E620). When the ET value is 00, the coded audio tape request byte (Bw Req.) is read (E607) as specified in Figure 2c. When the ET value is 01, a primary format request (E608) is read as specified in Figure 2c. The primary format is then specified according to the FMT code read. When the ET value is 10, the indication of the presence of a "PI" section (PI indic.) is read (E609) according to the format specified in Figure 2c or according to the format specified in Figure 7a. In the latter case, several indications can be read as shown later with reference to Figure 7a.
[0197] Figure 6d illustrates a method for reading the IVAS codec transport format according to the fourth embodiment described in Figure 5d.
[0198] As with figures 6a and 6b, we are only interested here in reading an additional secondary byte in E605.
[0199] In this embodiment, as in Figure 6a, the value of the ET field read in E606 defines the type of request or information. When the ET value is 00, the coded audio tape request byte (Bw Req.) is read (E607) as specified in Figure 2c. When the ET value is 10, the indication of the presence of a "PI" section (PI indic.) is read (E609) according to the format specified in Figure 2c or according to the format specified in Figure 7a. In the latter case, several indications can be read, as shown later with reference to Figure 7a. When the ET value is 01, E630 checks whether it is a first primary format request. In the case of a first request (O in E630), the primary format request is read (E608) as specified in Figure 2c. The primary format is then specified according to the FMT code read.
[0200] In the case of a second query with an ET=01 field (N in E630), following a first query with the same ET=01 value, this means that the second query contains secondary format information. This secondary format is then read (E631) by the value given to the sbFMT of the byte, as described with reference to Figure 5d.
[0201] In another embodiment, the current IVAS codec RTP payload format is modified to allow signaling, via an E-byte, that the audio data encoded and transported in the current packet was generated using IVAS split rendering, without requiring modification of the existing ToC byte definition. Furthermore, the E-byte indicating split rendering mode can also include additional information regarding the diegetic or non-diegetic nature of the audio stream, or even other rendering-related information (for example, that the input signal to the IVAS codec is actually binaural).
[0202] Figure 7a defines a modification of the existing E-byte (excluding the first E-byte reserved for CMR) indicating the presence of PI data with ET=10, to utilize the reserved bits. Bits 4 to 7 are defined as follows:
[0203] - S (1 bit): S=0 when the audio data encoded in the current packet is not generated by split rendering, S=1 when the audio data encoded in the current packet is generated by split rendering
[0204] - ID (2 bits): as in proposal S4-241961
[0205] - C (1 bit): indication of split rendering codec used (for example, 0 for LDLC or 1 for LC3plus)
[0206] - P (1 bit): indication of the presence of PI data
[0207] The order of these fields is only indicative and these fields can be interchanged. In variants the 2-bit ID field could be reduced to a 1-bit D field (where D=0 corresponds to ID=01 and D=1 to ID=10) - in this case bits 3 and 4 (P and C) are shifted to bits 4 and 5, the D field is for example at bit 6 and bit 3 becomes 'reserved'.
[0208] When the bit S=l, the BR field of the ToC bytes is limited to the rates allowed by the split rendering.
[0209] Thus, rather than using an E-byte with ET=10 without any additional data, solely to indicate the presence of PI data, this E-byte is extended here to also indicate the split rendering mode and associated information. This mode allows the reuse of the E-byte corresponding to the indication of the presence of PI data (which otherwise contains 5 unused bits) to avoid wasting a ToC byte to indicate the split rendering mode.
[0210] An IVAS-compatible terminal that implements the current IV AS payload could ignore bits 3 to 7 defined in the E-byte in Figure 6a, but it would also not be able to interpret the SDP parameters required for split rendering ("sr-dof", "sr-tc") and proposed in S4-241961; therefore, this embodiment illustrated in Figure 6a can be considered to have no negative impact on backward compatibility.
[0211] Thus, in this embodiment, mode information for audio data rendering is associated with an additional secondary byte (E-byte excluding the first E-byte reserved for the CMR) in the transport format header, specifically information about the existence of split rendering. In this embodiment, the information indicating an audio data rendering mode (split rendering (SR indic.)) is associated with the information indicating the existence of metadata related to audio data rendering (PI indic.) in the transport format header.
[0212] Since the PI section was defined in the IVAS payload to carry information related to "rendering" (rendering or playback of audio data) and to send requests or feedback, it is therefore appropriate to define here a new category of data "PI" (Processing Information) to indicate a split rendering request which corresponds well to a mode of rendering audio data.
[0213] This embodiment can be associated with the different embodiments described in Figures 5a, 5b, 5c, and 5d. Thus, the additional secondary byte in the IVAS codec transport format header allows for the specification of both a secondary audio data format and a specific rendering mode for this audio data, without altering the initial functions of these formats.
[0214] In a variant illustrated in Figure 7b, the E-byte type with ET=10 remains unchanged and indicates the presence of PI data (or section). Another E-byte type is defined with ET=11, where bit 3 is left as reserved and the remaining 4 bits are used for a general rendering indication, including split rendering. Since split rendering is a separate mode of IVAS, bit 3 is defined as:
[0215] - S (1 bit): S=0 when the audio data encoded in the current packet is not from split rendering, S=1 when the audio data encoded in the current packet is from a split rendering mode. Depending on the value of this bit S, two variants can be defined for the DESC (for Description) field corresponding to bits 4 to 7.
[0216] If S=1:
[0217] - C (1 bit): indication of split rendering codec used (LDLC or LC3plus)
[0218] - ID (2 bits) as in S4-241961
[0219] - In variants the 2-bit ID field could be reduced to a 1-bit D field (where D=0 corresponds to ID=01 and D=1 to ID=10) - in this case bits 4 and 5 are reserved and the C field can be shifted to bit 6 to leave the D field at bit 7.
[0220] If S=0:
[0221] - D (1 bit): indication whether the stream is of the diegetic type or not
[0222] - B (1 bit): Binaural format indication (only valid for stereo format). This variant does not allow combining the rendering indication (including split rendering) with certain embodiments described above, particularly the embodiment described in Figure 5a for specifying sub-formats, because it leaves no "free" E-byte type. However, this embodiment can be associated with the embodiments described in Figures 5b, 5c, and 5d. In this case, the request type ET=11 is modified as shown in Figure 7b.
[0223] The particular interest in reporting the binaural case, if the binaural flag B is 1 in the IVAS payload, is that the receiving terminal does not make a direct presentation on headphones and can "suppress" binauralization in case of playback on speakers.
[0224] In all variants of this embodiment, it is assumed that a split rendering query initially proposed in S4-241961 with an E-byte becomes an additional PI data type as illustrated in Figure 7c, where code 10101 ("bit type") is taken here as an example. The fields D, Y, P, R are those of Figure 3c (for the case ET=11). They are defined as follows:
[0225] - D (1 bit): Bit D controls whether the flow should be linked to head movement or not - Y (1 bit): Enables or disables position correction (pose) along the z-axis ("yaw").
[0226] - P (1 bit): Enables or disables position correction (pose) along the y-axis ("pitch").
[0227] - R (1 bit): Enabling or disabling position correction (pose) along the x-axis ("roll"). Figure 8 now describes an example of a hardware implementation of a communication device, for example a communication terminal implementing the data transport format generation method according to the invention, as described with reference to Figures 5a to 6c described above.
[0228] The TA terminal includes a storage space 11, for example a MEM memory, a processing unit 10 comprising a processor P, controlled by a computer program PG, stored in the memory 11 and implementing the transport method according to the invention.
[0229] At initialization, the code instructions of the program PG are, for example, loaded into a RAM memory (not shown) before being executed by the processor P of the processing unit 10. The program instructions can be stored on a storage medium such as a flash memory, a hard drive or any other non-transient storage medium.
[0230] The processing processor implements the transport process as described with reference to Figure 4, and according to the formats described with reference to Figures 5a to 6c, according to the PG program instructions.
[0231] The TA terminal includes a communication module 12 capable of receiving and transmitting real-time data to a communication network R and capable of reading SDP type signaling or configuration data 13 either through the signaling emitted in the network or in the terminal's memory 11.
[0232] The terminal includes a receiver module 14 capable of receiving and decoding a transport format as generated by a transmitter of another communication device and as described with reference to Figures 5a to 5d and / or 6a to 6c. For this purpose, it extracts from the fields of the RTP header, in particular the mode and / or secondary format information associated with a primary format for rendering audio data, present in an additional secondary byte (E-byte) of the transport format header.
[0233] The receiver module is also capable of managing a receive memory, decoding received frames and correcting any frame losses.
[0234] The terminal includes a transmitter module capable of encoding and generating RTP packets by adding protocol headers to generate a payload format corresponding to the encoding format. The transmitter also has a function for optionally adapting the transport (aggregation, redundancy, or other) as well as a function for transmitting packets.
[0235] Thus, the transmitter module implements a process for generating an audio data transport format during a real-time communication session with another communication device. For this, the data transport format for an immersive audio codec (IVAS) is such that secondary mode and / or format information associated with a primary format for rendering audio data is associated with an additional secondary byte (E-byte) of the transport format header.
[0236] Several embodiments of this data transport format are described with reference to figures 5a to 5d and / or 6a to 6c.
[0237] The term module can refer to a software component, a hardware component, or a set of hardware and software components. A software component itself corresponds to one or more computer programs or subprograms, or more generally, to any element of a program capable of implementing a function or set of functions as described for the modules in question. Similarly, a hardware component corresponds to any element of a hardware assembly capable of implementing a function or set of functions for the module in question (integrated circuit, smart card, memory card, etc.). A terminal is, for example, a telephone, a smartphone, a tablet, a computer, a home gateway, or a connected device. APPENDIX 1
[0238]
[0239] APPENDIX 2
[0240]
Claims
DEMANDS 1. Method for generating audio data transport format for an immersive audio codec (IVAS) during a real-time communication session between communication devices (A, B), wherein secondary format information associated with a primary format, for rendering audio data, is associated with an additional secondary byte (E-byte) of the transport format header.
2. A method according to claim 1, wherein the secondary format information is inserted into the transport format header by a dedicated request type in the secondary supplementary byte.
3. A method according to claim 1, wherein a value is allocated to a data field of the data format indicating the primary format to inform of the existence or not of additional secondary format information associated with the audio data.
4. A method according to claim 3, wherein when the value indicates the existence of additional secondary format information, this additional secondary format information is given in an additional byte following the byte indicating the primary format.
5. A method according to claim 1, wherein a data field whose value indicates the existence or not of additional secondary format information associated with the audio data is included in a data field of the secondary supplementary byte in addition to the information indicating the type of request of the supplementary byte.
6. Method according to claim 1, wherein the secondary format information is inserted into the header of the transport format by a repetition of the corresponding primary format query.
7. A method according to claim 1, wherein information indicating a mode of rendering audio data is further associated with information indicating the existence of metadata associated with the rendering of audio data from the transport format header.
8. Method according to claim 1, wherein information indicating a mode of rendering audio data is further inserted into the header of the transport format by a dedicated request type in the secondary supplementary byte.
9. Communication terminal comprising a processing circuit capable of implementing a method for generating an audio data transport format for an immersive audio codec (IVAS) during a real-time communication session with a communication device, wherein secondary format information associated with a primary format, for rendering audio data, is associated with an additional secondary byte (E-byte) of the transport format header.
10. Audio data transport format for an immersive audio codec (IVAS), in which secondary format information associated with a primary format, for rendering audio data, is associated with an additional secondary byte (E-byte) of the transport format header.
11. Computer program comprising code instructions for implementing the steps of the generation process according to any one of claims 1 to 8, when these instructions are executed by a processor.
12. A processor-readable storage medium storing a computer program containing instructions for executing the generation process according to any one of claims 1 to 8.