Method and apparatus for negotiating conversational immersive audio sessions

The method enhances IVAS codec negotiation by selecting optimal input formats and bitrates through SDP, addressing limitations in current codecs for immersive audio sessions, ensuring compatibility and reducing computational load.

JP2026509781APending Publication Date: 2026-03-25NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current conversational audio codecs, such as IVAS, lack mechanisms for negotiating input and output formats, rendering capabilities, and rate-adaptive processes, limiting flexibility and optimization for immersive audio sessions.

Method used

Implement a method for selecting preferred, mutually supported input formats through session description protocol (SDP) negotiation, allowing UEs to exchange and agree on optimal input formats and bitrate adaptations for immersive conversational speech codecs.

Benefits of technology

Enables flexible and optimized selection of input formats and bitrates for immersive audio sessions, ensuring compatibility with external renderers and reducing computational load by allowing UEs to agree on suitable audio bitstreams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509781000001_ABST
    Figure 2026509781000001_ABST
Patent Text Reader

Abstract

Negotiation for a conversational immersive audio session includes obtaining one or more supported immersive conversation codec input formats, providing one or more immersive conversation codec input formats as a list in a predetermined order, including immersive conversation codec input format attributes in a session description file, substituting the list of immersive conversation codec input formats into the input format attributes in the session description file, generating a session negotiation offer based on the session description file, and sending the session negotiation offer to the receiving user device. Non-limiting exemplary embodiments may include receiving a session negotiation offer, parsing one or more immersive conversation codec input formats from a session description file in the received session negotiation offer, obtaining one or more supported preferred input formats from the receiving user device, selecting one or more preferred input formats from the transmitting user device that are common with one or more supported preferred input formats from the receiving user device in the received session negotiation offer, substituting the selected one or more preferred input formats into the immersive conversation codec input format attribute in the session negotiation answer, and sending the session negotiation answer to the transmitting user device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary non - limiting embodiments generally relate to multimedia transport, and more specifically, to methods and apparatuses for negotiating a conversational immersive audio session.

Background Art

[0002] In a multimedia system, it is known to perform data compression and decoding.

Summary of the Invention

[0003] The above aspects and other features will be described in the following description in connection with the accompanying drawings.

Brief Description of the Drawings

[0004] [Figure 1] It is a block diagram showing one possible non - limiting system capable of implementing an exemplary embodiment. [Figure 2] It is a diagram showing an exemplary apparatus configured to implement the examples described herein. [Figure 3] It is a diagram showing an example of a non - volatile memory medium used to store instructions for implementing the examples described herein. [Figure 4] It is a diagram showing an exemplary method of a transmission apparatus based on the examples described herein. [Figure 5] It is a diagram showing an exemplary method of a receiving apparatus based on the examples described herein. [Figure 6] It is a diagram showing an example of a conversational immersive audio session between two parties.

Modes for Carrying Out the Invention

[0005] In this specification, methods and apparatuses for negotiating a conversational immersive audio session will be described.

[0006] 3GPP IVAS input format The Immersive Voice and Audio Services (IVAS) codec is an extension of the 3GPP EVS codec, intended for new immersive voice and audio services over 4G / 5G networks. Such immersive services include, for example, immersive voice and audio for virtual reality (VR). This multi-purpose audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. It is expected to support various input formats, including channel-based and scene-based inputs. It is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions. Currently, the IVAS codec standardization is expected to be completed in 2023 as part of Release 18.

[0007] Supported input formats for IVAS include stereo, multi-channel, object-based audio, scene-based audio, and MASA. Furthermore, several combinations can be supported through various means; for example, the Object with MASA (OMASA) composite format has been proposed. Additionally, a separate input format exists for binaural audio, which currently operates similarly to stereo input. For mono input, the 3GPP EVS codec is used. For a detailed algorithmic description of EVS, please refer to TS26.445.

[0008] Stereo input refers to an audio representation where two audio channels are assigned to the left and right audio channels.

[0009] Multi-channel (MC) input refers to an audio representation where each transport channel represents an audio signal for speakers surrounding the listener. IVAS supports surround formats 5.1 and 7.1, as well as surround formats 5.1.2, 5.1.4, and 7.1.4 with elevated speaker placement.

[0010] Object-based audio, or Independent Streams with Metadata (ISM) input, refers to an audio representation in which individual mono audio object streams are transmitted. In addition to the audio being transported, metadata describing the audio object is transmitted, which is expected to be the azimuth and elevation angles of the audio object (ISM).

[0011] Scene-based audio (SBA) input refers to an ambisonics-based audio representation. The ambisonics signal carries a representation of the audio scene, in which case the transport channels refer to the capture directions in a spherical domain. The first channel (W) represents omnidirectional capture, the incoming sound field from all directions. The next three channels (X, Y, Z) represent incoming sound from the corresponding spatial axis. These four channels form a first-order ambisonics (FOA) representation. Higher spatial accuracy can be achieved by increasing the number of capture directions with more channels. This increases the order of the ambisonics representation and is called higher-order ambisonics (HOA). Second-order ambisonics (HOA2) includes nine channels, and third-order (HOA3) includes sixteen channels. IVAS supports first-order, second-order, and third-order ambisonics.

[0012] MASA, or metadata-assisted spatial audio, refers to a parametric spatial audio representation defined for IVAS in Permanent Document IVAS-4 (Tdoc S4-221619). It uses an audio signal along with corresponding spatial metadata (including, for example, the direct-to-total energy ratio in the direction and frequency bands). A MASA stream can be obtained, for example, by capturing spatial audio using a mobile device's microphone, in which case the set of spatial metadata is estimated based on the microphone signal. MASA streams can also be obtained from other sources, such as specific spatial audio microphones (e.g., ambisonics), studio mixes (e.g., 5.1 mixes), or other content using appropriate format conversion. It is also possible to use the MASA tool within a codec for encoding multi-channel signals by converting a multi-channel signal to a MASA stream and then encoding that stream.

[0013] OMASA refers to MASA with additional object-based audio (1-4 objects). The object-based audio stream is supplied to the encoder as a separate stream from the MASA stream. OMASA is being proposed as part of IVAS, but is not yet officially part of the IVAS codec baseline.

[0014] IVAS bitrate Table 1 shows the currently supported bitrates for each IVAS input format. Table 2 shows the current continuous bitrate ranges for the IVAS input modes. It should be noted that the SBA, MC, MASA, and OMASA input modes are supported for a wide continuous bitrate range from 13.2 kbps (kilobits per second) to 512 kbps. The stereo input mode is supported up to 256 kbps. The supported bitrate range for ISM depends on the number of objects. The ISM input mode is supported up to 128 kbps (1 object), 256 kbps (2 objects), 384 kbps (3 objects), and 512 kbps (4 objects). (Exact values may change until the specification is finalized. Also, OMASA is currently under proposal and is not officially part of the IVAS codec baseline yet.)

[0015] Table 1 shows the supported input modes of IVAS and the supported bitrates for each mode.

[0016]

Table 1

[0017] Table 2 shows the IVAS input modes supported for the continuous bitrate range.

[0018]

Table 2

[0019] IVAS Output Format The IVAS output formats for IVAS include monaural, stereo, multi-channel (including custom speaker layouts), FOA, HOA2, HOA3, and binaural. Additionally, so-called pass-through operations, which enable MASA output for, for example, MASA input, are available. Binauralized audio can use default or custom HRIRs and room effects (BRIRs). Multi-channel output refers to rendering where multiple audio channels are rendered for the playback system (e.g., surround speaker settings 5.1, 7.1, 5.1.2, 5.1.4, or 7.1.4).

[0020] Scene-based audio output rendering (FOA, HOA2, HOA3) refers to rendering where the input stream is decoded and rendered to the corresponding ambisonic channels.

[0021] Binaural rendering renders the output binurally to the receiver through headphones. The binaural room output mode applies the room impulse response to the output signal.

[0022] RTP - Real-Time Transport Protocol RTP is intended for end-to-end real-time transfer of streaming media and provides provisions for jitter compensation and detection of packet loss and out-of-order delivery. RTP enables data transfer to multiple destinations via IP multicast or to a specific destination via IP unicast. Most RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols can also be used. RTP is used together with other protocols such as H.323 and the Real-Time Streaming Protocol RTSP.

[0023] The RTP specification describes two protocols, RTP and RTCP. RTP is used for the transfer of multimedia data, while its companion protocol (RTCP) is used for the periodic transmission of control information and QoS parameters.

[0024] RTP sessions are typically initiated between a client and a server, or between one client and another (or in a multi-party topology), using a signaling protocol such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols typically use the Session Description Protocol (RFC8866) to specify parameters for the session.

[0025] RTP Profile and Payload Format RTP is designed to transmit a large number of multimedia formats, enabling the transport of new formats without revising the RTP standard. For this purpose, information required by a specific use of the protocol is not included in the generic RTP header. RTP profiles can be defined for classes of use (e.g., audio, video). Associated RTP payload formats can be defined for media formats (e.g., specific video coding formats). Every instantiation of RTP in a particular use may require a profile and payload format specification.

[0026] The profile defines the codec used to encode the payload data and its mapping to the payload format code within the Payload Type (PT), which is the protocol field in the RTP header.

[0027] For example, an RTP profile for audio and video conferencing with minimal control is defined in RFC 3551. This profile defines a set of static payload type assignments and a dynamic mechanism for mapping between payload formats and PT values ​​using the session description protocol (SDP). The latter mechanism is used for newer video codecs, such as the RTP payload format for H.264 video defined in RFC 6184, or the RTP payload format for High Efficiency Video Coding (HEVC) defined in RFC 7798.

[0028] RTP-RTP session An RTP session is established for each multimedia stream. Audio and video streams can use separate RTP sessions, allowing the receiver to selectively receive components of a particular stream. The RTP specification also recommends the use of even port numbers for RTP and the following odd port numbers for the associated RTCP session. For applications that multiplex the protocol, a single port can be used for both RTP and RTCP.

[0029] Each RTP stream consists of RTP packets, and each RTP packet consists of a pair of RTP headers and payloads.

[0030] SDP - Session Description Protocol The Session Description Protocol (SDP) is a format for describing multimedia communication sessions for announcement and invitation purposes. Its primary use lies in supporting streaming media applications. While SDP does not deliver any media streams themselves, it is used between endpoints for negotiating network metrics, media types, bandwidth requirements, and other associated characteristics. This set of characteristics and parameters is called a session profile. SDP is extensible to support new media types and formats. SDP is widely adopted in the industry and is used for session initialization by various other protocols, such as SIP or WebRTC-related session negation.

[0031] The session description protocol describes a session as a group of fields in a text-based format, with one field per line. The form of each field is as follows:

[0032] [Table 3]

[0033] Here, <character>is a single, case-sensitive character. <value>This is structured text with a format that depends on the character. The value is typically UTF-8 encoded. Whitespace immediately on either side of the equals sign is not allowed.

[0034] A session description consists of three parts: a session description, a timing description, and a media description. Each description can contain multiple timing descriptions and media descriptions. Names are unique only within the associated syntactic structure.

[0035] The fields must appear in the order shown, and optional fields are marked with an asterisk.

[0036] [Table 4]

[0037] Time description (required)

[0038] [Table 5]

[0039] Media description (optional)

[0040] [Table 6]

[0041] The following is a sample session description from RFC4566. This session is initiated by user "jdoe" at IPv4 address 10.47.16.5. Its name is "SDP Seminar" and it includes extended session information ("Seminar on Session Description Protocol") along with a link and email address for contacting the person in charge, Jane Doe. This session is specified to last for 2 hours using an NTP timestamp, at a connection address specified as IPv4 224.2.17.12 with a TTL of 127 (indicating the address the client needs to connect to, or, if a multicast address is provided as in this case, the address it needs to subscribe to). The recipient of this session description is instructed to receive media only. Two media descriptions are provided, both using RTP audio / video profiles. The first is an audio stream on port 49170 using RTP / AVP payload type 0 (defined as PCMU by RFC3551), and the second is a video stream on port 51372 using RTP / AVP payload type 99 (defined as "dynamic"). Finally, there is an attribute that maps RTP / AVP payload type 99 to format h263-1998 with a clock rate of 90kHz. The RTCP ports for the audio and video streams on 49171 and 51373 are implicitly indicated, respectively.

[0042] [Table 7]

[0043] SDP-Attribute SDP uses attributes to extend the core protocol. Attributes can appear within session or media sections and are scoped accordingly as session-level or media-level. New attributes can be added to the standard by registering with IANA. A list of all registered attributes can be found at https: / / www.iana.org / assignments / sdp-parameters / sdp-parameters.xhtml#sdp-att-field. A media description can contain any number of "a=" lines (attribute fields) that are specific to the media description. Session-level attributes convey additional information that applies to the session as a whole, rather than to individual media descriptions.

[0044] Attributes are characteristics or values ​​such as the following:

[0045] [Table 8]

[0046] Examples of attributes defined in RFC8866 are "rtpmap" and "fmtp".

[0047] The "rtpmap" attribute maps the RTP payload type number (as used in the "m=" line) to the encoding name indicating the payload format being used. It also provides information about the clock rate and encoding parameters. Up to one "a=rtpmap:" attribute can be defined for each specified media format. Thus, it could look like this:

[0048] [Table 9]

[0049] In the example above, the media types are "audio / EVS" and "audio / L16".

[0050] The parameters added to the "a=rtpmap:" attribute should only be those necessary for the session directory to select the appropriate media to join the session. Codec-specific parameters should be added to other attributes, such as "fmtp".

[0051] The "fmtp" attribute allows parameters specific to a particular format to be communicated in a way that the SDP does not need to understand them. The format must be one of the formats specified for that media. The semicolon-separated format-specific parameters can be any set of parameters that need to be communicated by the SDP and are given unchanged to media tools using this format. A maximum of one instance of this attribute is allowed for each format. An example is shown below.

[0052] [Table 10]

[0053] For example, EVS defines the following ptime, maxptime, evs-mode-switch, hf-only, dtx, dtx-recv, max-red, channels, cmr, br, br-send, br-recv, bw, bw-send, bw-recv, ch-send, ch-recv, and ch-aw-recv.

[0054] Conversational audio codec session negotiation is currently limited to scenarios where rendering audio with different inputs is not possible. Typically, the audio input is limited to mono audio.

[0055] The following issues motivate the definition of new session negotiation parameters for IVAS.

[0056] 1. The upcoming IVAS standard will support a number of input formats. The input format has a clear impact in scenarios where the receiving end intends to use an external renderer that depends on a specific input format. For example, if an external renderer in the receiving UE depends on the MASA format for rendering, the receiving UE needs to negotiate the input format from the transmitting UE to be the MASA input format for the IVAS codec. Session negotiation for conversational audio codecs does not have a mechanism to indicate such functionality. In short, it is clearly necessary to reach an input format that can be agreed upon by both parties.

[0057] 2. The receiving UE may prefer a specific input format that is more suitable for a particular type of rendering of the output. For example, if the receiving UE intends to render IVAS output through a speaker system, while the MASA format is suitable for MASA-based head-track rendering via Nokia's proprietary external renderer, then a channel input format might be more appropriate. Thus, depending on the output scenario, the receiving UE may want to negotiate the most appropriate input format.

[0058] 3. It is necessary to indicate the renderer's output format so that the transmitting UE encoder can generate an IVAS bitstream that is optimized for a specific output format.

[0059] 4. It is necessary to inform the rate-adaptive process that the transmitting UE has the flexibility to switch between different input modes and bitrates.

[0060] The examples described herein address the above-mentioned problems.

[0061] The EVS codec (as presented in 3GPP TS26.445) operates only in mono mode and does not support other input or output formats. The codec can operate in EVS Primary or EVS AMR-WB IO mode. During session negotiation, the operating mode used at the start of the session is indicated by the evs-mode-swith parameter, with the allowed values ​​being 0 (EVS Primary) and 1 (EVS AMR-WB IO). This parameter reflects the mode used for the codec, not the input format, and the mode is mono in both cases. Therefore, a similar parameter cannot be used for IVAS that supports multiple input formats.

[0062] The EVS AMR-WB IO mode may have a mode-set parameter set during session negotiation. This parameter contains a list of supported operating modes for the codec during the session. The mode is selected from a table that indicates different bitrate operating modes for the EVS AMR-WB IO. In EVS, different codec modes are uniquely described by the bitrate used. For example, a bitrate of 8kbps is reserved only for the EVS primary mode, and the EVS AMR-WB IO does not have an operating mode at this bitrate. In IVAS, the same bitrate can operate for multiple different input formats, so a similar unique identification of operating modes based on bitrate is not possible.

[0063] The EVS codec does not support anything other than mono input, and the negotiation requirements related to multiple input formats and the resulting rate-adaptive negotiation techniques are not covered by prior art. Similarly, EVS or any other codec currently does not support defining output formats in the case of conversational audio session negotiation. Multi-mono operation is possible using multiple instances of the EVS encoder and decoder. For example, there is no standard mechanism for synchronizing two or more instances at the signal level.

[0064] The embodiments described herein relate to immersive speech and audio service codec session negotiation, which provides a method for selecting preferred, mutually supported input formats for an immersive conversational speech codec in order to realize the ability to select an input format that is optimal for at least one of an external renderer or a preferred output format. This is provided and performed by the transmitting UE and the receiving UE.

[0065] Transmitter UE • Retrieve one or more supported input formats • Sort input formats in a preferred order (for example, based on encoding computational complexity or UE's audio capture capabilities). • Include the immersive conversation codec input format attribute in the session description file. • In the session description file, assign a sorted list of immersive conversation codec input formats to the input format attribute. • Generate a session description offer • Send a session negotiation offer to the receiving UE.

[0066] Receiver UE • Receive a session negotiation offer • Parse the immersive conversation codec input format from the session description file received within the offer. • Obtain the preferred input format supported by the receiving UE. • From the received session description offer, select one or more preferred input formats that are common with the preferred input format of the receiving UE. • Include immersive conversation codec input format attributes • Substitute one or more preferred input formats into the Immersive Conversation Codec Input Format attribute. • Send the session negotiation answer to the sending UE.

[0067] The embodiments described herein relate to immersive speech and audio service codec session negotiation, which provides a method for selecting preferred mutually supported input formats for an immersive conversational speech codec in order to enable the function of selecting the optimal input format for the transmitting UE while serving an audio bitstream suitable for the receiving UE.

[0068] In one embodiment, the session description file includes, as attributes or media format parameters, input format indication, output format indication, and format switching for bitrate adaptation.

[0069] In one embodiment, the session negotiation description is represented as a session description protocol file (SDP).

[0070] In one embodiment, session negotiation is performed as a session description offer-answer model.

[0071] In another embodiment, if the response from the receiving UE includes a single input format, the transmitting UE bitrate adaptation is limited to that single input format.

[0072] In another embodiment, if the response from the receiving UE includes two or more input formats, the transmitting UE bitrate adaptation has the flexibility to use any of the agreed input formats.

[0073] In one embodiment, the codec format is indicated as an Immersive Speech and Audio Codec (IVAS) for Session Negotiation, having a bitstream restricted to an Immersive Speech Audio Codec bitstream and not including an Enhanced Speech Codec (EVS) bitstream for rate adaptation.

[0074] In another embodiment, the input format is included in the session description file as a media format parameter or attribute, along with the corresponding codec parameters, which are immersive audio and audio codec parameters.

[0075] The use of the EVS codec in the session description results in a fallback to TS26.445, which also covers EVS interoperability with AMR-WB IO mode. In the case of the IVAS payload format, IVAS is used as the codec name in the session description.

[0076] Table 3 shows that different IVAS input formats are assigned to the numerical value of the parameter `inf`. This parameter will be explained in more detail below.

[0077] Table 3 shows the IVAS input format and its assigned inf attribute value.

[0078] [Table 11]

[0079] Table 4 shows that different IVAS output formats are assigned to the numerical values ​​of the parameter `outf`. This parameter is explained in more detail below. The external output format indicates an (unrendered) output. For example, the receiver is using an external renderer. Passthrough is listed as a type of output format, but it is actually an operation of the IVAS codec that results in the same output format attributes as the input format, or in other words, the input stream format is preserved in the output. Similarly, external mode means that external rendering is required to listen to the audio, but it does not necessarily mean that an external renderer is used for that purpose.

[0080] Table 4 shows the IVAS output format and its assigned outf attribute value.

[0081] [Table 12]

[0082] Table 5 shows the specific operating modes for each IVAS output format, parameterized by the outf-specific-mode parameter. For example, an external renderer can use the outf-specific-mode parameter to indicate which output format it is using. Table 6 shows another embodiment in which the specific operating mode of the output format is included in the outf parameter, and in that case the outf-specific-mode parameter is not required. The outf-specific-mode parameter is described in more detail below.

[0083] Table 5 shows the IVAS output formats and their specific operating mode attribute, outf-specific-mode. XX indicates that the output format does not have a specific operating mode.

[0084] [Table 13]

[0085] Table 6 shows the IVAS output formats and their assigned outf attribute values. This table incorporates outf-specific-mode into the output format.

[0086] [Table 14]

[0087] `inf`: Indicates the input format functionality. If a range of input formats are supported, it is indicated by the first and last input formats in that range, separated by a hyphen (inf1-inf2). For multiple input formats that are individual rather than in a contiguous range, they are listed as comma-separated values ​​(inf1,inf2). Comma-separated values ​​are also used when the input formats are within a range, but the preferred order of those formats is not a default contiguous range. In both cases, the input formats are listed in a hyphen- or comma-separated list, in preferred order from the most preferred to the least preferred. `inf-send` and `inf-recv` are used when different input formats are used in both the send and receive directions. If the `inf` parameter is absent, all possible IVAS input formats are supported.

[0088] disable-inf-switch: A flag defined to restrict input format switching. The allowed values ​​are 0 and 1. If disable-inf-switch is 0 or does not exist, the sender can switch between negotiated IVAS input formats. If disable-inf-switch is 1, the sender is not allowed to switch IVAS input formats during a session.

[0089] `outf`: Indicates the output format functionality. If a range of output formats are supported, it is indicated by the first output format in that range and the last output format in that range, separated by a hyphen (outf1-outf2). For multiple output formats that are individual formats rather than a contiguous range, they are listed as comma-separated values ​​(out1,out2). Comma-separated values ​​are also used when the output formats are within a range, but the preferred order of those formats is not a default contiguous range. In both cases, the output formats are listed in a hyphen- or comma-separated list, in preferred order from the most preferred output format to the least preferred output format. `outf-send` and `outf-recv` are used when different output formats are used in both the send and receive directions. If the `outf` parameter is not present, all possible IVAS output formats are supported.

[0090] outf-specific-mode: Indicates a specific operating mode for the output format. If a range of specific operating modes is supported, it is indicated by the first and last modes of that range, separated by a hyphen (outf-specific-mode1-outf-specific-mode2). If there are multiple specific operating modes that are separate rather than a contiguous range, they are listed as comma-separated values ​​(outf-specific-mode1,outf-specific-mode2). Comma-separated values ​​are also used when specific operating modes are within a range, but the preferred order of those modes is not a default contiguous range. In both cases, the specific operating modes are listed in a hyphen- or comma-separated list, in preferred order from the most preferred mode to the least preferred mode. outf-specific-mode-send and outf-specific-mode-recv are used when different specific operating modes are used for both the transmit and receive directions, respectively. If outf-specific-mode does not exist, all possible specific input modes are supported. Some IVAS output formats may have only a single operating mode, in which case the outf-specific-mode parameter is redundant.

[0091] In the following examples, the parameters br, ptime, and maxptime follow the definitions presented in the EVS specification (3GPP TS26.445). In summary, the parameters represent the following:

[0092] br: Indicates the session bitrate, expressed in kilobits per second. This parameter can have a single value (br0) or a pair of bitrates separated by a hyphen (br1-br2), where br1 and br2 are used as the minimum and maximum bitrates, respectively.

[0093] ptime: Packet time, the length of time in milliseconds represented by the medium within the packet. In IVAS, ptime is set to 20ms.

[0094] maxptime: Indicates the maximum amount of media that can be encapsulated in each packet, in milliseconds. For frame-based codecs like IVAS, this time must be an integer multiple of the frame size (20ms in the case of IVAS).

[0095] The following Example 1 illustrates an example of SDP offer-answer negotiation for starting a session. The media line (m line) describes the port to be used for the session (49152). RTP / AVP refers to the audio and video RTP profile, and 96 is an indicator of the dynamic payload type. The type for payload number 96 is further described in the rtpmap and fmtp lines.

[0096] The rtpmap line indicates the use of the IVAS codec with a 16kHz timestamp clock frequency. The clock frequency used for IVAS has not yet been determined and may change before the standard is finalized. The EVS codec uses a 16kHz clock frequency, and the same value is used in the following examples. The timestamp is one of the fields in the fixed RTP header. The timestamp is incremented throughout the session and is reflected in the packet flow from sender to receiver. With a 20ms utterance frame block and a 16kHz timestamp clock frequency, the timestamp value is incremented by 320 for each consecutive frame block.

[0097] In this SDP offer, the fmtp line indicates the input formats supported by the sender in the inf parameter. The range 3-19 refers to the inf values ​​shown in Table 3. This range indicates that the sender supports all values ​​between 3 and 19, including 3 and 19. A bitrate of 512kbps is offered for this session.

[0098] The receiver sends the SDP answer to the sender with the fmtp line modified. The receiver has selected two preferred input formats for the sender to use (OMASA with 19=4 objects and MASA with 8=MASA) and sorted those formats in preferred order, where 19 is preferred over 8. The sender must use IVAS input format 19 during the session.

[0099] Example 1 is an example of an SDP offer-answer scenario, where the transmitter offers IVAS input modes 3-19, and the receiver answers with a list of preferred input modes (19,8) for the transmitter to use.

[0100] [Table 15]

[0101] Example 2 describes another SDP offer-answer scenario. The sender offers IVAS input modes 8 (MASA), 17 (OMASA with 2 objects), and 10 (ISM with 2 objects), and a bitrate range of 128–512 kbps. In this example, the receiver wants to use a single input mode across all available bitrates. However, only MASA and OMASA support the entire offered bitrate range from the offered input modes (ISM with 2 objects does not support bitrates higher than 256 kbps). The receiver sends an SDP answer indicating the preferred input modes 17 and 8 in that order. If the bitrate changes between the ranges negotiated at session time, the most preferred mode (17) should be used by the sender. However, the sender is not prohibited from switching to the use of input mode 8 (MASA) under certain circumstances. In one example scenario, the sender may unexpectedly face a heavy computational load and wish to use an input mode that is less computationally complex for processing. In this situation, the sender can switch to using any of the negotiated input modes, which are less computationally complex. If the receiver wants to explicitly prohibit switching input modes during a session, the receiver can select only a single input mode, as described in Example 3, or can include a disable-inf-switch parameter in the SDP answer, as described in Example 4.

[0102] Example 2 is an example of an SDP offer-answer scenario in which the transmitter offers IVAS input modes 8, 17, and 10 and a bitrate range of 128–512kbps. Input mode 10 does not support this entire bitrate range and may cause input mode switching if the bitrate changes during the session, so the receiver excludes input mode 10 from the answer.

[0103] [Table 16]

[0104] Example 3 describes another SDP offer-answer scenario. The sender offers IVAS input modes 8 (MASA), 17 (OMASA with 2 objects), and 10 (ISM with 2 objects) and a bitrate range of 128–512 kbps. In this example, the receiver wants to use a single input mode across all available bitrates. The receiver wants input mode 17 and includes only that input mode in the SDP answer.

[0105] Example 3 is an example of an SDP offer-answer scenario in which the transmitter offers IVAS input modes 8, 17, and 10, and the receiver selects only one input mode (17).

[0106] [Table 17]

[0107] Example 4 describes another SDP offer-answer scenario. The sender offers IVAS input modes 8 (MASA), 17 (OMASA with 2 objects), and 10 (ISM with 2 objects), and a bitrate range of 80–512 kbps. The receiver sends an answer to the sender indicating the preferred input modes 17 and 8 in that order. To prevent input mode switching during the session, the receiver also includes the parameter value disable-inf-switch=1 in the answer to restrict input mode switching between the two input modes. This set parameter prevents the sender from switching input modes, for example, in a scenario where the bitrate switches during the session.

[0108] Example 4 is an example of an SDP offer-answer scenario where the transmitter offers IVAS input modes 8, 17, and 10 and a bitrate range of 80–512kbps. The receiver desires input modes 17 and 8 in that order and includes those modes in the answer. The receiver also wants to avoid input mode switching during the session and includes the disable-inf-switch=1 parameter in the answer.

[0109] [Table 18]

[0110] Example 5 describes another SDP offer-answer scenario. The sender offers all possible IVAS input modes and a bitrate range of 13.2–512kbps. The receiver wants to support only the highest bitrate and indicates this in the answer as a bitrate range of 384–512kbps. This selected bitrate range also limits the availability of input modes, as not all offered modes support the highest bitrate. The receiver wants input modes 8 (MASA) and 17 (OMASA with 2 objects) in this order and indicates this in the SDP answer. Both preferred modes support the negotiated bitrate range.

[0111] Example 5 is an example of an SDP offer-answer scenario where the transmitter offers all possible IVAS input modes and a bitrate range of 13.2–512kbps. The receiver answers with a higher bitrate range (384–512kbps) that limits the availability of input modes. The receiver desires input modes 8 and 17, both of which support the negotiated bitrate range.

[0112] [Table 19]

[0113] Example 6 describes a different SDP offer-answer scenario. The sender offers all possible IVAS input modes and a bitrate range of 13.2–512kbps. The receiver wants to support only the highest bitrate and indicates this in the answer as a bitrate range of 384–512kbps. The receiver is using an external renderer (outf=15) with a binaural output setting (outf-specific-mode=binaural). In this example, the external renderer is specifically tuned for MASA type input, and the receiver wants input modes 8 (MASA) and 17 (OMASA with 2 objects) in that order and indicates this in the SDP answer.

[0114] Example 6 is an example of an SDP offer-answer scenario where the transmitter offers all possible IVAS input modes and a bitrate range of 13.2–512kbps. The receiver answers in the higher bitrate range (384–512kbps). The receiver desires input modes 8 and 17 because it is using an external renderer specifically tuned for MASA type inputs.

[0115] [Table 20]

[0116] In another embodiment, the disable-inf-switch parameter can be extended to allow more granular control over restricted mode switching. A parameter value of 0 or nonexistent means that the transmitter is allowed to switch between negotiated IVAS input formats. A value of 1 indicates that the transmitter is not permitted to switch between ungrouped input formats. For example, all multi-channel formats can be grouped under MC, and disable-inf-switch=1 would allow the transmitter to switch between different MC formats. In the context of the embodiments described herein, this restriction means that the transmitter cannot switch between input formats that are not in the same row in Table 3. That is, the transmitter is allowed to switch between different MC or SBA formats if these formats are negotiated. A disable-inf-switch parameter value of 2 indicates that switching within grouped input formats is also restricted, meaning the transmitter is not allowed to switch between any IVAS input formats, including different MC or SBA formats.

[0117] Example 7 demonstrates the use of an alternative version of the disable-inf-switch described above. The transmitter offers IVAS input formats 3-8 (MC mode and MASA) and bit rates 13.2-512. The receiver responds with the offered input mode and bit rate, and further includes the disable-inf-switch=1 parameter. This parameter indicates that the transmitter cannot switch between MC input formats (3-7) and MASA (8), but can switch between different MC input formats.

[0118] Example 7 is an example of an SDP offer-answer scenario in which the transmitter offers IVAS input modes 3-8 and a bitrate range of 13.2-512kbps. The receiver answers with the offered mode and bitrate but adds the parameter disable-inf-switch=1 (referring to an alternative implementation of the parameter described above). The added parameter restricts the transmitter from switching from MC input formats (3-7) to MASA (8), but allows the transmitter to switch between MC modes.

[0119] [Table 21]

[0120] Therefore, the embodiments described herein relate to IVAS standardization (standard specification), specifically to RTP payloads and signaling methods including session negotiation. Any of the embodiments described herein can be adopted from the 3GPP TS26.253 specification, "Codec for immersive voice and audio services - Detailed Algorithmic Description incl. RTP payload format and SDP parameter definitions."

[0121] Referring to Figure 1, this figure shows a block diagram of one possible non-limiting embodiment in which the embodiment can be carried out. A user device (UE) 110, a radio access network (RAN) node 170, and network elements 190 are shown, which can be used to communicate via VoLTE and 5G and Beyond. In the embodiment of Figure 1, the user device (UE) 110 wirelessly communicates with the radio network 100. The UE is a radio device that can access the radio network 100. The UE 110 includes one or more processors 120, one or more memories 125, and one or more transceivers 130 interconnected via one or more buses 127. Each of the one or more transceivers 130 includes a receiver Rx 132 and a transmitter Tx 133. The one or more buses 127 can be an address bus, a data bus, or a control bus and may include any interconnection mechanism, such as a series of wires on a motherboard or integrated circuit, optical fiber, or other optical communication equipment. The one or more transceivers 130 are connected to one or more antennas 128. One or more memories 125 contain computer program code 123. The UE 110 includes a module 140 which includes one or both of parts 140-1 and / or 140-2 that can be implemented in several ways. Module 140 may be implemented in hardware as module 140-1, such as being implemented as part of one or more processors 120. Module 140-1 may be implemented as an integrated circuit or by other hardware such as a programmable gate array. In another embodiment, module 140 may be implemented as computer program code 123 and implemented as module 140-2, which is executed by one or more processors 120. For example, one or more memories 125 and computer program code 123 can be configured by one or more processors 120 to cause the user device 110 to perform one or more of the operations described herein. The UE 110 communicates with the RAN node 170 via a radio link 111.

[0122] In this embodiment, the RAN node 170 is a base station that provides access to the radio network 100 to radio devices such as the UE 110. The RAN node 170 may also be a base station for 5G, also known as New Radio (NR). In 5G, the RAN node 170 may be an NG-RAN node defined as a gNB or ng-eNB. A gNB is a node that provides NR user plane and control plane protocol termination to the UE and is connected to the 5GC (e.g., network element 190) via an NG interface (e.g., connection 131). An ng-eNB is a node that provides E-UTRA user plane and control plane protocol termination to the UE and is connected to the 5GC via an NG interface (e.g., connection 131). An NG-RAN node may include a central unit (CU) (gNB-CU) 196 and multiple gNBs, including distributed units (DUs) (gNB-DUs), of which DU 195 is illustrated. Note that DU 195 may include or be coupled to a radio unit (RU) and be able to control the RU. gNB-CU196 is a logical node that hosts the Radio Resource Control (RRC), the SDAP and PDCP protocols of the gNB, or the RRC and PDCP protocols of the en-gNB, controlling the operation of one or more gNB-DUs. gNB-CU196 terminates the F1 interface connected to gNB-DU195. The F1 interface is illustrated as reference 198, which also shows links between remote elements and central elements of RAN node 170, such as between gNB-CU196 and gNB-DU195. gNB-DU195 is a logical node that hosts the RLC, MAC and PHY layers of the gNB or en-gNB, and its operation is partially controlled by gNB-CU196. One gNB-CU196 supports one or more cells. One cell may be supported by one gNB-DU195, or one cell may be supported / shared by multiple DUs under RAN sharing. The gNB-DU195 terminates the F1 interface 198 connected to the gNB-CU196.While DU195 is considered to include transceiver 160 as part of a RU, for example, it should be noted that some embodiments may have transceiver 160 as part of a separate RU connected to DU195, for example, under the control of DU195. RAN node 170 may be an eNB (Evolutionary NodeB) base station for LTE (Long-Term Evolution), or any other suitable base station or node.

[0123] RAN node 170 includes one or more processors 152, one or more memory 155, one or more network interfaces (N / WI / F) 161, and one or more transceivers 160, all interconnected by one or more buses 157. Each of the one or more transceivers 160 includes a receiver Rx162 and a transmitter Tx163. One or more transceivers 160 are connected to one or more antennas 158. One or more memory 155 contains computer program code 153. CU 196 may include a processor 152, one or more memory 155, and a network interface 161. DU 195 may include one or more memory and processors and / or other hardware of its own, but these are not shown in the illustration.

[0124] RAN node 170 includes module 150, which includes one or both of parts 150-1 and / or 150-2, which can be implemented in several ways. Module 150 may be implemented in hardware as module 150-1, such as being implemented as part of one or more processors 152. Module 150-1 may be implemented as an integrated circuit or by other hardware such as a programmable gate array. In another embodiment, module 150 may be implemented as module 150-2, which is implemented as computer program code 153 and executed by one or more processors 152. For example, one or more memories 155 and computer program code 153 can be configured by one or more processors 152 to cause RAN node 170 to perform one or more of the operations described herein. Note that the functionality of module 150 may be distributed, such as being distributed between DU 195 and CU 196, or may be implemented exclusively in DU 195.

[0125] One or more network interfaces 161 communicate over a network, such as via links 176 and 131. Two or more gNBs 170 can communicate, for example, using link 176. Link 176 may be wired, wireless, or both, and may implement, for example, an Xn interface for 5G, an X2 interface for LTE, or other suitable interfaces for other standards.

[0126] One or more buses 157 may be address buses, data buses, or control buses, and may include any interconnection mechanisms such as a series of wirings on a motherboard or integrated circuit, optical fibers or other optical communication equipment, or radio channels. For example, one or more transceivers 160 may be implemented as a remote radio head (RRH) 195 for LTE, or as a distributed unit (DU) 195 for a gNB embodiment for 5G, where other elements of the RAN node 170 may be physically located in a different location from the RRH / DU 195, and one or more buses 157 may be implemented in part as, for example, an optical fiber cable or other suitable network connection for connecting other elements of the RAN node 170 (e.g., a central unit (CU), gNB-CU 196) to the RRH / DU 195. Reference 198 also indicates those suitable network links.

[0127] A RAN node / gNB may include one or more TRPs to which the methods described herein can be applied. Figure 1 illustrates that RAN node 170 includes two TRPs, TRP51 and TRP52. RAN node 170 may host or include other TRPs not shown in Figure 1.

[0128] In NR, relay nodes are called integrated access and backhaul nodes. The mobile termination of an IAB node facilitates backhaul (parent link) connections. In other words, the mobile termination includes functions that support UE functions. The distributed unit of an IAB node facilitates so-called access link (child link) connections (i.e., backhaul for access link UEs and, in the case of multi-hop IAB, for other IAB nodes). In other words, the distributed unit is responsible for certain base station functions. IAB scenarios can follow a so-called segmented architecture in which a central unit hosts higher-layer protocols to the UE and terminates the control plane and user plane interfaces to the 5G core network.

[0129] It should be noted that while the description herein indicates that a “cell” performs a function, it should be obvious that the equipment forming a cell can perform those functions. A cell is part of a base station; that is, there can be multiple cells in one base station. For example, there can be three cells for a single carrier frequency and associated bandwidth, each cell covering one-third of a 360° area such that the coverage area of ​​a single base station covers an approximately elliptical or circular area. Also, each cell can correspond to a single carrier, and a base station can use multiple carriers. Thus, if there are three 120° cells for one carrier and two carriers, the base station has a total of six cells.

[0130] The wireless network 100 may include core network functions and may include one or more network elements 190 that provide connectivity to further networks such as telephone networks and / or data communication networks (e.g., the Internet) via one or more links 181. Such core network functions for 5G may include a location management function (LMF), and / or an access and mobility management function (AMF), and / or a user plane function (UPF), and / or a session management function (SMF). Such core network functions for LTE may include a mobility management entity (MME) / serving gateway (SGW) function. Such core network functions may also include a self-organizing / optimizing network (SON) function. It should be noted that these are merely examples of functions that may be supported by the network elements 190, and both 5G and LTE functions may be supported. The RAN node 170 is coupled to the network elements 190 via link 131. Link 131 may be implemented as, for example, an NG interface for 5G or an S1 interface for LTE, or another suitable interface for other standards. The network element 190 includes one or more processors 175, one or more memories 171, and one or more network interfaces (N / WI / F) 180 interconnected by one or more buses 185. One or more memories 171 contain computer program code 173. The computer program code 173 may include SON and / or MRO functions 172.

[0131] The wireless network 100 can implement network virtualization, which is the process of combining hardware and software network resources and network functions as a single software-based management entity or virtual network. Network virtualization often requires platform virtualization, which is combined with resource virtualization. Network virtualization can be classified into external virtualization, which combines many networks or parts of networks into a virtual unit, or internal virtualization, which gives network-like functionality to a software container on a single system. It should be noted that the virtualized entities resulting from network virtualization are still implemented to some extent using hardware such as processors 152 or 175 and memory 155 and 171, and such virtualized entities produce technical effects.

[0132] Computer-readable memories 125, 155, and 171 may be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, non-temporary memory, temporary memory, fixed memory, and removable memory. Computer-readable memories 125, 155, and 171 can be means for performing storage functions. Processors 120, 152, and 175 may be of any type suitable for the local technical environment and may include, in non-limiting examples, one or more general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), and processors based on multicore processor architectures. Processors 120, 152, and 175 can be means for performing functions such as controlling the UE 110, RAN node 170, network element 190, and other functions described herein.

[0133] In general, various exemplary embodiments of the user device 110 may include, but are not limited to, cellular phones such as smartphones, tablets, personal digital assistants (PDAs) with wireless communication capabilities, portable computers with wireless communication capabilities, image capture devices such as digital cameras with wireless communication capabilities, game devices with wireless communication capabilities, music storage and playback devices with wireless communication capabilities, internet-connected electronic devices including those that enable wireless internet access and browsing, tablets with wireless communication capabilities, head-mounted displays such as those that implement virtual / augmented / mixed reality, and portable units or terminals incorporating combinations of such functions. The UE 110 may also be a vehicle such as an automobile, or a UE mounted on a vehicle, such as a UAV such as a drone, or a UE mounted on a UAV. The user device 110 may also be a terminal device such as a mobile phone, mobile device, or sensor device, and the terminal device is a device used by or not used by a user.

[0134] UE110, RAN node 170, and / or network element 190 (and associated memory, computer program code, and modules) can be configured to implement (for example, in part) the methods and apparatus described herein, including methods and apparatus for negotiating conversational immersive audio sessions. Thus, the computer program code 123, module 140-1, module 140-2, and other elements / functions of UE110 shown in Figure 1 can implement the user equipment-related aspects of the embodiments described herein. Similarly, the computer program code 153, module 150-1, module 150-2, and other elements / functions of RAN node 170 shown in Figure 1 can implement the gNB / TRP-related aspects of the embodiments described herein. The computer program code 173 and other elements / functions of network element 190 shown in Figure 1 can be configured to implement the network element-related aspects of the embodiments described herein.

[0135] Figure 2 shows an exemplary hardware-implementable apparatus 200 configured to carry out the embodiments described herein. The apparatus 200 includes at least one processor 202 (e.g., FPGA and / or CPU), one or more memories 204 containing computer program code 205, the computer program code 205 having instructions for performing the methods described herein, and at least one memory 204 and the computer program code 205 are configured by at least one processor 202 to cause the apparatus 200 to implement circuits, processes, components, modules or functions (implemented by a control module 206) to carry out the embodiments described herein, including a method for negotiating a conversational immersive audio session. An offer 230 of the control module 206 can generate or receive an offer (e.g., an SDP offer), and an answer 240 can generate or receive an answer (e.g., an SDP answer). The memory 204 can be non-temporary memory, temporary memory, volatile memory (e.g., RAM) or non-volatile memory (e.g., ROM).

[0136] Apparatus 200 includes a display and / or I / O interface 208 which includes user interface (UI) circuits and elements that can be used to display aspects or situations of the methods described herein (for example, when one of the methods is performed or at a later point in time) or to receive input from a user, such as using a keypad, camera, touchscreen, touch area, one or more microphones, biometric recognition, one or more sensors, etc. In the case of immersive voice and audio, the embodiments described herein generally relate to devices having or connected to at least two microphones, for example, high-quality parametric spatial audio capture for MASA format generally uses at least three microphones.

[0137] The device 200 includes one or more communication interfaces (I / F) 210, such as a network (N / W) interface (I / F). The communication I / F 210 can be wired and / or wireless and can communicate over the Internet / other networks by any communication technology, including via one or more links 224. The communication I / F 210 may include one or more transmitters or one or more receivers.

[0138] The transceiver 216 includes one or more transmitters 218 and one or more receivers 220. The transceiver 216 and / or the communication I / F 210 may include standard, well-known components such as amplifiers, filters, frequency converters, (de)modulators, and encoder / decoder circuits, and one or more antennas, such as antenna 214, used for communication over the radio link 226.

[0139] The control module 206 of the device 200 includes one or both of parts 206-1 and / or 206-2, which can be implemented in several ways. The control module 206 may be implemented in hardware as control module 206-1, such as being implemented as part of one or more processors 202. The control module 206-1 may be implemented as an integrated circuit or by other hardware such as a programmable gate array. In another embodiment, the control module 206 may be implemented as control module 206-2, which is implemented as computer program code (having corresponding instructions) 205 and executed by one or more processors 202. For example, one or more memories 204 store instructions that, when executed by one or more processors 202, cause the device 200 to perform one or more of the operations described herein. Also, one or more processors 202, one or more memories 204, and exemplary algorithms encoded as instructions, programs, or code (e.g., as flowcharts and / or signaling diagrams) are means to cause the execution of the operations described herein.

[0140] The device 200 that performs the function of control 206 may correspond to UE 110, RAN node 170, or network element 190. Alternatively, since device 200 may be part of another node, such as a self-optimizing network (SON) node or a node in the cloud, device 200 and its elements do not have to correspond to any of the devices illustrated in Figure 1.

[0141] The device 200 may be distributed across the entire network (e.g., the Internet 28), including within the device 200 and between the device 200 and the UE 110, RAN node 170, or network element 190.

[0142] Interface 212 enables data communication and signaling between various items of the device 200, as shown in Figure 2. For example, interface 212 can be one or more buses, such as an address bus, a data bus, or a control bus, and may include any interconnection mechanism, such as a set of wiring on a motherboard or integrated circuit, optical fiber, or other optical communication equipment. Computer program code (e.g., instructions) 205, including control 206, may include object-oriented software configured to pass data or messages between objects in the computer program code 205. The device 200 does not have to include each of the features mentioned, or may include other features. Various components of the device 200 may be at least partially in a common housing 228, or subsets of various components of the device 200 may be at least partially in different housings, the different housings may include housing 228.

[0143] Figure 3 shows schematic diagrams of non-volatile memory media 300a (e.g., a computer / compact disc (CD) or digital versatile disc (DVD)) and 300b (e.g., a universal serial bus (USB) memory stick) that store instructions and / or parameters that, when executed by a processor, cause the processor to perform one or more steps of the methods described herein.

[0144] Figure 4 shows an exemplary method 400 performed by the transmitter, based on an exemplary embodiment described herein. In 410, the method includes obtaining one or more supported immersive conversation codec input formats. In 420, the method includes sorting the one or more immersive conversation codec input formats as a list sorted in a preferred order. In 430, the method includes including the immersive conversation codec input format attribute in a session description file. In 440, the method includes substituting the sorted list of immersive conversation codec input formats for the input format attribute in the session description file. In 450, the method includes generating a session negotiation offer based on the session description file. In 460, the method includes sending the session negotiation offer to the receiving user device. Method 400 may be performed by a transmitting device such as UE110, UE1 610-1, UE2 610-2, or device 200.

[0145] Figure 5 shows an exemplary method 500 performed by the receiving device based on an exemplary embodiment described herein. In 510, the method includes receiving a session negotiation offer. In 520, the method includes parsing one or more immersive conversation codec input formats from a session description file in the received session negotiation offer. In 530, the method includes obtaining one or more supported preferred input formats from the receiving user device. In 540, the method includes selecting one or more preferred input formats from the transmitting user device that are common with one or more supported preferred input formats from the receiving user device from the received session negotiation offer. In 550, the method includes substituting the selected one or more preferred input formats into the immersive conversation codec input format attribute in the session negotiation answer. In 560, the method includes sending the session negotiation answer to the transmitting user device. Method 500 may be carried out by a receiving device such as UE110, UE1 610-1, UE2 610-2, or device 200.

[0146] Figure 6 shows an example of a conversational immersive audio session between two parties, namely UE1 610-1 and UE2 610-2. The two UEs can negotiate an immersive conversational session via a suitable session negotiation mechanism via SIP / SDP, or via an SDP offer answer using a different signaling protocol (604, 606). Based on the UE's capabilities and preferences, a session offer is delivered from UE1 to UE2 via an SDP offer. The answer is provided by UE2 as an SDP answer. As a result of the agreed-upon session negotiation, RTP media delivery carrying an IVAS bitstream (608) as a payload is initiated between the two UEs (610-1, 610-2). A SIP server or webRTC signaling server (602) facilitates the session negotiation (604, 606).

[0147] In addition to the above Examples 1-7 which describe offer-answer negotiation scenarios, this specification provides and describes the following Examples (1-32).

[0148] (Example 1) A device comprising at least one processor and at least one memory for storing instructions, wherein the instructions, when executed by at least one processor, cause the device to perform at least one or more supported immersive conversation codec input formats; sort the one or more immersive conversation codec input formats as a list sorted in a preferred order; include immersive conversation codec input format attributes in a session description file; substitute the sorted list of immersive conversation codec input formats into the input format attributes in the session description file; generate a session negotiation offer based on the session description file; and send the session negotiation offer to a receiving user device.

[0149] (Example 2) The apparatus according to Example 1, wherein the session description file includes an input format indication, an output format indication, and a format switch for bitrate adaptation as attributes or media format parameters.

[0150] (Example 3) The apparatus according to any of Examples 1 to 2, wherein the session negotiation offer is represented as a session description protocol (SDP) file.

[0151] (Example 4) The apparatus according to any of Examples 1 to 3, wherein the sending of a session negotiation offer is performed as a session description offer answer model.

[0152] (Example 5) The apparatus according to any one of Examples 1 to 4, wherein when an instruction is executed by at least one processor, the apparatus causes the apparatus to at least receive a session negotiation answer from a receiving user device, and, if the session negotiation answer includes a single input format, to restrict the bitrate adaptation of the transmitting user device to a single input format.

[0153] (Example 6) The apparatus according to any one of Examples 1 to 5, wherein when an instruction is executed by at least one processor, the apparatus causes the apparatus to at least receive a session negotiation answer from a receiving user device, and, if the session negotiation answer explicitly disables input format switching, to restrict the bitrate adaptation of the transmitting user device to a single input format.

[0154] (Example 7) The apparatus according to Example 6, wherein input format switching is explicitly disabled by using the input format switching disable session description protocol (SDP) parameter.

[0155] (Example 8) The apparatus according to Example 7, wherein the SDP parameter for disabling input format switching includes disable-inf-switch.

[0156] (Example 9) The apparatus according to any one of Examples 1 to 8, wherein when an instruction is executed by at least one processor, the apparatus causes at least a session negotiation answer from a receiving user device, and if the session negotiation answer includes two or more input formats, the bitrate adaptation of the transmitting user device has the flexibility to use any input format from one or more agreed input formats.

[0157] (Example 10) The apparatus according to any one of Examples 1 to 9, wherein the codec format is indicated as an Immersive Speech and Audio Codec (IVAS) for a Session Negotiation Offer having a bitstream restricted to an Immersive Conversation Audio Codec bitstream, and the Session Negotiation Offer does not include an Enhanced Speech Codec (EVS) bitstream for rate adaptation.

[0158] (Example 11) The apparatus according to any one of Examples 1 to 10, wherein the input format is included in the description file as a media format parameter or as an attribute, along with the corresponding codec parameter, which is an immersive audio and audio codec parameter.

[0159] (Example 12) The apparatus according to any one of Examples 1 to 11, wherein one or more immersive conversation codec input formats are sorted based on encoding computation complexity.

[0160] (Example 13) The apparatus according to any one of Examples 1 to 12, wherein one or more immersive conversation codec input formats are sorted based on at least one audio capture function of the transmitting user device.

[0161] (Example 14) A device comprising at least one processor and at least one memory for storing instructions, wherein when an instruction is executed by at least one processor, the device causes the device to at least: receive a session negotiation offer; parse one or more immersive conversation codec input formats from a session description file in the received session negotiation offer; obtain one or more supported preferred input formats from a receiving user device; select one or more preferred input formats from a transmitting user device that are common with one or more supported preferred input formats from a receiving user device from the received session negotiation offer; assign the selected one or more preferred input formats to the immersive conversation codec input format attribute in a session negotiation answer; and send the session negotiation answer to the transmitting user device.

[0162] (Example 15) The apparatus according to Example 14, wherein the session description file includes an input format indication, an output format indication, and a format switch for bitrate adaptation as attributes or media format parameters.

[0163] (Example 16) The apparatus according to any of Examples 14 to 15, wherein the session negotiation offer is represented as a session description protocol (SDP) file and the session negotiation answer is represented as an SDP file.

[0164] (Example 17) The apparatus according to any one of Examples 14 to 16, wherein the reception of a session negotiation offer is performed as a session description offer answer model, and the transmission of a session negotiation answer is performed as a session description offer answer model.

[0165] (Example 18) The apparatus according to any of Examples 14 to 17, wherein if the session negotiation answer includes a single input format, the bitrate adaptation of the transmitting user equipment is limited to a single input format.

[0166] (Example 19) The apparatus according to any of Examples 14 to 18, wherein if the session negotiation answer explicitly disables input format switching, the bitrate adaptation of the transmitting user device is limited to a single input format.

[0167] (Example 20) The apparatus according to Example 19, wherein input format switching is explicitly disabled by using the input format switching disable session description protocol (SDP) parameter.

[0168] (Example 21) The apparatus according to Example 20, wherein the SDP parameter for disabling input format switching includes disable-inf-switch.

[0169] (Example 22) The apparatus according to any one of Examples 14 to 21, wherein, when the session negotiation answer includes two or more input formats, the bitrate adaptation of the transmitting user equipment has the flexibility to use any input format from one or more agreed input formats.

[0170] (Example 23) The apparatus according to any of Examples 14 to 22, wherein the codec format is indicated as an Immersive Speech and Audio Codec (IVAS) for session negotiation offers and session negotiation answers having a bitstream restricted to an Immersive Speech Audio Codec bitstream, and the session negotiation offers and session negotiation answers do not include an Enhanced Speech Codec (EVS) bitstream for rate adaptation.

[0171] (Example 24) The apparatus according to any one of Examples 14 to 23, wherein the input format is included in the session description file as a media format parameter or attribute, along with the corresponding codec parameter, which is an immersive audio and audio codec parameter.

[0172] (Example 25) The apparatus according to any one of Examples 14 to 24, wherein one or more immersive conversation codec input formats are sorted based on encoding computation complexity.

[0173] (Example 26) The apparatus according to any one of Examples 14 to 25, wherein one or more immersive conversation codec input formats are sorted based on at least one audio capture function of the transmitting user device.

[0174] (Example 27) A method comprising: obtaining one or more supported immersive conversation codec input formats; sorting the one or more immersive conversation codec input formats as a list sorted in a preferred order; including the immersive conversation codec input format attribute in a session description file; substituting the sorted list of immersive conversation codec input formats into the input format attribute in the session description file; generating a session negotiation offer based on the session description file; and sending the session negotiation offer to a receiving user device.

[0175] (Example 28) A method comprising: receiving a session negotiation offer; parsing one or more immersive conversation codec input formats from a session description file in the received session negotiation offer; obtaining one or more supported preferred input formats from the receiving user device; selecting one or more preferred input formats from the transmitting user device that are common with one or more supported preferred input formats from the receiving user device from the received session negotiation offer; substituting the selected one or more preferred input formats into the immersive conversation codec input format attribute in the session negotiation answer; and sending the session negotiation answer to the transmitting user device.

[0176] (Example 29) An apparatus comprising means for obtaining one or more supported immersive conversation codec input formats; means for sorting one or more immersive conversation codec input formats as a list sorted in a preferred order; means for including immersive conversation codec input format attributes in a session description file; means for substituting the sorted list of immersive conversation codec input formats into the input format attributes in the session description file; means for generating a session negotiation offer based on the session description file; and means for transmitting the session negotiation offer to a receiving user device.

[0177] (Example 30) An apparatus comprising: means for receiving a session negotiation offer; means for parsing one or more immersive conversation codec input formats from a session description file in the received session negotiation offer; means for obtaining one or more supported preferred input formats from a receiving user device; means for selecting one or more preferred input formats from a transmitting user device that are common with one or more supported preferred input formats from a receiving user device from the received session negotiation offer; means for substituting the selected one or more preferred input formats into the immersive conversation codec input format attribute in a session negotiation answer; and means for transmitting the session negotiation answer to a transmitting user device.

[0178] (Example 31) A computer-readable non-temporary program storage device that tangibly embodies a program of instructions executable by a computer to perform an operation, the operation comprising: obtaining one or more supported immersive conversation codec input formats; sorting one or more immersive conversation codec input formats as a list sorted in a preferred order; including immersive conversation codec input format attributes in a session description file; substituting the sorted list of immersive conversation codec input formats into the input format attributes in the session description file; generating a session negotiation offer based on the session description file; and transmitting the session negotiation offer to a receiving user device.

[0179] (Example 32) A computer-readable non-temporary program storage device that tangibly embodies a program of instructions executable by a computer to perform an operation, wherein the operation includes: receiving a session negotiation offer; parsing one or more immersive conversation codec input formats from a session description file in the received session negotiation offer; obtaining one or more supported preferred input formats from a receiving user device; selecting one or more preferred input formats from a transmitting user device that are common with one or more supported preferred input formats from a receiving user device from the received session negotiation offer; substituting the selected one or more preferred input formats into the immersive conversation codec input format attribute in a session negotiation answer; and sending the session negotiation answer to a transmitting user device.

[0180] When we refer to "computers," "processors," etc., please understand that this includes not only computers with different architectures such as single / multiprocessor architectures and sequential / parallel architectures, but also specialized circuits such as field-programmable gate arrays (FPGAs), application-specific circuits (ASICs), signal processing devices, and other processing circuits. When we refer to computer programs, instructions, code, etc., please understand that this includes software or firmware for programmable processors, programmable content for hardware devices such as instructions for processors, configuration settings for fixed-function devices, gate arrays, or programmable logic devices, etc.

[0181] As used herein, the terms “circuitry” and “circuit” and their variations may refer to any of the following: (a) hardware circuit embodiments, such as implementations in analog and / or digital circuitry; (b) (where applicable) (i) a combination of processors, or (ii) a part of processor / software including digital signal processors, software and one or more memories that cooperate to cause the device to perform various functions, such as circuitry, such as a microprocessor or part of a microprocessor that requires software or firmware for operation even if the software or firmware is not physically present. As further examples, the term “circuitry” as used herein may also cover a processor (or more processors) alone, or a part of a processor and / or embodiments of its accompanying software and / or firmware. The term “circuitry” may also cover, for example, a baseband integrated circuit or application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, cellular network device, or other network device, where applicable to the specific element. The terms "circuitry" or "circuit" are sometimes used to refer to a set of functions or processes used to carry out a method.

[0182] It should be understood that the above description is merely illustrative. Various alternative embodiments and modifications can be devised by those skilled in the art. For example, features described in various dependent claims can be combined with each other in any suitable combination. Furthermore, features can be selectively combined from the different embodiments described above to create new embodiments. Therefore, this description is intended to encompass all such alternative embodiments, modifications, and variations that fall within the scope of the appended claims.

[0183] The following acronyms and abbreviations, which may appear in this specification and / or in the drawings, are defined as follows: 3GPP 3rd Generation Partnership Project 4G: Fourth generation of broadband cellular network technology 5G (fifth-generation cellular network technology) 5GC 5G core network AMF (Access and Mobility Management Function) AMR-WB IO Interoperable Adaptive Multi-Rate Wideband ASIC: Application-Specific Integrated Circuit AVP (Audio-Video Protocol) BRIR (Binaural Room Impulse Response) CD Compact / Computer Disc CPU (Central Processing Unit) CR Carriage Return CU: Central unit or centralized unit DCT (Discrete Cosine Transform) DL Downlink DSP (Digital Signal Processor) DU (Distributed Unit) DVD (Digital Versatile Disc) eNB (evolved Node B) (e.g., LTE base station) EN-DC E-UTRAN New Radio-Dual Connectivity A node that provides NR user plane and control plane protocol termination for en-gNB UE and functions as a secondary node in EN-DC. E-UTRA, or evolved universal terrestrial radio access, is an LTE radio access technology. E-UTRAN E-UTRA Network EVS Enhanced Voice Service Interface between F1 CU and DU FOA (First-Order Ambisonics) FPGA (Field Programmable Gate Array) gNB is a base station for 5G / NR, i.e., it provides NR user plane and control plane protocol termination to the UE and is connected to the 5GC via the NG interface. H.2xx: A family of video coding standards in the ITU-T domain (e.g., H.264) H.323 is a standard that defines a protocol for providing audio-visual communication sessions over packet networks. HEVC (High Efficiency Video Coding) HOA (Higher-Order Ambisonics) HOA2 (Second-order higher-order ambisonics) HOA3 (Third-order higher-order ambisonics) HRIR (Head-related Impulse Response) IAB Integrated Access and Backhaul IANA (Internet Assigned Numbers Authority) I / F Interface IMS Instant Message Service I / O Input / Output IP Internet Protocol IPv4 Internet Protocol version 4 ISM (Independent Streams with Metadata) (i.e., a type of object-based audio) ITU (International Telecommunication Union) ITU-T ITU Telecommunications Standardization Sector IVAS (Immersive Voice and Audio Services) kbps (kilobits per second) Uncompressed audio data using L16 16-bit signed representation LF Line Feed LMF (Low-Mount Function) location management function LTE Long-Term Evolution MAC Media Access Control MASA Metadata-Assisted Spatial Audio MC Multichannel MME (Mobility Management Entity) MRO (Maintenance, Repair, and Overhaul) Mobility Robustness Optimization ng or NG new generation ng-eNB new generation eNB NG-RAN (New Generation Radio Access Network) NR new radio NTP (Network Time Protocol) Network OMASA object-based audio with MASA (composite input format) Pulse code modulation using the PCMU μ-law PDA (Personal Digital Assistant) PDCP (Packet Data Convergence Protocol) PHY physical layer PT payload type QoS (Quality of Service) RAN (Radio Access Network) RAM (Random Access Memory) RFC comment request RLC radio link control ROM (Read-only memory) RRC (Radio Resource Control) RTCP (Real-Time Transport Control Protocol) RTP (Real-time Transport Protocol) RTSP (Real-Time Streaming Protocol) RU (Radio Unit) Rx receiver SBA (Scene-Based Audio) SDP (Session Description Protocol) SGW (Serving Gateway) SIP session initiation protocol SMF session management function SON (Self-Organizing / Optimizing Network) Tdoc technical document TRP (Transmission Reception Point) TS Technical Specifications TTL (Time to Live) Tx transmitter UAV unmanned aerial vehicle UCS Universal Character Set UDP User Datagram Protocol UE (User Equipment) UI (User Interface) UMTS (Universal Mobile Telecommunication System) UPF User Plane Function URI Uniform Resource Identifier USB (Universal Serial Bus) UTF-8 UCS transformation format 8 UTRAN UMTS terrestrial radio access network VoLTE (Voice over LTE) voice service VR (Virtual Reality) WebRTC or WebRTC Web Real-Time Communication Network interfaces between X2 RAN nodes and between the RAN and the core network. Network interface between Xn NG-RAN nodes< / value> < / character>

Claims

1. At least one processor, A device including at least one memory for storing instructions, wherein when an instruction is executed by the at least one processor, the device provides at least, Obtaining one or more supported immersive conversation codec input formats, The one or more immersive conversation codec input formats are provided as a list in a preferred order, Include the immersive conversation codec input format attribute in the session description file, In the session description file, the list of immersive conversation codec input formats is assigned to the input format attribute, To generate a session negotiation offer based on the aforementioned session description file, Sending the aforementioned session negotiation offer to the receiving user device, A device that performs an action.

2. The apparatus according to claim 1, wherein the session description file includes an input format indication, an output format indication, and a format switch for bitrate adaptation as attributes or media format parameters.

3. The apparatus according to any one of claims 1 to 2, wherein the session negotiation offer is represented as a session description protocol (SDP) file.

4. The apparatus according to any one of claims 1 to 3, wherein the transmission of the session negotiation offer is performed as a session description offer answer model.

5. When the instruction is executed by the at least one processor, the device receives at least, Receiving a session negotiation answer from the receiving user device, If the session negotiation answer includes a single input format, the bitrate adaptation of the transmitting user equipment is limited to a single input format. The apparatus according to any one of claims 1 to 4, which causes the following to be performed.

6. When the instruction is executed by the at least one processor, the device receives at least, Receiving a session negotiation answer from the receiving user device, If the session negotiation answer explicitly disables input format switching, the bitrate adaptation of the transmitting user equipment will be limited to a single input format, The apparatus according to any one of claims 1 to 5, which causes the following to be performed.

7. The apparatus according to claim 6, wherein the input format switching is explicitly disabled by using an input format switching invalidation session description protocol (SDP) parameter.

8. When the instruction is executed by the at least one processor, the device receives at least, The receiving user device receives a session negotiation answer. The apparatus according to any one of claims 1 to 7, wherein, if the session negotiation answer includes two or more input formats, the bitrate adaptation of the transmitting user equipment has the flexibility to use any input format from one or more agreed-upon input formats.

9. The apparatus according to any one of claims 1 to 8, wherein the codec format is indicated as an Immersive Speech and Audio Codec (IVAS) for the session negotiation offer having a bitstream restricted to an Immersive Conversation Audio Codec bitstream, and the session negotiation offer does not include an Enhanced Speech Codec (EVS) bitstream for rate adaptation.

10. The apparatus according to any one of claims 1 to 9, wherein the input format is included in the session description file as a media format parameter or as an attribute, along with the corresponding codec parameter which is an immersive audio and audio codec parameter.

11. The apparatus according to any one of claims 1 to 10, wherein the one or more immersive conversation codec input formats are provided as the list in the preferred order based on encoding computation complexity.

12. The apparatus according to any one of claims 1 to 11, wherein the one or more immersive conversation codec input formats are provided as the list in the preferred order based on at least one audio capture function of the transmitting user device.

13. At least one processor, A device including at least one memory for storing instructions, wherein when an instruction is executed by the at least one processor, the device provides at least, Receiving a session negotiation offer, The process involves parsing one or more immersive conversation codec input formats from the session description file within the received session negotiation offer, To obtain one or more supported preferred input formats from the receiving user device, Selecting from the received session negotiation offer one or more preferred input formats from the transmitting user device that are common to one or more supported preferred input formats of the receiving user device, Within the session negotiation answer, the selected one or more preferred input formats are assigned to the immersive conversation codec input format attribute, The session negotiation answer is transmitted to the transmitting user device. A device that performs an action.

14. The apparatus according to claim 13, wherein the session description file includes an input format indication, an output format indication, and a format switch for bitrate adaptation as attributes or media format parameters.

15. The apparatus according to any one of claims 13 to 14, wherein the session negotiation offer is represented as a session description protocol (SDP) file, and the session negotiation answer is represented as an SDP file.

16. The apparatus according to any one of claims 13 to 15, wherein the reception of the session negotiation offer is performed as a session description offer answer model, and the transmission of the session negotiation answer is performed as a session description offer answer model.

17. The apparatus according to any one of claims 13 to 16, wherein, if the session negotiation answer includes a single input format, the bitrate adaptation of the transmitting user equipment is limited to a single input format.

18. The apparatus according to any one of claims 13 to 17, wherein if the session negotiation answer explicitly disables input format switching, the bitrate adaptation of the transmitting user device is limited to a single input format.

19. The apparatus according to claim 18, wherein the input format switching is explicitly disabled by using an input format switching disable session description protocol (SDP) parameter.

20. The apparatus according to any one of claims 13 to 19, wherein, if the session negotiation answer includes two or more input formats, the bitrate adaptation of the transmitting user device has the flexibility to use any input format from one or more agreed input formats.

21. The apparatus according to any one of claims 13 to 20, wherein the codec format is indicated as an Immersive Voice and Audio Codec (IVAS) for the Session Negotiation Offer and the Session Negotiation Answer, having a bitstream restricted to an Immersive Conversation Audio Codec bitstream, and the Session Negotiation Offer and the Session Negotiation Answer do not include an Enhanced Voice Codec (EVS) bitstream for rate adaptation.

22. The apparatus according to any one of claims 13 to 21, wherein the input format is included in the session description file as a media format parameter or as an attribute, along with the corresponding codec parameter which is an immersive audio and audio codec parameter.

23. The apparatus according to any one of claims 13 to 22, wherein the one or more immersive conversation codec input formats are provided as the list in the preferred order based on encoding computation complexity.

24. The apparatus according to any one of claims 13 to 23, wherein the one or more immersive conversation codec input formats are provided as the list in the preferred order based on at least one audio capture function of the transmitting user device.

25. Obtaining one or more supported immersive conversation codec input formats, The one or more immersive conversation codec input formats are provided as a list in a preferred order, Include the immersive conversation codec input format attribute in the session description file, In the session description file, the list of immersive conversation codec input formats is assigned to the input format attribute, To generate a session negotiation offer based on the aforementioned session description file, A method comprising sending the aforementioned session negotiation offer to the receiving user device.

26. Receiving a session negotiation offer, The process involves parsing one or more immersive conversation codec input formats from the session description file within the received session negotiation offer, To obtain one or more supported preferred input formats from the receiving user device, Selecting from the received session negotiation offer one or more preferred input formats from the transmitting user device that are common to one or more supported preferred input formats of the receiving user device, Within the session negotiation answer, the selected one or more preferred input formats are assigned to the immersive conversation codec input format attribute, A method comprising transmitting the session negotiation answer to the transmitting user device.

27. A means for obtaining one or more supported immersive conversation codec input formats, Means for providing the one or more immersive conversation codec input formats as a list in a preferred order, A means of including the immersive conversation codec input format attribute in the session description file, A means for substituting the list of immersive conversation codec input formats into the input format attribute in the session description file, A means for generating a session negotiation offer based on the aforementioned session description file, Means for transmitting the aforementioned session negotiation offer to the receiving user device, A device that includes this.

28. Means of receiving session negotiation offers, A means for parsing one or more immersive conversation codec input formats from the session description file in the received session negotiation offer, Means for obtaining one or more supported preferred input formats from the receiving user device, Means for selecting one or more preferred input formats from the transmitting user device that are common to one or more supported preferred input formats of the receiving user device from the received session negotiation offer, A means for substituting the selected one or more preferred input formats into the immersive conversation codec input format attribute within the session negotiation answer, An apparatus including means for transmitting the session negotiation answer to the transmitting user device.

29. A non-temporary program storage device readable by a computer, which tangibly embodies a program of instructions that can be executed by a computer in order to perform an action, wherein the action is Obtaining one or more supported immersive conversation codec input formats, The one or more immersive conversation codec input formats are provided as a list in a preferred order, Include the immersive conversation codec input format attribute in the session description file, In the session description file, the list of immersive conversation codec input formats is assigned to the input format attribute, To generate a session negotiation offer based on the aforementioned session description file, Sending the aforementioned session negotiation offer to the receiving user device, Includes non-temporary program storage devices.

30. A non-temporary program storage device readable by a computer, which tangibly embodies a program of instructions that can be executed by a computer in order to perform an action, wherein the action is Receiving a session negotiation offer, The process involves parsing one or more immersive conversation codec input formats from the session description file within the received session negotiation offer, To obtain one or more supported preferred input formats from the receiving user device, Selecting from the received session negotiation offer one or more preferred input formats from the transmitting user device that are common to one or more supported preferred input formats of the receiving user device, Within the session negotiation answer, the selected one or more preferred input formats are assigned to the immersive conversation codec input format attribute, A non-temporary program storage device, which includes transmitting the session negotiation answer to the transmitting user device.