Apparatus and method for enhanced voice communication
By dynamically adjusting microphone configurations and codec formats in response to calling mode changes, the system enhances voice communication quality and optimizes resource usage in communication terminals.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
Modern communication terminals face challenges in maintaining voice call quality and optimizing resource usage due to varying microphone configurations and capture capabilities, particularly when users switch between different calling modes, such as hands-free and handheld, without efficient codec renegotiation.
The system employs sensors to monitor calling mode changes and dynamically adjusts microphone configurations and codec formats by renegotiating codecs through SIP/SDP signaling, ensuring seamless transitions and efficient resource management.
This approach maintains call quality and optimizes resource usage by adapting to real-time changes in calling modes, reducing processing overhead and preserving battery life while ensuring high-quality audio experiences.
Smart Images

Figure JP2025039206_15052026_PF_FP_ABST
Abstract
Description
APPARATUS AND METHOD FOR ENHANCED VOICE COMMUNICATION
[0001] The present disclosure generally pertains to an apparatus and a method for enhanced voice communication.
[0002] Modern communication terminals increasingly support multi-channel audio capture, enabling rich spatial audio experiences.
[0003] In modern terminal devices, microphone configurations and capture capabilities vary, influencing voice communication quality and codec selection.
[0004] In one general aspect, the techniques disclosed in here feature: a communication apparatus comprising: control circuitry which, in operation, detects a change of a calling mode during a call, and if the change of the calling mode is detected, triggers a codec negotiation to switch from multi-channel codec to mono-channel codec or from mono-channel codec to multi-channel codec; and a transmitter which, in operation, transmits an update request to renegotiate a codec to be used after the change of the calling mode is detected.
[0005] It should be noted that general or specific disclosures may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any selective combination thereof.
[0006] With an apparatus and a method of the present disclosure, it is possible to achieve adaptation of the voice call efficiently while managing the codec and microphone setup to maintain call quality.
[0007] Figure 1 shows an example of signal exchange flow for codec negotiation.Figure 2 shows an example of communication stack layers.Figure 3 shows an example of SIP / SDP pre-call capability exchange.Figure 4 shows an example of SDP Answer.Figure 5 shows an example of SIP Acknowledgement.Figure 6 shows an example of sensor monitoring for mode change.Figure 7 shows an example of SIP Update Messages.Figure 8 shows an example of SDP Answer for the SIP Update.Figure 9 shows an example of a multi-party conference call for a spatial conferencing service.Figure 10 shows an example of a part of a communication apparatus.
[0008] Various embodiments of the present disclosure will be now described in detail with reference to the annexed drawings. In the following description, a detailed description of known functions and configurations has been omitted for clarity and conciseness.
[0009] This disclosure relates to an apparatus and a method for enhancing voice communication through terminal capability exchange and adaptive codec or coding format adjustments in response to real-time changes in calling mode.
[0010] To enhance spatial audio experience for 4G and future-generation services across conversational and / or non-conversational applications, systems can employ audio codecs supporting a range of formats, from mono up-to multi-channel coding formats can be employed to enhance the spatial audio experience after capability exchange between terminals. However, when users change the calling modes (e.g., from hands-free to handheld or vice versa), the system dynamically adjusts to the new calling mode by activating a relevant microphone configuration.
[0011] The present solution leverages sensors within the terminal devices to monitor calling mode events. Upon detecting a mode change, the terminal adapts by either renegotiating the codec via SIP / SDP signalling to switch to an appropriate codec or coding format or by modifying the audio input to multi-mono configuration without requiring codec renegotiation. By enabling real-time switching, the terminals efficiently manage resources, reduces the processing overhead and preserving battery life.
[0012] In modern communication terminals, microphone configurations and capture capabilities vary, influencing voice communication quality and codec selection. For instance, when a call is initiated between terminals with differing audio capture setups such as one connected to a Bluetooth headset (capable to receive mono to multi-channel / spatial audio, transmit binaural) and another terminal using on-device multi-microphones and stereo loudspeakers (capable to receive mono / stereo audio) --understanding each terminals capabilities is crucial to ensuring optimal audio quality. After this exchange, when a call starts in immersive calling mode like using 3GPP IVAS for enhanced spatial sound, switching to conventional mono due to change in calling mode in the ongoing call require better optimization, vice versa. Thus, the system must support a seamless transition in microphone configuration and renegotiating the codec or adjusting to a desired capture format based on the calling mode. The challenge lies in adapting the voice call efficiently while managing the codec and microphone setup to maintain call quality and optimize resource usage.
[0013] <Embodiment> The present solution involves negotiating an optimal microphone configuration based on device capabilities, selecting an appropriate voice codec or coding mode accordingly, monitoring calling mode changes, and dynamically re-negotiating microphone configuration settings and switching to the appropriate voice codec or coding format in response to user interactions.
[0014] The following implementation steps are illustrated through an example, assuming two terminal devices and a user behaviour scenario where a call mode switch from hands-free to handheld voice mode. Note that below steps can also apply to calling mode switch from handheld to hands free mode.
[0015] Terminal X: Equipped with a microphone array, this terminal supports multi-microphone capture and playback through built-in stereo speakers. It can encode Scene-Based Audio, MASA, or Stereo (and mono) and receive formats that can be rendered in Stereo through its loudspeakers.
[0016] Terminal Y: Connected to Bluetooth headphones, this terminal can capture binaural audio and supports playback via either built-in stereo speakers or headphones. It can encode binaural audio and receive any format that can be rendered as Stereo for loudspeakers or binaurally for headphones.
[0017] <Steps for Implementation:> Figure 1 illustrates an example of signal exchange flow for codec negotiation.
[0018] 1. Call Initiation, Device Mode detection, Capability Exchange
[0019] Initiation: The user dials a number, sending a call request over the 4G or next-generation network core.
[0020] Capability Exchange: As part of the call setup, the devices exchange capabilities via SIP / SDP, sharing details on:
[0021] This exchange allows both devices to share their capturing capabilities Identify and compile a list of device capabilities, including available microphones (e.g., internal, external, spatial arrays), speakers, headphone types (e.g., over-ear, in-ear, open-back), and audio channel capabilities (mono, stereo, spatial, etc.,).
[0022] Codec Options: Generate an optimal configuration list of codec lists and coding modes (example: AMR, AMR-WB, EVS (High Quality Voice), IVAS (Immersive Voice call))
[0023] Notify the network through SIP / SDP signalling, SIP signalling establishes the session with the other terminal, finalizing the setup.
[0024] Calling Mode Detection: At this initial stage, the device detects it’s in handsfree mode (the user connects the terminal with a Bluetooth headset, places the device on a table). The calling mode and relevant Quality of Experience (QoE) preferences for handsfree mode are bundled into the call request. Based on user usage patterns, even if the device has capabilities, the contents of the optimal configuration list may be limited by the user's usage patterns.
[0025] User usage patterns = Call states such as hands-free or handheld.
[0026] During the initial SIP INVITE, both terminals negotiate the use of suitable capture and rendering capabilities and corresponding codec option for the call to provide.
[0027] 2. QoS Profile Assignment with Guaranteed Bit Rate (GBR)
[0028] QoS with GBR: Based on the calling mode and microphone capability, codec, network bandwidth, the network assigns a Quality of Service (QoS) bearer with Guaranteed Bit Rate (GBR), ensuring that sufficient bandwidth is allocated for high-quality.
[0029] The GBR parameter secures a stable bitrate to maintain consistent, high-quality audio transmission, essential for immersive calls.
[0030] 3. Media Transfer
[0031] Voice Data Transmission: The initial voice data packets begin transmission over the configured bearer, ensuring prioritized handling of voice packets. The GBR parameter maintains the necessary bitrate, while the device microphone configuration provides a reliable audio experience.
[0032] QoE Monitoring: The network continuously tracks QoE metrics like packet loss, latency, and jitter. If these metrics fall below the threshold, the network can adjust, ensuring the user benefits from stable, high-quality audio.
[0033] 4. Calling Mode Event Monitoring (adaptation to new calling mode):
[0034] Calling Mode Both terminals continually monitor the calling mode using sensor input.
[0035] Detecting New Mode: Mid-call, the user hold the terminal closer to the ear switching to handheld mode. This change is detected using internal sensors to detect when the user switches to hand-held mode, and the device updates its mode.
[0036] Terminals can detect if the device switches from handheld mode (e.g., moving the phone away from the ear) to headphone / speaker mode (hands-free mode) using sensors available in the terminals.
[0037] Capability re-exchange - Monitor the status of device capabilities and determine the optimal codec mode based on changes in status, notifying the network. Changes in status = Hands-free / handheld, external speaker / built-in speaker / headphones, etc. - Based on changes in services, do the same. Changes in services = Mono / stereo / immersive, etc.
[0038] Codec Switching Upon Calling Mode Change: When the calling mode changes from hands-free to hands-handheld, the system dynamically triggers codec negotiation to switch from multi-channel / stereo to mono.
[0039] The source of the audio feed is also updated (e.g., switching from multi-microphone capture to a single bottom microphone for mono transmission).
[0040] Once a calling mode change is detected (e.g., from hands-free to handheld), the system needs to have a capability to support monaural capture by renegotiating the codec to monaural codec
[0041] or
[0042] maintaining the codec negotiation same but input signal can be changed to more than one-channel by splitting or duplicating the monaural input to multi-monaural.
[0043] For the latter case, there will be no change in the use of multi-channel / stereo codec but all input signals to the codec will be identical to each other. If the codec utilizes inter-channel correlations for encoding the multi-channel or stereo input, such multi monaural signal can be encoded efficiently. If supported, such identical multi monaural signal can be detected by the codec and the codec can switch to a monaural codec.
[0044] In the former case, such codec switching will be described in the following sections.
[0045] Steps for Codec Switching:
[0046] 1. Mode Change Detection: Upon detecting the mode change event (from stereo to mono), the system triggers codec renegotiation.
[0047] 2. Notify the Application Layer: The calling application is informed about the mode change. Figure 2 illustrates an example of communication stack layers.
[0048] 3. Send SIP UPDATE / REINVITE: Terminal X sends a SIP UPDATE or REINVITE message to renegotiate the codec.
[0049] SDP / Signalling Update for Codec Negotiation: - A re-negotiation occurs through SDP update embedded in a SIP re-INVITE or SIP UPDATE message to switch the codec (e.g., from stereo to EVS mono). - Once Terminal acknowledges the re-INVITE with SIP 200 OK, while maintaining the call the RTP stream is updated to reflect the new codec parameters, minimizing the transmission complexity while improving efficiency.
[0050] New QoE and QoS Profile: Updated QoE Requirements: In hands-free / hand-held mode, the user’s QoE preferences can be updated. QoS Profile Reconfiguration: The network adjusts the QoS profile for hands-free, reallocating GBR resources to match the new mode. For instance, if the Bluetooth headset has multi-microphone capability (binaural), the network enables the immersive feature and prioritizes stereo channel.
[0051] Network Slice Adaptation: Network slice dynamically adapts to the updated QoS requirements, reallocating GBR and other parameters for a seamless mode transition without dropping call quality.
[0052] NOTE: In the scenario of calling mode change detection (e.g., from hands-free mode to handheld mode), terminals can utilize one or more of the following sensors for event monitoring: - Proximity Sensor: Detects when the phone is close to the user’s ear, indicating a transition to handheld mode. - Accelerometer and Gyroscope: Detects movement and orientation changes, signalling a shift in the phone’s usage pattern. - Bluetooth Connectivity Status: Detects when a Bluetooth headset is connected or disconnected, which might trigger a hands-free mode change. - Microphone Activation: Tracks changes in the number of active microphones (e.g., multiple microphones for stereo or multi-channel, vs. a single mic for handheld mono mode).
[0053] 5. Real-Time Adaptation and QoE Optimization Continuous QoE Monitoring: Throughout the call, the network monitors QoE to dynamically maintain voice clarity and immersion, adjusting the GBR as needed.
[0054] Adaptations Based on Environment: If the user moves or if noise levels change, QoE adjustments are made. For instance, additional noise reduction could be applied in a noisy environment, leveraging multi-microphone capabilities when possible.
[0055] 6. Call Termination and Resource Release Termination Request: When the user ends the call, a termination request is sent over IMS, following SIP procedures.
[0056] Resource and QoS Release: The network releases the dedicated QoS bearer and network slice resources, including the GBR allocation, freeing up network capacity for other services. Call Quality Log: QoE and QoS performance data may be logged to optimize future calls with similar multi-microphone and immersive audio requirements.
[0057] <Use Cases> Note: Solution & Steps for Implementation explained above can be extended to below use cases
[0058] <Telephony Usage Scenarios (another example)> - Car Mode Switching
[0059] - Stereo and Immersive Telephony
[0060] <Example Implementation> Figure 3 illustrates an example of SIP / SDP pre-call capability exchange.
[0061] Explanation SDP Offer: m=audio 49170 RTP / AVP 97 98 99 100 0 8: Describes the media (audio) parameters, specifying that Terminal X can support RTP payload types97 (IVAS multi-channel), 98 (IVAS stereo / binaural), 99 (IVAS mono), 100 (AMR-WB) and legacy codecs (PCMU and PCMA for fallback). a=rtpmap: Maps the payload types to their respective codecs. - Payload type 97 is mapped to IVAS Multi-Channel at 32 kHz sampling rate with 6 channels. - Payload type 98 is mapped to IVAS Stereo at 32 kHz with 2 channels. - Payload type 99 is mapped to IVAS Mono at 32 kHz with 1 channel. - Payload type 100 is mapped to AMR-WB at 16 kHz with 1 channel. fmtp parameters: These are codec-specific parameters. For IVAS codecs, the maxplaybackrate defines the maximum playback rate, and maxplaybackchannels =2 ensures stereo playback for the multi-channel or stereo options. Pre-Call Capability Exchange: SDP Answer
[0062] Figure 4 illustrates an example of SDP Answer.
[0063] Explanation SDP Answer:
[0064] m=audio 49172 RTP / AVP 97 98: Terminal Y responds with its preferred codec list. It selects IVAS stereo (payload type 97) as the primary codec and IVAS mono (payload type 98) as a fallback option. a=rtpmap: Maps payload types 97 and 98 to their corresponding codecs, with the same sampling rate as Terminal X. fmtp parameters: Terminal Y echoes the codec-specific parameters, confirming it can handle stereo playback for IVAS stereo and mono playback for IVAS mono.
[0065] Figure 5 illustrates an example of SIP Acknowledgement.
[0066] The ACK confirms that Terminal X acknowledges the codec negotiation and is ready to begin media transmission using the agreed-upon codec (IVAS in this case).
[0067] Figure 6 illustrates an example of sensor monitoring for Mode Change.
[0068] The proximity sensor or accelerometer detects that the user is holding the phone, triggering the mode change event, which initiates codec switching.
[0069] Figure 7 illustrates an example of SIP Update Messages.
[0070] After detecting the mode change, Terminal X sends a SIP UPDATE message to renegotiate the codec.
[0071] Explanation: - Terminal X updates the codec list to offer IVAS Mono (100) and EVS Mono (96), optimizing for mono audio in handheld mode.
[0072] Figure 8 illustrates an example of SDP Answer for the SIP Update.
[0073] Terminal Y responds with an SDP answer accepting the codec switch.
[0074] Explanation: - Terminal Y accepts the codec switch to IVAS Mono (100), confirming the transition to mono audio in handheld mode.
[0075] <Spatial Conferencing Usage Scenarios> <Steps for Implementation:> The following example assumes spatial conferencing scenario with intermediary device (call server (MRFP)) that handles mixing, forwarding, and floor control functions and multiple endpoints with significantly different capabilities for instance in terms of supported codec, bitrate, audio bandwidth, capture and render. Figure 9 illustrates one typical realization of the spatial conferencing scenario. Some of the endpoints may be 3GPP UEs, some may be PSTN or generic VoIP clients. Among the 3GPP UEs, not all will necessarily support multi-channel codec like 3GPP IVAS but only legacy codecs like AMR, AMR-WB or EVS. There is also some capability variety among the IVAS-enabled UEs. Whilst other conferencing architectures are possible, the use of a centralized server architecture is practical when a diverse set of endpoints with very different capabilities participate in the same conference. Scenario involves negotiating an optimal microphone configuration based on device capabilities, selecting an appropriate voice codec or coding mode accordingly, monitoring calling mode changes for the endpoints, and dynamically re-negotiating microphone configuration settings for the endpoints with more than mono capturing and switching to the appropriate voice codec or coding format in response to user interactions.
[0076] - Server-based spatial voice conferencing
[0077] Figure 10 illustrated an exemplary configuration of a part of communication apparatus 10 according to an embodiment of the present disclosure. The communication apparatus 10 may include, for example, control circuitry 11, and a transmitter 12.
[0078] Control circuitry 11 controls a call processing. For example, control circuitry 11 detects a change of a calling mode during a call, and if the change of the calling mode is detected, triggers a codec negotiation to switch from multi-channel codec to mono-channel codec or from mono-channel codec to multi-channel codec.
[0079] Transmitter 12 transmits a signal (e.g., a message) relating to the call. For example, transmitter 12 transmits an update request to renegotiate a codec to be used after the change of the calling mode is detected.
[0080] According to the present Embodiment, it is possible to achieve adaptation of the voice call efficiently while managing the codec and microphone setup to maintain call quality and to optimize resource usage.
[0081] The disclosure of Japanese Patent Application No. 2024-196604, filed on November 11, 2024, including the specification, drawings and abstract, is incorporated herein by reference in its entirety.
[0082] This disclosure can be applied to an apparatus and a method or enhanced voice communication.
[0083] 10 Communication apparatus 11 Control circuitry 12 Transmitter
Claims
1. A communication apparatus comprising: control circuitry which, in operation, detects a change of a calling mode during a call, and if the change of the calling mode is detected, triggers a codec negotiation to switch from multi-channel codec to mono-channel codec or from mono-channel codec to multi-channel codec; and a transmitter which, in operation, transmits an update request to renegotiate a codec to be used after the change of the calling mode is detected.
2. The communication apparatus according to claim 1, wherein the calling mode includes a hands-free mode and a hand-held mode.
3. The communication apparatus according to claim 2, wherein if the detected change of the calling mode is a change from the hands-free mode to the hand-held mode, the triggered codec negotiation is for switching from the multi-channel codec to the mono-channel codec; and if the detected change of the calling mode is a change from the hand-held mode to the hands-free mode, the triggered codec negotiation is for switching from the mono-channel codec to the multi-channel codec.
4. The communication apparatus according to claim 1, wherein if the change of the calling mode is detected, the control circuitry requests update of a source of audio feed signal transmitted from a communication partner apparatus.
5. The communication apparatus according to claim 4, wherein if the detected change of the calling mode is a change from a hands-free mode to a hand-held mode, the requested update is a change from a multi-microphone capture signal to a single bottom microphone signal.
6. A communication method comprising: detecting a change of a calling mode during a call; if the change of the calling mode is detected during the call, triggering a codec negotiation to switch from multi-channel codec to mono-channel codec or from mono-channel codec to multi-channel codec; and transmitting an update request to renegotiate a codec to be used after the change of the calling mode is detected.
7. The communication method according to claim 6, wherein the calling mode includes a hands-free mode and a hand-held mode.
8. The communication method according to claim 7, wherein if the detected change of the calling mode is a change from the hands-free mode to the hand-held mode, the triggered codec negotiation is for switching from the multi-channel codec to the mono-channel codec; and if the detected change of the calling mode is a change from the hand-held mode to the hands-free mode, the triggered codec negotiation is for switching from the mono-channel codec to the multi-channel codec.
9. The communication method s according to claim 6, wherein if the change of the calling mode is detected, the control circuitry requests update of a source of audio feed signal transmitted from a communication partner apparatus.
10. The communication method according to claim 9, wherein if the detected change of the calling mode is a change from a hands-free mode to a hand-held mode, the requested update is a change from a multi-microphone capture signal to a single bottom microphone signal.