Receiver based dynamic audio suppression

Dynamic audio suppression methods negotiate listener preferences to enhance immersive audio experiences by selectively suppressing unwanted components, addressing the suboptimal issues in existing technologies.

WO2026101992A1PCT designated stage Publication Date: 2026-05-15DOLBY LABORATORIES LICENSING CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2025-11-05
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing immersive audio technologies fail to account for listener preferences in suppressing unwanted audio components, leading to suboptimal audio experiences by either collapsing ambiance or failing to suppress noise adequately.

Method used

Implement dynamic audio suppression methods that negotiate between capture and user devices based on listener preferences, using audio suppression processing information to selectively suppress unwanted audio content types.

Benefits of technology

Enhances audio quality by dynamically adjusting suppression levels based on listener preferences, ensuring an optimal immersive experience by maintaining desired audio components while reducing unwanted noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025054145_15052026_PF_FP_ABST
    Figure US2025054145_15052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods for dynamic audio suppression for captured audio, such as immersive audio in immersive voice and audio services (IVAS). One method comprises enabling a dynamic audio suppression capability based on a negotiated communication session and selecting a preferred audio content type based on user preferences for audio suppression in audio. The method further comprises sending audio suppression processing information (PI) data based on the preferred audio content, receiving audio bitstream and decoding the audio bitstream to obtain audio content from the audio bitstream. Wherein, a second audio content type of the audio content is suppressed relative to a first audio content type of the audio content when the first audio content type is selected. The method further comprises rendering the audio content in an audio format suitable for playing on the end user device.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No.: D24173WO01RECEIVER BASED DYNAMIC AUDIO SUPPRESSIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to US provisional Application No. 63 / 717,221, filed 06 November 2024 and US provisional application 63 / 911,082, filed 04 November 2025, all of which are incorporated herein by reference in their entirety.TECHNICAL FIELD

[0002] This application relates generally to audio signal processing and, more specifically, dynamically applying audio suppression based on negotiated communications between capture / sender and user / receiver devices based on capture / sender capabilities and user / receiver audio listening preferences.BACKGROUND

[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.

[0004] Voice and video encoder / decoder (codec) standards such as 3rdGeneration Partnership Project (3GPP) are recently focused on developing codecs for immersive audio service, such as for Immersive Voice and Audio Services or IV AS. The immersive audio services are expected to support a range of audio service capabilities, including but not limited to mono, stereo and fully immersive audio encoding, decoding and rendering.

[0005] A wide range of devices, including endpoints and network nodes, are expected to support immersive audio services (e.g., IVAS). Some example devices include but are not limited to: mobile devices such as smart phones and electronic tablets, personal computers, video and audio conferencing, home theater devices, television, vehicles such as cars, in-vehicle equipment such as in-vehicle infotainment systems or vehicle computers, video gaming devices, and extended reality (XR) devices, as well as other suitable devices. These devices, end points and network nodes can have various acoustic interfaces for sound capture, transmission, relay, and rendering.

[0006] IVAS codecs as standardized by the 3GPP support a variety of immersive audio formats, including but not limited to Ambisonics, Channel based Audio and Object Based Audio. The immersive audio formats offer an immersive experience to the listener devices that are capable ofDocket No.: D24173WO01 rendering the immersive audio formats. However, providing an immersive audio experience creates a number of challenges including but not limited to noise classification and cancellation and echo cancellation.

[0007] In some contemplated examples, a capture device may include audio processing for noise suppression and echo cancellation, which is intended to suppress unwanted audio components while maintaining the overall essence of the immersive capture. It may also be desirable to maintain phase and level differences between channels in an immersive capture to ensure audio quality to the listener.

[0008] It is with respect to the these and other considerations that the disclosure made herein is presented.SUMMARY

[0009] Techniques are generally described for audio signal processing and, more specifically, dynamically applying audio suppression based on negotiated communications between capture / sender and user / receiver devices based on capture / sender capabilities and user / receiver audio listening preferences.

[0010] Briefly stated, devices, systems and methods for dynamic audio suppression are described that may be applied to captured audio, such as immersive audio in immersive voice and audio services (IVAS). One example method comprises enabling a dynamic audio suppression capability based on a negotiated communication session and selecting a preferred audio content type based on user preferences for audio suppression in audio. The method further comprises sending audio suppression processing information (PI) data based on the preferred audio content, receiving audio bitstream and decoding the audio bitstream to obtain audio content from the audio bitstream. Wherein, a second audio content type of the audio content is suppressed relative to a first audio content type of the audio content when the first audio content type is selected. The method further comprises rendering the audio content in an audio format suitable for playing on the end user device.

[0011] Some voice communication signal chains make use of signal classification and noise suppression algorithms to enhance the desirable components (for e.g. speech component) of an input audio signal (e.g., input audio signal that is captured via microphones on a capture device) and suppress the undesirable components (e.g. noise) of the input audio signal. However, suppressing the undesirable audio components without taking a listener’s preference into account may result inDocket No.: D24173WO01 overall suboptimal experience for the listener. For example, if a capture device captures immersive audio that includes a directional speech component (e.g., located at front center) and a diffused noise component then suppressing the diffused noise completely may result in a collapsed ambiance and may not give the desired immersive experience to the listener. On the other hand, if the listener themselves is in a noisy environment, they may prefer maximum suppression of the undesirable components. Hence, the present disclosure considers that there is a need to take listener’s preference into account and provide improved methods and devices for performing suppression of unwanted audio content.

[0012] According to a first aspect there is provided a dynamic audio suppression method for captured audio content. The method comprises enabling, at a first device, a dynamic audio suppression capability based on a negotiated communication session between the first device and a second device, wherein the second device corresponds to a capture device that is configured to provide the captured audio in at least one of a variety of audio formats. The method further comprises selecting, at the first device, an audio suppression level based on user preferences for audio suppression, sending, from the first device to the second device, audio suppression processing information (PI) data based on the audio suppression level and receiving, at the first device, an audio bitstream from the second device. The method further comprises decoding, at the first device, the audio bitstream to obtain one or more audio signals from the audio bitstream, the one or more audio signals comprising a representation of the captured audio content, wherein, a second audio content type of the one or more audio signals is suppressed relative to a first audio content type of the one or more audio signals based on the audio suppression level and rendering, at the first device, the one or more audio signals in an audio format suitable for playing on an end user device.

[0013] Hereby, the first aspect enables the first device (e.g. a user / receiver device) to indicate to the second device (e.g. a capture / sender device) and adjust various parameters in the capture / sender device based on the listener’s preference with respect to the user / receiver device audio processing. The audio suppression level may be set manually by the listener or automatically and may vary dynamically over time. Using the audio suppression level it is possible to vary the level of audio suppression (such as noise suppression) that has been applied at the second device to the one or more audio signals of the audio bitstream. Accordingly, when the one or more audio signals of the audio bitstream are received at the first device, the second audio content is suppressed relative to theDocket No.: D24173WO01 first audio content type whereby it is not necessary for the first device to perform any further suppression.

[0014] According to a second aspect there is provided a dynamic audio suppression method for captured audio content. The method comprises enabling, at a second device, a dynamic audio suppression capability based on a negotiated communication session between a first device and the second device , wherein the second device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats, capturing, at the second device, the audio content wherein, the audio content contains a first audio content type and a second audio content type and receiving, at the second device, audio suppression processing information (PI) data from the first device, wherein the audio suppression processing information (PI) data indicates one or more of user preferences for an audio suppression level and user preferences for a preferred audio content type, wherein the preferred audio content type corresponds to one of the first audio content type and the second audio content type. The method further comprises performing, at the second device, audio suppression on the captured audio content based on the audio suppression processing information (PI) data to form suppressed audio content wherein a non-preferred audio content type is suppressed relative to the preferred audio content type, encoding, at the second device, the suppressed audio content using a codec to generate a bitstream and outputting, at the second device, the bitstream.

[0015] This second aspect enables the second device to perform audio suppression in a manner indicated by the audio suppression Processing Information (PI) data. Compared to the first device acting as a user / receiver device, the second device has direct access to captured (optionally immersive) audio content allowing the second device to perform enhanced audio suppression compared to performing audio suppression in the first device. For example, the first device may receive downmixed, compressed, quantized or otherwise distorted representation of the audio content whereby audio suppression may be less effective if performed at the first device.

[0016] According to a third aspect there is provided a dynamic audio suppression method for captured audio content. The method comprises enabling, at a first device, a dynamic audio suppression capability based on a negotiated communication session between the first device and a second device, wherein the second device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats, receiving, at the first device, an audio bitstream from the second device and decoding, at the first device, the audioDocket No.: D24173WO01 bitstream to obtain audio content from the audio bitstream and audio description (AD) processing information (PI) data indicating for each of the at least one audio signal of the audio content at least one active audio content type. The method further comprises obtaining processed audio content based on the AD PI data and rendering, at the first device, the processed audio content in an audio format suitable for playing on an end user device.

[0017] The Audio Description (AD) PI data may hereby be utilized by the first device to more accurately determine which audio content types are active in the audio content, which may be used to control audio suppression. The audio suppression may be performed in the first device, or in the second device with the first device selecting one of the active audio content types and indicating to the second device which audio content type is the preferred audio content type. It is also envisaged that the audio suppression is performed in both the first and second device.

[0018] According to a fourth aspect there is provided an audio encoding method for captured audio content. The method comprises enabling, at a second device, a dynamic audio suppression capability based on a negotiated communication session between the second device and a first device, wherein the second device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats. The method further comprises capturing, at the second device, audio content comprising at least one audio signal wherein, the audio contains two types of audio content, determining, based on the audio content, audio description (AD) data indicating a set of active audio content types and generating, AD processing information (PI) data based on the AD data. The method further comprises encoding, at the second device, the audio content and the AD PI data using a codec to generate a bitstream and outputting, at the second device, the bitstream.

[0019] By including AD PI data in the bitstream, the first device acting as a receiver device may more accurately distinguish which audio content types are active. Compared to the first device, the second device has direct access to captured (optionally immersive) audio content allowing the second device to more accurately identify which audio content types are active.

[0020] An additional aspect relates to an apparatus comprising a processor and a memory storing instructions configured to, when executed by the processor, cause the processor to perform the method of any of the first to fourth aspect .Docket No.: D24173WO01

[0021] A further aspect relates to a computer program product comprising instructions configured to, when executed by a processor, cause the processor to perform the method according to any one of the first to fourth aspect.

[0022] A still further aspect relates to a non-transitory computer readable medium comprising instructions configured to, when executed by a processor, cause the processor to perform the method according to any one of the first to fourth aspect.

[0023] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.

[0024] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associated drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims.DESCRIPTION OF THE DRAWINGS

[0025] Aspects of the present invention will be described in more detail with reference to the appended drawings, showing example embodiments.

[0026] Figure 1 illustrates two bytes used to signal an Audio Indicator (AID) value and Suppression Level Indicator (SLI) value, according to some implementations.

[0027] Figure 2 shows an example user interface (UI) for obtaining interactive user input related to the preferred audio content type and desired suppression level.

[0028] Figure 3 is a block diagram illustrating two devices performing an SDP negotiation process, according to some implementations.

[0029] Figure 4A is a block diagram illustrating a sender device according to some implementations.Docket No.: D24173WO01

[0030] Figure 4B is a flowchart illustrating a method for dynamic audio suppression that may be implemented by a sender device according to some implementations.

[0031] Figure 5A is a block diagram illustrating a receiver device according to some implementations.

[0032] Figure 5B is a flowchart illustrating a method for dynamic audio suppression that may be implemented by a receiver device according to some implementations.

[0033] Figure 6 is a block diagram illustrating a receiver device with automatic AID and SLI selection capability, according to some implementations.

[0034] Figure 7A is a block diagram illustrating a sender device with AD data extraction capability, according to some implementations.

[0035] Figure 7B is a flowchart illustrating a method that may be implemented by a sender device with AD data extraction capability, according to some implementations.

[0036] Figure 8 illustrates an AID byte that may be used to signal the AD data according to some implementations.

[0037] Figure 9A is a block diagram illustrating a receiver device configured to perform postprocessing based at least in part on the AD data provided in a forward bitstream.

[0038] Figure 9B is a flowchart illustrating a method that may be implemented by a receiver device configured to perform post-processing based at least in part on the AD data provided in a forward bitstream, according to some implementations.

[0039] Figure 10 illustrates a schematic block diagram of an example device or architecture that may be used to implement various aspects of the present disclosure.DETAILED DESCRIPTION

[0040] In the following description, numerous details are set forth, such as audio device configurations, timings, operations, and the like, in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to one skilled in the art that these specific details are merely examples and not intended to limit the scope of this application.Docket No.: D24173WO01

[0041] Any module, encoder, decoder or Tenderer described herein may be implemented in hardware or software. The same applies to each step in any method presented herein, which may be realized as a corresponding processing module, or vice versa.

[0042] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “at least both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0043] The Immersive Voice and Audio Services (IVAS) codec as standardized by the 3rdGeneration Partnership Project (3GPP) supports a variety of immersive audio formats (including but not limited to Ambisonics, Channel based and Object Audio) offering an immersive experience to a listener capable of rendering to an immersive format (e.g., multi-channel, binaural). Providing an immersive experience also makes it challenging for the audio processing in a capture device (such as noise suppression and echo cancellation) to suppress unwanted audio components while maintaining the overall essence of the immersive capture (for e.g., it may be desirable to maintain phase and level differences between channels). Moreover, in an immersive voice communication use case, if the captured immersive audio includes speech and background ambiance then suppressing the background ambiance completely may result in a collapsed scene which may not be desirable for the listener experiencing the audio content via a receiver device. However, this may depend on the service context and / or the expectation of the listener. In some cases, the listener may want to receive the clean speech content and accepts that some or all background ambience is suppressed, in other cases the listener may want to hear the ambience in addition to speech as this may contribute to aDocket No.: D24173WO01 more natural experience. Hence, there is a need to indicate listener or receiver’s preference with respect to the capture side audio processing.

[0044] While the aspects of the present disclosure may be implemented in an IVAS context for immersive audio it is envisaged that the aspects are not necessarily limited thereto. It is envisaged that the aspects of the present disclosure may be applicable to other types of audio content than immersive audio content and also applicable to any receiver device communicating with a sender device or any sender device communicating with a receiver device.

[0045] In an example implementation, a receiver device can indicate its preference with respect to the type of audio content.

[0046] Additionally or alternatively, a receiver device may indicate its preference with respect to the amount of audio suppression to be applied to unwanted audio content types also referred to as nonpreferred audio content types.

[0047] This indication may be transmitted as part of a IVAS Real Time Protocol (RTP) Payload as described below. The RTP payload may have a payload format as described in 3GPP technical specification 26.253 version 18.6.0, Annex A. IVAS RTP allows Processing Information (PI) data to be sent from a sender device to a receiver device.

[0048] Table 1 indicates a subset of types of forward PI data supported in the 3GPP technical specification 26.253 version 18.6.0. Here “forward” indicates that the PI data is transmitted from media sender device to media receiver device.Table 1Docket No.: D24173WO01

[0049] It is envisaged that the Dynamic Audio Suppression (DAS) PI data and audio description (AD) PI data described herein may be conveyed as PI data. Hereby, the DAS PI data and / or AD PI data may be added as additional PI types to future releases of 3GPP technical specification 26.253. Some of the PI data described herein may be conveyed from receiver device to sender device (e.g. as reverse PI data), e.g. to indicate the user preference from the receiver device. Here, “reverse” indicates that the PI data is transmitted from media receiver device to media sender device. Some of the PI data described herein may also be conveyed from the sender device to the receiver device (e.g. as forward PI data).

[0050] In an example implementation, the DAS PI data type, as indicated in Table 2, may be added to any preexisting PI data types to add DAS capabilities.Table 2

[0051] The DAS PI data describes at least one of the receiver device’s or listener’s preference with respect to the type of audio content (e.g., speech) that should be enhanced, e.g. the preferred audio content type(s) and the amount of suppression to be applied to the background noise. Background noise may be defined as the type of audio content that should be suppressed according to the receiver device’s preference, e.g. any audio content that is different from the preferred audio content type. Any audio content that is different from the preferred audio content type may be referred to as non-pref erred audio content types.

[0052] The background noise may comprise any audio which is different from the preferred type of audio content. For example, if the preferred type of audio content is speech, then the background audio content may be music, noise (e.g. white noise) and ambiance (e.g. the sound of traffic, birdsong and rain). In another example, if the preferred type of audio content is speech and music the background audio content may be noise (e.g. white noise) and ambiance (e.g. the sound of traffic, birdsong and rain). The background noise may be a predetermined type of audio content such as any non-speech audio content.

[0053] The size of DAS PI data according to some implementations is 2 bytes as indicated in Table 2 above and FIG. 1 illustrates an example 2 byte DAS PI data payload.Docket No.: D24173WO01

[0054] The DAS PI data payload contains an Audio Identifier (AID) byte and a Suppression Level Indicator (SLI) byte, as illustrated in FIG. 1. The value of a field of the AID byte may be non-zero for the audio identifier bits (V, M) and the reserved bits in AID may be set to 0, unless defined.

[0055] Likewise, there are four non-reserved bits of the SLI byte shown in FIG. 1 while the four remaining bits are reserved and set to 0, unless defined.

[0056] The AID bits indicate the user’s preferred audio content type, e.g. at least one of voice or speech (V) and music (M). In the example of FIG. 1 the AID byte has six reserved bits (which e.g. are set to zero) and two AID bits, a V bit for voice or speech and an M bit for music. In this example, when the V bit is set to one and the M bit is set to zero, this indicates that the preferred audio content type is speech. When the V bit is set to zero and the M bit is set to one, this indicates that the preferred type of audio content is music. When the V bit and the M bit are both set to zero, this indicates that there is no preferred audio content type. Notably, it is also possible to set both the V bit and the M bit to one which would indicate that both voice and music are preferred audio content types.

[0057] Other variants of the AID byte are also envisaged. For instance, one or more of the reserved bits may be used to signal additional preferred audio content types, such as ambiance. It is also envisaged that instead of signaling the preferred audio content types with the AID byte, the AID byte could be used to signal non-pref erred audio content types. The sender device may then be configured to treat any audio content type signaled by the AID byte as non-preferred and any other audio content type as preferred.

[0058] An SLI byte with four non-reserved bits, as shown in Table 3, allows specifying a desired degree of suppression for audio content types other than the preferred audio content type(s) specified by the AID, e.g. the non-preferred audio content types. The four non-reserved bits of the SLI byte of FIG. 1 indicate values from 0 to 15 wherein the expected level of audio suppression of the nonpreferred audio content type(s) is proportional to the indicated value. The expected level of audio suppression may be proportional to the value indicated by the SLI byte in a logarithmic or approximate logarithmic domain. That is, the level of audio suppression as a function of the SLI byte value follows a logarithmic or approximate logarithmic function. Examples of approximate logarithmic functions that may be used are the A-law and p-law of the G.711 standard.Docket No.: D24173WO01

[0059] For example, an SLI value of 0 indicates no audio suppression or minimum audio suppression and an SLI value of 15 indicates maximum audio suppression. Minimum audio suppression is the minimum audio suppression offered by the sender device and maximum audio suppression may be the maximum audio suppression offered by the sender device. The minimum and maximum suppression values may also be negotiated between sender device and receiver device, e.g. during session setup.Table 3

[0060] Other variants of the SLI byte with less than four non-reserved bits or more than four nonreserved bits are also envisaged for signaling values in a smaller or greater range than 0 to 15. For example, N bits may be non-reserved in the SLI byte wherein N is an integer that fulfills 2 < N < 8 and allows SLI values between 0 and 2N- 1 to be signaled.

[0061] The listener preference (also referred to as user preference) at the receiver device with respect to the type of audio content, e.g. used to determine the AID byte, may be determined by any appropriate method, including but not limited to: a prior stored user preference, an interactive selection by a user (e.g., selection of ‘Music’ or ‘Speech’ in the user interface of FIG. 2), or an automated process based on an algorithm (e.g., as a result of selection of ‘Auto’ in the user interface of FIG. 2).

[0062] The prior stored user preference may be a predetermined preferred audio content type. For example, the receiver may be associated with a conference system whereby the user has selected speech as the preferred content type during configuration of the conference system. As another example, the user may have selected speech and ambiance as the preferred content types for a smartphone when participating in video calls.

[0063] Interactive selection by the user may comprise the user selecting one or more preferred content types from e.g. a list presented on a user interface (UI). The interactive selection may be changed during streaming e.g. due to the user selecting one or more different preferred audio contentDocket No.: D24173WO01 types. The receiver device may accordingly convey an updated AID byte to the sender device in response to the user selecting a different set of one or more preferred audio content types.

[0064] FIG. 2 illustrates an example of a user interface (UI) 100 which may be present in a receiver device’s communication application. The UI 100 comprises a slider 101 for setting the desired suppression level, a list 102 of options for setting preferred audio content types, and a list 103 of options for selecting the desired bitrate.

[0065] The UI 100 allows the user to select “Speech” or “Music” as the preferred audio content type, e.g. the “Call preference”, from a list 102 of options. The UI 100 also presents the user with a third option “Auto” in the list 102. When “Speech” or “Music” is selected, speech or music is signaled as the preferred audio content type using an AID byte to the sender device. When “Auto” is selected the preferred audio content type is selected automatically using an automated process based on an algorithm. Implementations are also envisaged wherein instead of the user selecting the preferred audio content types, the UI 100 presents a list of different audio content types the user may select for suppression, e.g. allowing the user to select the non-preferred audio content type(s) instead of the preferred audio content type(s).

[0066] For example, an automated process based on an algorithm for selecting the preferred content type may be based on signal analysis methods or based on a service context.

[0067] An example signal analysis method may be a classifier model such as a Gaussian Mixture Model or GMM (see 3GPP TS 26.445 version 13.6.0 Release 13) or another classifier based analysis such as a machine learning or ML process.

[0068] In some implementations, the audio content received by the receiver device is processed with a classifier to determine if the audio content comprises speech, music, both or neither. The audio content received by the receiver device may comprise a mix of two or more audio content types such as speech, music, and noise. The classifier may be configured to determine to what extent the received audio content is dominated by each audio content type.

[0069] In response to the classifier model detecting that speech is the dominant audio content type, the preferred audio content type may be set to speech. In response to the classifier detecting that music is the dominant audio content type, the preferred content type may be set to music. In response to the classifier detecting that both speech and music are approximately equally dominant, ,Docket No.: D24173WO01 the preferred audio content type may be set to speech and music. In response to the classifier detecting that the audio content is dominated by an audio content type that is neither speech nor music (e.g., noise), no preferred audio content type may be set or a default preferred audio content type, such as speech, may be set. Hereby, with a classifier model, the receiver device is able to automatically select the preferred audio content type based on the audio content type that is determined to be most dominant in the received audio content.

[0070] An example service context may be based on a priori knowledge that the receiver device’s application has a voice focus (e.g. in a voice call or a conference call), whereby the preferred audio content type is speech, or that it has a music focus (e.g. in a music streaming application), whereby the preferred audio content type is music. For instance, the receiver may be processing audio associated with a video call in a smartphone whereby the preferred audio content type may be automatically set to speech and ambiance based on this context. For video calls it is often desirable to preserve the background ambience in addition to speech to provide better immersion for the video call participants.

[0071] The user preference with respect to the audio suppression level, used e.g. in determination of the SLI byte, may be determined by any appropriate method, included but not limited to: a prior stored user preference, an interactive selection by a user (e.g., moving the slider 101 titled “Noise” in the example UI 100 of FIG. 2), or an automated process based on an algorithm.

[0072] A prior stored user preference may be a selection of a suppression level the user has previously selected. The suppression level may be derived from other user parameters available to the receiver such as if speech enhancement or closed captions are activated, both of which may indicate that the user desires to make speech more intelligible whereby the user may benefit from a higher suppression level.

[0073] An interactive selection of the suppression level may involve the user entering a desired suppression level or selecting a desired suppression level from a list of suppression levels by suitable means (for example using a slider 101 as shown in FIG. 2). The suppression level may be selected from a list of options such as “OFF”, “LOW”, “MODERATE” or “HIGH” with each selection being associated with progressively higher suppression level. The suppression level may be selected using a slider 101 as shown in FIG. 2 with the user being able to set the suppression level anywhere between 0% and 100%.Docket No.: D24173WO01

[0074] The selected suppression level is transmitted to the sender device using the SLI byte. In the example SLI byte of FIG. 1, four bits of the SLI byte are used to signal the SLI level allowing 16 suppression levels to be signaled (or 15 levels and one “off’ state or minimum noise suppression state). The user preference of the suppression level may not necessarily map directly or one-to-one to one of the 16 SLI levels and in such cases the user preference may be mapped to one of the 16 SLI levels (e.g. the closest one) whereby the resulting SLI level is conveyed to the sender device.

[0075] If the SLI level is set to 0 this could be used as an indicator that no noise suppression is preferred or, alternatively, it could mean that minimum noise suppression is preferred. Wherein the minimum noise suppression may be the minimum suppression offered by the noise suppression algorithm of the sender device or a minimum suppression level that has been negotiated between sender device and receiver device. If the SLI level is set to its maximum level (e.g., 16) then it means maximum noise suppression is preferred, wherein maximum noise suppression may be the maximum suppression offered by the noise suppression algorithm of the sender device or a maximum suppression level that has been negotiated between sender device and receiver device.

[0076] The UI 100 further comprises a number of bitrate options 103 allowing the user to select a desired bitrate for the received audio content. The bitrates shown in FIG. 2 include 24.4 kbit / s, 64 kbit / s, 96 kbit / s, 160 kbit / s and 256 kbit / s, but this list of options is merely an example and the person skilled in the art will appreciate that different bitrate options may be presented instead. In some implementations, the selection of the bitrate influences the automatic selection of SLI level, as described below.

[0077] The bitrate may also be automatically determined based on network capacity, wherein the receiver device is configured to automatically request the highest bitrate possible based on the capacity of the network .

[0078] An example of automated processing according to an algorithm may be based on signal analysis methods that dynamically evaluate the signal levels (e.g., computing signal-to-noise levels, FFT or DFT processing, etc.) of the preferred audio content type relative to the non-preferred audio content types to update the suppression level of the user preference (e.g., continuously, periodically, or when triggered). For example, the suppression level for speech content may be automatically increased in response to the signal-to-noise ratio for speech being below a predetermined threshold.Docket No.: D24173WO01

[0079] In some implementations, noise in the environment in which the receiver device is located is measured (e.g. using a microphone) and used to automatically control the suppression level. If the noise level in the environment is high the suppression level may be increased and if the noise level of the environment is low the suppression level may be decreased.

[0080] To dynamically evaluate the signal levels a source separator may be employed in the receiver device which separates the preferred audio content type(s) from the received audio content so as to enable calculation of a signal-to-noise level.

[0081] In the UI 100 of FIG. 2, a user of a receiver device may set the user preferences manually, including but not limited to a bitrate of the audio, the audio suppression level, and the preferred content type (e.g., call preference). With respect to the preferred suppression level a user can request to increase or decrease the level of noise suppression by adjusting a slider bar (e.g., a slider with a value indicated in a range between 0% and 100%). With respect to the preferred type of audio content the user can select the content type using radio buttons based on the available content types from the capture device (e.g., a radio button for ‘Speech’, ‘Music’, etc.).

[0082] The available content types may be predetermined (e.g. speech / music / auto) or determined by a classifier of the receiver device. As a further option, the sender device may send audio description (AD) PI data to the receiver device indicating which audio content types are active, wherein the available content types are determined based on the AD PI data.

[0083] The sender device (e.g., capture device) receives a request from the receiver device, e.g. in form of an AID byte and adapts the tuning of the audio processing in the sender device accordingly. The classification of background noise at the sender device can be done based on the receiver device’s preference signaled using the V and M bits of the AID byte of FIG. 1.

[0084] To initiate DAS processing, two devices may negotiate their respective capabilities to establish that DAS processing is supported.

[0085] FIG. 3 is a block diagram illustrating a first device 10 and a second device 20 according to some implementations.

[0086] The first device 10 sends an SDP offer to the second device 20 and the second device 20 responsively sends an SDP answer to the first device 10.Docket No.: D24173WO01

[0087] As illustrated in FIG. 3, a first device 10 and second device 20 negotiate during the Session Description Protocol (SDP) negotiation to identify which SDP capabilities are supported by both devices 10, 20. In this example, the first device 10 may be a receiver device and the second device 20 may be a sender device. The SDP negotiation may comprise indicating if each device 10, 20 supports dynamic audio suppression (SDP type rdas of Table 2). Additionally, the SDP negotiation may comprise indicating if each device 10, 20 also supports any of the SDP types of Table 1.

[0088] The SDP negotiation may involve one of the devices 10, 20 sending an SDP Offer to the other device. In FIG. 3 the first device 10 sends an SDP Offer to the second device 20. The SDP Offer indicates one of more SDP types supported by the first device 10. In response, the second device 20 sends an SDP Answer back to the first device 10 wherein the SDP Answer indicates which out of the one or more SDP types of the SDP Offer the second device 20 supports. In response to both devices indicating support for DAS (see Table 2) during the SDP negotiation, an audio session with DAS is established. When the first and second device operate in an IVAS context the audio session is referred to as an IVAS session.

[0089] After the audio (IVAS) session is established, the described user preferences may be dynamically exchanged between the first and second devices 10, 20. For example, the first device 10 sends a 2 byte DAS PI data payload to the second device 20, to control the audio processing in the second device 20.

[0090] In general, each of the first and second device 10, 20 may act as either, or both, of a sender / capture device and receiver / user device. For example, in teleconferencing implementations the first device 10 may both send and receive audio content and the second device 20 may both send and receive audio content while PI data is exchanged between both devices 10, 20. In some examples, the protocol supports that either of the devices 10, 20 may initiate the exchange of information to indicate support for IVAS and the related parameters.

[0091] In an example implementation, a first device 10 (e.g. the receiver device) may send an SDP Offer indicating that the first device 10 supports DAS. Subsequently, the second device 20 may provide an SDP Answer indicating whether it supports DAS or not. Once both the first and second devices 10, 20 confirm that DAS is supported, the first device 10, at runtime during the audio session, may send backward PI data as part of the PI data in the RTP payload. The backward PI dataDocket No.: D24173WO01 comprises DAS data payload indicating the preferred audio content type and / or suppression level indicated (e.g. using the AID byte and / or SLI byte as previously described).

[0092] Once the audio session is established with the SDP negotiation (Offer and Answer), the first and second device 10, 20 start exchanging audio data, the second device 20 (acting as a sender device in this example) encodes audio using an IV AS encoder module and generates a real-time protocol (RTP) payload, which may include PI data, and transmits it to the first device 10 (acting as a receiver device in this example). The first device 10 decodes the IVAS RTP bitstream and renders the audio to desired format (e.g., loudspeaker, binaural, etc.). An example signal flow for IVAS with Dynamic Audio Suppression (DAS) is illustrated for a sender device in FIG. 4A and receiver device in FIG. 5A.

[0093] FIG. 4A is a block diagram illustrating a sender device 40a arranged according to some implementations.

[0094] The sender device 40a comprises a capture front-end module 41, an audio pre-processor module 42, an IVAS encoder module 43, an RTP packer module 44, an RTP depacker module 46, and a DAS data extractor module 45.

[0095] The capture front-end module 41 obtains as an input of one or more audio signals. The output of the capture front-end module 41 is provided as a first input to the audio pre-processor module 42. The output of the pre-processor module 42 is provided as an input to the IVAS encoder module 43, and the output of the IVAS encoder module 43 is provided to the RTP packer module 44. The output of the RTP packer module forms a forward IVAS RTP bitstream (an audio bitstream). The RTP depacker module 46 obtains as an input a backward IVAS RTP bitstream and the output of the RTP depacker module 46 is provided as input to the DAS data extractor module 45. The output of the DAS data extractor module 45 is provided as a second input to the audio preprocessor module 42.

[0096] In this example, the sender device 40a is also a capture device that comprises or communicates with microphones for capturing audio content (e.g., input from microphones).

[0097] The sender device 40a may comprise or be connected to one or more microphones, each microphone providing a separate input audio signal. In FIG. 4A the microphone(s) are provided separately from the sender device 40a, and the corresponding input audio signals are provided to theDocket No.: D24173WO01 capture front-end module 41. In some implementations, the input audio signal(s) obtained by the capture front-end module 41 are analogue audio signals whereby the capture front-end module 41 is configured to convert the analogue input audio signals to digital audio signals. To this end, the capture front-end module 41 may comprise an analogue to digital converter (ADC).

[0098] The audio signals output by the capture front-end module 41 form audio content comprising one or more audio signals. The audio content may be immersive audio content. The audio content is provided to the audio pre-processor module 42. The audio pre-processor module 42 is configured to process the one or more audio signals with an audio processing scheme. The audio processing scheme may comprise one or more of noise suppression, echo cancellation, and equalization. The audio processing scheme may be at least partially controlled by the backward DAS PI data provided in RTP packages of the backwards IVAS RTP bitstream from the receiver device.

[0099] As illustrated in FIG. 4A the sender device 40a obtains as an input the backwards IVAS RTP bitstream from the receiver device. The backwards IVAS RTP bitstream is provided to the RTP depacker module 46 which depacks the RTP packages to obtain the backwards PI data stored therein. The backwards PI data is provided to the DAS data extractor module 45, which extracts the DAS data from the backwards PI data. As indicated above, the DAS data may comprise an indicator of the preferred audio content type from the receiver device (e.g., an AID byte indicating preferred or non-preferred audio content types such as speech, music, etc.) and / or the suppression level (e.g. signaled using an SLI byte) to be applied to non-preferred audio content types. Information indicating one or more preferred (or non-preferred) audio content type and / or the suppression level is passed to the audio pre-processor module 42 whereby the processing scheme applied by the audio pre-processor module 42 is adapted based on this information. The pre-processor module 42 is configured to suppress audio content types different from the one or more preferred audio content types and / or wherein the suppression level indicates the amount of the suppression to be applied. For example, if the receiver device indicates speech as the preferred audio content type all other types of audio content will be suppressed by the audio pre-processor module 42 of the sender device 40a as per the suppression level indicated in the backwards PI data relating to DAS.

[0100] In some implementations, the sender device 40a comprises a classifier module configured to determine for each audio signal of the input audio signals a set of active audio content types for each audio signal. The set of active audio content types may comprise a single active audio content type or a plurality of active audio content types. It is envisaged that the set of active audioDocket No.: D24173WO01 content types may be empty, which may be the case for audio signals of the immersive audio content comprising silence or white noise, whereby the classifier does not detect any audio content type. Alternatively, if no audio content type other than the ambience type is detected, the ambiance audio content type may be assigned.

[0101] Based on the classification, the audio pre-processor module 42 may in some implementations be configured to perform suppression of the non-preferred audio content types on those audio signals having an identified audio content type other than the preferred audio content type.

[0102] In some implementations, the audio pre-processor module 42 is configured to perform more aggressive suppression of the non-preferred audio content types for those audio signals comprising identified audio content types other than the preferred audio content type and less aggressive suppression of non-preferred audio content types for audio signals also having a preferred audio content type.

[0103] In some implementations, the audio pre-processor module 42 is configured to omit audio signals not having the preferred audio content type and provide a fewer number of audio signals to the subsequent IVAS encoder module 43. As an example, in object based coding, objects containing non-preferred audio may be omitted by the pre-processor module 42 and do need to be coded by the IVAS encoder module 43 or transmitted to the receiver device, saving both computational resources and bandwidth. Alternatively, the audio pre-processor module 42 is configured to signal to the IVAS encoder module 43 which audio signals (or audio objects) do not include the preferred audio content type, whereby the IVAS encoder module 43 may be configured to refrain from encoding these audio signals or objects and / or refrain from encoding these audio signals or objects with a high bitrate.

[0104] Omission of audio signals not having the preferred audio content type is an aggressive form of audio suppression and, in some implementations, the pre-processor module 42 is configured to perform this type of audio suppression in response to the SLI indicating a maximum suppression level and / or a suppression level exceeding a predetermined threshold.

[0105] The output of the audio pre-processor module 42 is at least one pre-processed audio signal that includes the preferred audio content types and may also include the suppressed nonpreferred audio content types. The at least one pre-processed audio signal is provided to the IVASDocket No.: D24173WO01 encoder 43 configured to encode the at least one pre-processed audio signal. The IVAS encoder 43 may be configured to encode the at least one pre-processed audio signal as per 3GPP Technical Specification 26.253 to generate an IVAS bitstream. When a plurality of audio signals are provided to the IVAS encoder module 43, the IVAS encoder module 43 may be configured to encode each of the audio signals with a bitrate which depends on the number of audio signals. That is, the per-signal bitrate may be higher when fewer audio signals are provided and lower when more audio signals are provided. For example, the IVAS encoder module 43 may be configured to operate with a constant bitrate which is shared between the audio signals.

[0106] As described above, in some implementations the pre-processor module 42 may omit audio signals that do not have the preferred audio content type whereby the IVAS encoder module 43 obtains a reduced number of audio signals. This may allow encoding the remaining audio signals with a higher individual bitrate which increases the quality of the preferred audio content types at playback at the receiver device. The IVAS encoder module 43 may obtain information from the preprocessor module 42 indicating that one or more audio signals are to be encoded with a lower bitrate (or not encoded at all) which leaves more of the bitrate budget for encoding of the remaining audio signals.

[0107] The IVAS bitstream is subsequently packed into RTP packages by the RTP packer module 44. The RTP packer module 44 is configured to perform RTP packing. The RTP packer module 44 may perform RTP packing as described in 3GPP Technical Specification 26.253 version 18.6.0 Annex A to form a forward IVAS RTP bitstream. The forward IVAS RTP bitstream may then be transmitted to the receiver device.

[0108] FIG. 4B is a flowchart illustrating a method for dynamic audio suppression for captured audio content. The method may e.g. be implemented by the sender device 40a of FIG. 4A.

[0109] Example methods of FIG. 4B may be partitioned into steps, such as steps SI 1 - S16. The various steps may be described as operations, processes, methods, steps, acts, blocks or functions. The steps are not necessarily performed in the order indicated. In some implementations, one or more of the steps may be performed concurrently. Moreover, some example methods may include more or fewer steps than shown and / or described.

[0110] Processing may commence at step Sil titled “Enabling DAS capability”, which comprises enabling, at the sender device, a dynamic audio suppression capability based on aDocket No.: D24173WO01 negotiated communication session between the sender device and a receiver device, wherein the sender device corresponds to a capture device that is configured to provide the captured audio in a variety of audio formats. During the step Sil the DAS session negotiation described above may be performed.

[0111] Processing may continue to step S12 titled “Capturing audio content”. Step S12 comprises capturing, at the sender device, the audio content wherein for example the audio content contains a first audio content type and a second audio content type. For example, the audio content is captured using one or more microphones in communication with the capture front-end 41 described above.

[0112] Processing may continue to step S13 titled “Receiving audio suppression PI data”. Step S13 comprises receiving (e.g. by the RTP depacker module 46), at the sender device, audio suppression PI data from the receiver device, wherein the audio suppression PI data indicates one or more of user preferences for an audio suppression level and user preferences for a preferred audio content type, wherein the preferred audio content type corresponds to one of the first audio content type and the second audio content type. The PI data may be provided as RTP packets which are depacked and extracted by the RTP depacker module 46 and DAS data extractor module 45 described in connection with FIG. 4A. For example, the audio suppression PI data comprises an AID byte and / or an SLI byte as described above.

[0113] Processing may continue to step S14 titled “Performing audio suppression on audio content”. S14 comprises performing (e.g. by the audio pre-processing module 42), at the sender device, audio suppression on the captured audio content based on the audio suppression PI data to form suppressed audio content wherein a non-preferred audio content type is suppressed relative to the preferred audio content type. For example, if the PI data from the receiver device indicates speech as a preferred audio content type, the audio suppression involves suppression of non-speech audio content.

[0114] Processing may continue to step S15 titled “Encoding suppressed audio content to generate bitstream”. Step S15 comprises encoding (e.g. by the IVAS encoder 43), at the sender device, the suppressed audio content using a codec to generate a bitstream. For example, the IVAS codec is used wherein the suppressed audio content is encoded using an IVAS encoder.Docket No.: D24173WO01

[0115] Processing may continue to step S16 titled “Outputting bitstream”. Step S16 comprises outputting (e.g. by the RTP packer module 44), by the sender device, the bitstream. The bitstream may be temporarily stored and / or conveyed to the receiver device.

[0116] FIG. 5A is a block diagram illustrating a receiver device 50a according to some implementations.

[0117] The receiver device 50a comprises an RTP depacker module 51, an IVAS decoder module 52, audio post-processor module 53, playback front-end module 54, RTP packer module 55, audio preference type selector module 57 (also referred to as AID selector module), and audio suppression level selector module 56 (also referred to as SLI selector module).

[0118] The RTP depacker module 51 obtains as an input a forward IVAS RTP bitstream. The output of the RTP depacker module 51 is provided as input to the IVAS decoder 52 and an output of the IVAS decoder 52 is provided as an input to the audio post-processing module 53. An output of the audio post-processing module 53 is provided to an input of the playback front-end module 54 which outputs output audio for playback on loudspeakers / headphones. The AID selector module 57 obtains user input and provides an output that is used as a first input to the RTP packer module 55. The user input is also used as input to the SLI selector module 56, the output of which is provided as a second input to the RTP packer module 55. Optionally, an output of the AID selector module 57 is provided as input to the SLI selector module 56 in addition to, or instead of, the user input. The output of the RTP packer module 55 forms a backward IVAS RTP bitstream.

[0119] The receiver device 50a receives the forward IVAS RTP bitstream (e.g. from the sender device 40a of FIG. 4A) and unpacks it using the RTP depacker module 51 to obtain an IVAS bitstream. RTP depacker module 51 performs the inverse process of the RTP packer module 44 (see FIG. 4A). The IVAS bitstream is provided to the IVAS decoder module 52 which decodes the IVAS bitstream to obtain one or more decoded audio signals also referred to as decoded audio content. The decoded audio content is provided to the audio post-processor module 53 configured to perform an audio post-processing scheme. The audio post-processing scheme of the audio post-processor module 53 may comprise rendering the one or more decoded audio signals to the desired output format (e.g., 5.1.2, 7.1.4, binaural) depending on the playback device (end user device).

[0120] The one or more rendered audio signals may subsequently be provided to the playback front-end module 54, which provides the one or more rendered signals to theDocket No.: D24173WO01 corresponding playback devices (e.g., loudspeakers, headphones) for playback. Optionally, the playback front-end module 54 may comprise a digital to analogue converter (DAC) which converts the digitally coded (optionally rendered) audio signals to analogue signals suitable for providing to playback devices.

[0121] In FIG. 5A the receiver device 50a may also obtain user input indicating the user preference with respect to the preferred audio content type(s) (e.g. speech or music). For example, the user of the receiver device 50a has selected one or more content types from a pre-defined list of audio content types including but not limited to speech and music (e.g. see UI 100 of FIG. 2 illustrating how the selection may be presented to the user as a selectable call preference). Additionally or alternatively, the receiver device 50a may be configured to select the content type automatically or allow the user to select automatic content type selection as described below in connection with FIG. 6. In FIG. 5A, the AID selector module 57 obtains the user input and outputs the selected content type, e.g. in the form of an AID byte to the RTP packer module 55.

[0122] The receiver device 50a may obtain user input indicating user preference with respect to suppression level. Additionally or alternatively, the receiver device 50a may be configured to automatically determine the suppression level as described below.

[0123] In FIG. 5A the SLI module 56 obtains the user input and outputs the preferred suppression level, e.g. in the form of an SLI byte. As indicated above, in some implementations the suppression level selection is made using e.g. slider providing the user with a selection range spanning e.g. 0% to 100% with a discrete step length of 1% or even smaller whereas the SLI byte may be limited to indicating one of 16 discrete values (in implementations with 4 non-reserved bits). Hereby, the SLI selector module 56 may be configured to map the user input to an appropriate discrete SLI level.

[0124] Optionally, the output of the AID selector 57 may be provided as input to the SLI selector module 56 in addition to or instead of the user input being provided to the SLI selector module 56, as indicated with the dashed line in FIG. 5A going from the AID selector module 57 to the SLI selector module 56. This enables the selection of the suppression level to be made conditionally of the preferred audio content type selection. For example, this could be used to select the suppression level differently depending on the type of audio content that is preferred.Docket No.: D24173WO01

[0125] The suppression level and / or preferred audio content type is provided to the RTP packer module 55 which packages the SLI byte and / or AID byte into RTP packages (e.g. as per the RTP payload format described in 3GPP Technical Specification 26.253 version 18.6.0, Annex A) to form a backwards IVAS RTP bitstream. The backwards RTP IVAS bitstream is then transmitted to the sender device (see FIG. 4A).

[0126] FIG. 6 is a block diagram illustrating a receiver device 50b employing automatic selection of the preferred audio content type and suppression level.

[0127] The receiver device 50b comprises an RTP depacker module 51, an IVAS decoder module 52, an audio post-processor module 53, a playback front-end module 54, an RTP packer module 55, an audio preference type selector module 57 (also referred to as AID selector module), an audio suppression level selector module 56 (also referred to as SLI selector module), a classifier module 58 and a signal-to-noise (S / N) level calculator module 59.

[0128] The RTP depacker module 51 obtains as an input a forward IVAS RTP bitstream. A first output of the RTP depacker module 51 is provided as input to the IVAS decoder 52 and a second output of the RTP depacker module 51 is provided as a first input to the AID selector module 57. An output of the IVAS decoder 52 is provided as an input to the audio post-processing module 53. A first output of the audio post-processing module 53 is provided to an input of the playback front-end module 54 which outputs output audio for playback on loudspeakers / headphones. A second and third output of the audio post-processing module 53 is provided as input to the classifier module 58 and S / N calculator module 59, respectively. The AID selector module 57 obtains user input and provides a first output that is used as a second input to the S / N calculator module 59 and a second output that is used as a first input to the RTP packer module 55. The output of the S / N calculator module 59 is provided as input to the SLI selector module 56, which further obtains user input as a second input. The output of the SLI selector module 56 is provided as a second input to the RTP packer module 55. The output of the RTP packer module 55 forms a backward IVAS RTP bitstream.

[0129] The RTP depacker module 51, IVAS decoder module 52, audio post-processor module 53, and playback front-end module 54 may be equivalent to the corresponding modules of FIG. 5A. A difference compared to the receiver device 50a of FIG. 5A is that the receiver device 50b does not require a user input and instead selects the preferred audio content type andDocket No.: D24173WO01 suppression level, e.g. fully automatically, based on the audio output by the audio post-processor module 53.

[0130] In the example receiver device of FIG. 6, a classifier module 58 is used to determine one or more content types that are present in the decoded audio signal (and optionally which audio content type(s) is the dominating one) and provide information indicating the one or more content types (and optionally a confidence level for each content type) to the AID selector module 57. The AID selector module 57 may be configured to select the preferred content type to signal in the backward PI data based on the result of the classification, and optionally based on the confidence level at which each content type is detected by the classifier 58.

[0131] The AID selector module 57 may determine the selected content type based on the classification result or based partially on the classification result. As an example of the former, the AID selector module 57 may set the preferred audio content type(s) to be equal to the dominant content type(s) detected by the classifier module 58 (or optionally equal to the content types detected by the classifier module 58 with a confidence level exceeding a predetermined threshold). In some implementations, when determining the dominant content type(s), the classifier in the receiver device may take into account the level of audio suppression and the type of audio content that has been suppressed by the sender device and compensate for the effect of this suppression. This enables the determination to be made based on the relative strength of the different content types in the captured signal(s) of the sender device prior to audio suppression, rather than on the audio signals after audio suppression. As an example of the latter (e.g. the selected content type is based partially on the classification result), the AID selector module 57 may consider a previously set user preference, context information and / or interactive user input in addition to the classifier result. For instance, in a video call context speech and ambiance may be the preferred audio content types, but in response to the classifier 58 identifying (optionally with a very high confidence) music in the audio content, music may be automatically set as an preferred audio content type in addition to speech and ambiance or instead of ambiance.

[0132] The suppression level may be determined automatically using a signal-to-noise (S / N) level calculator module 59, which is configured to calculate the signal-to-noise level of the preferred audio content type(s) selected by the AID selector module 57 with respect to the other content types of the decoded audio signal. Hereby, to guide the S / N level calculator module 59 with respect to which content type(s) are preferred and which content types are to be treated as non-preferredDocket No.: D24173WO01 content types (noise) the S / N level calculator module 59 may be provided downstream of the AID selector module 57 as shown in FIG. 6. The signal-to-noise level calculator module 59 obtains the information indicating the preferred audio content type(s) from the AID selector 57 and calculates the signal-to-noise level for the preferred audio content type(s). The resulting signal-to-noise level is provided to the SLI selector module 56 which selects a suppression level based on the signal-to- noise level and signals the selected level using an SLI byte. For example, the SLI selector 56 may be configured to select a higher suppression level in response to the signal-to-noise level being lower or vice versa. Additionally, the SLI selector module 56 may obtain information indicating the noise, or at least the amount of noise, in the environment of the receiver device and select the suppression level based on the noise of the environment. For example, if the environment suffers from high levels of noise a higher SLI may be selected by the SLI selector module 56 compared to if the environment is less noisy.

[0133] In some implementations, the SLI selector 56 is configured to determine the SLI level based on the bitrate. For example, when a lower bitrate such as 24.4 kbit / s or 64 kbit / s is used a higher SLI level may be signaled compared to when a higher bitrate such as 160 kbit / s or 256 kbit / s is used, since more aggressive noise suppression is beneficial for intelligibility at lower bitrates. As an illustrative example, if the signal-to-noise level is X dB this may result in the signaled SLI level being 7 when the bitrate is 256 kbit / s. On the other hand, if the signal-to-noise level is X dB, but the bitrate is 24.4 kbit / s, this may result in the signaled SLI level being 10.

[0134] The preferred audio content type(s) determined by the AID selector 57 and the suppression level determined by the SLI selector 56 may be represented using an AID byte and an SLI byte which are provided to the RTP packer module 55 which packs this information as PI data into a backwards IVAS RTP bitstream conveyed to the sender device.

[0135] The receiver device 50b of FIG. 6 may operate in a fully automatic mode that does not require any user input. However, it is envisaged that user input may be provided and used to override the automatic operation or complement the automatic operation.

[0136] In the fully automatic mode, the AID selector module 57 is controlled by the classifier module 58 and the SLI selector module 56 is controlled by the signal-to-noise level calculator module 59 as well as the AID information output by the AID selector module 57.Docket No.: D24173WO01

[0137] In an AID automatic mode, the AID selector module 57 is controlled by the classifier module 58 and the SLI selector module 56 is controlled using user input. In this mode the user may not need to specify the selected content type (as this is automatically determined) and may specify the desired suppression level (e.g. by moving the slider 101 of the UI 100 shown in FIG. 2). In the AID automatic mode, the S / N level calculator module 59 may be deactivated or omitted. The AID automatic mode may correspond the “auto” setting in the list 102 of UI 100 of FIG. 2

[0138] In an SLI automatic mode, the AID selector module 57 is controlled by the user input and the SLI selector module 56 is controlled by the S / N level calculator module 59 and the AID information output by the AID selector module 57. In this mode the user may not need to specify the desired suppression level (as this automatically determined) and may specify the selected contenttype (e.g. by pressing a button in the UI 100 shown in FIG. 2). In the SLI automatic mode, the classifier module 58 may be deactivated or omitted. In the SLI automatic mode the bitrate may also be used by the SLI selector module 56 to set the SLI level, wherein the SLI selector module 56 may be configured to generally select higher SLI levels for lower bitrates or vice versa.

[0139] In a fully manual mode, the AID selector module 57 and SLI selector module 59 are both controlled using user input as described in connection with FIG. 5A. Optionally, the audio content types presented to the user for selection may be based on a classifier output.

[0140] FIG. 5B is a flowchart illustrating a method for dynamic audio suppression method for captured audio content. The method may be implemented by the receiver device 50a or 50b of FIG. 5A or 6.

[0141] Example methods of FIG. 5B may be partitioned into steps, such as steps S21 - S26. The various steps may be described as operations, processes, methods, steps, acts, blocks or functions. The steps are not necessarily performed in the order indicated. In some implementations, one or more of the steps may be performed concurrently. Moreover, some example methods may include more or fewer steps than shown and / or described.

[0142] Processing may commence at step S21 titled “Enabling DAS capability” comprises enabling, at the receiver device, a dynamic audio suppression capability based on a negotiated communication session between the receiver device and a sender device, wherein the sender device corresponds to a capture device that is configured to provide the captured audio in a variety of audio formats.Docket No.: D24173WO01

[0143] Processing may continue to step S22 titled “Selecting audio suppression level”. Step522 comprises selecting (e.g. by the SLI selector module 57), at the receiver device, an audio suppression level based on user preferences for audio suppression. The selection may be manual, automatic or semi-automatic as described above. It is envisaged that step S22 may be replaced with a step involving selecting (manually or automatically) one or more audio content types as preferred audio content or that step S22 may further comprise selecting (manually or automatically) one or more preferred audio content types in addition to selecting the audio suppression level.

[0144] Processing may continue to step S23 titled “Sending audio suppression PI data”. Step523 comprises sending (e.g. by the RTP packer module 55), from the receiver device to the sender device, audio suppression PI data based on the audio suppression level. The audio suppression PI data may comprise an SLI byte (and / or an AID byte).

[0145] Processing may continue to step S24 titled “Receiving audio bitstream”. Step S24 comprises receiving (e.g. by the RTP depacker module 51), at the receiver device, an audio bitstream from the sender device.

[0146] Processing may continue to step S25 titled “Decoding audio bitstream”. Step S25 comprises decoding (e.g. by the IVAS decoder 52), at the receiver device, the audio bitstream to obtain one or more audio signals from the audio bitstream, the one or more audio signals comprising a decoded representation of the captured audio content. The audio content carried by the bitstream and obtained by the receiver device after decoding of the bitstream has been processed by the sender device such that a second audio content type of the one or more audio signals is suppressed relative to a first audio content type of the one or more audio signals based on the audio suppression level.

[0147] Processing may continue to step S26 titled “Rendering audio signals”. Step S26 comprises rendering (e.g. by the playback front-end module 54), at the receiver device, the one or more audio signals in an audio format suitable for playing on an end user (e.g. playback) device.The end user device may e.g. be one or more loudspeakers (e.g. a set of headphones) associated with the receiver device.

[0148] In some implementations, the sender device may be configured to include audio description (AD) PI data in the forward IVAS RTP bitstream indicating the audio content types in the encoded audio content. This forward AD data may comprise at least one AID byte. The forward AD data may comprise a single AID byte representing all audio signals of the audio content or oneDocket No.: D24173WO01AID byte per audio signal. For example, if the number of audio signals is N, the forward AD data may comprise N AID bytes of AD data. As another example, a stereo pair of audio signals may be associated with a single AID byte.

[0149] Forward AD PI data capability may be negotiated during SDP negotiation between the sender device and receiver device as described above. Forward AD capability may form an SDP type that is separate from DAS. For example, Table 4 specifies the SDP type (titled faud) for forward AD capability. During SDP negotiation, the sender device may indicate to the receiver device that the sender device supports signaling of forward AD data. In response to the receiver device indicating that it supports AD PI data, the sender device may start sending forward AD PI data to the receiver device.Table 4

[0150] The forward AD data capability may be negotiated independently from DAS capability (see Table 2). It is understood that forward AD data capability may be used without DAS or vice versa.

[0151] In some implementations, forward AD data is negotiated in response to DAS being supported by both devices. That is, forward AD data may be made dependent on DAS. For example, a specific type of receiver device may require forward AD data to perform automatic selection of the preferred audio content type or to show appropriate options for audio content type selection on a UI.

[0152] In reference to FIG. 6, when AD PI data is present in the forward IVAS RTP bitstream, the AD PI data may be extracted by the RTP depacker module 51 and provided to the AID selector module 57 to partially or fully control the content type selection. For instance, the AID selector module 57 may be configured to select the audio content type based on one or more of the AD data, the classifier result and user input.

[0153] In some implementations, the AD data is used by the receiver device 50b to control which audio content types are shown on the UI 100 in FIG. 2 as an option for the user to select the preferred audio content type. Each audio content type indicated in the AD data for at least one audio signal may be shown on the UI 100 as a selectable option. Hereby, the user may be presented with aDocket No.: D24173WO01 customized list of audio content types to choose from, which reflects the actual audio content of the audio signals that are pre-processed at the sender device (as shown in FIG. 4).

[0154] FIG. 7 A is a block diagram illustrating an example sender device 40b with forward AD data extraction capability.

[0155] The sender device 40b comprises a capture front-end module 41, an audio preprocessor module 42, an IVAS encoder module 43, an RTP packer module 44, an RTP depacker module 46, a DAS data extractor module 45, a classifier module 47 and an AD data generator module 48.

[0156] The capture front-end module 41 obtains as an input one or more audio signals (e.g., from microphones). A first output of the capture front-end module 41 is provided as a first input to the audio pre-processor module 42 and a second output of the capture front-end module 42 is provided as input to the classifier module 47. A first output of the classifier module is provided as input to the AD data generator module 44 and a second output of the classifier module is provided as a second input to the audio pre-processing module 42. The output of the pre-processor module 42 is provided as an input to the IVAS encoder module 43, and the output of the IVAS encoder module 43 is provided to the RTP packer module 44 alongside the output of the AD data generator 48 which is also provided as a further input to the RTP packer module 44. The output of the RTP packer module 44 forms a forward IVAS RTP bitstream (an audio bitstream). The RTP depacker module 46 obtains as an input a backward IVAS RTP bitstream and the output of the RTP depacker module 46 is provided as input to the DAS data extractor module 45. The output of the DAS data extractor module 45 is provided as a third input to the audio pre-processor module 42.

[0157] The sender device 40b is similar to the sender device 40a of FIG. 4A but differs in that sender device 40b further comprises a classifier module 47 and AD data generator module 48. The classifier module 47 is configured to obtain the one or more audio signals output by the capture front end 71, and determine for each audio signal individually, one or more content types active in the audio signal.

[0158] In FIG. 7A, the classifier module 47 is depicted as separate from the audio preprocessor module 42 but it is envisaged that in some implementations, the classifier may form a part of the audio pre-processor module 42.Docket No.: D24173WO01

[0159] Information indicating the content type(s) (and optionally a classification confidence) is provided to the AD data generator 48 which generates AD data indicating the active content types to the RTP packer module 44. The AD data may, for each audio signal or for the entire audio content, comprise an AID byte as shown in FIG. 8, analogous to the AID byte of FIG. 1. The AID byte comprises 8 bits wherein six bits are reserved, and two bits are used to signal speech (voice) and / or music. Of course, this bit allocation is merely an example and one or more of the reserved bits may be used to enable signaling of additional content types.

[0160] The forward AD data extracted by the sender device 40b may be used to assist the receiver device in the determination of which types of audio content are available and may be selected as preferred audio content types. The forward AD data may be useful as the sender device 40b may have access to the undistorted, uncoded “raw” audio content during the capture process which may enable more accurate or robust content type classification.

[0161] As an alternative or addition to using the classifier module 47, the sender device 40b may also use context-based information of the capture environment to make this determination. For instance, there may be knowledge at the sender device 40b where the audio content or an audio signal is captured, e.g. in a conference room or a concert hall. In the first case, the sender device would indicate Voice (V) as the identified content with the AD data while it would indicate Music (M) in the second case. Captured video available at the sender side may also assist this determination.

[0162] The sender device 40b may convey the AD data in a forward PI data frame. Audio Description (AD) data may be provided in Processing Information (PI) data frames to convey the audio content type (e.g. speech / music / general audio) to the receiver device.

[0163] The size of AD PI data may vary from 1 to 12 bytes and depends on the IV AS format as described in Table 5 below. As shown, one byte (one AID byte) may be conveyed for stereo, scene based audio (SB A) and metadata-assisted spatial audio (MASA) indicating the active content types. For the IVAS format independent streams with metadata (ISM), an AID byte is provided for each stream. For a multi-channel (MC) IVAS format, one AID byte may be provided per channel. Alternatively, for a MC IVAS format, one AID byte may be provided for the center channel, and one AID byte may be provided for the remaining channels to assist with use cases benefiting from dialogue enhancement. For the combined formats objects with MASA (OMASA) and objects withDocket No.: D24173WO01SBA (OSBA), one AID byte is provided per object, and one AID byte is provided for the bed audio signal of MAS A or SBA.

[0164] In some implementations, the number of channels is limited to 12 (suitable for conveying 7.1.4 channel presentations). In some implementations, the number of objects is limited to four, wherein the AD PI data may be 4 bytes for the ISM format and 5 bytes for the OMASA format or OSBA format.

[0165] The AD PI data may also used for formats other than IVAS wherein the size of the AD PI data in general is N bytes, with N being the number of audio elements transmitted (objects and / or channels).Table 5: Audio Description PI data size

[0166] Each byte in AD PI data payload is an audio identifier (AID) byte that is defined as follows with reference to FIG. 8.

[0167] The AID byte comprises an 8 bit identifier, as described in FIG. 8, to specify type of audio that is being transmitted. This byte contains a V bit and an M bit, as defined in Table 6 and Table 7, that specifies whether audio contains speech and / or music. The reserved bits in AID may be set to 0 and may be ignored by a receiver device.Table 6: V field in AID identifier ByteDocket No.: D24173WO01Table 7: M field in AID identifier Byte

[0168] The classifier module 47 may in some implementations provide the classification results to the audio pre-processing module 42 which may omit one or more audio signals based on the classification results. Alternatively, the classification results are provided to the IVAS encoder module 43 which is configured to encode audio signals or objects carrying non-preferred audio content types with a reduced bitrate. The audio pre-processing module 42 or IVAS encoder module 43 may determine which audio content types are preferred and non-preferred by comparing the audio content types signaled in the DAS PI with the audio content types detected by the classifier module 47.

[0169] Thus, a receiver device may use the AID byte of the AD data obtained from the sender device in its determination of desired audio content versus undesired audio content either in an exclusive way or in an assisted way. In the latter case, the final determination is still done by the receiver, e.g. based on user preference and / or classification result, though assisted by the AID byte of the AD data obtained from the sender device.

[0170] Based on the AID byte conveyed by the sender device to the receiver device, the receiver may locally carry out noise suppression to reduce undesired audio components in the decoded audio signal. A possible advantage of this approach is that communication using reverse PI data may be avoided and computational resources to carry out noise suppression at the sender device may be saved. A further advantage is obtained in case the sender device transmits in a multi- or broadcast- session to multiple receiver devices with diverging preferences in terms of desirable versus undesirable audio components and / or suppression levels.

[0171] FIG. 7B is a flowchart illustrating a method for dynamic audio suppression for captured audio content. The method may be implemented by the sender device 50c of FIG. 7 A.

[0172] Example methods of FIG. 7B may be partitioned into steps, such as steps S31 - S36. The various steps may be described as operations, processes, methods, steps, acts, blocks or functions. The steps are not necessarily performed in the order indicated. In some implementations, one or more of the steps may be performed concurrently. Moreover, some example methods may include more or fewer steps than shown and / or described.Docket No.: D24173WO01

[0173] Processing may commence at step S31 titled “Enabling DAS capability”, which comprises enabling, at a sender device, a dynamic audio suppression capability based on a negotiated communication session between the sender device and a receiver device, wherein the sender device corresponds to a capture device that is configured to provide the captured audio content in any out of a variety of audio formats.

[0174] Processing may continue to step S32 titled “Capturing audio content”. Step S32 comprises capturing (e.g. with the capture front-end), at the sender device, audio content comprising at least one audio signal wherein, the audio contains two types of audio content.

[0175] Processing may continue to step S33 titled “Determining AD data”. Step S33 comprises determining (e.g. by the classifier module 47) , based on the audio content, audio description (AD) data indicating a set of active audio content types. To determine the AD data, the classifier module 47 of FIG. 7A may be used. In some implementations, the classifier module 47 processes each audio signal individually to determine which audio content types that are active in each audio signal.

[0176] Processing may continue to step S34 titled “Generating AD PI data”, which comprises generating (e.g. by the AD data generator module 48), at the sender device, AD processing information (PI) data based on the AD data. For example, the AD PI data comprises an AD byte for each of the at least one audio signal.

[0177] Processing may continue to step S35 titled “Encoding audio content and AD PI data to generate a bitstream”. Step S36 may comprise encoding (e.g. by the IVAS encoder 43), at the sender device, the audio content and the AD PI data using a codec (e.g. the IVAS codec) to generate a bitstream.

[0178] Processing may continue to step S36 titled “Outputting the bitstream”. Step S36 comprises outputting (e.g. by the RTP packer module 44), by the sender device, the bitstream. The bitstream may be temporarily stored and / or conveyed to a receiver device.

[0179] Yet another possibility is to carry out noise suppression both at the sender based on an AID indication obtained from one or several receivers while each receiver may suppress further undesired audio components that the sender has not (yet) suppressed.Docket No.: D24173WO01

[0180] FIG. 9 A is a block diagram illustrating an example receiver device 50c where the audio post-processor module 53 is configured to obtain the AD data from the forward IV AS RTP bitstream and perform suppression of non-preferred audio content types based on the AD data.

[0181] The example receiver device 50c comprises an RTP depacker module 51, an IVAS decoder module 52, an audio post-processor module 53, a playback front-end module 54, an RTP packer module 55, an audio preference type selector module 57 (also referred to as AID selector module), an audio suppression level selector module 56 (also referred to as SLI selector module), a classifier module 58 and a signal-to-noise (S / N) level calculator module 59.

[0182] The RTP depacker module 51 obtains as an input a forward IVAS RTP bitstream. A first output of the RTP depacker module 51 is provided as input to the IVAS decoder 52 and a second output of the RTP depacker module 51 is provided as a first input to the audio postprocessing module 53. An output of the IVAS decoder 52 is provided as a second input to the audio post-processing module 53. A first output of the audio post-processing module 53 is provided to an input of the playback front-end module 54 which outputs output audio for playback on loudspeakers / headphones. A second and third output of the audio post-processing module 53 is provided as input to the classifier module 58 and S / N calculator module 59 respectively. The AID selector module 57 obtains user input and provides a first output that is used as a second input to the S / N calculator module 59 and a second output that is used as a first input to the RTP packer module 55. The output of the S / N calculator module 59 is provided as input to the SLI selector module 56 which further obtains user input as a second input. The output of the SLI selector module 56 is provided as a second input to the RTP packer module 55. The output of the RTP packer module 55 forms a backward IVAS RTP bitstream.

[0183] For example, if, for an audio signal, the AD data indicates that speech is active, then the audio post-processor module 53 may perform suppression of any non-speech audio content in that audio signal. The post-processed audio signal is subsequently provided to the playback frontend 54.

[0184] The post-processed audio signal may also be provided to the classifier module 58 and / or S / N level calculator module 59 as shown in FIG. 9A. Despite the audio post-processor module 53 suppressing the non-preferred audio content types, the receiver device 50c may stillDocket No.: D24173WO01 convey backwards SLI and / or AID data to the sender device whereby the suppression of nonselected audio content types may be distributed between sender device and receiver device.

[0185] FIG. 9B is a flowchart illustrating a method for dynamic audio suppression for captured audio content. The method may be implemented by the receiver device 50c of FIG. 9A.

[0186] Example methods of FIG. 9B may be partitioned into steps, such as steps S41 - S45. The various steps may be described as operations, processes, methods, steps, acts, blocks or functions. The steps are not necessarily performed in the order indicated. In some implementations, one or more of the steps may be performed concurrently. Moreover, some example methods may include more or fewer steps than shown and / or described.

[0187] Processing may commence at step S41 titled “Enabling DAS capability” comprises enabling, at the receiver device, a dynamic audio suppression capability based on a negotiated communication session between the receiver device and a sender device, wherein the sender device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats.

[0188] Processing may continue to step S42 titled “Receiving audio bitstream”. Step S42 comprises receiving (e.g. by the RTP depacker module 51), at the receiver device, an audio bitstream from the sender device.

[0189] Processing may continue to step S43 titled “Decoding audio bitstream”. Step 43 comprises decoding (e.g. by the IVAS decoder 52), at the receiver device, the audio bitstream to obtain audio content and audio description (AD) PI data indicating for each of the at least one audio signal of the audio content at least one active audio content type. For example, the AD PI data comprises an AD byte associated with each audio signal.

[0190] Processing may continue to step S44 titled “Obtaining processed audio content”. Step S44 comprises obtaining (e.g. by the audio post-processing module 53 or playback front-end module 54) processed audio. In some implementations, the receiver device may at step S44 process one or more audio signals based on the AD PI data. The processing may comprise identifying one or more non-preferred audio content types as indicated by the AD PI data and suppressing the non-preferred audio content typesDocket No.: D24173WO01

[0191] Additionally or alternatively, step S44 may comprise determining processing information such as an AID and / or SLI byte based on the AD PI data, and sending the processing information to a sender device. At the sender device, the audio content is processed to yield processed audio content which is sent to the receiver device as an audio bitstream. The processed audio content is then obtained by the receiver device by decoding the audio bitstream to obtain the processed audio content.

[0192] Processing may continue to step S45 titled “Rendering the processed audio content”. Step S45 comprises rendering (e.g. by the playback front-end module 54), at the receiver device, the processed audio content in an audio format suitable for playing on an end user device.

[0193] FIG. 10 is a block diagram of a system 1100 for implementing the features and processes described in reference to FIGS. 1-9, according to one or more implementations.

[0194] FIG. 10 illustrates a schematic block diagram of an example system or device architecture that may be used to implement various aspects and processes described in reference to FIGS. 1-9, such as the sender devices, receiver devices, modules, encoders, and decoders according to one or more implementations described above. System 1100 includes but is not limited to servers and client devices, systems, System 1100 includes one or more server computers or any client devices, including but not limited to: call servers, user equipment, conference room systems, home theatre systems, virtual reality (VR) or extended reality (XR) gear, in-vehicle entertainment systems and immersive content ingestion devices. System 1100 includes any consumer devices, including but not limited to: smart phones, tablet computers, wearable computers, vehicle computers, game consoles, surround systems, kiosks, etc.

[0195] As shown, the system or device architecture 1100 includes central processing unit (CPU) 1101 which is capable of performing various processes such as those performed by the sender and / or receiver devices described above in accordance with a program stored in, for example, read-only memory (ROM) 1102 or a program loaded from, for example, storage unit 1108 to random access memory (RAM) 1103. The CPU 1101 may be, for example, an electronic processor 1101. In RAM 1103, the data required when CPU 1101 performs the various processes is also stored, as required. CPU 1101, ROM 1102, and RAM 71103 are connected to one another via bus 1104. Input / output interface 1105 is also connected to bus 1104.Docket No.: D24173WO01

[0196] The following components are connected to I / O interface 1105: input unit 1106, that may include a keyboard, a mouse, a touchscreen display (e.g. for displaying the UI 100 of FIG. 2) or the like; output unit 1107 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 1108 including a hard disk, or another suitable storage device; and communication unit 1109 including a network interface card such as a network card (e.g., wired or wireless).

[0197] In some implementations, input unit 1106 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats). The microphones may e.g. be connected to the capture front-end 41 described in connection with FIG. 4A and FIG. 7A above

[0198] In some implementations, output unit 1107 include systems with various numbers of speakers. Output unit 1107 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats). The speakers may be connected to the playback front-end 54 described in connection with FIG. 4B, FIG. 7B and FIG. 9A.

[0199] In some embodiments, communication unit 1109 is configured to communicate with other devices (e.g., via a network). Each of a sender device and receiver device may comprise a communication 1109 to enable two-way communication between the devices. Drive 1110 is also connected to I / O interface 1105, as required. Removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 1110, so that a computer program read therefrom is installed into storage unit 1108, as required. A person skilled in the art would understand that although apparatus 1100 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0200] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, theDocket No.: D24173WO01 computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 1109, and / or installed from the removable medium 1111, as shown in FIG. 10.

[0201] Generally, various example embodiments, devices, modules, encoders and decoders of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 1101 in combination with other components of FIG. 10), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, devices, modules, encoders, decoders, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0202] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0203] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non- transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a randomaccess memory (RAM), a read-only memory (ROM), an erasable programmable read-only memoryDocket No.: D24173WO01(EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD- ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0204] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0205] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims.

[0206] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0207] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein,Docket No.: D24173WO01 and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0208] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0209] This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

Claims

Docket No.: D24173WO01CLAIMS1. A dynamic audio suppression method for captured audio content, the method comprising: enabling, at a first device, a dynamic audio suppression capability based on a negotiated communication session between the first device and a second device, wherein the second device corresponds to a capture device that is configured to provide the captured audio in at least one of a variety of audio formats; selecting, at the first device, an audio suppression level based on user preferences for audio suppression; sending, from the first device to the second device, audio suppression processing information, PI, data based on the audio suppression level; receiving, at the first device, an audio bitstream from the second device; decoding, at the first device, the audio bitstream to obtain one or more audio signals from the audio bitstream, the one or more audio signals comprising a representation of the captured audio content; wherein, a second audio content type of the one or more audio signals is suppressed relative to a first audio content type of the one or more audio signals based on the audio suppression level; and rendering, at the first device, the one or more audio signals in an audio format suitable for playing on an end user device.

2. The dynamic audio suppression method according to claim 1, wherein the captured audio content is captured immersive audio content.

3. The dynamic audio suppression method according to claim 1 or claim 2, wherein the dynamic audio suppression method is for captured audio in immersive voice and audio services, IVAS.

4. The dynamic audio suppression method of any of the preceding claims, wherein the negotiated communication session includes: sending, from the second device to the first device, a set of capabilities associated with the second device, wherein the set of capabilities includes the dynamic audio suppression capability; receiving, at the first device, the set of capabilities associated with the second device;Docket No.: D24173WO01 identifying, at the first device, the dynamic audio suppression capability of the second device from the received set of capabilities; and sending, from the first device to the second device, an acknowledgment of support for dynamic audio suppression.

5. The dynamic audio suppression method of any of claims 1 - 3, wherein the negotiated communication session includes: sending, from the first device to the second device, a set of capabilities associated with the first device, wherein the set of capabilities includes the dynamic audio suppression capability; receiving, at the second device, the set of capabilities associated with the first device; identifying, at the second device, the dynamic audio suppression capability of the first device from the received set of capabilities; and sending, from the second device to the first device, an acknowledgment of support for dynamic audio suppression.

6. The dynamic audio suppression method of any of the preceding claims, wherein the captured audio content includes speech and background noise.

7. The dynamic audio suppression method of any of the preceding claims, wherein the variety of audio formats includes: Ambisonics, channel based audio, and object based audio.

8. The dynamic audio suppression method of any of the preceding claims, wherein user preferences for audio suppression are determined by prior stored user preferences for audio suppression.

9. The dynamic audio suppression method of any of claims 1 - 7, wherein user preferences for audio suppression are determined by interactive selection by a user.

10. The dynamic audio suppression method of any of claims 1 - 7, wherein user preferences for audio suppression are determined by automated processing based on signal analysis methods.Docket No.: D24173WO0111. The dynamic audio suppression method of claim 10, wherein the signal analysis methods are based on one or more of: a Discrete Fourier Transform, DFT, a Fast Fourier Transform, FFT, and a signal-to-noise level.

12. The dynamic audio suppression method of claim 10 or claim 11, wherein the signal analysis methods apply a classifier model to the audio content to identify content types for an audio identifier AID.

13. The dynamic audio suppression method of claim 12, wherein the classifier model includes one or more of a Gaussian Mixture Model, GMM, or a Machine Learning Model, ML,.

14. The dynamic audio suppression method of any of claims 10 - 13, wherein the signal analysis methods dynamically evaluate the signal levels of a selected type of audio content relative to a nonselected type of audio content to update a suppression level indicator, SLI.

15. The dynamic audio suppression method of claim 14, wherein the update of the SLI is performed continuously, periodically, or when triggered.

16. The dynamic audio suppression method of any of claims 1 - 7, wherein user preferences for audio suppression are determined by automated processing based on service context.

17. The dynamic audio suppression method of claim 16, wherein the service context is based on either a voice focus or a music focus.

18. The dynamic audio suppression method of any of the preceding claims, further comprising: selecting, at the first device, a preferred audio content type based on user preferences for audio suppression in one or more audio signals; wherein the audio suppression PI data is further based on the preferred audio content type, and wherein the first audio content type is the preferred audio content type.Docket No.: D24173WO0119. The dynamic audio suppression method of claim 18, wherein the audio suppression PI data describes the user’s preference with respect to one or more types of audio content that should be enhanced.

20. The dynamic audio suppression method of claim 18, wherein the audio suppression PI data describes the user’s preference with respect to one or more types of audio content that should be suppressed.

21. The dynamic audio suppression method of claim 19, wherein the audio suppression PI data includes one byte of data that indicates the user preferences with respect to the one or more types of audio content that are preferred.

22. The dynamic audio suppression method of claim 21, wherein the data payload of the audio suppression PI data contains an Audio Identifier, AID, byte based on the preferred audio content type(s).

23. The dynamic audio suppression method of claim 22, wherein audio content types other than those specified by the Audio Identifier, AID, field are considered as preferred audio components.

24. The dynamic audio suppression method of claim 20, wherein the audio suppression processing information, PI, data includes one byte of data that indicates the user preferences with respect to the type of audio content that is non-preferred and should be suppressed.

25. The dynamic audio suppression method of any of the preceding claims, wherein the audio suppression processing information, PI, data comprises a Suppression Level Indicator, SLI, byte that is based on the audio suppression level and indicating a desired degree of suppression.

26. The dynamic audio suppression method of claim 25, wherein the Suppression Level Indicator, SLI, byte comprises N bits for signaling the indicator value, wherein 2 < N < 8, and wherein the indicator value ranges from 0 to 2N- 1, wherein the expected amount of audio suppression is proportional to the indicator value with 0 indicating one of: no audio suppression and minimum audio suppression, and 2N- 1 indicating maximum audio suppression.Docket No.: D24173WO0127. The dynamic audio suppression method of any of claims 25 - 26, wherein the values of the Suppression Level Indicator, SLI, are associated with an approximate logarithmic domain.

28. The dynamic audio suppression method of any of claims 18 - 27, wherein the audio bitstream comprises encoded audio description, AD PI data indicating for the audio content at least one active audio content type, the method further comprises: decoding, at the first device, the audio bitstream to obtain the AD PI data; wherein the preferred audio content type is further based on the AD PI data.

29. The dynamic audio suppression method of claim 28, further comprising: performing, at the first device, audio suppression on the audio content based on the AD PI data prior to rendering the audio content.

30. The dynamic audio suppression method of any of claims 18 - 29, wherein the audio content includes the first audio content type and the second audio content type.

31. The dynamic audio suppression method of any of claims 18 - 29, wherein the audio content comprises the first audio content type.

32. A dynamic audio suppression method for captured audio content, the method comprising: enabling, at a second device, a dynamic audio suppression capability based on a negotiated communication session between a first device and the second device, wherein the second device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats; capturing, at the second device, the audio content wherein, the audio content contains a first audio content type and a second audio content type; receiving, at the second device, audio suppression processing information, PI, data from the first device, wherein the audio suppression PI data indicates one or more of user preferences for an audio suppression level and user preferences for a preferred audio content type, wherein the preferred audio content type corresponds to one of the first audio content type and the second audio content type;Docket No.: D24173WO01 performing, at the second device, audio suppression on the captured audio content based on the audio suppression PI data to form suppressed audio content wherein a non-preferred audio content type is suppressed relative to the preferred audio content type; encoding, at the second device, the suppressed audio content using a codec to generate a bitstream; and outputting, at the second device, the bitstream.

33. The dynamic audio suppression method of claim 32, wherein the captured audio content comprises a plurality of audio signals, the method further comprising: determining, for each audio signal, a set of active audio content types; wherein performing audio suppression on the captured audio content comprises selecting audio signals having a set of active audio content types that includes the preferred audio content type, and wherein encoding the suppressed audio content comprises encoding the selected audio signals having a set of active audio content types that include the preferred audio content type.

34. The dynamic audio suppression method of claim 33, wherein an encoding bitrate per audio signal is based on the number of selected audio signals.

35. A dynamic audio suppression method for captured audio content, the method comprising: enabling, at a first device, a dynamic audio suppression capability based on a negotiated communication session between the first device and a second device, wherein the second device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats; receiving, at the first device, an audio bitstream from the second device; decoding, at the first device, the audio bitstream to obtain audio content from the audio bitstream and audio description, AD, processing information, PI, data indicating for each of the at least one audio signal of the audio content at least one active audio content type; obtaining processed audio content based on the AD PI data; and rendering, at the first device, the processed audio content in an audio format suitable for playing on an end user device.Docket No.: D24173WO0136. The dynamic audio suppression method according to claim 35, wherein obtaining the processed audio content comprises performing audio suppression based on the AD PI data on the audio content at the first device.

37. The dynamic audio suppression method according to claim 35 or claim 36, wherein obtaining the processed audio content comprises performing audio suppression based on the AD PI data on the audio content at the second device.

38. The dynamic audio suppression method according to claim 37, wherein obtaining the processed audio content based on the AD PI data comprises: sending, from the first device to the second device, audio suppression PI data based on the AD PI data; receiving, at the first device, an audio bitstream from the second device; and decoding, at the first device, the audio bitstream to obtain the processed audio content.

39. An audio encoding method for captured audio content, the method comprising: enabling, at a second device, a dynamic audio suppression capability based on a negotiated communication session between the second device and a first device, wherein the second device corresponds to a capture device that is configured to provide the captured audio content in at least one of a variety of audio formats; capturing, at the second device, audio content comprising at least one audio signal wherein, the audio contains two types of audio content; determining, based on the audio content, audio description, AD, data indicating a set of active audio content types; generating, AD processing information, PI, data based on the AD data; encoding, at the second device, the audio content and the AD PI data using a codec to generate a bitstream; and outputting, at the second device, the bitstream.

40. The audio encoding method for captured audio content according to claim 39, wherein the audio content comprises a plurality of audio signals, and wherein determining AD data comprises determining a set of active audio content types for each audio signal.Docket No.: D24173WO0141. The audio encoding method for captured audio content according to claim 40, AD PI data comprises one AID byte for each audio signal of the plurality of audio signals.

42. The audio encoding method for captured audio content according to any of claims 39 - 41, wherein determining a set of active audio content types comprises performing audio content type classification.

43. An apparatus comprising a processor and a memory storing instructions configured to, when executed by the processor, cause the processor to perform the method of any one of claims 1 - 42.

44. A computer program product comprising instructions configured to, when executed by a processor, cause the processor to perform the method according to any one of claims 1 - 42.

45. A non-transitory computer readable medium comprising instructions configured to, when executed by a processor, cause the processor to perform the method according to any one of claims 1 - 42.