A method for adaptive speech enhancement based on speech experience
The adaptive speech enhancement method addresses the limitations of static gain application by using dynamic gains based on speech experience measures, enhancing speech clarity and reducing artifacts across varying audio conditions.
Patent Information
- Application Number
- PCT/US2025/040375
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-04
- Filing Date
- 2025-08-01
- Publication Date
- 2026-02-12
AI Technical Summary
Existing speech enhancement methods apply static modifications to speech and background signals, which can lead to unnatural sounding mixes and fail to adapt to varying audio conditions, affecting the speech experience across different audio signals.
An adaptive method that determines dynamic gains based on speech experience measures for each audio segment, splitting the gains into boosting and attenuating components to improve speech clarity and reduce artifacts.
Enhances speech experience by dynamically adjusting gains, reducing the risk of introducing audible artifacts and improving speech clarity in diverse audio scenarios.
Smart Images

Figure IMGF000014_0001 
Figure IMGF000015_0001 
Figure IMGF000015_0002
Abstract
Description
A METHOD FOR ADAPTIVE SPEECH ENHANCEMENT BASED ON SPEECH EXPERIENCE CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from U.S. Provisional Application Ser. No.63 / 679,988, filed on 6 August 2024, and European Patent Application No 24204858.5, filed on 4 October 2024, each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates to a method and audio processing system for performing adaptive speech processing. BACKGROUND
[0003] In the field of audio processing, speech enhancement (SE) (often also referred to as dialogue enhancement (DE)) is a technique that is applied to improve the end-user’s experience of speech when consuming media content such as TV, films, podcasts etc.
[0004] SE benefits from having access to separate speech and non-speech (background) signals or stems. Separate speech and non-speech signals may be available in production as separate signals, so called stems. However, in many scenarios only a complete mix, with speech and background content combined, is available. For this case, the speech can be separated from the background in an unguided fashion, for example using Deep Learning algorithms.
[0005] The input signals, be it separate speech and non-speech stems, or a complete mix, can be provided in many different spatial configurations, e.g., mono, multi-channel (e.g. stereo, 5.1, 7.1.4), or object based, or a combination of channels and objects (as in e.g. the Atmos format), or High-Order Ambisonics (HOA) or any combination of the aforementioned spatial configurations. Typically, the spatial configuration is preserved throughout the SE processing. For example, if the input is a complete stereo mix, an unguided separation typically produces a stereo speech estimate signal and a stereo background estimate signal.
[0006] With speech and background separated, modifications of the speech and / or the background can be done with the goal of improving the clarity of the speech, improving the speech intelligibility or decreasing listening effort for the end-user etc., while typically preserving the spatial configuration.
[0007] A common modification is to apply a static (time-invariant) gain to the speech, for example amplifying the speech by 12 dB. Alternatively, a static gain can be applied to thebackground, for example attenuating the background by 12 dB. In both these cases, a change in the speech-to-non-speech-ratio (SNR) of +12 dB is introduced. SUMMARY
[0008] A drawback with existing solutions is that the desired modifications to speech and non-speech background are selected once and then applied in a static manner going forward, even if the need for such modifications has ended. However, speech content could be clean with no interfering background in one segment and in the next segment, the background could severely interfere with the speech. It may seem innocuous to apply a speech boost in the former case (no interfering background), however, that can lead to an unnatural sounding mix of speech and background. In other words, applying a static boost to all speech may not be subjectively preferred. While a static boost may work for some types of audio signals it will not work across a wide variety of audio signals.
[0009] To overcome at least some of the drawbacks of existing solutions an improved audio processing method is proposed.
[0010] According to a first aspect there is provided an audio processing method comprising obtaining a first and second sequence of time-aligned audio segments, wherein the first sequence of audio segments comprises speech audio content and the second sequence of audio segments comprises non-speech audio content. The method further comprises for each time-aligned pair of segments in the first and second sequence, obtaining at least one speech experience measure (SEM) and determining, based on the at least one speech experience measure, a first gain and a second gain, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain.
[0011] In some implementations, determining the first and second gain comprises determining a gain level for improving the speech experience of the pair of segments based on the at least one speech experience measure, obtaining a split ratio and partitioning the gain level into the first gain and the second gain, based on the split ratio.
[0012] In some implementations, the method further comprises applying the first gain to the segment of the first sequence to generate a modified first segment and applying the second gain to the segment of the second sequence to generate a modified second segment, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain. The method further comprises outputting a segment of mixed audio content comprising a mix of the modified first and second segment.
[0013] Hereby, by obtaining at least one speech experience measure (associated with the first segment and second segment) the gain level can be determined adaptively based on the atleast one SEM to achieve a dynamic gain modification which takes the properties of the audio content in the segments into account. Furthermore, by splitting the gain level into two modifications gains, wherein one modification gain is applied as a boosting gain and the other modification gain is applied as an attenuating gain adds the benefit of reducing the overall risk of introducing audible undesirable audio artifacts that otherwise may appear if the full gain level is used as boosting gain or attenuating gain applied to only one of the first segment and the second segment.
[0014] Alternatively to splitting a gain level into the first and second gain, the first and second gain may be determined using a respective mapping function, each mapping the SEM to a respective gain.
[0015] The first sequence of segments may be segments originating from the separated speech of a speech / non-speech separation process, a separate speech stem of a multi-stem production or a speech enhancement processing of a mix input audio signal. The segments of the first sequence therefore comprise speech content and, in the case of the first sequence being the result of speech / non-speech separation processing the first sequence of segments may be clean speech with little to no non-speech content. However, it is understood that the first sequence of segments may also comprise some non-speech content, in addition to the speech content. For example, the speech / non-speech separation process may be imperfect whereby some non-speech content remains in the first sequence of segments.
[0016] The second sequence of segments may be segments originating from the separated non-speech of a speech / non-speech separation process, a non-speech stem of a multi-stem production or a mix input audio signal. The segments of the second sequence therefore comprise non-speech content and, in the case of the second sequence being the result of speech / non-speech separation processing the second sequence may be substantially free from speech with little to no speech content. However, it is understood that the second sequence of segments may also comprise some speech content, in addition to the non-speech content. For example, the speech / non-speech separation process may be imperfect whereby some speech content remains in the second sequence of segments.
[0017] In some implementations, the first sequence of segments is more speech heavy compared to the second sequence of segments. For example, a speech to non-speech ratio between a speech energy measure and a non-speech energy measure is higher for the first sequence of segments compared to the second sequence of segments.
[0018] As an alternative to applying the first and second gain to the first and second segment it is envisaged that the first and second gain are encoded into a bitstream together with the first and second segment and / or the mix segment from which the first and second segmentwere extracted. The bitstream may then be transmitted to a receiving device, which decodes the bitstream to access the first and second gain and the audio content of the first and second segment or the mix segment, whereby the receiving device performs the application of the gains (and optionally separates the mix segment into a first and second segment with a speech / non- speech separator in the receiving device). A benefit with including the first and second gain in a bitstream for later application is that the receiving device may determine if and to what extent the gains are to be applied.
[0019] According to a second aspect there is provided an audio processing method comprising receiving, by a receiving device, a bitstream and processing, by the receiving device, the bitstream to obtain a first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, a first and second gain. The method further comprises for each time-aligned pair of segments in the first and second sequence applying the first gain to the segment of the first sequence to generate a modified first segment and applying the second gain to the segment of the second sequence to generate a modified second segment. Wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain. The method further comprises outputting a segment of mixed audio content comprising a mix of the modified first and second segment.
[0020] Additionally, it is generally difficult to recreate the original, unmodified, segments from the modified segments resulting from the processing whereby the original mix segment(s) may be transmitted alongside the first and second gain to the receiving device. The receiving device may then separate the speech from the mix segments to generate the first and second segments and recreate the modified segments (by applying the gains) while at the same time the original segments are available if needed.
[0021] In some implementations, the bitstream comprises encoded segments of a mix signal and wherein processing the bitstream comprises decoding the bitstream to obtain the segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment.
[0022] According to a third aspect there is provided an apparatus comprising a processor and a memory, configured to perform the method according to the first aspect.
[0023] According to a fourth aspect there is provided a computer-readable storage medium storing the computer program according to the first aspect.
[0024] According to a fifth aspect there is provided an audio processing method comprising obtaining a first and second sequence of time-aligned input audio segments, whereinthe first sequence of input audio segments comprises speech audio content and the second sequence of input audio segments comprises non-speech audio content. The method further comprises, for each time-aligned pair of input audio segments in the first and second sequence, obtaining at least one speech experience measure and determining at least one dynamic range compression (DRC) parameter based on the at least one speech experience measure.
[0025] In some implementations, the method further comprises performing dynamic range compression based on the at least one DRC parameter on at least one of the segments of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence.
[0026] That is, based on the at least one speech experience measure, at least one DRC parameter is adaptively controlled. For example, the DRC parameter is the dynamic range compression ratio which may be controlled so as to increase the perceived “presence” of audio content associated with either the first or the second segment when the speech experience measure indicates that the speech experience is poor, or otherwise in need of improvement. It is hereby no longer necessary to set a static DRC parameter which is used for all signal types since the DRC parameter will be dynamically adjusted based on the at least one speech experience measure. The at least one DRC parameter may be one or more of a DRC ratio, a DRC threshold, a make-up gain, a DRC attack time constant, and a DRC release time constant.
[0027] According to a sixth aspect there is provided an audio processing method comprises receiving, by a receiving device, a bitstream, processing, by the receiving device, the bitstream to obtain a first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, a at least one dynamic range compression, DRC, parameter The method further comprises for each time-aligned pair of segments in the first and second sequence performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence.
[0028] That is, as an alternative to performing dynamic range compression processing directly, it is again envisaged that at least one DRC parameter is encoded into a bitstream together with the first and second segment and / or the mix segment from which the first and second segment were extracted. The bitstream may then be transmitted to a receiving device, which decodes the bitstream to access the at least one DRC parameter and the audio content ofthe first and second segment or the mix segment, whereby the receiving device performs the DRC processing based on the at least one DRC parameter. As described above a benefit with this approach is that the receiving device is given more flexibility in terms of deciding if and when to apply the DRC processing while the receiving device also has access to the original (non-DRC processed) segment(s) which may be used by other sub-systems in the receiving device.
[0029] In some implementations, the bitstream comprises encoded segments of a mix signal and wherein processing the bitstream comprises decoding the bitstream to obtain the segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment.
[0030] That is, the receiving device may receive the mix segments in the bitstream and separates the mix segment into the first and second segment using its own speech / non-speech separation process. Alternatively, the first and second segments are included in the bitstream directly, wherein the receiving device does not need to perform its own speech / non-speech separation process.
[0031] According to a seventh aspect there is provided an apparatus comprising a processor and a memory, configured to perform the method according to the fourth aspect.
[0032] According to an eighth aspect there is provided a computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to the fourth aspect. DESCRIPTION OF THE DRAWINGS
[0033] Aspects of the present disclosure will be described in more detail with reference to the appended drawings, showing exemplary embodiments.
[0034] Figure 1 is a block diagram illustrating an audio processing system according to some implementations.
[0035] Figure 2A is a block diagram illustrating a speech enhancement processor with intrusive gain modification according to some implementations.
[0036] Figure 2B is a flowchart describing a method for processing two audio segments with the speech enhancement processor of FIG.2A.
[0037] Figure 2C is a block diagram illustrating a speech enhancement processor configured to apply modification gains to segments of the mix input audio signal and segments of the speech audio signal.
[0038] Figure 3 is graph illustrating an example of how various processing parameters can be adjusted based on at least one SEM represented as a variable X.
[0039] Figure 4 is a graph illustrating an example of how the split ratio can be adjusted based on a measure of the temporal stationarity of a non-speech segment or mix segment.
[0040] Figure 5A is a block diagram illustrating a speech enhancement processor with intrusive DRC modification according to some implementations.
[0041] Figure 5B is a flowchart describing a method for processing two audio segments with the speech enhancement processor of FIG.5A.
[0042] Figure 6 is a block diagram illustrating an audio processing system for non- intrusive gain modification.
[0043] Figure 7 is a block diagram illustrating a speech experience measure generator according to some implementations.
[0044] Figure 8 is a block diagram illustrating non-intrusive gain modification according to some implementations.
[0045] Figure 9 is a block diagram illustrating non-intrusive DRC modification according to some implementations.
[0046] Figure 10 is a block diagram illustrating a system for processing a plurality of mix input audio signals according to some implementations.
[0047] Figure 11 illustrates a schematic block diagram of an example device or architecture that may be used to implement embodiments of the invention. DETAILED DESCRIPTION
[0048] FIG.1 is a block-diagram illustrating an audio processing system 10 according to some implementations. The audio processing system 10 comprises a speech experience measure generator 1, a parameter mapper 2 and a speech enhancement processor 3. The purpose of the audio processing system 10 is to dynamically control the speech enhancement processing performed by the speech enhancement processor 3 based on at least one speech experience measure (SEM) extracted by the SEM generator 1.
[0049] The audio processing system 10 is generally tasked with improving the speech experience and providing an output audio signal with generally better speech experience compared to the input audio signal. Improving the speech experience may involve boosting or otherwise amplifying the speech to make it louder and / or more intelligible. However, there are more aspects to speech experience than just loudness and intelligibility, and in some implementations the speech experience may benefit from a local attenuation of speech to e.g. avoid large or rapid changes in speech loudness.
[0050] The speech enhancement processor 3 is configured to obtain a mix input audio signal MIN and generate a modified mix audio signal MM as an output. The mix input audio signal MIN may be a single channel audio signal or a single audio object, an audio signal with multiple channels (e.g. stereo, 5,1, 7.1.4), an audio signal with multiple audio objects, or a signal comprising at least one channel and at least one object (e.g. an Atmos signal). An audio object is an audio stem which has an associated spatial position which may vary with time.
[0051] For example, the mix input audio signal MIN is a single channel audio signal comprising both speech audio content and non-speech audio content. With non-speech audio content it is meant any audio content which is not related to human speech, such as music, effects or noise. The non-speech audio content may also be referred to as background audio content.
[0052] As another example, the mix input audio signal MIN may be a multi-channel, multi-stem or multi-object signal comprising at least two channels or stems, at least two audio objects or at least one channel or stem and at least one audio object. For instance, the mix input audio signal MIN comprises a center channel and one or more surround channels of a multi- channel audio presentation or a plurality of audio objects wherein each channel or object comprises speech content, non-speech content or mixtures thereof.
[0053] As yet another example, the mix input audio signal MIN may comprise a plurality of content specific stems, such as a speech stem (often labeled as a dialogue stem) and one or more non-speech stems (such as a music and effects stems). Each stem signal is provided in a spatial configuration, i.e., the stem can be, e.g., a mono signal, a multi-channel signal, an object based signal, or a signal that is a combination of channels and objects. Content specific stems are sometimes used as the soundtrack of a movie.
[0054] The speech enhancement processor 3 obtains the mix input audio signal MIN as an input and performs at least one form of speech enhancement processing to generate a modified mix audio signal MM as the output. The modified mix audio signal MM may comprise portions, or all, of the audio content of the mix input audio signal MIN but in the modified mix audio signal MM the audio content has been rebalanced or otherwise modified to enhance the speech experience.
[0055] The speech enhancement processor 3 may separate a single channel input mix MIN into a speech signal and non-speech signal whereby it individually modifies the speech signal and / or the non-speech signal prior to recombining the two (modified) signals to form the modified mix audio signal MM with improved speech experience. The modifications implemented by the speech enhancement processor 3 may involve gain adjustment or dynamic range compression.
[0056] For a multi-channel or multi-object input mix MIN each channel or object of the input mix MIN may comprise a mix of speech and non-speech content whereby one or more (e.g. each) channel and / or object is processed with the speech enhancement processor 3 to enhance the speech experience of the one or more channels and / or objects. By enhancing the speech experience of one or more (e.g. each) channel and / or object the speech experience of the multi-channel mix as a whole may be enhanced.
[0057] The processing performed by the speech enhancement processor 3 is governed by at least one processing parameter ^, wherein the at least one processing parameter ^ is based on at least one speech experience measure (SEM). The at least one processing parameter ^ may for example dictate a modification gain to be applied and / or a DRC ratio to be applied as will be described below. In general, the speech enhancement processor 3 may be seen as a function ^ which produces an output ^^^^given an input ^^^and at least one processing parameter ^. That is:^^^^ = ^^^^^, ^^, … , ^^^. ^1^Accordingly, the speech enhancement processing applied in the speech enhancement processor 3 can be controlled via at least one processing parameter ^.
[0058] To obtain the at least one processing parameter ^, the speech experience measure generator 1 determines at least one speech experience measure (SEM) and provides the at least one SEM to the parameter mapper 2 which maps the at least one SEM to at least one processing parameter ^. The at least one processing parameter ^ is then in turn provided to the speech enhancement processor 3. In general the SEM generator 1 may output multiple SEMs and the parameter mapper 2 may output multiple processing parameters ^ meaning that the parameter mapper 2 may in general implement a multi-dimensional function ℎ which maps one or more SEMs (represented using a vector ^⃗) to one or more processing parameters ^ (represented usinga vector ^⃗) such that ^⃗ = ℎ^^^⃗. Examples of processing parameters include a gain level ^, asplit ratio ^, a subband specific split ratio ^^^^, a scaling factor ^ and a DRC ratio ^ as described below. That is, the parameter mapper 2 may implement a unique mapping function for each processing parameter. For example, the parameter mapper 2 may map a first set of one or more SEMs to the split ratio ^ (typically a value between 0 and 1) with a first mapping function and map a second set of one or more SEMs to the gain level ^ (typically a value between 0 and 15 dB) using a second mapping function where the first and second set of SEMs are mutually exclusive, partially overlapping or fully overlapping.
[0059] To make the audio processing dynamic with time, the speech enhancement processor 3, the parameter mapper 2 and the SEM generator 1 may all operate on a segment-to-segment basis wherein segments of the involved audio signals are processed in sequence. For example, the SEM generator 1 extracts the at least one SEM once for each segment allowing the processing of the speech enhancement processor 3 to be updated for each segment. Hereby, any audio signals described herein may be subdivided into a sequence of segments. The segments may be of the same length or of different lengths. The segments may be partially overlapping in time or non-overlapping. For example, each segment is between 5 and 500 ms long.
[0060] The SEM generator 1 obtains as an input at least one of the mix input audio signal M, auxiliary data (e.g. metadata) associated with the mix input audio signal M, a speech signal MS carrying speech content separated from the mix input audio signal MIN and a non-speech signal MNS separated from the mix input audio signal M. Based on this input, the SEM generator 1 determines at least one SEM associated with the speech experience of each segment of the mix input audio signal M. The at least one SEM indicates a property of the speech experience and generally a SEM is said to be higher if it is associated with speech which is clear, easy to perceive, highly intelligible or louder than other non-speech content and a SEM is said to be lower if the SEM is associated with speech that is less clear, hard to perceive, has poor intelligibility or is less loud compared to other non-speech content. As described below, the one or more SEM may be extracted using a variety of methods, for example the SEM is computed based on the ratio between the energy of speech separated from an original audio signal (or segments thereof) and the energy of the original audio signal (or segments thereof).
[0061] In some implementations, the at least one SEM may be smoothed or averaged over multiple segments (e.g. smoothed or averaged over time scales ranging from 0.5-10s) since too rapidly varying SEMs may lead to too rapid control of the speech enhancement processor 3 which may lead to undesired acoustic artifacts.
[0062] The at least one SEM may be a discrete or continuous value which evolves from segment to segment over the course of the mix input audio signal M. The at least one SEM may be based on properties of the speech content separated from the mix input audio signal MIN and / or non-speech separated from the mix input audio signal M. Additionally or alternatively, the at least one SEM may further be based on auxiliary data (e.g. metadata) associated with the mix input audio signal. For example, the auxiliary data may indicate the presence of interference in a playback environment, or the auxiliary data may comprise metadata associated with the mix input audio signal M.
[0063] For example, a segment of the mix input audio signal MIN having comparatively loud non-speech content (e.g. noise) may be associated with a lower SEM value since non- speech content (such as noise) may impair the speech experience. Acoustic noise present in the environment during playback may also be seen as something that degrades the speech experiencewhereby a SEM associated with playback interference may be lower when noise is present in the playback environment. As further examples, artifacts or degradations in the signal processing path may be determined based on the transmission codec, type of playback device (phone, tablet, TV, set-top box), transducer type (headphones, loudspeakers), playback configuration (e.g., 5.1, stereo downmix of 5.1, virtualized stereo based on 5.1), acoustic path distortions (colorization, reverberation). For instance, if the transducer type is loudspeakers a transducer SEM value may be lower compared to if the transducer type is headphones whereby the transducer SEM is higher.
[0064] Without loss of generality, it will in the following be assumed that a high SEM value is associated with clear, high-quality speech and that a low SEM value is associated with poor-quality speech which may have low intelligibility. While an opposite definition of a SEM is possible (with a low SEM indicating clear and high-quality speech and a high SEM indicating poor quality speech), it is understood that this would be technically equivalent.
[0065] In some examples, the at least one SEM is a value defined on the range 0 to 1, with 0 indicating poor-quality speech which may have low intelligibility and 1 indicating clear and / or high-quality speech. A “high” SEM may in this example be a SEM between 0.5 and 1 wherein a “low” SEM may be a SEM between 0 and 0.5 although other definitions of a “high” or “low” SEM are envisaged.
[0066] Generally, if the SEM generator 1 determines one or more SEMs that have a high value this may be an indicator that less aggressive speech enhancement processing is to be performed in the speech enhancement processor 3 compared to the situation where one or more SEMs that have a low value are determined.
[0067] In some implementations, the at least one SEM is both time and frequency dependent and determined by the SEM generator 1 for each frequency band of a plurality of frequency bands of the mix input audio signal M.
[0068] In some implementations, the at least one SEM comprises a speech-to-non-speech ratio (SNR) extracted by the SEM generator 1. To extract the SNR, the SEM generator 1 may obtain the segments of the mix input audio signal MIN and separate each segment into a speech segment MS and non-speech segment MNS. For example, the SEM generator 1 may utilize a speech / non-speech separator to separate a segment of the mix input audio signal MIN into a speech segment MS and non-speech segment MNS. The speech separator of the SEM generator 1 may be identical to the speech separator used in the speech enhancement processor in some implementations.
[0069] Hereby, it is envisaged that in implementations where the SEM generator 1 uses a separated speech and non-speech content and the speech enhancement processor 3 uses speechand non-speech content, a single speech / non-speech separator may be used. The single speech / non-speech separator extracts the separated speech and non-speech content wherein the separated speech content and non-speech content is provided on the one hand to the SEM generator 1 (e.g. for calculation of the SNR) and other hand to the speech enhancement processor 3 (e.g. for determining modification gains and application of the modification gains to the speech and non-speech content respectively, based on the at least one SEM) as will be described in connection to FIG.2. By using the same speech / non-speech separator the computational efficiency is enhanced.
[0070] In some implementations, the separated speech segments MS and non-speech segments MNS may already be available (e.g. when the mix input audio signal MIN comprises a speech stem and a non-speech stem, or, when the mix input audio signal MIN has been separated into speech and non-speech segments elsewhere in the SE processing, e.g., the output of a common speech / non-speech separator may be used by both the SEM generator and the SE processor) wherein the SEM generator 1 does not need to perform the speech / non-speech separation.
[0071] The SEM generator 1 then determines the energy of a speech segment MS and the energy of an associated time-aligned non-speech segment MNS for each of a plurality of time- frequency tiles that span each segment MS, MNS. Each time-frequency tile is associated with a time slot ^ and a frequency band ^. Subsequently, the SEM generator 1 determines the ratio of the speech content energy in the speech segment MS to the non-speech content energy of the non-speech segment MNS, for each time-frequency tile.
[0072] There are several options for calculating this ratio, referred to as the time andfrequency selective speech-to-non-speech ratio (SNR), which may be denoted ^^^^^, ^^. InFIG.2A the SEM generator 1 obtains the MS and MNS audio segments and calculates the SNR as E(MS) / E(MNS) with “E” indicating the energy of a segment. The energy is typically proportional to the square of the absolute value of the signal values (amplitude or spectral coefficient magnitude). A drawback with an SNR calculated as E(MS) / E(MNS) is however that, when the non-speech audio segments MNS are extracted in the speech / non-speech separator 31 as MNS = MIN – MS, there is a risk that E(MNS) underestimates the energy of non-speech components since some non-speech content gets mistakenly classified as speech and appears in the MS segments. This leads to an underestimation of E(MNS) and an overestimation of the SNR which, when the SNR is used as a SEM, in turn may lead to undesirable misguiding of the speech enhancement process.
[0073] An alternative method for calculating the SNR is to instead calculate the SNR as SNR = E(MS) / [E(MIN) – E(MS)]. Compared to calculating the SNR as E(MS) / E(MNS), thisapproach has the advantage of reducing the impact of any residual non-speech content present in the MS segments, whereby the non-speech energy metric becomes more robust against the presence of non-speech residuals. If the SNR is calculated as E(MS) / [E(MIN) – E(MS)] the SEM generator 11 may not need to obtain the MNS segments at all and instead obtains the MIN segments as will be described below, in connection with FIG.2C.
[0074] In the following, unless specified otherwise, it is understood that either of these two methods for calculating the SNR may be used.
[0075] In some implementations, the time and frequency selective ^^^^^, ^^ is outputteddirectly by the SEM generator 1 whereby the time and frequency selective ^^^^^, ^^ is useddirectly as the at least one SEM. The speech enhancement processing in the speech enhancementprocessor 3 may then be based on ^^^^^, ^^ whereby the speech enhancement processing isperformed at the same time-frequency tiling as that of the time and frequency selective^^^^^, ^^.
[0076] In some implementations, the SEM generator 1 extracts a broadband SNR for each segment as the at least one SEM. The broadband SNR may be determined by computing the frequency- and / or time-weighted energy of the speech segment MS and the frequency- and / or time-weighted energy of the non-speech segment MNS, and then computing the ratio between the frequency-weighted and / or time-weighted speech segment MS energy and the frequency- weighted and / or time-weighted non-speech segment MNS energy.
[0077] Optionally, perceptual filtering, e.g., A- or B-weighting, or K-weighting as described in ITU-R BS.1770: Algorithms to measure audio programme loudness and true-peak audio level, 2023, hereby incorporated in it its entirety by reference, is applied to the speech segment and the non-speech segment prior to energy computations.
[0078] The SNR may be computed in time-frequency tiles and aggregated into a time- varying scalar broadband SNR referred to as ^^^^^^. For example, the aggregation is performed in accordance with^^^^^^ = !^"^, ^^^^^^^ + "^, ^^^2^^,!^"^, ^^ is a time and frequency dependent weight, * specifies the analysis time window, and ^specifies the set of analysis frequencies.
[0079] In some cases, the dynamic range of the SNR is limited (e.g., to the range -15 to+15 dB) wherein a proper selection of weights !^"^, ^^ yields a metric of speech intelligibility(e.g., the Articulation Index). In one example, the weights !^"^, ^^ have a perceptual weightingalong frequency, e.g., A-weighting, represented by a vector !^^^^, and a simple triangularwindow along time, represented by a vector !^^"^^ wherein !^"^, ^^ is then the outer productof !^^"^^ and !^^^^.
[0080] The above weights !^"^, ^^ may be static over time but it is also envisaged that insome cases weights may be time varying and adapted based on the signal properties of the speech segment MS, non-speech segment MNS or the mix input audio segment M. For example, in some implementations, SNR values are gated out in time-frequency tiles with very low speech energy (relative to adjacent time-frequency tiles).
[0081] Another type of aggregation alternative to the aggregation presented in equation 2 is^^^^^^ = $^ m∈&i,'n∈( !^"^, ^^^^^^^ + "^, ^^ ^3^2 and 3^^^^^^ =$^ m∈&a,'x∈( !^"^, ^^^^^^^ + "^, ^^. ^4^anda or a of channels and objects, the computation of the speech energy, in e.g. a TF-tile, can be done by computing the per-channel and per-object energies in that TF-tile for each channel and object, and then summing the energies over all the channels and objects that constitute the speech signal. Similarly, the non-speech energy in a TF-tile can be computed by first computing the per- channel and per-object energies in that TF-tile, and then summing the energies.
[0083] The multiple channels and / or audio objects may still be processed individually even if the SEM is based on a combination of properties across multiple channels or objects, as described in more detail in connection with FIG.10.
[0084] Any of the above SNR examples may be provided as the at least one SEM to the parameter mapper 2 which maps the at least one SEM to at least one processing parameter ^ that will influence the operation of the speech enhancement processor 3. Since the SEM (e.g. the SNR) changes dynamically over time, from segment to segment or even from time-frequency tile to time-frequency tile, the speech enhancement processor 3 will also be controlled dynamically based on the one or more SEMs and the processing parameters ^ extracted therefrom.
[0085] FIG.2A illustrates a speech enhancement processor 3 implementing so-called intrusive speech enhancement processing. In this example, a SEM generator 11 and a parameter mapper 21 are integrated in the speech enhancement processor 3 meaning that the speech enhancement processor 3 acts as a complete audio processing system 10. The SEM generator 11 and parameter mapper 21 operate analogously (and may be identical) to the (external) SEMgenerator 1 and parameter mapper 2 of FIG.1. In the implementation of FIG.2A the SEM generator 11 is configured to determine at least one SEM (e.g. the SNR) based on separated speech content and non-speech content.
[0086] The speech enhancement processor 3 of FIG.2A obtains the mix input audio signal M. The mix input audio signal MIN comprises a sequence of mix audio segments and the speech enhancement processor 3 separates each mix audio segment into a speech audio segment MS and a non-speech audio segment MNS using a speech / non-speech separator 31. The speech enhancement processing is performed by the gain modification stage 34 comprising gain modification units 34a, 34b which apply a boosting gain to one of the speech audio segment MS and non-speech audio segment MNS and an attenuating gain to the other one of the speech audio segment MS and non-speech audio segment MNS, forming gain modified speech segments and non-speech segments MMS, MMNS which are mixed at mixing stage 35 to form a modified mix audio segment MM. In some implementations, the gain modification stage 34 and mixing stage 35 are implemented in a receiving device (e.g. in a decoder) and the speech / non-speech separator 31, SEM generator 11 and parameter mapper 21 are implemented in a separate transmitting device (e.g. in an encoder).
[0087] Hereby, the same speech / non-speech separator 31 extracts the speech audio segment MS and non-speech audio segment MNS that in turn can be used by both the SEM generator 11 and the parameter mapper 21.
[0088] Alternatively, the speech enhancement processor 3 directly obtains a sequence of speech audio segments MS and non-speech audio segments MNS (e.g. the mix input audio signal comprises multiple stems, a speech stem and a non-speech stem) whereby the speech / non-speech separator 31 can be omitted.
[0089] The separation performed by the speech / non-speech separator 31 may be any type of speech / non-speech separation. For example, the speech / non-speech separator applies a time domain filter, a frequency domain filter or employs a neural network that predicts a separation gain mask for separating speech content from non-speech content on a segment-to-segment basis.
[0090] The speech audio segments MS and / or the non-speech audio segments MNS are in this implementation provided to the SEM generator 11 which is configured to determine at least one SEM based on the speech audio segment MS and / or the non-speech audio segment MNS. For example, the SEM generator 11 determines one of the SNR measures as described above. The at least one SEM extracted by the SEM generator 11 is mapped to at least one processing parameter ^ by the parameter mapper 21. In this embodiment, the at least one processing parameter ^ is two processing parameters, a first and second modification gain ^1, ^2 and thesemodification gains ^1, ^2 are provided to the first and second gain modification units 34a, 34b. The first modification unit 34a also obtains the speech audio segments MS and the second modification unit 34b also obtains the non-speech audio segments MNS.
[0091] Based on the first and second modification gain ^1, ^2, the gain modification units 34a, 34b applies a corresponding gain to the speech segment MS and non-speech segment MNS respectively. Since the at least one SEM is determined for each segment, and therefore may vary on a segment-to-segment basis, the gains applied by each gain modification unit 34a, 34b will in general vary from segment-to-segment yielding a dynamic speech enhancement for improving the speech experience by rebalancing the speech content with respect to the non-speech content.
[0092] It is noted that the gain modification units 34a, 34b may operate in a time domain or a frequency domain whereby the modification gains ^1, ^2 are applied in the time domain or frequency domain. For example, the gain modification units 34a, 34b operate on time-frequency tile representations of the speech segment MS and non-speech segment MNS respectively.
[0093] In the following, the at least one SEM will be described with a single variable X. However it is understood that the parameter X can represent a single SEM or a (e.g. weighted) combination of multiple SEMs. For example, X is the SNR in dB or X is the SNR in dB weighted with a playback environment SEM.
[0094] In some implementations, to find the modification gains ^1 and ^2, the parameter mapper uses an individual mapping function for each gain. That is, ^1 is found using a first mapping function f1(X) which maps the parameter X to a first modification gain ^1 and a second mapping function f2(X) which maps the parameter X to the second modification gain ^2. Notably, the same parameter X (i.e. the SEM or combination of SEMs) is used to determine both modification gains ^1 and ^2. For example, each mapping function f1, f2implements a functionin accordance with equation 5 below, with individual values of ^23, ^456, ^&7, ^'5$8,^^^^9:;88<=, *0, *1, *2, *3, with ^1 interpreted as an amplification and ^2 as an attenuation.
[0095] In some implementations to find the modification gains ^1 and ^2 to be applied, X is first mapped to gain level G (e.g. expressed in dB) according to a single mapping function. One exemplary mapping function that maps X to the gain level G is ^23 , X > *35^values of the thresholds are chosen to fit the one or more SEMs used as X and they satisfy *0 <*1 < *2 < *3. The mapping function described in equation 5 is illustrated in FIG. 3, and asseen in FIG.3, when X is equal to or exceeds T3 the gain level ^ is very low and equal to GHQwhich represents the gain level ^ to be used when the speech experience is already good. For example, GHQ = 0 dB meaning that no gain adjustment may be performed when the speech experience is already satisfactory.
[0096] Between T2 and T3, the gain level ^ is gradually increasing as X decreases, up to amaximum gain level of ^ = ^456. The reason for introducing the upper bound ^456 is that toohigh a gain level risks making artifacts in the speech / non-speech separation audible. Even in case there are no speech / non-speech separation artifacts, the upper bound ^456can be useful to reduce the risk of excessive speech / non-speech modification which can be objectionable to a listener. As the X decreases below T2 and approaches T1, the gain level ^ is gradually reduced and reaches ^'5$8at T1. ^'5$8may be 0 dB or > 0 dB. For example, it may be perceptuallybeneficial to attempt to carefully boost speech segments MS even when X is below T1 whereby^'5$8 may be selected as > 0 dB . The gain level ^ is then constant between T0 and T1 and as Xdecreases below T0 it is assumed that the at least one SEM are so low, that likely no speech is present whereby speech enhancement processing is not needed, and the gain level is set to 0 dB or a minimum amount of gain ^^^^9:;88<=.
[0097] It is also noted that the function shown in FIG.3 is merely exemplary. Generally, a bell shaped function over X is desirable whereby comparatively less aggressive speech enhancement (lower gain level ^) processing is performed for the extreme values of X (at the low extreme end there is no speech activity; at the high extreme end the background does not interfere with the speech) and comparatively more aggressive speech enhancement processing is performed for intermediate values of X (in this region we expect SE to be useful). However,while FIG. 3 shows that ^'5$8> ^^^^ :;88<=> ^23this is merely exemplary and ^'5$8,^^^^ :;88<=, and ^23 may be defined differently and may even be the same, e.g. equal to 0 dB.Another example of a mapping function based on equation 5 is when ^'5$8 = ^^^^ :;88<= =^456, i.e., we modify the speech and background to the maximum amount specified by ^456 atvery low values of the SEM and also when there is no speech activity. Yet another example of amapping function is to use ^23 = ^456, i.e., we modify the speech and background to themaximum amount specified by ^456even at very high values of the SEM (where it may not be necessary in order to make speech more intelligible but it may still be desirable in order to improve the overall experience).
[0098] In some implementations, the gain level ^ is smoothed across two or more segments to avoid sudden gain level changes. The smoothed gain level in a segment ^ is denoted Gsmooth(t) wherein the smoothing of the instantaneous gain level G as c^ROORST^LMNN , G > ^:4^^^=^t − 1^cLMNNOP = Q OP^W ^6^c XYXRLX^LMNNOP , otherwise^:4^^^=^t^ = ^1 − E:4^^^=^G + E:4^^^=GLMNNOP^t − 1^ ^7^wherein the attack and release coefficients c^WXYXRLX^ ^ROORST^LMNNOPand cLMNNOPinfluence the attack / release behavior of the smoothing. In some is slower compared to the ack, meaning that c^WXYattXRLX^LMNNOPis greater cLMNNOP.
[0099] The gain level G is then partitioned or “split” by the parameter mapper 2 into twomodification gains, a first modification gain ^1 and a second modification gain ^2 wherein|^1| + |^2| = ^. The first modification gain ^1 is applied to the speech segments Ms by thefirst modification unit 34a and the second modification gain G2 is applied to the non-speech segments MNS by the second modification unit 34b. At least one of the first modification gain ^1 and the second modification gain ^2 is a boosting (amplifying) gain and the other one of the first modification gain ^1 and the second modification gain ^2 is an attenuating gain.
[0100] In one example the first modification gain ^1 is a boosting gain whereby the first modification unit 34a boosts the speech segments MS and the second modification gain ^2 is an attenuating gain whereby the second modification unit 34b attenuates the non-speech segment MNS.
[0101] The smoothing described in relation to equations 6 and 7 may be done after the gain level has been split into ^1 and ^2. That is, the smoothing may be applied individually to ^1 and ^2, which allows using individual attack and release coefficients for ^1 and ^2. Similarly, in applications with two individual mapping functions, the smoothing may also be done individually on the speech and non-speech sequence of gains using individual attack and release coefficients.
[0102] The partitioning of the gain level ^ into the first and second modification gain is governed by a static or dynamic split ratio ^ which indicates how the gain level ^ is to be split into the first and second gain ^1, ^2.
[0103] For example, the split ratio ^ is defined on the range [0, 1] wherein the partitioningmay be performed asG1 = ^1 − ^^^G2 = −^^ ^8^wherein ^, ^1 and ^2 are expressed in decibels and the negative sign on ^2 indicates that it is an attenuating gain whereas ^1 is a positive boosting.
[0104] Alternatively, if ^1 is an attenuating gain and ^2 is a boosting gain the partitioningmay be performed as^^ = −^^^7 = ^1 − ^^^ ^9^in accordance with equation 9 also could be used to keep the speech level stable over time, which is one type of speech experience enhancement.
[0105] The split ratio ^ will determine the overall behavior of the gain modification performed by the gain modification units 34a, 34b. To illustrate this, two extreme cases arepresented. For example, assuming that the partitioning of equation 8 is used, if ^ = 0 the non-speech segments MNS will remain at a constant level (adjusted with ^2 = 0 dB) while thespeech segments MS are boosted by the amount dictated by the first modification gain ^1 = ^.For a bell-shaped (SEM to gain level) mapping function, and assuming no speech / non-speech separation artifacts, only speech will be perceived to be modified, and as the SEM increases from intermediate values to high values, the amount of speech modification decreases. Depending on the details of the mapping, the change in speech modification as SEM decreases may not be perceivable.
[0106] Conversely, if ^ = 1 (assuming that the partitioning of equation 8), the speechsegments MS will remain at a constant level while the non-speech segments MNS are boosted bythe amount dictated by the second modification gain ^2 = ^. For a bell-shaped mappingfunction, only the background will be perceived to be modified. At low values of the SEM the background will be unmodified (i.e., no or very little attenuation). As the SEM increases the amount of background modification increases (more attenuation of the background). As the SEM increases past intermediate values toward the high end, the background modification is decreased.
[0107] Artifacts that may be present in the speech to non-speech separation used to form the speech segments MS and the non-speech segments MNS will manifest themselves differently depending on the value of s. It is beyond the scope of this presentation to discuss these effects; however, the differences are relatively small compared to the sensitivity to mis-timed modifications discussed in the next paragraph. More critical to the robustness to separation artifacts is the behavior of the mapping function as the SEM decreases.
[0108] Especially disruptive, and more critical to the choice of s, are mis-timed gain modifications (e.g. gain modifications occurring mid-utterance, mid-word, mid-vowel, or just after a speech onset) applied by the first and second modification units 34a, 34b. A sudden (mis- timed) change of the gain level can result in a sudden perceived unnatural change in level of the speech and / or the non-speech content. Assuming that a sudden change in gain level ^ poses an equal risk of unnatural level fluctuations in the speech segments MS as in the non-speechsegments MNS, a static value of the split ratio s may be set to ^ = 0.5.
[0109] However, speech and non-speech content may exhibit different sensitivity to sudden imposed level changes and the sensitivities may vary over time. Thus, in some implementations the split ratio s is determined dynamically for each pair of a speech segment MS and non-speech segment MNS. The dynamic split ratio ^ may be based on at least one of: at least one SEM and a temporal stationarity measure.
[0110] For example, assuming the partitioning of equation 8 it may be beneficial to adapt the split ratio ^ by setting ^ closer to 0 for segments pairs where the content of the non-speech segment MNS is (temporally) stationary and where sudden changes to ^ risk leading to unnatural level fluctuations. On the other hand ^ can be set closer to 0.5, or closer to 1, when the content of the non-speech segment MNS is more non-stationary since sudden changes to the gain level G are then more likely to be masked by the non-stationarity of the non-speech segment MNS.
[0111] The split ratio ^ can further be interpolated between these stationarity measures so as to vary smoothly with the stationarity measure. In FIG.4 a graph illustrating an exemplary mapping between the split ratio s and the stationarity measure of the non-speech segments MNS is shown. As seen, the split ratio s may be a monotonically non-increasing function with increasing stationarity measure of the non-speech segment MNS.
[0112] To this end, the processing parameter mapper 21 may e.g. obtain two SEMs for each pair of speech and non-speech segments MS, MNS. These two SEMS are SNR and a measure of the stationarity of the non-speech segment MNS so as to dynamically determine ^ based on the SNR and the split ratio S for portioning ^ into ^1 and ^2 based on the stationary measure.
[0113] Determining the level of stationarity of the non-speech content can be done in several ways and various stationarity measures are known to those skilled in the art. An example method is to compute the variance of the full- or sub-band energy in a context window. The time context window comprises a plurality of time slots. For example, the variance of the energy or loudness of each tile of a time-frequency tile representation of a segment of audio content can be calculated and the level of stationarity determined based on the variance. It is understood that acomparatively larger variance will be associated with a lower level of stationarity and that a comparatively smaller variance will be associated with a higher level of stationarity.
[0114] Additionally or alternatively, the split ratio ^ is determined based on the at least one SEM (e.g. the SNR). As the energy level of the speech content decreases relative to the non- speech content, (i.e. as the SNR decreases) the speech separation process used to separate the mix input segments MIN into speech and non-speech segments MS, MNS comes under increased strain, and many speech separation processes tend to output a weaker (lower loudness) speech estimate as this strain becomes more severe. This can partially be compensated for by decreasing the split ratio ^, assuming equation 8 is used, i.e., allocate more of the modification to the speech, as the strain increases or, in other words, by decreasing the split ratio ^ as e.g. the SNR (or one or more other SEMs) decreases.
[0115] The split ratio S may hereby vary on a segment-to-segment basis whereby the split ratio s for at least one pair of time aligned segments is such that both ^1 and ^2 are non-zero. For example the split ratio for at least one pair of time aligned segments may be in the range of 2% to 98%, or in the range of 5% to 95%.
[0116] The gains ^1, ^2 applied by the modification units 34a, 34b may be broadband gains that are applied to the full frequency band of the speech segments MS and the non-speech segments MNS respectively. However, it is also envisaged that the parameter mapper 21 determines a gain level ^ and a split ratio ^ individually for a plurality of frequency bands for each segment whereby the modification gains ^1, ^2 are also determined individually for each frequency band of each segment. The modification units 34a, 34b may then apply the frequency band specific modification gains ^1, ^2 individually for each frequency band. The at least one SEM, and also the stationarity metric of the non-speech segments MNS, may be determined individually in a plurality of frequency bands to allow the speech enhancement processor 3 to operate independently on individual frequency bands.
[0117] Additionally or alternatively, the modification units 34a, 34b may be configured to apply a dynamic EQ filter wherein a scaling of the EQ filter is based on at least one SEM. That is, each modification unit 34a, 34b is configured to apply a plurality of subband gains e^^^ to acorresponding plurality of ^ ≥ 2 frequency bands of the speech segments MS and the non-speech segments MNS respectively. For example, each segment may comprise a DFT representation wherein ^ indicates a frequency range of the DFT. The nominal gains are to be applied to the segments, for example in a suitable frequency transform domain, such as in the discrete Fourier transform (DFT) domain. The details of DFT based filtering are well-known to those skilled in the art.
[0118] The subband gains e^^^may be scaled based on the at least one SEM, wherein the at least one SEM may be a broadband SEM or frequency selective SEM.
[0119] In general, EQ adjustment comprises applying an EQ filter to attenuate and / or boost the speech and / or non-speech segments MS, MNS in a frequency selective manner. Traditionally, an EQ filter is a time invariant, and linear, filter but in the implementations envisaged here the EQ filters applied by each gain modification unit 34a, 34b may be dynamically scaled and / or otherwise modified on a segment-by-segment basis based on the at least one SEM.
[0120] The subband gains e^^^may be adjusted with a factor ^, to yield adjusted subbandgains e ∗ ^^^ wherein the factor ^ is based on the at least one SEM, again described as X, bymultiplying the subband gains with ^ which yields:e ∗ ^^^ = ^^^^e^^^, ^ = 1, … , h ^10^where the gains are assumed to be in decibels. The factor ^ is typically between 0 and 1, and a high value of X maps to a low value of ^ and a low value of X maps to a high value of ^. However, it may be beneficial to decrease ^ as the SEM value X decreases beyond some threshold and again reference is made to FIG.3 showing how ^ may be adjusted as X varies.
[0121] Notably, modification of the subband gains e^^^ in accordance with equation 10 is analogous to the determination the gain level ^ based on the at least one SEM. While ^ is a single broadband gain applied to all frequencies of the segments MS, MNS, the adjusted nominal gains according to equation 10 are frequency selective.
[0122] Based on the at least one SEM and / or the non-speech segment MNS stationaritymeasure the subband gains e^^^ may be split to form two sets of subband gains e^^^^ ande 7^^^ using a split ratio s similar to the gain splitting described above. The split ratio s may bethe same for each subband or determined individually for each subband yielding ^^^^. The two sets of sets of subband gains e^^^^ and e7^^^ are then applied by the gain modification units 34a, 34b in manner analogous to the application of the single (broadband) gains ^1 and ^2 described above, with one set of subband gains being a boosting gain and the other set of subband gains being an attenuating gain.
[0123] In some implementations, the EQ filter is a non-parametric filter specified by a setof nominal frequency selective gains eiNMjiRY^^^, ^ = 1, … , h, wherein k is a frequency indexindicating frequency subbands of each segment. For example, nominal frequency selective gains eiNMjiRY^^^ specify a bandpass filter. In such implementations, the subband gains are given by the nominal gains eiNMjiRY^^^ and the processing is made dynamic by determining the split ratio s or ^^^^ and / or the scaling factor ^ based on the at least one SEM and / or the non-speech segment MNS stationarity measure for each pair of time aligned segments.
[0124] In some implementations, the nominal gains eiNMjiRY^^^define a parametric filter which is defined by at least one filter parameter defining the relative level of the nominal gains (i.e. the spectral shape of the filter). Hereby the filter may further be adjusted by adjusting the filter parameter based on the at least one SEM.
[0125] In some implementations, the frequency specific gains e^^^ define a parametric high pass shelving filter which is defined by the shelf gain acting as a filter parameter. The shelf gain of the high pass shelving filter is controlled based on the at least one SEM represented as X. For example, as shown in FIG.3, the shelf gain increases as X decreases from T3 towards T2, the shelf gain decreases as X decreases from T2 to T1 and so on, giving a mapping between X and shelf gain similar to that between X and gain level ^ or between X and the scaling factor ^. It is also envisaged that other types of mapping between shelf gain and X can be applied depending on e.g. the type of application. For example, X may be mapped to the shelf gain with a non-decreasing function of X or a non-increasing function of X.
[0126] There are many examples of high-pass shelving filters that may be used. In some cases, the filter is expressed using a Z-transform transfer function and one exemplary high-pass shelving filter expressed with a transfer function is: m^ tan n!< o + ^ + n ^ tan n!< o − ^ o l9^:k l^ = 2 : m : 2 :^ ^11^^:expressed in radians per second (i.e. p< = 2q^<⁄ ^: ). For example, the cut-off frequency is 4 kHz,the sampling frequency is 48 kHz (giving p< = 2q 4⁄ 48 = q / 6) and the shelf gain ^: variesbetween 1 dB and 6 dB. The frequency specific gains e^^^of this filter may be obtained by sampling the transfer function k^l^at frequencies ^ at the center of the ^ subbands. The frequency specific gains e^^^ may then be applied directly to a frequency domain representation of the respective audio segments.
[0127] Alternatively, is understood that frequency selective gains forming filter may be applied in the time domain. As explained above, filters can be expressed as transfer function in the Z-transform domain whereby the transfer function can be sampled to yield frequency specific gains e^^^ that may be applied to e.g. a time-frequency tile representation of the respective segment. To apply the EQ filter in the time domain an autoregressive moving-average, ARMA, filter corresponding to the Z-transform transfer model can be extracted. If t^l^ and ^^l^ denote the Z-transforms of the time domain output and input u^v^ and ^^v^ of the filter any transfer function k^l^ is by definitionk^l^ = w^x^<=> t^l^ = k^l^^^l^ ^12^which corresponds to a time domain linear difference equation of: z{u^v^ + z^u^v − 1^ = |{ + |^^^v − 1^ ^13^which can be rewritten as |+ | ^^v − 1^ − z u^v − 1^u^v^ = { ^ ^z ^14^{=skilledappreciate that time domain filter implementations other than ARMA filters mayFor example, time domain filter implementations with greater numerical stability may be used.
[0128] Hereby, it is understood that any application of an EQ filter to audio segments can be performed in the frequency domain or in the time domain.
[0129] In some examples, the cut-off frequency ^<is kept constant as X varies but it is also envisaged that the cut-off frequency also varies as a function of X.
[0130] It is also envisaged that other types of parametric filters could be used instead of the high pass shelving filter of equation 11 and / or other parameters of parametric filters such as Q-value and / or filter gain, are adjusted based on the at least one SEM.
[0131] FIG.2B is a flowchart illustrating a method for processing audio which e.g. may be performed by the speech enhancement processor 3 of FIG.2A.
[0132] At step S1 a first and second audio segment is obtained. The first and second audio segment may be speech audio segment MS and non-speech audio segment MNS as extracted by a speech / non-speech separator 31. The first and second audio segment may further be segments directly obtained from a multi-stem input audio signal, for example the first audio segment is a segment of a speech stem signal of the multi-stem input audio signal and the second audio segment is a segment of a non-speech stem signal (e.g. a music and effects stem) of the multi- stem input. As a further example, the first segment is a segment of the mix input audio signal which has undergone speech enhancement processing MSE and the second segment is a segment of the input mix audio signal MIN as will be described below, in connection to the non-intrusive processing.
[0133] The first and second segment are time-aligned meaning that they correspond to the same time interval of the (original) mix input audio signal M.
[0134] The method then goes to step S2 involving obtaining at least one SEM. For example, the at least one SEM is extracted by the SEM generator 11 or included as metadata together with the mix input audio signal M.
[0135] At step S3, the gain level G is determined based on the at least one SEM. The parameter mapper 21 may determine the gain level G based on the at least SEM.
[0136] The method then goes to step S4 involving obtaining the split ratio S indicatinghow the gain level is to be partitioned into the two modification gains ^1, ^2. As describedabove the split ratio ^ may be static, whereby obtaining the split ratio ^ may simply involve accessing the split ratio ^. Alternatively, the parameter mapper 21 determines the split ratio ^ based on at least one of at least one SEM and a stationarity metric of the second segment (e.g. a non-speech segment MNS) as described above.
[0137] With the split ratio ^ and gain level ^, the method goes to step S5 comprisingportioning (splitting) the gain level into two modification gains ^1, ^2 wherein one of thesegains is a boosting gain and the other gain is an attenuating gain. In step S6 the modification gains are applied to the first and second segment respectively so as to form gain modified segments MMMS, MMNS, which are mixed together to form a gain modified mix audio segment MM. Finally, at step S7 the modified mix audio segment MM is output. Hereby the method has achieved dynamic speech enhancement by extracting and applying modification gains in a dynamic manner, based on the at least one SEM.
[0138] In some implementations, the gain application and mixing of steps S6 and S7 is instead performed in a receiving device whereas steps S1-S5 are performed in a transmitting device. To this end, the method may comprise a further optional step after step S5, and prior to step S6, comprising encoding the modification gains into a bitstream and transmitting the bitstream to a receiving device.
[0139] FIG.2C shows an alternative speech enhancement processor 3a that is similar to the speech enhancement processor of FIG.2A with the difference that speech enhancement processor 3a does not utilize dedicated non-speech audio segments MNS.
[0140] The speech enhancement processor 3a obtains the mix input audio signal MIN and extracts for each mix audio signal segment a speech audio segment MS using the speech separator 31’. Compared to the speech / non-speech separator 31 of FIG.2A, the speech separator 31’ only extracts and outputs speech segments MS as opposed to both speech MS and non- speech segments MNS. Hereby, extraction of non-speech segments MNS, for example as MNS = MIN – MS, is not necessary wherein speech separator 31’ may be computationally less complex compared to the speech / non-speech separator 31.
[0141] The speech segments MS are provided to a SEM generator 11 which determines at least one SEM based on the speech segments MS. The at least one SEM may be configured to determine an SNR as the ratio between the energy of the speech segments, i.e. E(MS), and the energy of non-speech content e.g. calculated as E(MIN) - E(MS). To this end, the SEM generator11 may first obtain the speech segments MS as input and determine the energy E(MS). The SEM generator 11 then obtains the input mix segments MIN and determines E(MIN) – E(MS) describing the non-speech energy whereby the SNR is extracted as E(MS) / [E(MIN) – E(MS)]. Accordingly, the SEM generator 11 may not need dedicated non-speech segments and may operate using only MS and MIN segments.
[0142] A benefit with determining the SNR as E(MS) / [E(MIN) – E(MS)] is therefore that no dedicated MNS segments need to be extracted while also this SNR is more robust against the presence of non-speech residuals in the speech segments MS, as previously described.
[0143] The SEM is provided to the parameter mapper 21 which maps the SEM to two gains ^1 and ^2 using two mapping functions f1(X), f2(X) or splitting, as described above in connection with FIG.2A. However, in the speech enhancement processor 3a of FIG.2C no dedicated non-speech segments are available and the gains G1 and G2 are instead to be applied to the mix input audio segments MIN and the speech audio segments MS by the gain modification stage 34.
[0144] It has been realized that a gain modified mix audio segment MM equivalent to that of FIG.2A may be generated by the gain modification stage 34 and the mixing stage 35 using the mix input segments MIN, the speech segments MS and the gains ^1 and ^2. Morespecifically, the gain modified mix audio segment MM may be formed as^^ = ^^1 − ^2^^^ + ^2 × ^^^. ^15^This is illustrated in FIG.2C wherein the first gain modification module 34a applies a gain of ^1 – ^2 to the MS segment to form MMS and wherein the second gain modification module 34b applies a gain of ^2 to the MIN segment to form MMNS wherein MMS and MMNS are mixed in the mixing stage 35. The resulting MM segment generated by the speech enhancement processor 3a is equivalent the MM segment generated by the speech enhancement processor 3 of FIG.2A for implementations wherein MNS = MIN – MS. For reference, in FIG.2A, the MM segments are extracted from MS and MNS as: ^^ = ^1 × ^^ + ^2 × ^^^ ^16^
[0145] Accordingly, as is understood from FIGS.2A and 2C, the speech enhancement processor 3, 3a may be implemented with or without a dedicated sequence of non-speech segments, i.e. with MS segments and MNS segments or using MS segments and MIN segments.
[0146] Turning to FIG.5A, a speech enhancement processor 3' is shown for performing DRC processing instead of gain modification. The speech enhancement processor 3' is similar to the speech enhancement processor 3 of FIG.2A with the difference being that the gain modification stage 34 comprising gain modification modules 34a, 34b has been replaced with DRC processing stage 36 comprising DRC processing modules 36a, 36b.
[0147] While FIG.5A shows two DRC processing modules 36a, 36b, one DRC processing module 36a for the speech segments MS and one DRC processing module 36b for the non- speech segments MNS. DRC is in some implementations performed only on speech segments MS by the first DRC processing module 36a (and the second DRC processing module 36b can be omitted) or performed only on non-speech segments MNS by the second DRC processing module unit 36b (and the first DRC processing module 36a can be omitted). Alternatively, DRC processing is performed on both the speech segments MS and the non-speech segments MNS. That is, in general, the DRC processing may be performed on either the speech segments MS or the non-speech segments MNS or on both the speech segments MS and the non-speech segments MNS.
[0148] DRC is a well-established technique to control level and amplitude of an audio signal. It is often used in general audio processing (e.g. in music production) to make the perceived loudness more constant over time.
[0149] DRC tends to increase the perceived “presence” of the dynamic range compressed content whereby, to e.g. improve the speech experience, DRC may be applied to the speech segments MS only. However, DRC may also be applied to the non-speech segments MNS to make the loudness of the non-speech content more constant over time which also may improve the speech experience.
[0150] At least one of the first and second DRC processing unit 36a, 36b may apply DRC with a compression ratio R that is based on at least one SEM extracted by the SEM generator 11. For example, the compression ratio R is determined by the parameter mapper 21 which converts the at least one SEM to a compression ratio R which then is provided to the first and / or second DRC processing unit 36a, 36b which applies DRC with the compression ratio R. The DRC parameters will accordingly be time-varying , since the applied compression ratio R may vary from segment to segment, as the at least one SEM varies.
[0151] In implementations where both DRC processing units 36a, 36b are used the parameter mapper 21 may determine two DRC ratios R1 and R2 which are provided to, and applied by, a respective one of the DRC processing units 36a, 36b, as shown in FIG.5A.
[0152] In some implementations, the DRC processing performed by the first and / or second DRC processing unit 36a, 36b comprises application of an instantaneous compression gain e^^^that fulfills ^− 1^^ − ^ > *DRC processing unit 36a, 36b, T is a DRC threshold (in dB), and E is the signal energy (in dB)of the speech segment MSand / or non-speech segment MNSprovided to the first and / or second DRC processing unit 36a, 36b.
[0153] The instantaneous DRC gain e^^^may be applied directly to the speech and / or non-speech segments MS, MNS to form corresponding modified segments MMS, MMNS.
[0154] Alternatively, as the instantaneous DRC gain e^^^obtained using equation 17 is determined for each segment MS, MNS, and may fluctuate rapidly, the instantaneous DRC gain e^^^may be smoothed using attack and release processing (with time-constants tattackand trelease) to form a smoothed DRC gain, whereby the smoothed DRC gain is applied to the speech and / or non-speech segments MS, MNS to form DRC modified segments MMS, MMNS.
[0155] It is envisaged that DRC parameters other than the ratio R may be adjusted based on the at least one SEM in addition to, or as an alternative to, adjustment of the ratio R. For example, multiple DRC parameters (such as the ratio R, the threshold T and / or the smoothing time-constants tattackand trelease) may be adjusted at the same time based on the at least one SEM and / or the signal energy E of the speech or non-speech segments MS, MNS.
[0156] As mentioned above, the ratio R for the DRC processing may be adjusted based on the at least one SEM. In FIG.3 an example of how the DRC ratio R may be adjusted based on at least one SEM (expressed as the variable X) is shown. As seen the compression ratio R varies from 1 (meaning no compression) to 10 (strong compression) so as to favor low or no compression for high values of X (i.e. already good speech experience), favor relatively strong compression when X is between T1 and T3 and favor relatively weak (but not non-existent) compression when X is below T1. However, this is merely one exemplary mapping function between X and it is understood that the compression ratio R may be determined with a different mapping function and / or that the minimum and maximum compression ratios (i.e.1 and 10) shown in fig.3 could be selected differently.
[0157] In some implementations, a make-up gain Mg is applied to the segments MS, MNS subject to DRC processing to compensate for the likely loss of total energy as the DRC ratio R increases. For example, the make-up gain Mg is adjusted according to a curve similar to that shown in fig.3 favoring higher make-up gains for higher applied DRC ratio R. While the DRC gain e^^^is applied with very fine granularity (e.g. applied for each sample of each segment or at least updated multiple times during a segment) the make-up gain Mg is applied with a coarser granularity (e.g. once of each segment).
[0158] In some implementations, the DRC threshold T is static. Alternatively, the DRC threshold T is adjusted based on the segment energies (or loudness) of the speech and / or non- speech segments MS, MNS, accumulated over a time window (ideally positioned centered around the current segment that is being processed) ranging from the order of seconds up to theduration of the complete soundtrack. A DRC threshold adaptation block 37 may be used to determine the accumulated signal energy or loudness over a plurality of speech and / or non- speech segments and adjusts the DRC threshold T based on the determined energy or loudness. For example, the DRC threshold T increases with increasing accumulated signal energy or signal loudness and decreases with decreasing accumulated signal energy or signal loudness.
[0159] Irrespective of the type of DRC processing applied, the DRC processing modules 36a, 36b output DRC modified speech and non-speech segments MMS, MMNS to a mixing stage 35 which mixes the modified segments MMS, MMNS to form the output mix audio segment MM. In some implementations, the DRC processing stage 36 and mixing stage 35 are implemented in a receiving device and the speech / non-speech separator 31, SEM generator 11 and parameter mapper 21 (and optional DRC threshold adaptation blocks 37) are implemented in a transmitting device.
[0160] FIG.5B is a flowchart illustrating a method for processing audio which e.g. may be performed by the speech enhancement processor 3' of FIG.5A.
[0161] Similar to the method of FIG.2A, the method involves obtaining a first and second segment at step S1 and step S2 comprising obtaining at least one SEM associated with the first and second segment (or the segment of the mix input audio signal from which the first and second segment has been extracted).
[0162] Since the speech enhancement processor 3' of FIG.5A relates to dynamic range adjustment the method of FIG.5B then goes to step S31 comprising determining at least one DRC parameter based on the at least one SEM. The at least one DRC parameter is at least one of a DRC ratio, a DRC threshold, a make-up gain, a DRC attack time constant, and a DRC release time constant.
[0163] In some implementations, the step S31 comprises determining a DRC parameter for each respective DRC processing unit 36a, 36b. For example, step S31 comprises determining two DRC ratios R1 and R2, where the first DRC ratio R1 is conveyed to the first DRC processing unit 36a processing the speech segments MS and the second DRC ratio R2 is conveyed to the second DRC processing unit 36b processing the non-speech segments MNS. The DRC ratio(s) R may e.g. be calculated by the parameter mapper 2 based on the at least one SEM.
[0164] The method then goes to step S41 comprising performing DRC processing of at least one of the first and second segment using at least one of the first and second DRC processing modules 36a, 36b with the DRC parameter(s). Hereby, at least one DRC modified segment MMS, MMNS is obtained at the output of the DRC processing modules 36a, 36bwhereby the DRC modified segment MMS, MMNS is mixed with another DRC modified segment or a non-DRC modified segment to form a DRC modified mix output segment MM.
[0165] The method then goes to step S71, involving outputting the DRC modified mix segment MM.
[0166] The DRC processing method of FIG.5B and the gain modification processing method of FIG.2B may be used separately or combined. For example, the gain modification of both the first and second segment may be performed first whereafter DRC modification of at least one of the gain modified segments is performed, or vice versa. Optionally, the gain modification involves dynamic EQ application using subband gains e^^^, as described above, when combined with the DRC modification.
[0167] The above speech enhancement processors of FIG.2A and FIG.5A are examples of intrusive audio processing schemes wherein the gain modification or DRC processing is applied to separated speech and / or non-speech segments MS, MNS which may only be available inside the speech enhancement processors. With intrusive audio processing schemes it is meant that the internal processing flow of the speech enhancement processors 3, 3' is accessible.
[0168] On the other hand, in some implementations, the speech enhancement processor 3, 3’ is a “black box” which obtains a mix input audio signal MIN and produces a speech enhanced mix signal MSE as an output. For example, while some speech enhancement processors used in audio engineering offers access to the internal processing, some speech enhancement processors offer no possibility of external control to modify or otherwise adjust the speech enhancement processing occurring inside the speech enhancement processor.
[0169] In some implementations, the DRC processing and mixing of steps S41 and S71 are instead performed in a receiving device whereas steps S1, S2 and S31 are performed in a transmitting device. To this end, the method may comprise a further optional step after step S31, and prior to step S41, comprising encoding the at least one DRC parameter into a bitstream and transmitting the bitstream to a receiving device.
[0170] In FIG.6 a block diagram with a “black box” speech enhancement processor 3'' is shown. For the speech enhancement processor 3'' intrusive processing or any adaptive control is not possible and / or desirable. The speech enhancement processor 3'' obtains a mix input audio signal MIN and outputs a speech enhanced mix audio signal MSE. The details of the speech enhancement processing performed by the speech enhancement processor 3'' is irrelevant but generally the speech enhancement processor 3'' performs some type of processing to improve the speech experience.
[0171] However, since the separated speech and non-speech segments (if at all present) inside the speech enhancement processor 3'' are not accessible it is not possible to perform thesame type of gain modification or DRC processing described above. Nevertheless, a non- intrusive version of the gain modification and / or DRC modification may still be performed even though the inner workings of the processing of the speech enhancement processor 3'' is static or not known.
[0172] In FIG.6 the speech enhancement processor 3'' obtains a sequence of mix input audio segments MIN and outputs a sequence of speech enhanced mix audio segments MSE. The speech enhanced audio segments MSE comprises speech content but may also comprise non- speech content (e.g. attenuated compared to the mix input audio segments M).
[0173] The non-intrusive audio processing system 10' further comprises a SEM generator 1'. The SEM generator 1' is similar to the SEM generator 1 of FIG.1 or the SEM generator 11 described in connection to the intrusive processing implementations of FIG.2A and FIG.5A however the SEM generator 1' may not have immediate access to separated speech segments and non-speech segments as is the case for SEM generator 11.
[0174] The SEM generator 1' may hereby determine the at least one SEM based on the mix input audio segments MIN (carrying a mix of speech and non-speech content).
[0175] To determine the at least one SEM it is envisaged that the SEM generator 1' comprises a speech / non-speech separator as shown in FIG.7 whereby speech and non-speech segments are extracted from the mix input audio segments allowing e.g. the SNR to be extracted in a manner analogous to equations 2, 3 or 4 above to be used as the at least one SEM. Since the separated outputs of a speech / non-speech separator (if at all present) inside the non-intrusive speech enhancement processor 3'' are not available the SEM generator 1' may comprise its own speech / non-speech separator to enable e.g. the SNR to be determined as the at least one SEM.
[0176] Additionally or alternatively, at least one of the mix input audio segments M, auxiliary data (e.g. metadata) associated with the mix input audio segments MIN and the speech enhanced mix audio segments MSE output by the non-intrusive speech enhancement processor 3'' is used by the SEM generator 1 to extract at least one SEM. Notably, it is possible for the SEM generator to obtain the speech enhanced mix audio segments MSE as an estimate of the speech content and calculate an estimate of the non-speech content as the difference MIN - MSE.
[0177] The at least one SEM is provided to a parameter mapper 2 that maps the at least one SEM to at least one processing parameter ^ which is provided as control information to gain modification stage 44. The gain modification stage 44 may operate completely analogously to the gain modification stage 34 of FIG.2A. Alternatively, if DRC processing is desired the gain modification stage 44 is replaced with the DRC modification stage 34 of FIG.5A.
[0178] For example, in FIG.8 the details of the gain modification stage 44 are shown. In the non-intrusive processing system 10' the speech enhanced mix segments MSE from the speechenhancement processor 3'' are available and these are provided to the first gain modification module 44a. Segments of the (unprocessed) mix input audio signal MIN are provided to the second gain modification module 44b. The first and second gain modification modules 44a, 44b may now operate analogously to the gain modification modules 34a, 34b of FIG.2A and apply a boosting gain and attenuating gain respectively as dictated by the determined gain level and split ratio. The gain adjusted segments MGMSE, MGM output by each of the gain modification modules 44a, 44b are provided to the mixer 45 which mixes segments to form a modified mix output segment MM.
[0179] That is, the processing applied by each gain processing module 44a, 44b may be equivalent to the processing applied by the gain modification modules 34a, 34b used in intrusive audio processing systems. The main difference is that here, in the non-intrusive implementation, the segments input to the modification module 44 are speech enhanced mix audio segments MSE instead of speech segments MS from a speech / non-speech separator and the original input mix audio segments MIN instead of non-speech segments MNS from the speech / non-speech separator.
[0180] For example, the parameter mapper 2 may determine a non-intrusive gain level ^^^based on the at least one SEM and subsequently split the non-intrusive gain level ^^^into a firstand second non-intrusive gain ^^^^ and ^^^7 such that |^^^^| + |^^^7| = ^^^ in accordance witha non-intrusive split ratio ^^^(which in turn is based on at least one SEM or a stationarity measure of the mix input audio segments M). The first non-intrusive gain ^^^^is provided to the first gain modification unit 44a and applied to the speech enhanced mix segments MSE (coming from the “black box” speech enhancement processor) and the second non-intrusive gain ^^^7is provided to the second gain modification unit 44b and applied to the mix input audio segments M. One of the first and second non-intrusive gain ^^^^and ^^^7is used as boosting gain and the other one is used as an attenuating gain. For example, ^^^^is a boosting gain and ^^^7is an attenuating gain.
[0181] A benefit of non-intrusive speech enhancement processing is the structural simplicity of the solution which allows e.g. any speech enhancement processor 3'' to be used without knowledge of the internal configuration of the speech enhancement processor 3''. An advantage of the intrusive speech enhancement processing on the other hand is the added control of the behavior in intermediate points of operation, between low SEM values and high SEM values.
[0182] In some implementations, gain modification stage 44 is implemented in a receiving device, which receives a bitstream with the modification gains ^1, ^2 and audio content in the form of either a segment pair (e.g. an MSE / MIN pair as shown or an MS / MNS pair) or a mixsignal such as MIN (which can be separated into an MSE / MIN pair or an MS / MNS pair). Hereby, the application of the modification gains ^1, ^2 is performed in the receiving device which is separate from the transmitting device which has extracted the modification gains ^1, ^2 and optionally the segment pairs.
[0183] Turning to FIG.9, a DRC processing stage 46 for non-intrusive implementations is shown. The DRC processing stage is similar to the DRC processing stage 36 used in the intrusive DRC processing of FIG.5A. The difference is that the non-intrusive DRC processing stage 46 is configured to obtain speech enhanced mix audio segments MSE instead of speech segments MS from a speech / non-speech separator and the original input mix audio segments MIN instead of non-speech segments MNS from the speech / non-speech separator. Additionally, while the DRC processing stage 46 shows two DRC modification units 46a, 46b it is envisaged that only one DRC modification unit 46a, 46b is used if DRC processing of only one of the speech enhanced mix audio segments MSE and mix input audio segments MIN is to be performed.
[0184] When both DRC modification units 46a, 46b are used the parameter mapper extracts two DRC ratios R1 and R2 for each of the DRC modification units 46a, 46b respectively based on at least one SEM. The parameter mapper may similarly determine a make-up gain Mg1, Mg2 and / or threshold T1, T2 for each DRC modification units 46a, 46b based on at least one SEM.
[0185] The first and / or second DRC modification unit 46a, 46b obtains the input segments (MSE or M) and performs DRC processing governed by at least one of the DRC ratio R, make- up gain Mg and DRC threshold T. The DRC ratio R, make-up gain Mg1, Mg2 and / or threshold T are determined adaptively and based on the at least one SEM. The first and / or second DRC modification unit then outputs a respective DRC modified segment MDMSE, MDM which is provided to the mixer 45 which mixes the segments to form the modified mix output segments MM.
[0186] In some implementations, DRC processing stage 46 including mixing stage 45 is implemented in a receiving device, which receives a bitstream with at least one DRC parameter (e.g. a single DRC ratio R1 or two DRC ratios R1, R2) and audio content in the form of either a segment pair (e.g. an MSE / MIN pair as shown or an MS / MNS pair) or a mix signal such as MIN (which can be separated into an MSE / MIN pair or an MS / MNS pair). Hereby, the DRC processing based on the at least one DRC parameter is performed in the receiving device which is separate from the transmitting device which has extracted DRC parameter and optionally the segment pairs.
[0187] In the above, various examples of speech experience measures, SEMs, have been presented. Additional examples of SEMs include state of the art methods for determining speech experience measures (or dialogue experience measures as they are also called).
[0188] For example, the partial loudness of speech in the presence of noise (or non-speech content) can be determined in frequency bands and accumulated over equivalent rectangular bandwidths (ERBs) yielding a scalar measure of loudness in sones as described in Moore, B. C., Glasberg, B. R., & Baer, T. (1997), A Model for the Prediction of Thresholds, Loudness, and Partial Loudness, Journal of the Audio Engineering Society, 45(4), 224-240, hereby incorporated by reference in its entirety, as well as in Glasberg, B. R., & Moore, B. C. (2002), A Model of Loudness Applicable to Time-Varying Sounds. Journal of the Audio Engineering Society, 50(5), 331-342, hereby incorporated by reference in its entirety.
[0189] As another example, objective dialogue quality metrics may be used as a SEM. For example, the state-of-the-art objective dialogue quality metrics as proposed in ITU, 2001, Recommendation ITU-R BS.1387-1: Method for objective measurements of perceived audio quality, Rix, A., Beerends, J., M, H., & Hekstra, A., 2001, Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs, hereby incorporated by reference in its entirety, can be used.
[0190] As yet another example the dialogue quality metrics of Taal, C. H., Hendriks, R. C., Heusdens, R., & Jensen, J., 2010, a short-time objective intelligibility measure for time frequency weighted noisy speech can be used, or the dialogue quality metrics of Hines, A., Skoglund, J., Kokaram, A., & Harte, N., 2012 ViSQOL: The Virtual Speech Quality Objective Listener, can be used or the dialogue quality metrics of Reddy, C. K., Gopal, V., & Cutler, R, 2022, DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors can be used as a SEM. The above mentioned references are hereby incorporated by reference in their entirety.
[0191] Further examples of SEMs include SEMs that can be extracted from user data associated with the audio content (and optionally video content) of the mix input audio signal M.
[0192] Direct user feedback (e.g. events related to a users’ effort to improve speech intelligibility) may be used as a user data SEMs. As a first example, overall volume adjustment may be used as a user data SEM where volume increases in general, and in particular volume increases coinciding with speech content, and / or volume decreases in general, and in particular coinciding with non-speech content, result in a lower valued user data SEM. As a second example, enabling or disabling subtitles may result in a lower or higher valued user data SEM. Here in case of enabling subtitles the SEM is increased and in the case of disabling subtitles the SEM is decreased meaning that when subtitles are enabled it is assumed that less aggressivespeech experience improvement is needed. On the other hand, implementations are also envisaged where enabling subtitles may decrease the SEM and disabling subtitles increase the SEM based on the assumption that a user enabling subtitles means that the speech experience is worse than desired. As a third example, a user performing a rewind followed by (re)playing of the audio content may be associated with a lower valued used data SEM. The above three examples of user data SEMs are merely exemplary and may be used separately or combined together or with other SEMs.
[0193] Static user preferences may be used as one or more user preference SEMs. As a first example, an explicit user preference indicating a desired level of speech enhancement (high, medium, low or off) may be used as a user preference SEM wherein a user preference SEM indicating more speech enhancement is associated with lower SEM value which will bias the processing towards introducing more speech enhancement. As a second example, audiologist hearing measurements may be available, wherein one or more user preference SEMs are determined based on the audiologist hearing measurements. As a third example, biological information such as age and / or gender for the user may be available, wherein at least one user preference SEM is extracted from the biological information. Specifically, an older user may be assumed to prefer improved speech intelligibility and more speech enhancement processing wherein a higher age is mapped to a lower SEM which may result in boosted speech. As a fourth example, a user’s hearing characteristics may be derived from user preferences or calibration data such as thresholds of hearing, a personalized Head Related Transfer Function (PHRTF) or EQ settings which in turn may be used to extract at least one user preference SEM. As a fifth example, a user’s utterances related to their experience during playback may be mapped to a SEM using a chatbot and / or natural language processing, NLP, model.
[0194] Further examples of SEMs include metadata SEMs that can be extracted from video or scene metadata associated with the audio content (e.g. when the audio content is the soundtrack of a movie). As a first example, scenes featuring certain actresses or actors known to speak in a less intelligible fashion (e.g. murmuring or soft speaking style) may be associated with a lower SEM value. As a second example, metadata indicating the studio that has produced the content can be used to extract at least one SEM. Particularly, some studios may be known to mix speech at low levels whereby the content of such studios may be associated with a lower valued SEM. As a third example, metadata indicating the genre may be used to extract at least one SEM.
[0195] Further examples of SEMs include playback environment SEMs based on the information regarding the playback environment. As an example, the noise level in the playbackenvironment is measured wherein the playback environment SEM is assumed to be lower when there is more noise and higher when there is less noise.
[0196] All SEMs do not need to be computed at the same time or in the same location. For example, signal-derived SEMs may be computed at the head end of the signal chain (e.g., at content generation or encoding time) and be transmitted together with the signal. These signal- derived SEMs may then be combined with SEMs that are computed later and in a different location (e.g., at the time of playback on the playback device).
[0197] For example, one or several SEM values are included as metadata that is carried along the mix input audio signal. At the receiving end, the one or several SEMs in the metadata may be modified or combined with one or more SEMs available at the receiving end (e.g. user preferences, playback configuration and environmental noise) and then used to control the speech enhancement processing. This allows for use of a single bitstream to cater all speech enhancement needs while maintaining the flexibility to adapt the speech enhancement processing based on the particulars that are specific to the receiving end.
[0198] Each SEM may be represented with a discrete or continuous value. For example, the SNR of equations 2, 3, or 4 is one example of a SEM which can be determined given a speech segment and a non-speech segment. Another example of a SEM is a playback configuration SEM which indicates playback configuration. For example if the playback configuration is headphones the playback configuration SEM may assume a first discrete value, if the playback configuration is TV loudspeakers the playback configuration SEM may assume a second discrete value and if the playback configuration is mobile phone loudspeakers the playback configuration SEM may assume a third discrete value.
[0199] One or more SEMs (continuous and / or discrete) may also be combined to form combined SEMs. For example a combined SEM may be based on the SNR penalized with a playback configuration SEM. For example, the original SNR may be maintained if the playback configuration SEM indicates headphones, decreased if the playback configuration SEM indicates TV loudspeakers and further decreased if the playback configuration SEM indicates mobile phone loudspeakers.
[0200] In the below, four implementation examples will be described to further highlight how the audio processing system could operate at runtime.
[0201] In a first implementation example, the playback configuration SEM is used to control intrusive speech enhancement as shown in FIG.2A or FIG.5A. If the playback configuration is headphones the playback configuration SEM assumes a first discrete value, if the playback configuration SEM is TV loudspeakers the playback configuration SEM assumes a second discrete value and if the playback configuration is mobile phone loudspeakers theplayback configuration SEM assumes a third discrete value. If the playback configuration SEM is the first value this may be an indicator that no speech enhancement processing is needed and the DRC ratio R (in case of DRC modification) is set to 1 and / or the gain level ^ (in case of gain modification) is set to 0 dB. If the playback configuration SEM is the second value this may be an indicator that comparatively more speech enhancement processing is needed wherein e.g. the gain level is set to 9 dB and / or the DRC ratio R is kept at 1. If the playback configuration SEM is the third value this may be an indicator that comparatively even more speech enhancement processing is needed wherein e.g. the gain level is increased to 12 dB and / or the DRC ratio R is set to 10.
[0202] In a second implementation example, a continuous SEM (e.g. the SNR) is used together with the discrete playback configuration SEM. The continuous SNR represented as a parameter X which in turn is mapped to a gain level G e.g. in accordance with the graph of FIG. 3, wherein the maximum gain level is ^456. As the parameter X varies from segment to segment, so does the gain level G and / or DRC ratio R for modifying the speech and non-speech segments according to the graph of FIG.3. However, based on the playback configuration SEM the maximum gain level ^456is adjusted. For example, if the playback configuration SEM is the first value, ^456may be 0 dB, if the playback configuration SEM is the second value, ^456may be 9 dB and if the playback configuration SEM is the third value, ^456may be 15 dB. Additionally or alternatively, the maximum DRC ratio R is also modified based on the playback configuration SEM wherein the maximum DRC ratio R is 1 for the first and second value of the playback configuration SEM and 10 for the third value of the playback configuration SEM.
[0203] In a third implementation example, two continuous SEMs are combined, SEM A and SEM B. The combined SEM may e.g. be determined as a weighted sum, the minimum or the maximum of SEM A and SEM B. For example SEM A is an SNR determined for each segment as the ratio between the energy of the speech segment and the energy of the mix input signal segment and SEM B is an SNR determined as the ratio between the energy of a separated speech segment and the energy of measured environmental noise at the playback side.
[0204] In a fourth implementation example the two continuous SEMs, SEM A and SEM B are mapped individually to a same processing parameter (e.g. the gain level G). That is, SEM Ais mapped to the gain level using a function ^5 giving a first instance ^5 of the gain level as ^5=^5(SEM A) and SEM B is mapped to the gain level using a function ^^ giving a second instance^^ of the gain level as ^^=^^ (SEM B). A final value of the gain level ^< may then bedetermined using a function ^<that determines ^<as function of ^5and ^^as ^<= ^<(^5, ^^). The function ^<may e.g. be the average, minimum or maximum.
[0205] Additionally, it is noted that prior to providing an audio segment to a SEM generator for extraction of at least one SEM, the audio segment may be pre-processed. The pre- processing may involve applying a head related transfer function (HRTF) and / or a gain. The HRTF may be configured to modify the segment so as to more accurately represent the audio content as it will be experienced by the listener. For a channel based presentation, the spatial relationship between the canonical listening position and the canonical loudspeaker position associated with an audio segment may be used to form an appropriate HRTF for application during pre-processing.
[0206] FIG.10 shows a block diagram illustrating a system for processing a multi-element audio presentation comprising a plurality of audio channels and / or audio objects. In general, a multi-element audio presentation may be described as comprising an arbitrary number K of mix input audio signals. Each of the K mix input audio signals may be associated with an audio channel or an audio object. The system of FIG.10 may be used to process all audio signals of an audio presentation or a subset of the audio signals.
[0207] For example, the system may be used to process all audio signals of a 5.1 presentation or only a subset of the audio signals, such as the audio signals associated with the LRC audio channels. Since the surround channels typically comprise background or non-speech content it may be assumed that these channels do not affect the speech intelligibility whereby these channels may be left out of processing in some implementations.
[0208] Alternatively, two or more instances of the system of FIG.10 may be used in parallel, e.g. one system is used to process the LRC audio channels, and another system is used to process the surround channels of a 5.1 presentation.
[0209] The system of FIG.10 may generally be divided into an analysis part comprising the optional pre-processing modules 391-1, … 391-K, 392-1, … 392-K and a SEM generator 11, and a processing part comprising the gain mappers 21-1, …, 21-K and gain application and mixing modules 48-1, … 48-K.
[0210] Each mix input audio signal is first processed with a speech separator or speech / non-speech separator to yield, for each mix input audio signal, two audio signals, a speech audio signal and an auxiliary audio signal. The auxiliary audio signal may be a non- speech audio signal or a mix input audio signal as described above (see e.g. FIGS.2A and 2C).
[0211] Each pair of a speech audio signal and an auxiliary audio signal is provided to a SEM generator 11 which generates a SEM based on the received audio signals. Many alternatives for determining a SEM have been described above and may be applied by the SEM generator 11. For example, the SEM generator 11 may determine the same based on one or more of the K mix input audio signals (not shown)and / or based on the K pairs of a speech audio signaland auxiliary audio signal extracted from the K mix input audio signals (as illustrated in FIG. 10).. For example, the SEM extractor 11 determines K SNR values, one for each speech / auxiliary audio signal pair and forms a single SEM as a weighted combination of the K number of SNR values.
[0212] The SEM generated by the SEM generator 11 is provided to each of K gain mappers 21-1, 21-2, … 21-K each configured to map the SEM to a pair of gains for application to the respective speech and auxiliary audio signal in the subsequent gain application and mixing module 48-1, 48-2, 48-K. For example, each gain mapper 21-1, 21-2, … 21-K may be configured to utilize two mapping functions, a first mapping function for mapping the SEM to the first gain and a second mapping function for mapping the SEM to the second gain. As another example, each gain mapper may be configured to determine a gain level and splitting ratio, wherein the two gains are determined based on the gain level and the splitting ratio as described in connection with FIG.2A.
[0213] Each gain mapper 21-1, 21-2, …, 21-K outputs two gains, a first gain referred to as the speech gain ^^,^, ^^,7, ^^,^that is applied to the segments of the speech audio signal and a second gain ^^^,^, ^^^,7, ^^^,^referred to as a non-speech gain that is applied to the segments of the auxiliary audio signal (i.e. the mix input audio signal or a non-speech audio signal extracted therefrom).
[0214] Each gain application and mixing module 48-1, 48-2, 48-K may correspond to module 44 of FIG.8. It is also envisaged that in addition to, or as an alternative to, modifying the gain of the audio signals the processing part of the system in FIG.10 may be configured to perform DRC adjustment as described in connection with FIG.5A. For example, each gain mapper 21-1, 21-2, … 21-K is replaced with a DRC parameter mapper configured to map the SEM to at least one DRC parameter which is applied to the audio signals by a DRC processing stage replacing each gain application and mixing module 48-1, 48-2, 48-K. The DRC processing stage is described e.g. in connection with FIG.9.
[0215] Notably, the gains determined by one gain mapper may in general be different from the gain gains determined by another gain mapper depending e.g. on the details of the gain mapping function(s) utilized by each gain mapper 21-1, 21-2, … 21-K. Hereby, the processing part of the system of FIG.10 is configured to process each of the K mix input audio signals (or the speech auxiliary components extracted therefrom) individually using techniques described in connection with FIGS.8-9 wherein the gain or DRC parameters applied to each mix input signal are controlled using the same SEM.
[0216] In some implementations, one or more audio signals are subject to pre-processing using pre-processing modules 391-1, 391-1, …, 391-K, 392-1, 392-1, …, 392-K whereby it isthe pre-processed audio signals which are input to the SEM generator 11 whereby the SEM will ultimately be based on the pre-processed version of the audio signals.
[0217] The pre-processing performed by each pre-processing module 391-1, 391-1, …, 391-K, 392-1, 392-1, …, 392-K may involve applying a gain to the respective audio signal. The gain is used to control to what extent the audio signal will influence the SEM generation and audio signals from audio objects or channels deemed more important, or more likely to contain speech, may be provided with a higher gain whereas remaining audio signals are provided with a lower gain. For example, in a 5.1 presentation the audio signals associated with the front left, center and front right channel may be provided with a higher gain and the surround channels provided with a lower gain since speech is most likely present in the LRC channel triplet. Alternatively, in some applications it may be beneficial in a 5.1 presentation to apply higher surround channel gains and lower front channel gains to reflect the fact that the same sound played from the surround loudspeakers appears louder than when played from the front loudspeakers.
[0218] Additionally or alternatively, the pre-processing performed by each pre-processing module 391-1, 391-1, …, 391-K, 392-1, 392-1, …, 392-K may involve applying a HRTF to each audio signal to provide audio signals more closely resembling the audio content actually perceived by the listener to the SEM generator 11. For example, in a channel based presentation a HRTF between the canonical loudspeaker position and a listener positioned at the intended listening position of the presentation may be used to form an audio signal more accurately representing what the listener will actually hear.
[0219] While the above examples relate to channel-based presentations it is understood that the same may be applied for audio objects. For example, audio objects associated with object metadata indicating speech, or indicating a higher importance, may be provided with a higher gain during pre-processing. As another example, position metadata of the audio object and position data associated with the listener may be used to determine appropriate HRTFs for application during pre-processing.
[0220] In one example implementation, the channels of a multi-channel presentation are separated into a first subset and a second subset. The audio signals of the channels of the first subset are processed with the analysis part and processing part of the system of FIG.10. The audio signals of the second subset are however only subject to processing with the processing part, wherein the gain applied to the background audio signals of the second subset is based on the gain applied to the background audio signal of the first subset. For example, the first subset comprises one or more of the front left, center and front right channel and the second subsetcomprises one or more surround or height channels of a 5.0, 5.1, 5.1.2, 5.1.4, 7.1.7.1.2, or 7.1.4 representation.
[0221] Hereby, the channels of the first subset which are likely to contain speech are used for extracting the SEM controlling the speech enhancement processing. However, if the channels of the second subset are used as is, without modification, the non-speech background audio will suffer from spatial distortion since it is suppressed in the front channels but left unaffected in the surround channels. Hereby, by determining a non-speech gain for application to the second subset, based on the non-speech gain for the first subset, it is possible to mitigate the spatial distortion. The non-speech gain applied to the second subset may be equal to one of, or an average of, the non-speech gains applied to the first subset.
[0222] FIG.11 shows a schematic block diagram of an example electronic device or architecture 200 (e.g., an apparatus 200) suitable for implementing example embodiments of the present disclosure. Architecture 200 includes but is not limited to servers and client devices, systems, modules and methods as described in reference to FIGs.1A and 1B. As shown, the architecture 200 includes central processing unit (CPU) 201 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 202 or a program loaded from, for example, storage unit 208 to random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201, which may include one or more processor cores, and in some examples the processor 201 may be multiple processors. In RAM 203, the data used when CPU 201 performs the various processes is also stored, as required. CPU 201, ROM 202 and RAM 203 are connected to one another via bus 204. Input / output (I / O) interface 205 is also connected to bus 204.
[0223] The following components are connected to I / O interface 205: input unit 206, that may include a keyboard, a mouse, or the like; output unit 207 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 208 including a hard disk, or another suitable storage device; and communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).
[0224] In some implementations, input unit 206 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0225] In some implementations, output unit 207 include systems with various number of speakers. Output unit 207 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0226] In some embodiments, communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I / O interface 205, as required.Removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 210, so that a computer program read therefrom is installed into storage unit 208, as required. A person skilled in the art would understand that although apparatus 200 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0227] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 209, and / or installed from the removable medium 211, as shown in FIG.11.
[0228] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 201 in combination with other components of FIG.11), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, a processor and / or other computing device(s), which may include control circuitry. While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0229] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0230] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instructionexecution system, apparatus, or device. The machine-readable medium may be a machine- readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0231] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to one or more processors of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by one or more processors of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0232] Various aspects may be appreciated from the following enumerated example embodiments (EEEs): EEE 1. An audio processing method comprising: obtaining a first and second sequence of time-aligned audio segments, wherein the first sequence of audio segments comprises speech audio content and the second sequence of audio segments comprises non-speech audio content; for each time-aligned pair of segments in the first and second sequence: - obtaining at least one speech experience measure; - determining, based on the at least one speech experience measure, a first gain and a second gain, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain. EEE 2. The method according to EEE 1, wherein determining the first and second gain comprises:- determining a gain level for improving the speech experience of the pair of segments based on the at least one speech experience measure; - obtaining a split ratio; - partitioning the gain level into the first gain and the second gain, based on the split ratio. EEE 3. The method according to EEE 1, wherein determining the first and second gain comprises: - determining the first gain using a first mapping function, mapping the at least one speech experience measure to a first gain; and - determining the second gain using a second mapping function, mapping the at least one speech experience measure to a second gain. EEE 4. The method according to any of EEEs 1 - 3, further comprising: for each time-aligned pair of segments in the first and second sequence: - applying the first gain to the segment of the first sequence to generate a modified first segment; - applying the second gain to the segment of the second sequence to generate a modified second segment, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment. EEE 5. The method according to any of EEEs 1 - 3, further comprising: for each time-aligned pair of segments in the first and second sequence: - determining a third gain as the difference between the first gain and the second gain; - applying the third gain to the segment of the first sequence to generate a modified first segment; - applying the second gain to the segment of the second sequence to generate a modified second segment, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment, wherein the first sequence of audio segments comprises speech audio content from a speech separation and / or speech enhancement processing of a mix signal and the second sequence of audio segments comprises segments from the mix signal. EEE 6. The method according to any of EEEs 1 - 3, further comprising:encoding the first gain and second gain into a bitstream, transmitting the bitstream to a receiving device. EEE 7. The method according to EEE 6, further comprising: encoding the first and second sequence of time aligned segments into the bitstream. EEE 8. The method according to EEE 6, further comprising: obtaining a sequence of segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment; and encoding a mix segment associated with the first and second gain into the bitstream. EEE 9. The method according to any of EEEs 6 – 8, further comprising: receiving, by the receiving device, the bitstream; processing, by the receiving device, the bitstream to obtain the first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, the first and second gain; for each time-aligned pair of segments in the first and second sequence: applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment. EEE 10. The method according to any of the preceding EEEs, wherein the first gain is the boosting gain and the second gain is the attenuating gain. EEE 11. The method according to any of EEEs 1 - 9, wherein the first gain is the attenuating gain and the second gain is the boosting gain.EEE 12. The method according to any of the preceding EEEs, when depending on EEE 2, wherein obtaining the split ratio comprises: determining the split ratio for the segment pair based on the speech experience measure for the segment pair. EEE 13. The method according to EEE 12, wherein the split ratio indicates a ratio of the gain level that is to be assigned to the second gain, and wherein the speech experience measure assumes one of at least two values, a high value and a low value, wherein the split ratio for the high value of the speech experience measure is larger than the split ratio for the low value of the speech experience measure. EEE 14. The method according to EEE 12, wherein the split ratio indicates a ratio of the gain level that is to be assigned to the second gain, and wherein the speech experience measure assumes one of at least two values, a high value and a low value, wherein the split ratio for the low value of the speech experience measure is larger than the split ratio for the high value of the speech experience measure. EEE 15. The method according to any of the preceding EEEs, wherein the split ratio indicates a ratio of the gain level that is to be assigned to the second gain, the method further comprises: determining, for each segment in the second sequence of audio segments a stationarity measure, the stationarity measure indicating to what extent the content of the audio segment in the second sequence is stationary; determining the split ratio for each segment pair based on the stationarity measure, wherein the split ratio for segment pairs comprising a second sequence segment with stationary audio content is lower compared to the split ratio for segment pairs comprising a second sequence segment with non-stationary audio content. EEE 16. The method according to any of the preceding EEEs, wherein the speech experience measure assumes a value on the range [T2, T3] with T2 indicating lower speech experience compared to T3, and wherein the gain level is determined by evaluating a mapping function mapping the range [T2, T3] to a gain level in the range [GHQ, Gmax], wherein the mapping function maps the gain level to GHQ when the at least one speech experience measure is T3 and maps the gain level to Gmax when the speech experience measure is T2, wherein T2 < T3.EEE 17. The method according to EEE 16, wherein the mapping function is a monotonically decreasing function for the gain level on the range [T2, T3]. EEE 18. The method according to EEE 16 or EEE 17 wherein the speech experience measure assumes a value on the range [T1, T3], wherein T3 > T2 > T1 and wherein the mapping function maps the gain level to Gfade at a speech experience measure of T1, wherein Gmax > Gfade and / or Gmax > GHQ. EEE 19. The method according to EEE 18, wherein the mapping function is a monotonically increasing function for the gain level on the range [T1, T2]. EEE 20. The method according to any of the preceding EEEs, wherein each audio segment comprises a plurality of frequency bands and wherein the gain level comprises a plurality of subband gains, each associated with a frequency band of said plurality of frequency bands, and wherein partitioning the gain level into first gain and a second gain, based on the split ratio, comprises partitioning each subband gain into a first subband gain and a second subband gain. EEE 21. The method according to EEE 20, wherein the subband gains form an equalizer, EQ, filter. EEE 22. The method according to EEE 21, wherein the method further comprises modifying the EQ filter based on the at least one speech experience measure to form a modified EQ filter. EEE 23. The method according to EEE 22, wherein modifying the EQ filter comprises scaling the plurality of subband gains with a scaling factor, wherein the scaling factor is based on the at least one speech experience measure. EEE 24. The method according to EEE 22 or EEE 23, wherein the EQ filter is a parametric high- pass shelving filter and wherein modifying the EQ filter comprises adjusting the subband gains to adjust a shelf gain of the EQ filter, based on the at least one speech experience measure.EEE 25. The method according to any of the preceding EEEs, wherein the step of applying the first gain to the segment of the first sequence and applying the second gain to the segment of the second sequence is performed in a time domain or in a frequency domain. EEE 26. The method according to any of the preceding EEEs, wherein a speech to non-speech ratio between a speech energy measure and a non-speech energy measure is higher for the first sequence of segments compared to the second sequence of segments. EEE 27. The method according to any of the preceding EEEs, wherein the first sequence of audio segments comprises speech audio content from a speech separation and / or speech enhancement processing of a mix signal and the second sequence of audio segments comprises non-speech audio content from the mix signal, or wherein the first sequence of audio segments comprises speech audio content from a speech stem of a multi-stem mix signal and the second sequence of audio segments comprises non-speech audio content from a non-speech stem of the multi-stem mix signal. EEE 28. The method according to any of the preceding EEEs, wherein the first sequence of audio segments comprises speech audio content from a speech separation processing of a mix signal and the second sequence of audio segments comprises non-speech audio content from the mix signal, the method further comprising: obtaining a sequence of segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment. EEE 29. The method according to any of the preceding EEEs, wherein the first sequence of audio segments comprises speech audio content from a speech enhancement processing of a mix signal and the second sequence of audio segments comprises non-speech audio content from the mix signal, the method further comprising: obtaining a sequence of mix segments of the mix signal, each mix segment comprising speech audio content mixed with non-speech audio content; processing each segment of the mix signal with a speech enhancement system to form speech enhanced mix segments, the speech enhancement system being configured to extract a speech enhanced mix segment with increased speech experience given an input mix segment;wherein the speech enhanced mix segments constitute the first sequence of segments, and wherein the sequence of mix segments constitutes the second sequence of segments. EEE 30. The method according to any of EEEs 1 - 26, wherein the first sequence of audio segments comprises speech audio content from a speech stem of a multi-stem mix signal and the second sequence of audio segments comprises non-speech audio content from a non-speech stem of the multi-stem mix signal. EEE 31. The method according to any of the preceding EEEs, wherein the at least one speech experience measure comprises at least one of: a speech-to-non-speech ratio associated with the time-aligned pair of segments, a speech-to-non-speech ratio associated with a pre-processed version of the time-aligned pair of segments, playback configuration, direct user feedback, static user preferences, playback environment data, and metadata associated with the first and / or second sequence of audio segments. EEE 32. The method according to any of the preceding EEEs, wherein the first and second sequence of time-aligned audio segments are extracted from a first audio channel or object, the method further comprising: obtaining a second audio channel or object; extracting from the second audio channel or object a third and fourth sequence of time-aligned audio segments, wherein the third sequence of audio segments comprises speech audio content and the fourth sequence of audio segments comprises non-speech audio content; determining, based on the at least one speech experience measure, a third gain and a fourth gain, wherein one of the third gain and fourth gain is an attenuating gain and the other one of the third gain and fourth gain is a boosting gain; applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment; outputting a segment of mixed audio content comprising a mix of the modified first and second segment.applying the third gain to the segment of the third sequence to generate a modified third segment; applying the fourth gain to the segment of the fourth sequence to generate a modified fourth segment; and outputting a segment of mixed audio content comprising a mix of the modified third and fourth segment. EEE 33. The method according to EEE 32, wherein the at least one speech experience measure is determined based on the first audio channel or object and the second audio channel or object. EEE 34. The method according to any of the preceding EEEs, wherein the first and second sequence of time-aligned audio segments are extracted from a first audio channel or object, the method further comprising: obtaining a third audio channel or object; extracting from the third audio channel or object a fifth and sixth sequence of time- aligned audio segments, wherein the fifth sequence of audio segments comprises speech audio content and the sixth sequence of audio segments comprises non-speech audio content; applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified sixth segment; outputting a segment of mixed audio content comprising a mix of the modified first and second segment; determining a second attenuating gain based on the attenuating gain of the first gain or second gain; applying the second attenuating gain to a segment of the sixth sequence to generate a modified fourth segment; outputting a segment of mixed audio content comprising a mix of the fifth segment and the modified sixth segment. EEE 35.The method according to EEE 34, wherein the first channel or object is an audio channel selected from a group consisting of a front left channel of an audio presentation, a front right channel of an audio presentation and a center channel of an audio presentation, and wherein the second channel or object is a surround channel of an audio presentation.EEE 36. An audio processing method comprising: receiving, by a receiving device, a bitstream; processing, by the receiving device, the bitstream to obtain a first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, a first and second gain; for each time-aligned pair of segments in the first and second sequence: applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment. EEE 37. The method according to EEE 36, wherein the bitstream comprises encoded segments of a mix signal and wherein processing the bitstream comprises: decoding the bitstream to obtain the segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment. EEE 38. An apparatus comprising a processor and a memory, configured to perform the method according to any of the preceding EEEs. EEE 39. A computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to any of EEEs 1-37. EEE 40. A computer-readable storage medium storing the computer program product according to EEE 38. EEE 41. An audio processing method comprising:obtaining a first and second sequence of time-aligned input audio segments, wherein the first sequence of input audio segments comprises speech audio content and the second sequence of input audio segments comprises non-speech audio content; for each time-aligned pair of input audio segments in the first and second sequence: obtaining at least one speech experience measure; determining at least one dynamic range compression, DRC, parameter based on the at least one speech experience measure. EEE 42. The method according to EEE 41, further comprising: for each time-aligned pair of input audio segments in the first and second sequence: performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence; and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence. EEE 43. The method according to EEE 42, further comprising: encoding the at least one DRC parameter into a bitstream, transmitting the bitstream to a receiving device. EEE 44. The method according to EEE 43, further comprising: encoding the first and second sequence of time aligned segments into the bitstream. EEE 45. The method according to EEE 43, further comprising: obtaining a sequence of segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment; and encoding a mix segment associated with the first and second gain into the bitstream. EEE 46. The method according to any of EEEs 43 – 45, further comprising:receiving, by the receiving device, the bitstream; processing, by the receiving device, the bitstream to obtain the first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, the at least one DRC parameter; for each time-aligned pair of segments in the first and second sequence: performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence; and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence. EEE 47. The method according to any of EEEs 35 - 39, wherein the at least one DRC parameter is at least one of: a DRC ratio, a DRC threshold, a make-up gain, a DRC attack time constant, and a DRC release time constant. EEE 48. The method according to any of EEEs 42 - 47, wherein performing dynamic range compression based on the DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence comprises: performing dynamic range compression of the segment of the first sequence based on the DRC parameter to form a first DRC segment, and outputting the segment of mixed audio content comprising a mix of the first DRC segment and the segment of the second sequence. EEE 49. The method according to any of EEEs 42 - 47, wherein performing dynamic range compression based on the DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence comprises: performing dynamic range compression of the segment of the second sequence based on the DRC parameter to form a second DRC segment, and outputting the segment of mixed audio content comprising a mix of the second DRC segment and the segment of the first sequence.EEE 50. The method according to any of EEEs 42 - 47, wherein performing dynamic range compression based on the DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence comprises: performing dynamic range compression of the segment of the first sequence to form a first DRC segment and performing dynamic range compression of the second sequence to form a second DRC segment, and outputting the segment of mixed audio content comprising a mix of the first DRC segment and the second DRC segment. EEE 51. The method according to EEE 50, wherein determining a dynamic range compression, DRC, parameter based on the at least one speech experience measure comprises: determining two DRC compression parameters based on at least one speech experience measure, a first DRC parameter and a second DRC parameter, wherein the first DRC segment is formed based on the first DRC parameter and the second DRC segment is formed based on the second DRC parameter. EEE 52. The method according to any of EEEs 42 - 51, wherein the at least one speech experience measure comprises at least one of: a speech-to-non-speech ratio associated with the time-aligned pair of segments, a speech-to-non-speech ratio associated with a preprocessed version of the time- aligned pair of segments; direct user feedback, static user preferences, playback environment data, and metadata associated with the first and / or second sequence of audio segments. EEE 53. An audio processing method comprising: receiving, by a receiving device, a bitstream; processing, by the receiving device, the bitstream to obtain a first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, a at least one dynamic range compression, DRC, parameter; for each time-aligned pair of segments in the first and second sequence:performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence; and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence. EEE 54. The method according to EEE 53, wherein the bitstream comprises encoded segments of a mix signal and wherein processing the bitstream comprises: decoding the bitstream to obtain the segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment. EEE 55. An apparatus comprising a processor and a memory, configured to perform the method according to any of EEEs 41-54. EEE 56. A computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to any of EEEs 51-54. EEE 57.A computer-readable storage medium storing the computer program according to EEE 56.
Claims
1. CLAIMS 1. An audio processing method comprising: obtaining a first and second sequence of time-aligned audio segments, wherein the first sequence of audio segments comprises speech audio content and the second sequence of audio segments comprises non-speech audio content; for each time-aligned pair of segments in the first and second sequence: - obtaining at least one speech experience measure; - determining, based on the at least one speech experience measure, a first gain and a second gain, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain.
2. The method according to claim 1, wherein determining the first and second gain comprises: - determining a gain level for improving the speech experience of the pair of segments based on the at least one speech experience measure; - obtaining a split ratio; - partitioning the gain level into the first gain and the second gain, based on the split ratio.
3. The method according to claim 1, wherein determining the first and second gain comprises: - determining the first gain using a first mapping function, mapping the at least one speech experience measure to a first gain; and - determining the second gain using a second mapping function, mapping the at least one speech experience measure to a second gain.
4. The method according to any of claims 1 - 3, further comprising: for each time-aligned pair of segments in the first and second sequence: - applying the first gain to the segment of the first sequence to generate a modified first segment; - applying the second gain to the segment of the second sequence to generate a modified second segment, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment.
5. The method according to any of claims 1 - 3, further comprising: for each time-aligned pair of segments in the first and second sequence: - determining a third gain as the difference between the first gain and the second gain; - applying the third gain to the segment of the first sequence to generate a modified first segment; - applying the second gain to the segment of the second sequence to generate a modified second segment, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment, wherein the first sequence of audio segments comprises speech audio content from a speech separation and / or speech enhancement processing of a mix signal and the second sequence of audio segments comprises segments from the mix signal.
6. The method according to any of claims 1 - 3, further comprising: encoding the first gain and second gain into a bitstream, transmitting the bitstream to a receiving device.
7. The method according to claim 6, further comprising: encoding the first and second sequence of time aligned segments into the bitstream.
8. The method according to claim 6, further comprising: obtaining a sequence of segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment; and encoding a mix segment associated with the first and second gain into the bitstream.
9. The method according to any of claims 6 – 8, further comprising: receiving, by the receiving device, the bitstream;processing, by the receiving device, the bitstream to obtain the first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, the first and second gain; for each time-aligned pair of segments in the first and second sequence: applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment.
10. The method according to any of the preceding claims, wherein the first gain is the boosting gain and the second gain is the attenuating gain.
11. The method according to any of claims 1 - 9, wherein the first gain is the attenuating gain and the second gain is the boosting gain.
12. The method according to any of the preceding claims, when depending on claim 2, wherein obtaining the split ratio comprises: determining the split ratio for the segment pair based on the speech experience measure for the segment pair.
13. The method according to claim 12, wherein the split ratio indicates a ratio of the gain level that is to be assigned to the second gain, and wherein the speech experience measure assumes one of at least two values, a high value and a low value, wherein the split ratio for the high value of the speech experience measure is larger than the split ratio for the low value of the speech experience measure.
14. The method according to claim 12, wherein the split ratio indicates a ratio of the gain level that is to be assigned to the second gain, and wherein the speech experience measure assumes one of at least two values, a high value and a low value, wherein the split ratio for the low value of the speech experience measure is larger than the split ratio for the high value of the speech experience measure.
15. The method according to any of the preceding claims, wherein the split ratio indicates a ratio of the gain level that is to be assigned to the second gain, the method further comprises: determining, for each segment in the second sequence of audio segments a stationarity measure, the stationarity measure indicating to what extent the content of the audio segment in the second sequence is stationary; determining the split ratio for each segment pair based on the stationarity measure, wherein the split ratio for segment pairs comprising a second sequence segment with stationary audio content is lower compared to the split ratio for segment pairs comprising a second sequence segment with non-stationary audio content.
16. The method according to any of the preceding claims, wherein the speech experience measure assumes a value on the range [T2, T3] with T2indicating lower speech experience compared to T3, and wherein the gain level is determined by evaluating a mapping function mapping the range [T2, T3] to a gain level in the range [GHQ, Gmax], wherein the mapping function maps the gain level to GHQ when the at least one speech experience measure is T3and maps the gain level to Gmaxwhen the speech experience measure is T2, wherein T2 < T3.
17. The method according to claim 16, wherein the mapping function is a monotonically decreasing function for the gain level on the range [T2, T3].
18. The method according to claim 16 or claim 17 wherein the speech experience measure assumes a value on the range [T1, T3], wherein T3 > T2 > T1 and wherein the mapping function maps the gain level to Gfade at a speech experience measure of T1, wherein Gmax > Gfade and / or Gmax > GHQ.
19. The method according to claim 18, wherein the mapping function is a monotonically increasing function for the gain level on the range [T1, T2].
20. The method according to any of the preceding claims, wherein each audio segment comprises a plurality of frequency bands and wherein the gain level comprises a plurality of subband gains, each associated with a frequency band of said plurality of frequency bands, and wherein partitioning the gain level into first gain and a second gain, based on the split ratio, comprises partitioning each subband gain into a first subband gain and a second subband gain.
21. The method according to claim 20, wherein the subband gains form an equalizer, EQ, filter.
22. The method according to claim 21, wherein the method further comprises modifying the EQ filter based on the at least one speech experience measure to form a modified EQ filter.
23. The method according to claim 22, wherein modifying the EQ filter comprises scaling the plurality of subband gains with a scaling factor, wherein the scaling factor is based on the at least one speech experience measure.
24. The method according to claim 22 or claim 23, wherein the EQ filter is a parametric high-pass shelving filter and wherein modifying the EQ filter comprises adjusting the subband gains to adjust a shelf gain of the EQ filter, based on the at least one speech experience measure.
25. The method according to any of the preceding claims, wherein the step of applying the first gain to the segment of the first sequence and applying the second gain to the segment of the second sequence is performed in a time domain or in a frequency domain.
26. The method according to any of the preceding claims, wherein a speech to non-speech ratio between a speech energy measure and a non-speech energy measure is higher for the first sequence of segments compared to the second sequence of segments.
27. The method according to any of the preceding claims, wherein the first sequence of audio segments comprises speech audio content from a speech separation and / or speech enhancement processing of a mix signal and the second sequence of audio segments comprises non-speech audio content from the mix signal, or wherein the first sequence of audio segments comprises speech audio content from a speech stem of a multi-stem mix signal and the second sequence of audio segments comprises non-speech audio content from a non-speech stem of the multi-stem mix signal.
28. The method according to any of the preceding claims,wherein the first sequence of audio segments comprises speech audio content from a speech separation processing of a mix signal and the second sequence of audio segments comprises non-speech audio content from the mix signal, the method further comprising: obtaining a sequence of segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment.
29. The method according to any of the preceding claims, wherein the first sequence of audio segments comprises speech audio content from a speech enhancement processing of a mix signal and the second sequence of audio segments comprises non-speech audio content from the mix signal, the method further comprising: obtaining a sequence of mix segments of the mix signal, each mix segment comprising speech audio content mixed with non-speech audio content; processing each segment of the mix signal with a speech enhancement system to form speech enhanced mix segments, the speech enhancement system being configured to extract a speech enhanced mix segment with increased speech experience given an input mix segment; wherein the speech enhanced mix segments constitute the first sequence of segments, and wherein the sequence of mix segments constitutes the second sequence of segments.
30. The method according to any of claims 1 - 26, wherein the first sequence of audio segments comprises speech audio content from a speech stem of a multi-stem mix signal and the second sequence of audio segments comprises non-speech audio content from a non-speech stem of the multi-stem mix signal.
31. The method according to any of the preceding claims, wherein the at least one speech experience measure comprises at least one of: a speech-to-non-speech ratio associated with the time-aligned pair of segments, a speech-to-non-speech ratio associated with a pre-processed version of the time-aligned pair of segments, playback configuration, direct user feedback, static user preferences,playback environment data, and metadata associated with the first and / or second sequence of audio segments.
32. The method according to any of the preceding claims, wherein the first and second sequence of time-aligned audio segments are extracted from a first audio channel or object, the method further comprising: obtaining a second audio channel or object; extracting from the second audio channel or object a third and fourth sequence of time-aligned audio segments, wherein the third sequence of audio segments comprises speech audio content and the fourth sequence of audio segments comprises non-speech audio content; determining, based on the at least one speech experience measure, a third gain and a fourth gain, wherein one of the third gain and fourth gain is an attenuating gain and the other one of the third gain and fourth gain is a boosting gain; applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment; outputting a segment of mixed audio content comprising a mix of the modified first and second segment. applying the third gain to the segment of the third sequence to generate a modified third segment; applying the fourth gain to the segment of the fourth sequence to generate a modified fourth segment; and outputting a segment of mixed audio content comprising a mix of the modified third and fourth segment.
33. The method according to claim 32, wherein the at least one speech experience measure is determined based on the first audio channel or object and the second audio channel or object.
34. The method according to any of the preceding claims, wherein the first and second sequence of time-aligned audio segments are extracted from a first audio channel or object, the method further comprising: obtaining a third audio channel or object; extracting from the third audio channel or object a fifth and sixth sequence of time- aligned audio segments, wherein the fifth sequence of audio segments comprises speech audio content and the sixth sequence of audio segments comprises non-speech audio content;applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment; outputting a segment of mixed audio content comprising a mix of the modified first and second segment; determining a second attenuating gain based on the attenuating gain of the first gain or second gain; applying the second attenuating gain to a segment of the sixth sequence to generate a modified sixth segment; outputting a segment of mixed audio content comprising a mix of the fifth segment and the modified sixth segment.
35. The method according to claim 34, wherein the first channel or object is an audio channel selected from a group consisting of a front left channel of an audio presentation, a front right channel of an audio presentation and a center channel of an audio presentation, and wherein the second channel or object is a surround channel of an audio presentation.
36. An audio processing method comprising: receiving, by a receiving device, a bitstream; processing, by the receiving device, the bitstream to obtain a first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, a first and second gain; for each time-aligned pair of segments in the first and second sequence: applying the first gain to the segment of the first sequence to generate a modified first segment; applying the second gain to the segment of the second sequence to generate a modified second segment, wherein one of the first gain and second gain is an attenuating gain and the other one of the first gain and second gain is a boosting gain, and outputting a segment of mixed audio content comprising a mix of the modified first and second segment.
37. The method according to claim 36, wherein the bitstream comprises encoded segments of a mix signal and wherein processing the bitstream comprises: decoding the bitstream to obtain the segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment.
38. An apparatus comprising a processor and a memory, configured to perform the method according to any of the preceding claims.
39. A computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to any of claims 1-37.
40. A computer-readable storage medium storing the computer program product according to claim 38.
41. An audio processing method comprising: obtaining a first and second sequence of time-aligned input audio segments, wherein the first sequence of input audio segments comprises speech audio content and the second sequence of input audio segments comprises non-speech audio content; for each time-aligned pair of input audio segments in the first and second sequence: obtaining at least one speech experience measure; determining at least one dynamic range compression, DRC, parameter based on the at least one speech experience measure.
42. The method according to claim 41, further comprising: for each time-aligned pair of input audio segments in the first and second sequence: performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence; andoutputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence.
43. The method according to claim 42, further comprising: encoding the at least one DRC parameter into a bitstream, transmitting the bitstream to a receiving device.
44. The method according to claim 43, further comprising: encoding the first and second sequence of time aligned segments into the bitstream.
45. The method according to claim 43, further comprising: obtaining a sequence of segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment; and encoding a mix segment associated with the first and second gain into the bitstream.
46. The method according to any of claims 43 – 45, further comprising: receiving, by the receiving device, the bitstream; processing, by the receiving device, the bitstream to obtain the first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, the at least one DRC parameter; for each time-aligned pair of segments in the first and second sequence: performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence; and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence.
47. The method according to any of claims 35 - 39, wherein the at least one DRC parameter is at least one of: a DRC ratio, a DRC threshold, a make-up gain, a DRC attack time constant, and a DRC release time constant.
48. The method according to any of claims 42 - 47, wherein performing dynamic range compression based on the DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence comprises: performing dynamic range compression of the segment of the first sequence based on the DRC parameter to form a first DRC segment, and outputting the segment of mixed audio content comprising a mix of the first DRC segment and the segment of the second sequence.
49. The method according to any of claims 42 - 47, wherein performing dynamic range compression based on the DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence comprises: performing dynamic range compression of the segment of the second sequence based on the DRC parameter to form a second DRC segment, and outputting the segment of mixed audio content comprising a mix of the second DRC segment and the segment of the first sequence.
50. The method according to any of claims 42 - 47, wherein performing dynamic range compression based on the DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence comprises: performing dynamic range compression of the segment of the first sequence to form a first DRC segment and performing dynamic range compression of the second sequence to form a second DRC segment, and outputting the segment of mixed audio content comprising a mix of the first DRC segment and the second DRC segment.
51. The method according to claim 50, wherein determining a dynamic range compression, DRC, parameter based on the at least one speech experience measure comprises: determining two DRC compression parameters based on at least one speech experience measure, a first DRC parameter and a second DRC parameter, wherein the first DRC segment is formed based on the first DRC parameter and the second DRC segment is formed based on the second DRC parameter.
52. The method according to any of claims 42 - 51, wherein the at least one speech experience measure comprises at least one of: a speech-to-non-speech ratio associated with the time-aligned pair of segments, a speech-to-non-speech ratio associated with a preprocessed version of the time- aligned pair of segments; direct user feedback, static user preferences, playback environment data, and metadata associated with the first and / or second sequence of audio segments.
53. An audio processing method comprising: receiving, by a receiving device, a bitstream; processing, by the receiving device, the bitstream to obtain a first and second sequence of time-aligned audio segments and, for each time-aligned pair of segments, a at least one dynamic range compression, DRC, parameter; for each time-aligned pair of segments in the first and second sequence: performing dynamic range compression based on the at least one DRC parameter on at least one of the segment of the first sequence and the segment of the second sequence to form at least one DRC segment associated with the first or second sequence; and outputting a segment of mixed audio content comprising a mix of the at least one DRC segment associated with the first or second sequence and a segment associated with the other one of the first and second sequence.
54. The method according to claim 53, wherein the bitstream comprises encoded segments of a mix signal and wherein processing the bitstream comprises:decoding the bitstream to obtain the segments of the mix signal, the mix signal segments comprising speech audio content mixed with non-speech audio content; and processing each mix segment with a separator for speech separation to form the first and second sequence of segments, wherein the separator is configured to separate each segment into a speech content segment and non-speech content segment.
55. An apparatus comprising a processor and a memory, configured to perform the method according to any of claims 41-54.
56. A computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to any of claims 51-54.
57. A computer-readable storage medium storing the computer program according to claim 56.
Citation Information
Patent Citations
Enhancing intelligibility of speech content in an audio signal
EP3149730B1
Ratio of Speech to Non-Speech Audio such as for Elderly or Hearing-Impaired Listeners
US20100106507A1
Method and Apparatus for Maintaining Speech Audibility in Multi-Channel Audio with Minimal Impact on Surround Experience
US20110054887A1