Apparatus, method, computer program and bitstream for quality control and / or enhancement of audio scene

By obtaining the short-term intensity difference between speech and background components through an audio analyzer, the problem of speech being difficult to keep up with due to excessive background noise is solved. This enables audio scene quality control and enhancement with low computational complexity, adapts to the characteristics of different audio content, and improves intelligibility and auditory impression.

CN121925703APending Publication Date: 2026-04-24FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2024-07-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing audio mixing techniques are unable to effectively solve the problem of excessive background noise causing speech to be difficult to keep up with, leading to audience fatigue and frustration. Furthermore, traditional methods cannot achieve a good trade-off between quality control and enhancement of the speech and background parts.

Method used

An audio analyzer is used to obtain the short-term intensity difference between the speech and background parts of the audio content. Through short-term intensity measurement and analysis, a quality control report is provided to improve the intelligibility and auditory impression of the audio scene. The audio scene is manipulated and enhanced using low computational complexity and low-level features.

Benefits of technology

It effectively improves the intelligibility of audio scenes with low computational complexity. Through short-term intensity measurement and analysis, it enhances the quality control and enhancement of audio scenes, adapts to the characteristics of different audio content, and supports personalized audio improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925703A_ABST
    Figure CN121925703A_ABST
Patent Text Reader

Abstract

Embodiments according to the present invention include apparatuses, methods, computer programs and bitstreams for quality control and / or enhancement of audio scenes. Embodiments in accordance with the present invention relate to devices and methods for quality control and enhancement of audio scenes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Field ‌

[0002] Embodiments of the present invention include devices, methods, computer programs, and bitstreams for quality control and / or enhancement in audio scenarios.

[0003] Background Technology ‌

[0004] For audio mixing in television (broadcast or streaming), it is frequently reported that the speech is difficult to follow due to excessive background noise. Excessive background music and sound effects in audio mixing can mask the dialogue in the foreground, leading to viewer fatigue and frustration [1]. This is a problem known for decades [2]. Modern audio codec systems, such as Next Generation Audio (NAG) systems, offer the ability to provide technical solutions at the user end. These solutions include dynamic range compression (DRC) and the possibility of personalized speech levels, also known as dialogue enhancement [3]. However, these traditional methods still fail to produce satisfactory results.

[0005] Therefore, it is desirable to obtain a concept that achieves a better trade-off between the quality of an audio scene with a speech component and a background component (e.g., in the form of auditory impressions, such as intelligibility), the computational efficiency of providing, representing, encoding, decoding, and / or rendering the scene, and the computational complexity of the concept.

[0006] This is achieved through the subject matter of the independent claims of this application.

[0007] Further embodiments of the present invention are defined by the subject matter of the dependent claims of this application.

[0008] Summary of the Invention ‌

[0009] Embodiments of the invention (e.g., according to a first aspect) include an audio analyzer, for example for supporting audio production, post-production, or quality control stages, wherein the audio analyzer is configured to acquire (e.g., receive) audio content (e.g., audio scenes, such as in a format commonly used in audio production), the audio content including speech portions and background portions.

[0010] Alternatively, the audio analyzer may be configured, for example, to acquire audio content including a speech portion and a background portion, such as acquiring a "final mix" in which the speech portion and the background portion are combined, or to acquire separate signals representing the speech portion and the background portion of the audio content, respectively.

[0011] Furthermore, the audio analyzer is configured to determine the short-term intensity difference (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity difference values ​​describing multiple short-term intensity differences at different frequencies or frequency ranges) between the speech portion of audio content (e.g., in the sense of "dialogue" type audio content) and the background portion of the audio content (e.g., music and / or sound effects, such as diffused sound, such as stadium ambiance). Alternatively or additionally, the audio analyzer is configured to determine short-term intensity information of the speech portion of the audio content.

[0012] Furthermore, the audio analyzer is configured to provide a representation of short-term intensity differences (e.g., a single short-term intensity difference value for each part of the audio content, or a set of short-term intensity differences describing multiple different frequencies or frequency ranges, such as a coarse quantization representation of short-term intensity differences, such as quantization to only 2 quantization steps, or quantization to only 3 quantization steps, or quantization to only 4 quantization steps) and / or a representation of short-term intensity information of the speech portion as an analysis result, such as as a quality control report, or the audio analyzer is configured to derive analysis results (e.g., binary, ternary, or quaternary values, or information about key segments (e.g., segments in an audio scene where a particular signal characteristic of at least one audio component (audio portion) does not meet one or more expected (predefined) criteria) from short-term intensity differences (e.g., using comparisons between short-term intensity differences and one or more thresholds, or using comparisons between short-term intensity difference values ​​describing multiple different frequencies or frequency ranges and corresponding (e.g., frequency-dependent) associated thresholds) and / or from short-term intensity information of the speech portion).

[0013] The inventors recognized that the short-term intensity difference between the speech and background components, as well as the short-term intensity information of the speech component of audio content, could be a key indicator and adjustment tool for audio scene quality (e.g., regarding auditory impressions, especially intelligibility).

[0014] It is recognized that analysis of such short-term intensity measures (e.g., information or differences, such as absolute or relative values) allows for the identification of key segments of an audio scene (e.g., whether the quality of the audio scene is critical). The inventors further recognize that the “criticality” of such audio segments can be determined even based on the relationship between the short-term intensity of the speech portion (e.g., the acoustic foreground (e.g., also known as dialogue, though not limited to human conversation)) and the short-term intensity of the background portion, or based on the absolute short-term intensity of the speech portion.

[0015] In particular, the inventors recognized that short-term intensity metrics can provide a profound measure of the required listening effort and intelligibility, wherein said metrics can be obtained with low computational cost. Furthermore, based on such metrics, the inventors recognized that manipulation of an audio scene can be performed in a direct manner, such as improving, for example, locally (e.g., temporally local, spatially local, for example within a specific frequency range, for example in a portion of the time-frequency domain) the intensity difference or ratio, for example in the form of energy ratio or loudness difference.

[0016] Furthermore, the audio scene can be directly modified based on the measured characteristics of the audio signal, for example, by changing the audio mix, or the corresponding modifications can be provided or captured as metadata elements that can be applied in the receiving and / or playback devices. Therefore, it is further recognized that audio enhancement based on short-term intensity metrics can be efficiently provided as bitstream elements in the form of metadata, thus allowing for good flexibility in the inventive concept.

[0017] Furthermore, the embodiments do not rely on the analysis of higher-order cognitive factors, and therefore can be performed independently of the influence of factors such as unfamiliar vocabulary or accents or language fluency (although the embodiments may optionally include such analyses additionally). Therefore, audio scene enhancement can be performed more effectively by using factors that do not primarily or even specifically depend on individual listener characteristics or even the listener's current cognitive state.

[0018] In addition, using short-term intensity metrics allows for the provision of statistical data, such as, in particular, statistics on the number, severity, and temporal location of key segments. This enables differentiated improvements for different aspects and / or parts of the audio scene, and even allows—though not necessarily requires—user-specific settings, such as allowing the respective end user to set a preferred level of intelligibility.

[0019] According to embodiments of the invention (e.g., according to the first aspect), short-term intensity differences and / or short-term intensity information include a time resolution of no more than 3000 milliseconds (e.g., short-term loudness, e.g., according to EBU TECH 3341), or a time resolution of no more than 1000 milliseconds, or a time resolution of no more than 400 milliseconds [e.g., for instantaneous loudness, e.g., according to EBU TECH 3341], or a time resolution of no more than 100 milliseconds, or a time resolution of no more than 40 milliseconds, or a time resolution of no more than 20 milliseconds. Alternatively, short-term intensity differences and / or short-term intensity information include a time resolution between 3000 milliseconds and 400 milliseconds.

[0020] The inventors recognized that this absolute temporal resolution allows for the provision of conclusive analytical results with fine temporal granularity.

[0021] According to embodiments of the present invention (e.g., according to a first aspect), short-term intensity information of short-term intensity differences and / or speech portions includes a temporal resolution of one audio frame, or short-term intensity information of short-term intensity differences and / or speech portions includes a temporal resolution of two audio frames, or short-term intensity information of short-term intensity differences and / or speech portions includes a temporal resolution of no more than 10 audio frames.

[0022] The inventors recognized that a temporal resolution that depends on the frame size can achieve a good trade-off between the granularity of the analysis results, computational cost, and conclusivity.

[0023] According to embodiments of the invention (e.g., according to a first aspect), short-term intensity difference is short-term loudness difference, such as the short-term loudness difference between speech or dialogue and background, or instantaneous loudness difference, such as the instantaneous loudness difference between speech or dialogue and background. Alternatively or additionally, short-term intensity information of the speech portion is short-term loudness, such as the short-term loudness of speech or dialogue, or instantaneous loudness, such as the instantaneous loudness of speech or dialogue.

[0024] The inventors recognized that loudness difference, or loudness information, could allow for conclusive analytical results with low computational cost.

[0025] According to embodiments of the invention (e.g., according to a first aspect), short-term intensity difference is a short-term energy ratio, such as the short-term energy ratio between speech or dialogue and background, or an instantaneous energy ratio, such as the instantaneous energy ratio between speech or dialogue and background, and / or the short-term intensity information of the speech portion is short-term energy or instantaneous energy.

[0026] The inventors recognized that energy ratios or energy information could allow for conclusive analytical results with low computational cost.

[0027] Furthermore, the determination of loudness or energy metric can be achieved in a direct manner, thus limiting the additional complexity of the method of the present invention.

[0028] According to embodiments of the invention (e.g., according to the first aspect), short-term intensity difference and / or short-term intensity information are low-level features of audio content, such as those that do not take into account the temporal correlation between different parts of the audio content, and / or, for example, those that do not take into account the meaning of the speech portion of the audio content and / or the vocabulary of the speech portion of the audio content, and / or the accent of the speech portion of the audio content, and / or the speed of the speech portion of the audio content, and / or the fluency of the speech portion of the audio content, and / or the sentence complexity of the speech portion of the audio content, and / or the transmission speed of the speech portion of the audio content, and / or the phoneme pronunciation of the speech portion of the audio content, and / or the ambiguity of the speech portion of the audio content, and / or the unclear dialogue of the speech portion of the audio content, and / or the cognitively relevant intelligibility of the audio content.

[0029] Therefore, audio enhancement can be performed with low computational complexity, for example, independent of complex analysis of higher-order cognitive factors. This can be particularly advantageous because higher-order cognitive factors can be highly individualized, based on which general audio improvements (e.g., improving the audio experience for a broad audience) cannot be performed (or can only be performed in a complex manner). Thus, the embodiments allow for reduced computation by enabling the provision of “average” improvements rather than a large number of highly individualized improvement options.

[0030] According to embodiments of the invention (e.g., according to the first aspect), the audio analyzer is configured to provide analysis results based on features beyond intensity characteristics of the speech portions of the audio content, such as temporal correlation between different portions of the speech content; for example, meaning of the speech portions of the audio content, and / or vocabulary of the speech portions of the audio content, and / or accent of the speech portions of the audio content, and / or speed of the speech portions of the audio content, and / or fluency of the speech portions of the audio content, and / or sentence complexity of the speech portions of the audio content, and / or delivery speed of the speech portions of the audio content, and / or phoneme pronunciation of the speech portions of the audio content, and / or ambiguity of the speech portions of the audio content, and / or unclear dialogue of the speech portions of the audio content; for example, beyond intensity characteristics; for example, beyond intensity metrics.

[0031] It is recognized that the method of the present invention allows for the provision of analysis results for audio scene enhancement based, for example, solely on intensity metrics, thus allowing for low complexity in implementing the method. Furthermore, intensity metrics can be readily obtained from existing frameworks, thus facilitating the integration of the inventive concepts.

[0032] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to provide analysis results based solely on one or more features of the audio content (e.g., based on a short-term intensity difference metric, and optionally also based on an absolute intensity metric), which can be modified by scaling (e.g., intensity scaling, level scaling, or energy scaling) one or more portions of the audio content, such as the speech portion and / or the background portion.

[0033] It is recognized that scene modification through scaling allows for a good trade-off between computational cost and quality improvement in scene enhancement, thus allowing analysis work to be focused only on those features that can be modified through scaling, thereby reducing the computational cost of analysis.

[0034] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to separate acquired audio content (e.g., "final mix") into a speech portion of the audio content and a background portion of the audio content, for example, using source separation.

[0035] This can be beneficial in certain audio scenarios, enabling the acquisition of analytical results that can be used to effectively improve the scene. In particular, it may help determine the intensity ratio or difference between the acoustic foreground and background, and thus manipulate it.

[0036] According to embodiments of the invention (e.g., according to the first aspect), an audio analyzer is configured to determine or estimate the intensity of the speech portion of the acquired audio content (e.g., "final mix") and the intensity of the background portion of the acquired audio content (e.g., "final mix"), for example, without actually separating the speech portion and the background portion of the audio content, for example, using a "measurement tool".

[0037] The inventors recognized that estimations using intensity measures allow for a good trade-off between the accuracy of analytical results and computational costs.

[0038] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to provide metadata of audio content and / or encoded audio content as an analysis result, wherein the metadata can, for example, control modifications to the audio content.

[0039] It is recognized that providing metadata about audio content and / or encoded audio content as part of the analysis results may be computationally efficient, thereby reducing the computational load of subsequent processing steps.

[0040] According to embodiments of the invention (e.g., according to the first aspect), the audio analyzer is configured to provide analysis results in character-encoded form, such as as readable text, or in XML format, or in comma-separated values ​​(CSV) format, and / or in binary form.

[0041] It is recognized that this character encoding format allows the analysis of the present invention to be represented in a small number of bits and in a straightforward manner for subsequent processing (and thus, for example, easy to implement and further process, such as easy to integrate into existing frameworks).

[0042] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to provide visualization of analysis results, for example, in the form of plotting analysis results over time, wherein, for example, coloring is determined based on short-term level differences.

[0043] The inventors recognize that the inventive method for determining the analysis results can be explained in a straightforward manner, thus simplifying the determination of subsequent individual settings (e.g., parameterization) for scene enhancement (e.g., regarding the desired level of understandability), for example, for the corresponding content creator.

[0044] According to embodiments of the invention (e.g., according to the first aspect), an audio analyzer is configured to acquire short-term intensity differences and / or short-term intensity information of speech portions based on the local power of one or more audio signals, or based on the local power of multiple portions of audio content. Alternatively, the audio analyzer is configured to acquire short-term intensity differences and / or short-term intensity information of speech portions based on short-term or instantaneous loudness (e.g., using different time window sizes) according to ITU-R BS.1770 and EBU R 128 or variations thereof.

[0045] This method of determining short-term intensity measures allows for a good trade-off between the conclusiveness of the measures and the computational cost of determining them. Furthermore, the determination according to ITU-R BS.1770 and EBU R 128 allows for simple integration into the corresponding frameworks.

[0046] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to acquire short-term intensity differences and / or short-term intensity information of speech portions based on (e.g., using) one or more filtered portions of audio content, wherein the filtering may, for example, be configured to mimic the frequency-selective sensitivity of the human ear.

[0047] Filtering can allow manipulation of the audio scene to improve the conclusiveness of short-term metrics, for example, by mimicking the frequency-selective sensitivity of the human ear.

[0048] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to use a loudness calculation model to obtain short-term intensity differences and / or short-term intensity information of speech portions, the calculation model being adapted, for example, to use linear or nonlinear processing of time intervals of the audio content to derive short-term intensity values ​​of portions of the audio content.

[0049] This achieves a good trade-off between the conclusiveness of short-term intensity measurements and the computational cost.

[0050] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to use one or more AI-based short-term intensity estimates to obtain short-term intensity differences and / or short-term intensity information of speech portions.

[0051] The inventors recognized that artificial intelligence, such as using neural networks, can be trained to effectively provide short-term intensity measurements.

[0052] According to embodiments of the invention (e.g., according to the first aspect), an audio analyzer is configured to combine multiple portions of audio content (e.g., multiple portions of audio content of a speech type; e.g., multiple speech portions of audio content originating from different speakers; e.g., multiple portions of audio content of a background type) to obtain short-term intensity differences and / or to obtain short-term intensity information of the speech portions. Alternatively or additionally, the audio analyzer is configured to combine multiple audio signals of audio content (e.g., multiple audio signals of audio content of a speech type; e.g., multiple speech portions of audio content originating from different speakers; e.g., multiple audio signals of audio content of a background type; e.g., based on weighted combination; e.g., short-term intensity of the combined portions; e.g., short-term intensity of the combined audio signals) to obtain short-term intensity differences and / or to obtain short-term intensity information of the speech portions.

[0053] It is recognized that combining parts of audio content (e.g., segments, aspects) can be advantageous in order to provide a sufficient information basis for inventive analysis, thereby preventing analysis based on isolated parts of the content that contain only insufficient information about their relationship to other parts of the scene.

[0054] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to determine one or more key segments of audio content (e.g., segments whose speech intelligibility or comprehensibility is considered threatened, for example, in terms of start / end / duration / one bit per frame) based on short-term intensity differences and / or short-term intensity information based on speech portions, wherein key segments are portions of the audio content with lower intensity of speech portions, for example, below a threshold in an absolute sense and / or lower relative to the background portion, and wherein the analysis results include information about one or more key segments.

[0055] It is recognized that the method of the present invention allows for improvements to audio scenes, particularly regarding segments containing key relationships between speech and background portions.

[0056] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to determine one or more key segments based on a comparison of short-term intensity differences with a single threshold and / or multiple thresholds and / or based on a comparison of short-term intensity information of a speech segment with a single threshold and / or multiple thresholds.

[0057] It is recognized that, based on the method of the present invention, the use of—for example, easily implemented—thresholds can allow analysis results to be obtained with low computational complexity. Furthermore, using multiple thresholds can allow key segments to be categorized according to their severity. Categorization regarding severity can be simplified to finding settings for corresponding subsequent audio improvements, such as for different target groups (e.g., healthy, mild hearing loss, etc.).

[0058] According to embodiments of the invention (e.g., according to a first aspect), the audio analyzer is configured to use different thresholds for different portions of the audio content (e.g., different time portions), such as different portions of a corresponding speech portion, such as different portions of a corresponding background portion. Alternatively or additionally, the audio analyzer is configured to use different thresholds for different types of the audio signal of the audio content (e.g., speech type, such as background type; such as the type of acoustic scene of the audio content).

[0059] It is recognized that the threshold concept of the present invention can be adjusted, for example, in terms of the value and number of thresholds, to adapt to the specific circumstances of the audio scene or audio content, thereby improving the conclusivity of the analysis results.

[0060] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to use one or more frequency-dependent thresholds, such as those based on the frequency-selective sensitivity of the human ear or other psychoacoustic models; for example, using one threshold per frequency band; for example, to provide a signal when a predetermined number of thresholds are exceeded; for example, a signal indicating that the intelligibility of the speech portion and the background portion is at risk or no longer satisfied.

[0061] Therefore, this allows the analysis to be adapted to the characteristics of human hearing.

[0062] According to embodiments of the invention (e.g., according to the first aspect), the audio analyzer is configured to use artificial intelligence (e.g., using a neural network; e.g., using manual classification as training) to adjust one or more thresholds.

[0063] This can provide an effective way to selectively adjust thresholds, for example, for a target audience and / or specific scenario characteristics.

[0064] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to perform inference of a neural network to determine one or more key segments of audio content, wherein a key segment is a portion of the audio content in which the intensity of a speech portion is locally lower in an absolute sense and / or relative to a background portion, and wherein the analysis results include information about one or more key segments.

[0065] It is recognized that using neural networks can allow for the efficient delivery of analytical results.

[0066] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to determine information about one or more key segments, such as statistical information, for example, statistics: start, end, duration, quantity, severity and / or criticality (e.g., regarding the level of intelligibility and / or listening effort required to understand the key segment) and / or criticality level (e.g., which may be associated with the level of intelligibility or listening effort required to understand the key segment) and / or time position. Furthermore, the analysis results include said information.

[0067] Therefore, differentiated analytics results can be provided to allow for targeted improvements to audio content, such as in a goal-oriented manner.

[0068] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to use two or more states to determine information about the severity and / or criticality of one or more key segments (e.g., about the level of intelligibility and / or listening effort required to understand the key segments).

[0069] This allows for the provision of analytical results with a good trade-off between complexity and granularity.

[0070] According to embodiments of the invention (e.g., according to the first aspect), the audio analyzer is configured to provide analysis results in the form of binary analysis results, for example, indicating whether features of the audio content meet conditions (e.g., the audio signal is easy to understand) or not (e.g., the audio signal is not easy to understand; for example, the audio signal is difficult to understand). Alternatively, the audio analyzer is configured to provide analysis results in the form of ternary analysis results, for example, distinguishing three different classification levels of features, three different grades of features, or three different value ranges of features, for example, indicating whether the audio content or portions thereof are comprehensible with low, medium, and / or high listening effort. Alternatively, the audio analyzer is configured to provide analysis results in the form of quaternary analysis results, for example, indicating the degree to which features of the audio content meet conditions; for example, distinguishing four different classification levels of features, four different grades of features, or four different value ranges of features.

[0071] This allows for the provision of analysis results with a good trade-off between complexity and granularity.

[0072] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to determine short-term intensity differences and / or short-term intensity information of speech portions as an approximation of the intelligibility or listening effort of the audio content.

[0073] Therefore, low-level features can be used to approximate high-level cognitive features, which allows for a good trade-off between the conclusiveness of the analysis and the computational workload.

[0074] According to embodiments of the invention (e.g., according to a first aspect), the audio analyzer is configured to determine additional quality control information based on the audio content, such as loudness compliance and peak level.

[0075] This additional information can allow for improvements to subsequent scene enhancements.

[0076] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to use at least one absolute or relative threshold on one or more audio components (e.g., portions of audio content) or one or more groups of audio components as one or more desired criteria, for example, to identify key segments, wherein the absolute threshold is related to the short-term intensity of, for example, one or more selected audio components (e.g., speech portions of audio content) or groups of audio components, and wherein the relative threshold is related to the short-term intensity difference between audio components (e.g., between speech portions of audio content and background portions of audio content) or groups of audio components.

[0077] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to combine multiple audio signals of the same or similar type to form component groups.

[0078] This can allow for an improved information basis for subsequent analysis, for example, by combining information about the relationships between combined signals, rather than analyzing them individually, such as allowing the determination or use of relative thresholds.

[0079] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to combine multiple audio signals of the same or similar type to form component groups based on the importance of multiple audio signals, wherein the importance of the audio signals is manually set or determined based on the speech portion contained in the signal, and / or based on the contribution of each audio signal to the final mix, taking into account audio masking effects and characteristics of the human auditory system, to combine multiple audio signals of the same or similar type to form component groups.

[0080] This allows for the analysis of the importance of a given audio signal or group of audio signals. Therefore, it allows for the introduction of parameters to emphasize certain aspects of the audio content, such as its intelligibility.

[0081] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to acquire average speech information of audio content, compare short-term intensity information (e.g., as absolute (short-term) intensity of speech) with the average speech information to obtain a comparison result (e.g., a quantitative comparison result, such as information about the ratio or difference between the short-term intensity information and the average speech information), and derive an analysis result based on the comparison result that includes information about the deviation between the short-term intensity information and the average speech information. Therefore, the analysis result may optionally be a comparison result, or, for example, an interpretation thereof, such as indicating the severity of the deviation.

[0082] According to embodiments of the invention (e.g., according to a first aspect), average speech information includes information about at least one of average speech level, average speech intensity, or average speech loudness of the audio content.

[0083] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to determine average speech information based on the average (e.g., integral) of audio content (e.g., audio program) over a predetermined time interval.

[0084] According to embodiments of the invention (e.g., according to a first aspect), the audio analyzer is configured to provide analysis results based on a combined evaluation of short-term intensity difference and deviation of short-term intensity information from average speech information.

[0085] According to embodiments of the invention (e.g., according to a first aspect), an audio analyzer is configured to: determine information about local speech levels as short-term intensity information; determine average speech information based on the average speech loudness over the entire audio content, or the average speech loudness over a time period at least ten times the duration for which the short-term intensity difference or short-term intensity information is determined; compare the local speech levels with the average speech information to obtain a comparison result; and derive an analysis result based on an evaluation of the comparison result with respect to a threshold, for example, to classify the severity of the deviation.

[0086] According to embodiments of the present invention (e.g., according to the first aspect), the audio analyzer is configured to determine information about the evolution of short-term intensity difference over time between a speech portion of audio content and a background portion of audio content, and / or the audio analyzer is configured to determine information about the evolution of short-term intensity information over time between a speech portion of audio content, and the audio analyzer is configured to derive analysis results from the information about the evolution of short-term intensity difference over time and / or from the information about the evolution of short-term intensity information over time between a speech portion (102,622).

[0087] The absolute intensity of speech can be an important complement to the short-term loudness difference between speech and the background, and according to embodiments, analysis results are provided based on this loudness difference. In particular, according to some embodiments, the variation of the absolute intensity of speech over time is of interest. In other words, some embodiments (e.g., particularly those discussed above) can be based on the idea of ​​detecting whether the speech level (e.g., indicated by short-term intensity) deviates too much from the average speech level (e.g., indicated by average speech information) over the entire program (or time interval of audio content), for example, even without considering its relationship to the background.

[0088] Therefore, embodiments may optionally include corresponding short-term intensity information or changes in short-term intensity differences over time. Another inventive concept according to embodiments is that a critical segment is detected if the absolute level of the speech deviates locally from the speech loudness integrated over the entire program (or, for example, a portion of the program) to a certain threshold (e.g., 10 LU). Specifically, according to embodiments, a combined analysis of absolute speech loudness deviation and short-term speech loudness relative to the background can be performed.

[0089] Furthermore, the embodiments include an audio analyzer (100, 600, 800), wherein the audio analyzer is configured to acquire audio content (101, 601, 631, 1001, 1201, 1241) of an audio scene including a speech portion and a background portion; wherein the audio analyzer is configured to determine a short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content, and / or wherein the audio analyzer is configured to determine short-term intensity information (112, 612, 1112) of the speech portion of the audio content, and wherein the audio analyzer is configured to derive an analysis result (102, 622) from the short-term intensity difference and / or from the short-term intensity information of the speech portion, in order to provide information about key segments of the audio scene for which a particular signal characteristic of at least one audio component in the audio scene does not meet one or more predefined criteria.

[0090] For example, the ultimate goal of the analysis could be the concept of key segments (which may lead to increased listening effort (e.g., identifying segments that require increased listening effort), or more generally, “not meeting one or more expected criteria”), and thus the enhancement of the audio content for such segments.

[0091] Embodiments of the invention (e.g., according to the second aspect) include an audio analyzer, for example, supporting audio production, post-production, or quality control stages, wherein the audio analyzer is configured to acquire (e.g., receive) audio content comprising a speech portion and a background portion (e.g., an audio scene, for example, in a format commonly used in audio production), for example, acquiring a “final mix” of the speech portion and the background portion, or acquiring separate signals representing the speech portion and the background portion of the audio content, respectively. In addition, the audio analyzer includes a neural network configured to derive quality control information based on the audio content (e.g., a representation of short-term intensity differences; e.g., a representation of the short-term intensity of a speech portion; e.g., a single short-term intensity difference for each audio content portion, or a set of short-term intensity differences describing multiple different frequencies or frequency ranges; e.g., a coarse quantization representation of short-term intensity differences, e.g., quantized to only 2 quantization steps, or only to only 3 quantization steps, or only to only 4 quantization steps; e.g., information about key segments, e.g., information about segments of an audio scene where the specific signal characteristics of at least one audio component (audio portion) do not meet one or more expected (predefined) criteria).

[0092] The inventors recognize that the embodiments according to the first aspect can be implemented using neural networks, or their respective functions can be implemented using neural networks. Therefore, the analysis results may correspond to quality control information, or the quality control information may include the analysis results. Therefore, embodiments according to the second aspect of the invention may individually or in combination include any features, functions, and / or details of embodiments according to the first aspect of the invention. Optionally, embodiments according to the first aspect can be used to train a corresponding neural network. Once trained, the neural network may perform well in terms of processing speed of audio content, for example, compared to an implementation according to the first aspect without a neural network.

[0093] According to embodiments of the present invention (e.g., according to the second aspect), a neural network is configured (e.g., trained) to acquire a representation of the short-term intensity difference between the speech portion of an audio content and the background portion of the audio content (e.g., a single short-term intensity difference value for each audio content portion, or a set of short-term intensity differences describing multiple different frequencies or frequency ranges, e.g., a coarse quantization representation of the short-term intensity difference, e.g., quantized to only 2 quantization steps, or only to only 3 quantization steps, or only to only 4 quantization steps) and / or a representation of the short-term intensity information of the speech portion of the audio content as an analysis result. Furthermore, the audio analyzer is configured to provide the representation of the short-term intensity difference and / or the representation of the short-term intensity information of the speech portion of the audio content as quality control information.

[0094] According to embodiments of the invention (e.g., according to the second aspect), the neural network is trained using an audio analyzer according to embodiments (e.g., according to the first aspect), wherein the audio analyzer and the neural network are provided with the same input, and wherein, for example, the neural network is tuned and / or optimized to approximate the corresponding output of the audio analyzer.

[0095] According to embodiments of the invention (e.g., according to a second aspect), an audio analyzer is configured to separate the speech portion and the background portion of audio content in order to provide the separated speech portion and / or background portion to a neural network, for example, performing source separation before performing inference using the neural network; for example, performing preprocessing.

[0096] According to embodiments of the invention (e.g., according to a second aspect), the audio content includes a speech portion and a background portion in a combined manner (e.g., a mixed manner); for example, in the form of a combined audio signal composed of the speech portion and the background portion. Furthermore, the audio analyzer is configured to provide the speech portion and the background portion to the neural network in a combined manner (e.g., as a "final mix").

[0097] According to embodiments of the invention (e.g., according to the second aspect), the audio analyzer is configured to provide the speech portion and the background portion to the neural network as separate signals, thus, for example, providing the speech portion and the background portion separately.

[0098] According to embodiments of the present invention (e.g., according to a second aspect), the audio analyzer includes a first neural network for deriving quality control information based on an audio mix of audio content (e.g., including speech and background portions in an interleaved manner), and the audio analyzer includes a second neural network for deriving quality control information based on the speech and background portions of the audio content provided as separate information entries.

[0099] It is recognized that combining the analysis of audio content as a whole (e.g., speech + background) with its separate forms can increase the conclusivity of the results. Using neural networks can allow for limiting computational costs, despite this dual approach.

[0100] According to embodiments of the invention (e.g., according to the second aspect), the audio analyzer includes an end-to-end detector (e.g., a module, such as one based on artificial intelligence or a neural network) to directly detect key segments (e.g., key segments of audio content) from one or more input signals.

[0101] It has been recognized that neural networks can be trained to provide information about key segments in one step.

[0102] According to embodiments of the invention (e.g., according to the second aspect), an end-to-end detector (e.g., a module) is configured to switch between two sub-modules (e.g., a first sub-module configured to detect key segments in a mixed signal representation of audio content and a second sub-module configured to detect key segments in a separate signal representation of audio content) based on the input format type.

[0103] Therefore, the analysis of this invention can be performed in a manner optimized for the corresponding input format.

[0104] Embodiments of the invention (e.g., according to a third aspect) include an audio processor for processing audio content, wherein the audio processor is configured to acquire (e.g., receive) audio content comprising a speech portion and a background portion. Furthermore, the audio processor is configured to determine short-term intensity differences (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity difference values ​​describing short-term intensity differences at multiple different frequencies or frequency ranges) between the speech portion and the background portion of the audio content. Alternatively or additionally, the audio processor is configured to determine the short-term intensity of the speech portion. Furthermore, the audio processor is configured to modify audio content based on short-term intensity differences and / or based on the short-term intensity of the speech portion (e.g., to (optionally selectively) improve speech intelligibility (e.g., for portions of audio content with relatively low short-term intensity differences), for example, by (optionally selectively) modifying the speech portion of the audio content and / or the background portion of the audio content and / or parameter information of the audio content (e.g., gain values ​​or processing parameters) (e.g., for portions of audio content with relatively low short-term intensity differences), for example, thereby (optionally selectively) increasing the intensity difference between the speech portion of the modified audio content and the background portion of the modified audio content compared to the original intensity difference (e.g., for portions of audio content with relatively low short-term intensity differences), or wherein the audio processor, for example, is configured to, based on short-term intensity differences... Metadata information about the audio content is determined based on the difference in intensity and / or the short-term intensity of the speech portion, such that the relationship between the speech portion and the background portion of the audio content can be modified based on the metadata information and / or the speech (e.g., in the form of speech portions) can be modified or enhanced (e.g., by modifying its intensity and / or by applying frequency-related filtering) (e.g., by boosting its absolute level, e.g., by compressing it, e.g., by applying frequency-related filtering) (e.g., regardless of its relationship with the background) [e.g., in cases where the speech in some portions is lower in an absolute sense rather than lower relative to the background; this can, e.g., be processed by the decoder; this can, e.g., be important for pure speech segments (e.g., with little or no background), e.g., but also applies to cases where it may be considered better to apply compression or equalization to the full mix than to rebalance].

[0105] In addition, the audio processor can, for example, be configured to provide a file or stream as a result of modification of the audio content, wherein the file or stream includes unmodified audio content and metadata information.

[0106] Alternatively, the audio processor is configured to determine metadata information about the audio content based on short-term intensity differences and / or based on the short-term intensity of the speech portion, and to provide a file or stream (e.g., an audio master file) comprising (e.g., unmodified; e.g., uncompressed) audio content and metadata information (e.g., wherein the audio processor is configured to, for example, change or modify the audio content or add additional information to the audio content based on short-term intensity differences and / or based on the short-term intensity of the speech portion, such as keeping the audio content itself unchanged but adding additional metadata information, wherein, for example, the audio processor is configured to modify the audio content based on short-term intensity differences and / or based on the short-term intensity of the speech portion by using changes to the representation of the audio content input to the audio processor, or by adding additional metadata information to the representation of the audio content input to the audio processor).

[0107] An audio processor according to embodiments of the present invention may optionally include functions discussed in the context of an audio analyzer according to the first and / or second aspects, for example, relating to processing audio content in order to determine short-term intensity metrics. Therefore, an audio processor according to the third aspect may include, individually or in combination, any or all functions, details, and / or features discussed in the context of an audio analyzer according to the first and / or second aspects.

[0108] Furthermore, the inventors recognize that the audio processor of the present invention can allow modification of audio content (e.g., direct modification) and / or provide metadata information about the audio content (based on which, for example, subsequent modifications can be performed).

[0109] Therefore, the inventors recognized that, based on short-term intensity metrics, metadata information could be incorporated into the stream in addition to direct audio improvements, allowing for audio improvements at the decoder end or end-user end.

[0110] According to embodiments of the invention (e.g., according to a third aspect), an audio processor is configured to (e.g., in a time-varying manner) modify the short-term intensity difference between the speech portion of an audio content and the background portion of the audio content to obtain a processed version of the audio content. Alternatively or additionally, the audio processor is configured to (e.g., in a time-varying manner) modify the short-term intensity of the speech portion of the audio content to obtain a processed version of the audio content.

[0111] The inventors recognized that modifying this intensity metric could allow for improvements in audio content, resulting in good outcomes, particularly in intelligibility, and with limited computational cost.

[0112] According to embodiments of the invention (e.g., according to a third aspect), an audio processor is configured to modify acquired audio content at a time resolution not exceeding 3000 milliseconds (e.g., short-term loudness, e.g., according to EBU TECH 3341), or at a time resolution not exceeding 1000 milliseconds, or at a time resolution not exceeding 400 milliseconds (e.g., for instantaneous loudness, e.g., according to EBU TECH 3341), or at a time resolution not exceeding 100 milliseconds, or at a time resolution not exceeding 40 milliseconds, or at a time resolution not exceeding 20 milliseconds, in order to acquire a processed version of the audio content. Alternatively, the audio processor is configured to modify the acquired audio content at a time resolution between 3000 milliseconds and 400 milliseconds.

[0113] According to embodiments of the present invention (e.g., according to a third aspect), an audio processor is configured to modify acquired audio content at a time resolution of one audio frame, or at a time resolution of two audio frames, or at a time resolution of no more than 10 audio frames, in order to obtain a processed version of the audio content.

[0114] According to embodiments of the invention (e.g., according to a third aspect), an audio processor is configured to scale the speech portion and / or the background portion of the acquired audio content in order to obtain a processed version of the audio content, optionally such that, for example, the audio mix is ​​directly altered.

[0115] According to embodiments of the invention (e.g., according to a third aspect), an audio processor is configured to provide or modify metadata (e.g., gain information describing one or more gains of one or more different portions of audio content to be applied to the audio decoder side, such as affecting the scaling of the speech portion of the acquired audio content and / or the scaling of the background portion of the acquired audio content, such as on the audio decoder side) in order to obtain a processed version of the audio content, optionally causing, for example, to indirectly change the audio mix.

[0116] It is recognized that the inventive method for solving intensity metrics allows for improvements in the quality of audio scenes with limited impact on scene representation, i.e., by altering or providing scene metadata. Therefore, scene improvements can be performed with low computational effort.

[0117] According to embodiments of the invention (e.g., according to a third aspect), an audio processor is configured to determine metadata information about audio content based on short-term intensity differences and / or based on the short-term intensity of speech portions, such that the relationship between the speech portions and background portions of the audio content can be modified based on the metadata information and / or the speech portions can be modified based on the metadata information, for example, speech (e.g., by modifying its intensity and / or by applying frequency-related filtering, e.g., enhancement; e.g., making the speech enhanced (e.g., increasing its absolute level, e.g., compressing it, e.g., by applying frequency-related filtering); e.g., disregarding the relationship with the background, e.g., where the speech in some portions is lower in an absolute sense rather than lower relative to the background; this can, e.g., be processed by a decoder; this can, e.g., be important for pure speech segments (e.g., with little or no background), e.g., but also applicable to situations where it might be considered better to apply compression or equalization to the entire mix than to rebalance.

[0118] Furthermore, the audio processor is configured to provide modified audio content, including audio content (e.g., in an unmodified manner) and metadata information, such that the modified audio content includes the original audio content or at least its unmodified audio signal, as well as corresponding metadata information for manipulating the audio content or audio signal, such as the relationship between the speech portion and the background audio portion.

[0119] According to embodiments of the present invention (e.g., according to a third aspect), an audio processor is configured to format metadata information according to the audio data frame rate of the audio content.

[0120] Therefore, the implementation can allow for seamless integration into existing frameworks.

[0121] According to embodiments of the invention (e.g., according to a third aspect), the audio processor is configured to separate the speech portion and the background portion of the audio content.

[0122] According to embodiments of the present invention (e.g., according to a third aspect), the audio processor includes an audio analyzer according to any of the embodiments discussed above (e.g., according to a first and / or second aspect), wherein the audio processor is configured to modify audio content based on the analysis results.

[0123] According to embodiments of the present invention (e.g., according to a third aspect), an audio processor includes an audio analyzer according to any of the embodiments discussed above (e.g., according to a first and / or second aspect), wherein the audio processor is configured to modify or generate metadata information based on the results of the audio analyzer (e.g., based on the analysis results).

[0124] According to embodiments of the invention (e.g., according to a third aspect), the audio processor includes an audio analyzer according to any of the embodiments discussed above (e.g., according to a first and / or second aspect), wherein the audio processor is configured to store metadata information aligned with audio data in a file or stream (e.g., in an audio master file or stream). Therefore, the embodiments can allow seamless integration into existing frameworks, for example, adapting to the specific requirements of the corresponding file or stream format.

[0125] Embodiments of the invention (e.g., according to the fourth aspect) include a bitstream provider, for example, for providing an audio bitstream or for providing a transport stream, wherein the bitstream provider is configured to include an encoded representation of audio content (e.g., an encoded representation of audio content including a speech portion and a background portion) and quality control information (e.g., quality control metadata) into the bitstream (e.g., an audio bitstream including both the encoded representation of the audio content and the quality control information, or a transport bitstream including the audio bitstream and the quality control information; wherein, for example, the quality control information may be embedded in the descriptor of the MPEG-2 transport stream or in the file format frame of ISOBMFF [wherein, for example, the quality control information may be provided in a bitstream syntax element suitable for (conventional) devices that do not evaluate or process quality control information]).

[0126] Quality control information may include metadata information as discussed in the embodiments of the third aspect of the invention and / or information about the analysis results as discussed in the embodiments of the first and / or second aspects of the invention (e.g., as a basis for metadata information or separately).

[0127] Therefore, a bitstream provider according to embodiments of the present invention may include functions discussed in the context of an audio analyzer according to the first and / or second aspects, for example, relating to processing audio content in order to determine short-term intensity metrics, and / or may include functions discussed in the context of an audio processor according to the third aspect of the present invention.

[0128] In other words, the bitstream provider according to the embodiments of the fourth aspect may include, individually or in combination, any or all of the functions, details and / or features discussed in the context of the audio analyzer according to the first and / or second aspects and / or in the context of the audio processor according to the third aspect of the invention.

[0129] It has been recognized that, for example, based on the analysis of this invention, quality control information can be provided in the bitstream, which enables efficient improvement of audio scenes, thereby imposing little or limited additional load on the corresponding bitstream and achieving good enhancement results.

[0130] According to embodiments of the invention (e.g., according to the fourth aspect), quality control information enables or supports decoder-side (e.g., selective) modification of the relationship between the intensity of the speech portion of the audio content and the background portion of the audio content. Alternatively or additionally, the quality control information enables or supports decoder-side modification (e.g., enhancement) of the speech portion (e.g., speech) (e.g., making the speech portion modifiable based on the quality control information; e.g., by modifying its intensity and / or by applying frequency-related filtering; e.g., enhancement; e.g., making the speech enhanced (e.g., making its absolute level boosted, e.g., compressing it, e.g., by applying frequency-related filtering); e.g., disregarding the relationship with the background; e.g., in cases where the speech in certain portions is lower in an absolute sense rather than lower relative to the background; e.g., this can be processed by the decoder; e.g., this may be important for pure speech segments (e.g., with little or no background), e.g., but also applicable to cases where it may be considered better to apply compression or equalization to the entire mix than to rebalance).

[0131] According to embodiments of the invention (e.g., according to the fourth aspect), quality control information enables or supports decoder-side (e.g., selective) improvement of the speech intelligibility of the speech portion of audio content, for example, in the presence of background portions of audio content that reduce speech intelligibility.

[0132] According to embodiments of the invention (e.g., according to the fourth aspect), quality control information selectively (e.g., in a time-related manner) enables and disables decoder-side improvements to the speech intelligibility of the speech portion of audio content (e.g., in the presence of background portions of audio content that reduce speech intelligibility; for example, although the quality control information may not actually actively enable decoder-side improvements, instead, it may indicate (e.g., selectively) where decoder-side improvements are meaningful or appropriate).

[0133] Alternatively, the quality control information indicates which segments of the audio content, decoder-side improvements to the speech intelligibility of the speech portion of the audio content (e.g., in the presence of background portions of the audio content that reduce speech intelligibility) are permissible.

[0134] It has been recognized that the analysis of audio content in this invention allows for the acquisition of information defining constraints on scene modifications to prevent modifications that would degrade the speech and background components.

[0135] According to embodiments of the invention (e.g., according to the fourth aspect), the quality control information includes information about key time portions (e.g., segments) of the audio content (e.g., specific information; e.g., specific flags or specific quantitative values) (e.g., information about; e.g., information for signaling notification; e.g., information for description) (e.g., signals for signaling key time portions of the audio content and / or information indicating the criticality of different time portions of the audio content, or information indicating whether a time portion is critical).

[0136] According to embodiments of the invention (e.g., according to the fourth aspect), the quality control information includes information indicating which segments of the audio content, under obstructed listening conditions (e.g., in a noisy listening environment, or, for example, in the presence of an unstable auditory environment, or, for example, in the presence of a hearing impairment by the listener, or, for example, in the presence of listener fatigue), the decoder-side improvement of the speech intelligibility of the speech portion of the audio content is considered recommendable, such as dedicated information; for example, dedicated flags or dedicated quantitative values.

[0137] Therefore, it can facilitate decisions regarding subsequent adjustments to audio content. In particular, by providing such information, user-specific adjustments can be made. Specifically, ordinary end users (e.g., those without specific knowledge in the field of audio processing) can be able to make informed decisions about whether to implement the corresponding modifications.

[0138] According to embodiments of the invention (e.g., according to the fourth aspect), the quality control information includes information indicating whether a portion of the audio content contains a speech intelligibility metric (e.g., a single numerical value describing the speech intelligibility of a speech portion) or a speech intelligibility-related feature (e.g., a short-term intensity difference between the speech portion of the audio content and the background portion of the audio content, or the level of the speech portion of the audio content) that is predeterminedly related to one or more thresholds (e.g., a single threshold or multiple thresholds, e.g., associated with different frequencies, e.g., greater than, equal to, or less than one or more thresholds) (e.g., specific information; e.g., a specific flag or a single specific quantitative value).

[0139] Therefore, metadata can be provided that provides differentiated information about the quality of audio content in terms of speech intelligibility.

[0140] According to embodiments of the invention (e.g., according to the fourth aspect), quality control information includes information (e.g., specific information; e.g., specific flags or a single specific quantitative value) that quantitatively describes speech intelligibility-related features of a portion of audio content (e.g., short-term intensity difference between the speech portion of the audio content and the background portion of the audio content, or the level of the speech portion of the audio content).

[0141] Having numerical metrics allows for the quantification of the severity of intelligibility problems, and thus the quantification of the “intensity” of the modifications required. Furthermore, this can allow for the categorization of audio scenarios for specific audiences (e.g., people with hearing impairments).

[0142] According to embodiments of the invention (e.g., according to the fourth aspect), quality control information includes information indicating whether an audio scene is considered difficult to understand or easy to understand (e.g., specialized information; e.g., specialized flags or a single specialized quantitative value).

[0143] According to embodiments of the invention (e.g., according to the fourth aspect), quality control information includes information indicating whether an audio scene is considered comprehensible with low listening effort, medium listening effort, or high listening effort (e.g., specific information; e.g., specific flags or a single specific quantitative value).

[0144] According to embodiments of the invention (e.g., according to a fourth aspect), the quality control information includes information indicating segments in the audio content where the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content is less than or equal to a threshold. Alternatively or additionally, the quality control information includes information indicating segments in the audio content where the short-term intensity of the speech portion of the audio content is less than or equal to a threshold.

[0145] It has been recognized that, based on the analytical method of the present invention, audio scenes or audio content can be selectively classified or categorized regarding quality (e.g., ease of understanding or difficulty of understanding) for certain temporal portions, spatial portions, and certain frequency intervals. Therefore, differentiated manipulation of audio content can be performed based on this knowledge.

[0146] According to embodiments of the invention (e.g., according to the fourth aspect), the quality control information describes the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content (e.g., a time resolution not exceeding 3000 milliseconds, for example, short-term loudness, for example, according to EBU TECH 3341, or a time resolution not exceeding 1000 milliseconds, or a time resolution not exceeding 400 milliseconds, for example, for instantaneous loudness, for example, according to EBU TECH 3341, or a time resolution not exceeding 100 milliseconds, or a time resolution not exceeding 40 milliseconds, or a time resolution not exceeding 20 milliseconds; or a time resolution between 3000 milliseconds and 400 milliseconds). Alternatively or additionally, the quality control information describes the short-term intensity of the speech portion of the audio content.

[0147] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to add quality control information (e.g., quality control metadata) to pre-existing metadata (e.g., such metadata from the production session, e.g., metadata already present in the encoded representation of the audio content).

[0148] Therefore, seamless integration with existing frameworks can be provided.

[0149] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to adjust processing parameters (e.g., filter coefficients, gain values, etc.) for decoding the audio content based on the short-term intensity difference between the speech portion and the background portion of the audio content (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity differences describing multiple different frequencies or frequency ranges) and / or based on the short-term intensity of the speech portion, for example, to implicitly signal key time portions; for example, to implicitly trigger decoder-side improvements in speech intelligibility.

[0150] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to include quality control information in the extended payload of the bitstream stream (e.g., in payloads that can be enabled and disabled, for example, using flags or list entries indicating the presence (or absence) of the payload; for example, in payloads defined as optional; for example, MHAS packets in the case of MPEG-H 3D audio codecs, and / or bitstream extension elements, such as usacExtElement, in the case of MPEG-H 3D audio codecs and / or USAC resp. xHE-AAC audio codecs; for example, conforming to XHE-AAC and / or MPEG-H 3D audio).

[0151] Therefore, these methods allow for scenario enhancements of the present invention with limited impact on existing frameworks. Furthermore, additional computational and bandwidth costs can be kept to a minimum.

[0152] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider includes an audio analyzer according to any of the previously discussed embodiments (e.g., according to the first and / or second aspects), wherein the bitstream provider is configured to determine quality control information based on the analysis results, and / or the bitstream provider includes an audio processor according to any of the embodiments discussed above (e.g., according to the third aspect).

[0153] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to format quality control information into quality control metadata packets (e.g., aligned with the audio frame rate).

[0154] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to encapsulate quality control metadata (e.g., quality control metadata packets) in packets and insert the packets into the audio bitstream during encoding.

[0155] Similarly, this method allows for scenario enhancements of the invention to be implemented with limited impact on existing frameworks. Furthermore, additional computational and bandwidth costs can be kept to a minimum.

[0156] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is implemented using a neural network, wherein the neural network is configured to receive a representation of audio content (e.g., a “final mix” or multiple audio signals representing different parts of the audio content, such as an audio signal representing the speech portion of the audio content and an audio signal representing the background portion of the audio content) and provide quality control information based thereon. Furthermore, the neural network is trained using training audio scenarios (i.e., audio content) that are labeled (e.g., categorized) in terms of speech intelligibility (e.g., difficult to understand / easy to understand; e.g., understandable with low listening effort / understandable with medium listening effort / understandable with high listening effort), wherein, for example, the parameters of the neural network are adjusted such that the output provided by the neural network in response to the training audio scenario approximates the corresponding label associated with the corresponding training audio scenario.

[0157] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is implemented using a neural network, wherein the neural network is configured to receive a representation of audio content (e.g., a “final mix” or multiple audio signals representing different parts of the audio content, such as an audio signal representing the speech portion of the audio content and an audio signal representing the background portion of the audio content), and provide quality control information based thereon, wherein the neural network is trained using an audio analyzer according to any of the above embodiments (e.g., according to the first aspect), and wherein the audio analyzer is configured to provide reference quality control information for the training of the neural network based on multiple audio scenes (i.e., audio content).

[0158] According to embodiments of the invention (e.g., according to the fourth aspect), the quality control information includes clarity information metadata, or accessibility enhancement metadata, or speech transparency metadata, or speech enhancement metadata, or intelligibility metadata, or content description metadata (e.g., which may be related to intelligibility; wherein the content description metadata may be, for example, low-level and audio-oriented); or local enhancement metadata, or signal description metadata.

[0159] It has been recognized that the present invention’s analysis based on audio content can provide the metadata elements discussed above, thereby allowing for in-depth analysis and improvement of audio content.

[0160] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to include quality control information (e.g., audio quality control metadata qcInfo() and information about audio quality control metadata, such as qcInfoCount, the number of data structures qcInfo() carrying audio quality control metadata present in the MHAS packet, and / or information about when the audio quality control metadata should be applied (e.g., qcInfoActive), and / or information about under what circumstances the audio quality control metadata should be applied, and / or information about under what circumstances the corresponding decoder can choose to apply the audio quality control metadata, and / or information indicating whether the audio quality control metadata (e.g., qcInfo()) is associated with a specific (e.g., a single) audio element or with an audio scene defined by a combination of audio elements (qcInfoType)) into the MHAS packet (e.g., into the MPEG-H 3D audio stream packet, e.g., into a dedicated MHAS packet, which may, e.g., contain only quality control information).

[0161] Therefore, seamless integration with existing frameworks can be achieved.

[0162] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured to provide audio quality control metadata (e.g., qcInfo()) and information about the audio quality control metadata (wherein the audio quality control metadata may, for example, quantitatively describe speech intelligibility-related features of a portion of the audio content, and / or where the audio quality control metadata may, for example, include information as defined in any of the embodiments previously discussed (e.g., according to the fourth aspect). In addition, information about audio quality control metadata includes descriptions of how many data structures (qcInfo()) carrying audio quality control metadata exist in the MHAS package, and / or descriptions of audio quality control metadata or indications of when audio quality control metadata should be applied (e.g., qcInfoActive), and / or descriptions of audio quality control metadata or indications of under what circumstances audio quality control metadata should be applied (e.g., providing information on criteria that can be used for decoder-side decisions), and / or descriptions of audio quality control metadata or indications of under what circumstances the corresponding decoder or renderer can choose to apply audio quality control metadata (e.g., providing information on criteria that can be used for decoder-side decisions), and / or information about audio quality control metadata (e.g., qcInfoType) indicating audio quality control metadata (e.g., qcInfoType). For example, whether qcInfo() is associated with a specific (e.g., a single) audio element or with an audio scene defined by a combination of audio elements; and / or information about the audio quality control metadata (e.g., qcInfoType) indicates the type of audio content to which the audio quality control metadata is associated (e.g., a single audio element or, for example, an aggregation of audio elements or, for example, multiple audio elements); and / or information about the audio quality control metadata (e.g., qcInfoType) indicates which type of audio content the audio quality control metadata can be applied to manipulate audio elements of the corresponding type; and / or information about the audio quality control metadata includes an identifier (e.g., mae_groupID, e.g., mae_groupPresetID) indicating which audio element or group of audio elements (e.g., a combination) the corresponding audio quality control metadata is associated with.

[0163] The above implementation allows for particularly efficient audio content enhancement.

[0164] According to embodiments of the invention (e.g., according to the fourth aspect), a bitstream provider is configured, for example, within an MHAS packet, to provide audio quality control information with a first granularity (e.g., with a first temporal granularity, or with a separate association with a specific audio element) and audio quality control information with a second granularity (e.g., with a first temporal granularity; e.g., without an association with a specific audio element). Therefore, the provision of quality control information can be adapted to apply specific constraints, such as the bandwidth currently available for the audio stream.

[0165] According to embodiments of the invention (e.g., according to the fourth aspect), the bitstream provider is configured, for example, within an MHAS packet, to provide multiple different audio quality control metadata (e.g., lists) (e.g., different qcInfo() data structures) associated with different audio elements and / or different combinations of audio elements, and optionally to provide extended audio quality control metadata, which may be, for example, general to all audio elements and combinations of audio elements.

[0166] Therefore, differentiated metadata information can be provided to selectively improve relevant aspects of the audio scene.

[0167] Embodiments of the invention (e.g., according to the fifth aspect) include an audio decoder for providing a decoded audio representation (e.g., one or more decoded audio signals) based on an encoded media representation (e.g., an encoded audio representation; e.g., based on an audio bitstream including quality control information; e.g., based on a transport stream including an audio bitstream and quality control information and possible additional media information (such as a video bitstream); e.g., based on a bitstream including encoded data comprising one or more audio signals containing at least two different audio types (e.g., different portions of audio content), which can be characterized as, for example, dialogue and / or narration and / or background and / or music and / or effects; wherein, for example, at least two different audio types can be contained in at least two different audio signals, e.g., a stereo channel signal with music and effects and a mono dialogue audio object; or wherein, for example, at least two different audio types can be contained in the same one or more audio signals, e.g., a stereo full main channel containing a mixture of music and effects with dialogue).

[0168] In addition, the audio decoder is configured, for example, to use a bitstream parser to obtain (e.g., extract) quality control information (e.g., quality control metadata) from the encoded media representation, wherein the quality control information may, for example, be extracted from the audio bitstream or from a transport stream that includes quality control information and the audio bitstream as separate information.

[0169] In addition, the audio decoder is configured to provide a decoded audio representation based on quality control information.

[0170] Therefore, a decoder provider according to embodiments of the present invention may include functions discussed in the context of an audio analyzer according to the first and / or second aspects, such as those relating to processing audio content to determine short-term intensity metrics, and / or may include functions discussed in the context of an audio processor according to the third aspect of the present invention. Thus, embodiments according to the fifth aspect may include, individually or in combination, any or all functions, details, and / or features discussed in the context of an audio analyzer according to the first and / or second aspects and / or in the context of an audio processor according to the third aspect of the present invention.

[0171] Furthermore, the audio decoder according to the embodiments can form a counterpart to the bitstream provider of the present invention, and therefore can include the corresponding (decoder-side) features, functions, and details disclosed in the context of the bitstream provider according to the fourth aspect. Additionally, the audio decoder of the present invention can be additionally configured to perform the rendering of an audio scene, and therefore can include any or all of the rendering functions previously discussed.

[0172] According to embodiments of the invention (e.g., according to the fifth aspect), the encoded media representation includes a representation of audio content comprising a speech portion and a background portion, and the audio decoder is configured to receive quality control information comprising at least one of the following: information for modifying the relationship between the intensity of the speech portion of the audio content and the background portion of the audio content; information for modifying, e.g., enhancing, the speech portion (e.g., speech) of the audio content (e.g., by modifying its intensity and / or by applying frequency correlation filtering) (e.g., such that the speech portion can be modified based on the quality control information (e.g., enhancement; e.g., such that the speech is enhanced (e.g., such that its absolute level is increased, e.g., such that it is compressed, e.g., by applying frequency correlation filtering); e.g., disregarding the relationship with the background); e.g., in cases where there are portions of speech that are lower in an absolute sense rather than lower relative to the background; e.g., this can be processed by the decoder; e.g., this may be important for pure speech segments (e.g., with little or no background), e.g., but also applicable to cases where it may be considered better to apply compression or equalization to the full mix than to rebalance. Information for improving the speech intelligibility of the speech portion of audio content; information for selectively enabling and disabling improvements in the speech intelligibility of the speech portion of audio content; information indicating which segments of the audio content are permissible for improvements in the speech intelligibility of the speech portion of audio content; information indicating key time segments of the audio content; information indicating which segments of the audio content are considered recommendable for improvements in the speech intelligibility of the speech portion of audio content under hearing impairment conditions; information indicating whether a portion of the audio content contains a speech intelligibility metric or speech intelligibility-related feature that is predetermined to one or more threshold values; information indicating whether an audio scene is considered difficult or easy to understand; information indicating whether an audio scene is considered comprehensible with low, medium, or high listening effort; information indicating segments in the audio content where the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content is less than or equal to a threshold; and / or information indicating segments in the audio content where the short-term intensity of the speech portion of the audio content is less than or equal to a threshold. Furthermore, the decoder is configured to provide decoded audio information based on these (e.g., one or more of the information discussed above).

[0173] Embodiments of the invention allow for the provision, on the one hand, differentiated information regarding the presence and classification of problematic portions or segments of audio content; on the other hand, optionally, information or instructions on how to mitigate or overcome such problems (e.g., regarding intelligibility). Therefore, the embodiments allow for good flexibility, as decisions regarding whether to change or modify the audio content can be made based on the selection of different evaluation information (e.g., one or more) and user-specific constraints (e.g., hearing impairment). Furthermore, based on additional information for modifying and / or improving the content, the computational workload on the decoder side can remain limited despite the option to change the audio content.

[0174] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to perform speech enhancement based on quality control information (e.g., increasing the intensity difference between the speech portion of the audio content and the background portion of the audio content; e.g., increasing the intensity of at least one segment or fragment of the speech portion).

[0175] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to selectively perform speech enhancement on segments of audio content indicated by quality control information. Therefore, the additional computational cost can be kept low.

[0176] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to selectively perform speech enhancement on segments of audio content (e.g., speech segments relative to background portions) that indicate difficulty in understanding based on quality control information.

[0177] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to receive control information defining whether speech enhancement should be performed, such as user input, and the audio decoder is configured to activate and deactivate speech enhancement according to the control information defining whether speech enhancement should be performed and optionally also according to quality control information (e.g., selectively).

[0178] Implementations may allow for global improvements to audio content or local improvements to audio characteristics. Depending on the audio content, changes (e.g., modifications, such as improvements) can be made in a highly targeted manner, so that, for example, the audio parts that have already achieved the listener's expected effect (e.g., in a quiet scene, the background is more important) are not changed, but only the problematic parts are changed (e.g., a conversation between two people that is masked by background noise).

[0179] In particular, it is recognized that the manipulation of audio content by the present invention is particularly effective in solving the intelligibility problem.

[0180] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to receive control information defining interaction with an audio scene, such as user input (e.g., adjusting the level of a portion of audio content (e.g., the level of an audio object) or adjusting the position of an audio object; wherein the audio content may, for example, include at least a portion of a description of the audio scene), and the audio decoder is configured to (e.g., selectively) activate and deactivate speech enhancement based on the control information defining interaction with the audio scene (and optionally also based on quality control information). Alternatively or additionally, the audio decoder is configured to adjust one or more parameters of the speech enhancement based on the control information defining interaction with the audio scene and optionally also based on quality control information.

[0181] Therefore, the embodiments allow for processing interactive AR and / or VR (Augmented Reality / Virtual Reality) scenes, where how the audio scene evolves may not be predetermined. For example, introducing additional background noise can trigger intelligibility improvements according to the embodiments to enhance the dialogue portion of the audio scene. Because operations can optionally be performed only on low-level features, computational workload can be kept low, thereby allowing even challenging time constraints for rendering, such as real-time rendering, to be met.

[0182] According to an embodiment of the invention (e.g., according to the fifth aspect), the audio decoder is configured to adjust one or more parameters of the speech enhancement according to quality control information (e.g., the intensity relationship (e.g., intensity ratio) between the speech portion of the audio content and the background portion of the audio content is adjusted by the speech enhancement to (e.g., increase (or decrease) the degree).

[0183] Therefore, the embodiments allow for fine-tuning of the level of speech enhancement, for example, depending on the severity of the acoustic problem (e.g., intelligibility problem).

[0184] According to embodiments of the present invention (e.g., according to the fifth aspect), the audio decoder is configured to acquire (e.g., receive) information about the listening environment (e.g., information about the intensity of noise in the listening environment; e.g., information about background noise in the listening environment; e.g., information about the type of listening environment (e.g., public place, living room, car, etc.); e.g., information about the location (e.g., absolute location) of one or more listeners; e.g., information about the location of one or more listeners in the listening environment; e.g., information about conditions affecting the listener's attention in the listening environment; e.g., information about lighting conditions in the listening environment; e.g., information about motion in the listening environment; e.g., information about visual stimuli in the listening environment; e.g., information about time, e.g., dynamic, time-varying information of the listening environment). Furthermore, the audio decoder is configured to acquire (e.g., based on information about the listening environment and based on quality control information (e.g., at a time resolution not exceeding 3000 milliseconds, e.g., for short-term loudness, e.g., according to EBU TECH 3341, or at a time resolution not exceeding 1000 milliseconds, or at a time resolution not exceeding 400 milliseconds, e.g., for instantaneous loudness, e.g., according to EBU TECH 3341). TECH3341, or at a time resolution of no more than 100 milliseconds, or at a time resolution of no more than 40 milliseconds, or at a time resolution of no more than 20 milliseconds, or at a time resolution between 3000 milliseconds and 400 milliseconds, or at a time resolution of one audio frame, or at a time resolution of two audio frames, or at a time resolution of no more than 10 audio frames, determines whether to perform speech enhancement (e.g., such that speech enhancement is selectively performed on segments in the audio content where quality control information indicates that the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content is insufficient (e.g., less than a threshold, which can be determined, for example, by the audio decoder based on information about the listening environment), or, for example, such that if the audio decoder determines, based on information about the listening environment, that speech enhancement should be performed on key segments, speech enhancement is selectively performed on segments in the audio content where quality control information indicates that key segments (e.g., segments with a relatively low short-term intensity difference between the speech portion of the audio content and the background portion of the audio content).

[0185] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to acquire (e.g., receive) information about the listening environment (e.g., information about the intensity of noise in the listening environment; e.g., information about background noise in the listening environment; e.g., information about the location of one or more listeners in the listening environment; e.g., information about conditions in the listening environment that affect the listener's concentration; e.g., information about lighting conditions in the listening environment; e.g., information about motion in the listening environment; e.g., information about visual stimuli in the listening environment). Furthermore, the audio decoder is configured to adjust one or more parameters of the speech enhancement based on quality control information and based on the information about the listening environment (e.g., the degree to which the intensity relationship (e.g., intensity ratio) between the speech portion and the background portion of the audio content is increased by the speech enhancement).

[0186] Therefore, the implementation not only allows users to adapt to specific needs, but also allows users to adapt to specific needs of their surrounding environment, thereby enabling the auditory experience to be optimized for users in a highly personalized way.

[0187] According to embodiments of the present invention (e.g., according to the fifth aspect), the audio decoder is configured to acquire (e.g., receive) user input (e.g., information about the user's speech intelligibility requirements, or information about the user's concentration or cognitive load or focus (e.g., concentration), or information about the user's hearing impairment, or information from a hearing aid, or information from one or more other devices or sensors). Furthermore, the audio decoder is configured to perform speech enhancement based on user input and quality control information (e.g., at a time resolution of no more than 3000 milliseconds [e.g., short-term loudness, e.g., according to EBU TECH 3341], or at a time resolution of no more than 1000 milliseconds, or at a time resolution of no more than 400 milliseconds (e.g., for instantaneous loudness, e.g., according to EBU TECH 3341), or at a time resolution of no more than 100 milliseconds, or at a time resolution of no more than 40 milliseconds, or at a time resolution of no more than 20 milliseconds, or at a time resolution between 3000 milliseconds and 400 milliseconds, or at a time resolution of one audio frame, or at a time resolution of two audio frames, or at a time resolution of no more than 10 audio frames). (e.g., selectively performing speech enhancement for segments in the audio content where the quality control information indicates that the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content is insufficient (e.g., less than a threshold defined by user input), or, for example, selectively performing speech enhancement for segments in the audio content where the quality control information indicates that the key segment is.)

[0188] The adjustments to the low-level characteristics for audio enhancement according to the present invention allow for keeping computational workload low, thereby enabling additional user input without significantly increasing hardware requirements.

[0189] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to acquire (e.g., receive) system-level information (e.g., information about system settings; e.g., information about dialogue enhancement options; e.g., information about settings for hearing impairments; e.g., information about settings for visual impairments). Furthermore, the audio decoder is configured to perform speech enhancement based on user input and quality control information (e.g., at a time resolution of no more than 3000 milliseconds (e.g., short-term loudness, e.g., according to EBU TECH 3341), or at a time resolution of no more than 1000 milliseconds, or at a time resolution of no more than 400 milliseconds (e.g., for instantaneous loudness, e.g., according to EBU TECH 3341), or at a time resolution of no more than 100 milliseconds, or at a time resolution of no more than 40 milliseconds, or at a time resolution of no more than 20 milliseconds, or at a time resolution between 3000 milliseconds and 400 milliseconds, or at a time resolution of one audio frame, or at a time resolution of two audio frames, or at a time resolution of no more than 10 audio frames). This enhancement is configured to selectively perform speech enhancement on segments in the audio content where the quality control information indicates that the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content is insufficient (e.g., less than a threshold defined by user input), or, for example, selectively perform speech enhancement on segments in the audio content where the quality control information indicates that the segment is a key segment.

[0190] This can simplify the handling of the highly flexible adjustment options provided by the embodiments, for example, for end users.

[0191] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to acquire (e.g., receive) information about one or more sound reproduction devices (e.g., sound transducers) (e.g., information indicating whether an internal sound transducer (e.g., a speaker) or an external sound transducer (e.g., a speaker or headphones) of the device is used to reproduce audio content; e.g., information about the location of one or more sound transducers used to reproduce audio content). Furthermore, the audio decoder is configured to adjust one or more parameters of speech enhancement based on quality control information and based on information about the one or more sound reproduction devices (e.g., the degree to which the intensity relationship (e.g., intensity ratio) between the speech portion and the background portion of the audio content is increased by the speech enhancement).

[0192] Furthermore, the efficiency of audio analysis and audio content modification (e.g., based on intensity metrics) allows for the optimization of audio rendering by incorporating characteristics of the consumer's corresponding sound system.

[0193] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to acquire one or more of the following system-level information: information about system settings, such as preferred language settings, dialogue enhancement option settings, hearing impairment settings, and visual impairment settings; information about user settings, such as object level adjustment settings; for example, object position adjustment settings; information about the environment, such as sensors from acoustic monitoring of the environment, or sensors from optical monitoring of the environment, or from position sensors; and / or information about one or more additional devices, such as one or more sound transducers. Furthermore, the audio decoder is configured to perform one or more of the following functions based on quality control information and system-level information: determining (e.g., whether) critical segments exist in the audio content, for example, in any audio signal, and require improvement (wherein, for example, the quality control information may indicate the degree of criticality of the audio content, and where the system-level information may be used to determine whether quality improvement (e.g., speech enhancement) is required for the indicated degree of criticality); determining the level and / or intensity of the quality improvement to be applied (wherein, for example, the quality control information may describe the quality level of the audio content without quality improvement, and where, for example, the system-level information may describe the expected or required quality level); and deriving the quality control information required by the audio decoder to enhance the audio quality of one or more critical segments to improve intelligibility and / or reduce listening effort.

[0194] According to embodiments of the invention (e.g., according to the fifth aspect), quality control information (e.g., derived quality control information obtained from quality control information obtained from encoded media representation) includes one or more of the following: one or more gain sequences that need to (e.g., should or will) be applied to one or more audio signals or one or more portions of audio content in an audio scene; information about which signals or portions of the audio content should (e.g., or can) be processed to improve one or more key segments; and / or information about the duration of one or more key segments.

[0195] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to apply quality control information (e.g., derived quality control information obtained from quality control information obtained from the encoded media representation) to obtain a quality-enhanced version of the audio content.

[0196] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to apply quality control information (e.g., derived quality control information obtained from quality control information obtained from encoded media representation) to audio content (e.g., applied to different portions of the audio content, or applied to one or more audio signals representing the audio content) to obtain a quality-enhanced (e.g., in terms of intelligibility) version of the audio content.

[0197] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to perform filtering to obtain a quality-enhanced version of audio content, and the audio decoder is configured to determine one or more filter coefficients for filtering based on quality control information (e.g., quality control information obtained from the encoded media representation).

[0198] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to perform filtering to obtain a quality-enhanced version of audio content, and the audio decoder is configured to determine one or more filter coefficients for filtering based on quality control information (e.g., quality control information obtained from the encoded media representation).

[0199] It is recognized that the quality control information defined according to the embodiments can allow for the efficient determination of filter coefficients to improve audio content. Furthermore, existing filtering architectures can be reused in implementing the audio enhancement of this invention.

[0200] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to base its information on system-level information (e.g., information about system settings; e.g., information about dialogue enhancement options; e.g., information about settings for hearing impairments; e.g., information about settings for visual impairments), and / or on information about one or more sound reproduction devices (e.g., sound transducers) (e.g., information indicating whether to use internal sound transducers (e.g., speakers) or external sound transducers (e.g., speakers or headphones) to reproduce audio content; e.g., information about the location of one or more sound transducers used to reproduce audio content), and / or on information about the listening environment (e.g., information about noise intensity in the listening environment; e.g., information about background noise in the listening environment; e.g., information about the sound source). Information about the type of listening environment (e.g., public place, living room, car, etc.); information about the location of one or more listeners (e.g., absolute location); information about the location of one or more listeners in the listening environment; information about conditions in the listening environment that affect the listener's attention; information about lighting conditions in the listening environment; information about motion in the listening environment; information about visual stimuli in the listening environment; information about time, such as dynamic, time-varying information about the listening environment), and / or information about user input (e.g., information about the user's speech intelligibility requirements, or information about the user's attention or cognitive load or focus, or information about the user's hearing impairment) to determine one or more filter coefficients for filtering.

[0201] Therefore, the implementation allows for the combination of multiple individual factors to provide, for example, the best listening experience.

[0202] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to apply filtering to one or more output signals of a decoder core, or wherein the audio decoder is configured to apply filtering to one or more rendered audio signals obtained by rendering using the output signals of the decoder core, for example, to the final rendered output.

[0203] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to trigger time-frequency modifications, such as modifications to the audio content in a frequency-dependent and time-varying manner, based on quality control information and optionally also on user input and / or information about the listening environment.

[0204] It is recognized that enhancements can be performed particularly effectively in the time-frequency domain based on the quality control information of the present invention.

[0205] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to detect one or more key segments present in at least one audio signal contained in an encoded media representation (e.g., detecting one or more key segments present in at least one audio signal contained in an encoded media representation that require improvement under one or more current system settings (e.g., taking into account one or more current system settings), wherein the current system settings may, for example, include information about one or more user selections and / or information about the environment; wherein, for example, a segment is considered key if the reproduction of at least two different audio types (e.g., dialogue (speech) and background) results in an increase in the user's listening effort; wherein, for example, the audio decoder may be configured to identify key segments based on quality control information).

[0206] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to decode encoded audio data and process one or more detected key segments, for example, key segments of one or more audio signals obtained by decoding encoded audio data.

[0207] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to process one or more detected key segments to optionally improve the audio quality of the audio scene and / or optionally reduce the user's listening effort.

[0208] According to embodiments of the invention (e.g., according to the fifth aspect), quality control information includes metadata associated with an audio scene, which contains information about key segments present in at least one audio signal contained in an audio stream. Furthermore, an audio decoder is configured to process information about key segments present in at least one audio signal contained in an encoded media representation (e.g., an audio stream), and at least one additional piece of information from system-level or from other metadata in the encoded media representation (e.g., an audio stream), to determine whether the key segments present in at least one audio signal contained in the encoded media representation (e.g., an encoded audio stream) can be improved. Additionally, the audio decoder is configured to decode the encoded audio stream contained in or constituting the encoded media representation, and, when it is determined that the key segments present in at least one audio signal contained in the audio stream can be improved, use the information about the key segments present in the at least one audio signal to improve the audio quality of the entire audio scene.

[0209] Therefore, the decoder of this invention can automatically determine whether signal enhancement is needed. This avoids the computational burden of unnecessary signal processing. Furthermore, the increased autonomy of the decoder reduces the complexity for end users.

[0210] According to embodiments of the invention (e.g., according to the fifth aspect), information about key segments (e.g., which may be included in an encoded media representation) includes at least one parameter associated with the short-term intensity of an audio signal or part of an audio signal in an audio scene (e.g., the speech portion of audio content), or at least one parameter associated with the short-term intensity difference between two or more audio types contained in the audio scene (e.g., between the speech portion of the audio content of the audio scene and the background portion of the audio content of the audio scene).

[0211] According to embodiments of the invention (e.g., according to the fifth aspect), information about key segments (e.g., which may be included in the encoded media representation) includes at least one of the following parameters: information about which audio signals contain key segments; information about which audio signals need to be processed to improve the key segments (it is possible that not all audio signals contain key segments); one or more gain sequences to be applied to one or more audio signals (e.g., containing a time resolution of no more than 3000 milliseconds [e.g., short-term loudness, e.g. according to EBU TECH 3341], or a time resolution of no more than 1000 milliseconds, or a time resolution of no more than 400 milliseconds [e.g., for instantaneous loudness, e.g. according to EBU TECH 3341]). [TECH3341], or a time resolution of no more than 100 milliseconds, or a time resolution of no more than 40 milliseconds, or a time resolution of no more than 20 milliseconds, or a time resolution between 3000 milliseconds and 400 milliseconds, or a time resolution of one audio frame, or a time resolution of two audio frames, or a time resolution of no more than ten audio frames] (e.g., to improve speech intelligibility) [but, for example, if decoder-side enhancement of speech intelligibility is desired, it may be applied only to one or more audio signals, for example, as a result of a decoder-side decision, which may be based, for example, on user input, listening environment information, etc.); information about the start, and / or end and / or duration of at least one key segment; short-term intensity values ​​associated with at least one audio signal; short-term intensity differences associated with at least two audio types (e.g., at least two different parts of audio content, which may be characterized as, for example, dialogue and / or narration and / or background and / or music and effects).

[0212] Recognize that the above methods allow for the efficient representation or indication of key segments.

[0213] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to evaluate quality control information contained in an MHAS packet (e.g., an MPEG-H 3D audio stream packet, such as a dedicated MHAS packet that may contain only quality control information, for example) including audio quality control metadata (qcInfo() and information about the audio quality control metadata, such as qcInfoCount, how many data structures carrying audio quality control metadata exist in the MHAS packet, and / or information about when the audio quality control metadata should be applied (e.g., qcInfoActive), and / or information about under what circumstances the audio quality control metadata should be applied, and / or information about under what circumstances the corresponding decoder can choose to apply the audio quality control metadata, and / or information indicating whether the audio quality control metadata (e.g., qcInfo()) is associated with a specific (e.g., a single) audio element or with an audio scene defined by a combination of audio elements (qcInfoType)).

[0214] Therefore, the implementation allows for seamless integration into existing frameworks.

[0215] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to evaluate audio quality control metadata (e.g., qcInfo()) and information about the audio quality control metadata. Optionally, the audio quality control metadata may, for example, quantitatively describe speech intelligibility-related features of a portion of the audio content, and / or the audio quality control metadata may, for example, include information defined in one of the embodiments discussed above (e.g., according to the fourth aspect). Furthermore, the information about the audio quality control metadata describes how many data structures qcInfo() carrying the audio quality control metadata exist in the MHAS package, and / or describes or provides an indication of when the audio quality control metadata should be applied (e.g., qcInfoActive), and / or describes or provides an indication of under what circumstances the audio quality control metadata should be applied (e.g., providing information about standards that can be used for decoder-side decision-making), and / or describes or provides an indication of under what circumstances the corresponding decoder or corresponding renderer can choose to apply the audio quality control metadata (e.g., providing information about standards that can be used for decoder-side decision-making), and / or information about the audio quality control metadata, such as qcInfoType, indicating the audio quality control metadata... For example, qcInfo(), associated with a specific (e.g., a single) audio element or with an audio scene defined by a combination of audio elements; and / or information about audio quality control metadata, such as qcInfoType, indicating the type of audio content associated with the audio quality control metadata (e.g., a single audio element, an aggregation of audio elements, or multiple audio elements); and / or information about audio quality control metadata, such as qcInfoType, indicating which type of audio content the audio quality control metadata can be applied to in order to manipulate audio elements of the corresponding type; and / or information about audio quality control metadata containing identifiers, such as mae_groupID, e.g., mae_groupPresetID, indicating which audio element or group of audio elements (e.g., a combination) the corresponding audio quality control metadata is associated with.

[0216] According to embodiments of the invention (e.g., according to the fifth aspect), an audio decoder is configured to evaluate (e.g., within an MHAS package) audio quality control information having a first granularity (e.g., having a first temporal granularity, or associated individually with a specific audio element) and audio quality control information having a second granularity (e.g., having a first temporal granularity; e.g., not associated with a specific audio element). Therefore, the computational workload and / or the quality of audio rendering can be scalable.

[0217] According to embodiments of the invention (e.g., according to the fifth aspect), the audio decoder is configured to evaluate (e.g., within an MHAS package) multiple (e.g., lists) different audio quality control metadata (e.g., different qcInfo() data structures) associated with different audio elements and / or different combinations of audio elements, and optionally also to evaluate extended audio quality control metadata (e.g., possibly generalizable to all audio elements and combinations of audio elements). Therefore, quality control information can be provided in an audio element-specific manner, e.g., such that corresponding quality control metadata is associated with corresponding audio elements. This can allow for improved rendering quality of the audio scene.

[0218] Embodiments of the present invention (e.g., according to a first aspect) include a method for analyzing audio content, the method comprising: acquiring (e.g., receiving) audio content comprising a speech portion and a background portion (e.g., acquiring a “final mix” of the speech portion and the background portion, or acquiring separate signals representing the speech portion and the background portion of the audio content, respectively); determining a short-term intensity difference (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity difference values ​​describing multiple short-term intensity differences of different frequencies or frequency ranges) between the speech portion of the audio content (e.g., in the sense of “dialogue” type audio content, such as speech, and / or narration, and / or audio description) and the background portion of the audio content (e.g., music and / or effects, and / or diffuse sound, and / or stadium atmosphere) and the background portion of the audio content (e.g., music and / or effects, and / or diffuse sound, and / or stadium atmosphere); and / or determining the short-term intensity difference between the speech portion of the audio content (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity difference values ​​describing multiple short-term intensity differences of different frequencies or frequency ranges); and / or determining the short-term intensity difference between the speech portion of the audio content (e.g., a single short-term intensity difference value describing the short-term intensity difference ... Short-term intensity information of the speech portion, and providing a representation of short-term intensity differences (e.g., a single short-term intensity difference for each part of the audio content, or a set of short-term intensity differences describing short-term intensity differences of multiple different frequencies or frequency ranges; e.g., a coarse quantization representation of short-term intensity differences, such as quantization to only 2 quantization steps, or quantization to only 3 quantization steps, or quantization to only 4 quantization steps) and / or a representation of short-term intensity information of the speech portion as an analysis result (e.g., as a quality control report), or deriving analysis results (e.g., binary, ternary, or quaternary values) from short-term intensity differences (e.g., using comparisons between short-term intensity differences and one or more thresholds, or using comparisons between short-term intensity differences describing short-term intensity differences of multiple different frequencies or frequency ranges and corresponding (e.g., frequency-dependent) associated thresholds) and / or from short-term intensity information of the speech portion.

[0219] An embodiment of the invention (e.g., according to the second aspect) includes a method for analyzing audio content, the method comprising: acquiring (e.g., receiving) audio content comprising a speech portion and a background portion (e.g., acquiring a “final mix” of the speech portion and the background portion, or acquiring separate signals representing the speech portion and the background portion of the audio content, respectively), and using a neural network to derive quality control information based on the audio content (e.g., a representation of short-term intensity differences; e.g., a representation of the short-term intensity of the speech portion; e.g., a single short-term intensity difference for each portion of the audio content, or a set of short-term intensity differences describing multiple short-term intensity differences at different frequencies or frequency ranges; e.g., a coarse quantization representation of the short-term intensity differences, e.g., quantized to only 2 quantization steps, or quantized to only 3 quantization steps, or quantized to only 4 quantization steps).

[0220] Embodiments of the present invention (e.g., according to a third aspect) include a method for processing audio content, the method comprising: acquiring (e.g., receiving) audio content, wherein the audio content comprises a speech portion and a background portion; determining a short-term intensity difference between the speech portion of the audio content and the background portion of the audio content (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity difference values ​​describing short-term intensity differences of multiple different frequencies or frequency ranges); and / or determining the short-term intensity of the speech portion, and modifying the audio content based on the short-term intensity difference and / or based on the short-term intensity of the speech portion (e.g., to (selectively) improve speech intelligibility (e.g., for portions of the audio content having a relatively low short-term intensity difference); for example, by (selectively) modifying the speech portion of the audio content and / or the background portion of the audio content and / or parameter information of the audio content (e.g., gain values ​​or processing parameters) (e.g., for portions of the audio content having a relatively low short-term intensity difference); for example, thereby (selectively) increasing the intensity difference between the modified speech portion of the audio content and the modified background portion of the audio content compared to the original intensity difference (e.g., for portions of the audio content having a relatively low short-term intensity difference)).

[0221] Embodiments of the present invention (e.g., according to the fourth aspect) include a method for providing a bitstream (e.g., for providing an audio bitstream or for providing a transport stream), the method comprising: incorporating an encoded representation of audio content (e.g., an encoded representation of audio content including a speech portion and a background portion) and quality control information (e.g., quality control metadata) into the bitstream (e.g., an audio bitstream containing the encoded representation of the audio content and the quality control information, or a transport bitstream containing the audio bitstream and the quality control information; wherein, for example, the quality control information may be embedded in a descriptor of an MPEG-2 transport stream or a file format frame of an ISOBMFF).

[0222] Embodiments of the present invention (e.g., according to the fifth aspect) include a method for providing a decoded audio representation (e.g., one or more decoded audio signals) based on an encoded media representation (e.g., an encoded audio representation; e.g., an audio bitstream containing quality control information; e.g., a transport stream containing an audio bitstream and quality control information, and possibly additional media information such as a video bitstream), the method comprising: obtaining (e.g., extracting) quality control information (e.g., quality control metadata) from the encoded media representation (e.g., using a bitstream parser) [wherein, e.g., the quality control information may be extracted from the audio bitstream or from a transport stream containing quality control information and the audio bitstream as separate information]; and providing the decoded audio representation based on the quality control information.

[0223] Embodiments of the present invention include a computer program for performing any of the methods discussed above (e.g., according to the first, second, third, fourth and / or fifth aspects) when run on a computer.

[0224] The methods described above can be based on the same considerations as the corresponding audio analyzers, audio processors, bitstream providers, and audio decoders mentioned above. Incidentally, the corresponding methods can be accomplished through all the features and functions described individually or in combination within the context of the corresponding audio analyzers, audio processors, bitstream providers, and audio decoders.

[0225] According to embodiments of the present invention (e.g., according to the sixth aspect), a bitstream (e.g., an audio bitstream or a transmission bitstream containing an audio bitstream) is included, the bitstream comprising: an encoded representation of audio content (e.g., an encoded representation of audio content including a speech portion and a background portion); and quality control information (e.g., within an audio bitstream containing an encoded representation of audio content and quality control information, or within a transmission bitstream containing an audio bitstream and quality control information).

[0226] The bitstream according to the embodiments may be the result of a bitstream provider according to the fourth aspect, for example, particularly having any features of an audio analyzer according to the first and / or second aspects and / or having any features of an audio processor according to the third aspect.

[0227] Therefore, the bitstream according to the embodiments may include any features, functions and / or details disclosed in the context of the audio analyzer and / or audio processor discussed above, and in particular the bitstream provider discussed above.

[0228] Furthermore, the bitstream according to the embodiments can be the input to the decoder according to the fifth aspect. Therefore, the bitstream of the present invention can include the corresponding features, functions, and / or details disclosed in the context of the decoder of the present invention.

[0229] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information enables or supports decoder-side modification (e.g., selective modification) of the relationship between the intensity of the speech portion of the audio content and the background portion of the audio content. Alternatively or additionally, the quality control information enables or supports decoder-side modification (e.g., enhancement) of the speech portion (e.g., by modifying its intensity and / or by applying frequency-dependent filtering, e.g., making the speech portion modifiable [e.g., enhanced] based on the quality control information; e.g., making the speech enhanced (e.g., making its absolute level increased, e.g., making it compressed, e.g., by applying frequency-dependent filtering); e.g., disregarding the relationship with the background) (e.g., in cases where there are portions of speech that are lower in an absolute sense than lower relative to the background; e.g., this can be processed by the decoder; e.g., this may be important for pure speech segments [e.g., with little or no background], e.g., but also applicable to cases where it may be considered better to apply compression or equalization to the full mix than to rebalance).

[0230] According to embodiments of the invention (e.g., according to the sixth aspect), quality control information enables or supports decoder-side improvements (e.g., selective improvements) in the speech intelligibility of the speech portion of audio content, for example, in the presence of background portions of audio content that reduce speech intelligibility.

[0231] According to embodiments of the invention (e.g., according to the sixth aspect), quality control information selectively (e.g., in a time-related manner) enables and disables decoder-side improvements to the speech intelligibility of speech portions of audio content (e.g., in the presence of background portions of audio content that reduce speech intelligibility; for example, although the quality control information may not actually actively enable decoder-side improvements, it may instead (e.g., selectively) indicate where decoder-side improvements to the speech intelligibility of speech portions of audio content are permissible for which segments of the audio content (e.g., in the presence of background portions of audio content that reduce speech intelligibility).

[0232] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes information about key time portions (e.g., segments) of the audio content (e.g., specific information; e.g., specific flags or specific quantitative values) (e.g., relevant information; e.g., information for signaling notification; e.g., information for description) (e.g., information for signaling key time portions of the audio content and / or information indicating the degree of criticality of different time portions of the audio content, or information indicating whether a time portion is critical).

[0233] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes information indicating the decoder-side improvement of the speech intelligibility of the speech portion of the audio content under obstructed listening conditions (e.g., in a noisy listening environment, or in the presence of an unstable listening environment, or in the presence of a hearing impairment by the listener, or in the presence of listener fatigue), for which segments of the audio content are considered recommendable (e.g., specific information; e.g., specific flags or specific quantitative values).

[0234] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes information (e.g., specific information; e.g., a single numerical value describing the speech intelligibility of a speech portion) or speech intelligibility-related features (e.g., short-term intensity difference between the speech portion of the audio content and the background portion of the audio content, or the level of the speech portion of the audio content) that indicate whether a portion of the audio content contains a speech intelligibility measurement (e.g., a single numerical value describing the speech intelligibility of a speech portion) that is in a predetermined relationship (e.g., greater than, equal to, or less than one or more thresholds) with respect to one or more thresholds (e.g., a single threshold or a single specific quantitative value).

[0235] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes information (e.g., specific information; e.g., specific flags or a single specific quantitative value) that quantitatively describes speech intelligibility-related features of a portion of the audio content (e.g., short-term intensity difference between the speech portion of the audio content and the background portion of the audio content, or the level of the speech portion of the audio content).

[0236] Embodiments of the invention allow for the provision of a measure of speech intelligibility, for example, even through the analysis and processing of low-level audio characteristics such as intensity values.

[0237] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes information indicating whether an audio scene is considered difficult or easy to understand (e.g., specialized information; e.g., specialized flags or a single specialized quantitative value).

[0238] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes information indicating whether an audio scene is considered comprehensible with low listening effort, medium listening effort, or high listening effort (e.g., specific information; e.g., specific flags or a single specific quantitative value).

[0239] Therefore, the analysis of this invention allows for the classification of audio scenarios, making audio improvements available even to ordinary end users.

[0240] According to embodiments of the present invention (e.g., according to a sixth aspect), the quality control information includes information indicating segments in which the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content is less than or equal to a threshold; and / or the quality control information includes information indicating segments in which the short-term intensity of the speech portion of the audio content is less than or equal to a threshold.

[0241] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information describes the short-term intensity difference between the speech portion of the audio content and the background portion of the audio content (e.g., a time resolution not exceeding 3000 milliseconds, for example, for EBU short-term loudness; or a time resolution not exceeding 1000 milliseconds; or a time resolution not exceeding 400 milliseconds (e.g., for instantaneous loudness); or a time resolution not exceeding 100 milliseconds; or a time resolution not exceeding 40 milliseconds; or a time resolution not exceeding 20 milliseconds; or a time resolution between 3000 milliseconds and 400 milliseconds). Alternatively or additionally, the quality control information describes the short-term intensity of the speech portion of the audio content.

[0242] According to embodiments of the invention (e.g., according to the sixth aspect), the bitstream includes information for adjusting processing parameters (e.g., filter coefficients, gain values, etc.) for decoding the audio content based on short-term intensity differences between the speech portion and the background portion of the audio content (e.g., a single short-term intensity difference value describing the short-term intensity difference, or a set of short-term intensity differences describing multiple different frequencies or frequency ranges) and / or based on the short-term intensity of the speech portion, for example, to implicitly indicate key time portions; for example, to implicitly trigger decoder-side improvements in speech intelligibility.

[0243] According to embodiments of the invention (e.g., according to the sixth aspect), the bitstream includes an extended payload (e.g., a payload that can be enabled and disabled, for example, using flags or list entries indicating the presence (or absence) of the payload; for example, a payload defined as optional, for example, an MHAS packet in the case of the MPEG-H 3D audio codec, and / or a bitstream extension element in the case of the MPEG-H 3D audio codec and / or the USAC resp. xHE-AAC audio codec, for example, usacExtElementType; for example, conforming to XHE-AAC and / or MPEG-H; for example, an mpegh3daConfigExtension() extension element or a usacConfigExtension() extension element, or a usacExtElement() extension element), and the extended payload contains quality control information.

[0244] This allows for seamless integration into existing frameworks. Furthermore, the method of this invention allows for the simple integration of information into the payload, thereby reducing the amount of integration work required.

[0245] According to embodiments of the invention (e.g., according to the sixth aspect), the bitstream includes quality control meta-data packets, into which quality control information is formatted, for example, aligned with the audio frame rate.

[0246] According to embodiments of the invention (e.g., according to the sixth aspect), the bitstream contains data packets, and quality control metadata (e.g., quality control metadata packets) is encapsulated within the data packets.

[0247] These methods allow for simplified subsequent parsing of the bitstream.

[0248] According to embodiments of the invention (e.g., according to the sixth aspect), the quality control information includes proficiency information metadata, or the quality control information includes accessibility enhancement metadata, or the quality control information includes speech transparency metadata, or the quality control information includes speech enhancement metadata, or the quality control information includes intelligibility metadata, or the quality control information includes content description metadata (e.g., this may be related to intelligibility; wherein the content description metadata may be low-level and audio-oriented), or the quality control information includes local enhancement metadata, or the quality control information includes signal descriptive metadata.

[0249] It should be noted that in the above discussion, the embodiments are presented in an order of different aspects; however, features, functions, and details of an embodiment of one aspect may be implemented individually or in combination in the same, similar, or corresponding manner in an embodiment according to another aspect. The order of different aspects is for the purpose of facilitating understanding of the embodiments of the present invention.

[0250] Attached Figure Description ‌

[0251] The accompanying drawings are not necessarily drawn to scale; their main purpose is usually to illustrate the principles of the invention. In the following description, various embodiments of the invention are described in conjunction with the following drawings, wherein:

[0252] Figure 1 A schematic diagram of an audio analyzer according to an embodiment of the present invention is shown;

[0253] Figure 2 A schematic diagram of an audio analyzer incorporating a neural network according to an embodiment of the present invention is shown;

[0254] Figure 3 A schematic diagram of an audio processor according to an embodiment of the present invention is shown;

[0255] Figure 4 A schematic diagram of a bitstream provider according to an embodiment of the present invention is shown;

[0256] Figure 5 A schematic diagram of an audio decoder according to an embodiment of the present invention is shown;

[0257] Figure 6 A schematic diagram of an audio analyzer with optional features according to an embodiment of the present invention is shown;

[0258] Figure 7 A schematic example visualization of short-term loudness difference (ST-LD) for quality control (QC) according to an embodiment of the present invention is shown;

[0259] Figure 8 A schematic diagram of an audio analyzer for the optional determination of short-term intensity differences is shown according to an embodiment;

[0260] Figure 9 A schematic diagram of an audio analyzer with a detector according to an embodiment of the present invention is shown;

[0261] Figure 10 A schematic diagram of an audio analyzer with two detectors according to an embodiment of the present invention is shown;

[0262] Figure 11 A schematic diagram of an audio processor with optional features according to an embodiment of the present invention is shown;

[0263] Figure 12 A schematic diagram of a second audio processor with optional features according to an embodiment of the present invention is shown;

[0264] Figure 13 A schematic diagram of a third audio processor with optional features according to an embodiment of the present invention is shown;

[0265] Figure 14 A schematic diagram of a bitstream provider with optional features according to an embodiment of the present invention is shown;

[0266] Figure 15 A schematic diagram of an audio decoder with optional features according to an embodiment of the present invention is shown;

[0267] Figure 16 A schematic diagram of an audio decoder with filters according to an embodiment of the present invention is shown;

[0268] Figure 17 A schematic diagram of a bitstream provider with a multiplexer according to an embodiment of the present invention is shown;

[0269] Figure 18 A schematic diagram of an audio decoder with an optional demultiplexer according to an embodiment of the present invention is shown;

[0270] Figure 19 An example of the syntax for MHASPacketPayload() according to an embodiment of the present invention is shown;

[0271] Figure 20 An example of the value of MHASPacketType according to an embodiment of the present invention is shown; and

[0272] Figure 21 An example of the syntax for audioQualityControlInfo() according to an embodiment of the present invention is shown.

[0273] Detailed Implementation ‌

[0274] Even if they appear in different figures, in the following description, the same or equivalent elements or elements having the same or equivalent functions are represented by the same or equivalent reference numerals.

[0275] In the following description, numerous details are set forth to provide a more comprehensive explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detailed form to avoid obscuring embodiments of the invention. Furthermore, unless otherwise specifically stated, features of the different embodiments described herein can be combined with each other.

[0276] Figure 1 A schematic diagram of an audio analyzer according to an embodiment of the present invention (e.g., according to a first aspect) is shown. Figure 1 An audio analyzer 100 is shown, including a short-term intensity determiner 110 and an analysis result provider 120. The audio analyzer 100 is configured to acquire audio content 101 comprising a speech portion and a background portion, and to use the short-term intensity determiner 110 to determine a short-term intensity metric 112, such as the short-term intensity difference between the speech portion and the background portion of the audio content and / or short-term intensity information of the speech portion of the audio content. Furthermore, the audio analyzer 100 is configured to use the analysis result provider 120 to provide a representation of the short-term intensity difference and / or a representation of the short-term intensity information of the speech portion as an analysis result 102, or to derive the analysis result 102 from the short-term intensity difference and / or from the short-term intensity information of the speech portion.

[0277] Optionally, the determiner 110 may include a filter or filtering unit configured to obtain a short-term intensity metric based on one or more filtered portions of the audio content 101.

[0278] Figure 2A schematic diagram of an audio analyzer incorporating a neural network according to an embodiment of the present invention (e.g., according to a second aspect) is shown. Figure 2 An audio analyzer 200 is shown, configured to acquire audio content 101 comprising a speech portion and a background portion, and including a neural network configured to derive quality control information 202 based on the audio content 101. The quality control information 202 may correspond to (e.g., may be similar to or even identical to) [the original text / image ... Figure 1 The difference in analysis result 102 is that it was obtained based on artificial intelligence.

[0279] Note again here that it conforms to... Figure 2 The audio analyzer can be included in Figure 1 Any feature discussed in the context of [the context].

[0280] For example, quality control information 202 may be a representation of short-term intensity differences, and / or a representation of short-term intensity of speech portions, and / or a single short-term intensity difference for each portion of audio content, and / or a set of short-term intensity differences describing short-term intensity differences of multiple different frequencies or frequency ranges, and / or information about key segments, such as segments of an audio scene where the specific signal characteristics of at least one audio component (audio portion) do not meet one or more expected (predefined) criteria.

[0281] Figure 3 A schematic diagram of an audio processor according to an embodiment of the present invention (e.g., according to a third aspect) is shown. Figure 3 An audio processor 300, including a short-term intensity determiner 110, is shown. It is configured to acquire audio content 101 comprising a speech portion and a background portion. Furthermore, the audio processor 300 is configured to use the short-term intensity determiner 110 to determine the short-term intensity difference between the speech portion and the background portion of the audio content, and / or to determine the short-term intensity of the speech portion, as indicated by short-term intensity metric 112. Additionally, the audio processor 300 is configured to use a modifier 310 to modify the audio content 101 based on the short-term intensity difference and / or based on the short-term intensity of the speech portion. Therefore, as... Figure 3 As shown, modified audio content 312 can be provided. As an example, the modification may include scaling of at least a portion of the audio content. As another example, modifier 310 can be configured to determine metadata information based on short-term intensity metric 112 to provide modified audio content 310 including audio content 101 and metadata information.

[0282] As an optional alternative, the audio processor 300 can be configured as follows: Figure 3As shown, a metadata provider 320 is used to determine metadata information 321 about the audio content based on short-term intensity differences and / or short-term intensity based on speech portions, and a file or stream provider 330 (which is also optional) is used to provide a file or stream 332, such that the file or stream includes the audio content 101 and the metadata information 321.

[0283] Furthermore, as an optional feature, for example, using metadata provider 320 and / or modifier 310, the audio processor can be configured to provide or modify metadata to obtain a processed version 312 of the audio content. Thus, alternatively, metadata already existing in the audio content can be modified or additional metadata elements can be added. Therefore, metadata provider 320 can provide metadata information 321 to the corresponding modifier 310 (where, for example, provider 330 may or may not exist).

[0284] Specifically, the modified audio content may include a scaled version of the speech portion of the acquired audio content 101 and / or a scaled version of the background portion of the acquired audio content 101.

[0285] Furthermore, it should be noted that the embodiments are not limited to... Figure 3 Two alternative options are shown. Therefore, as an example, metadata information 321 (based on which intensity metrics can be modified) can be determined and provided to be included in a file or stream 332, or provided as part of modified audio content (e.g., 312). Specifically, the file or stream provider 330 can be configured to format the metadata information according to the audio data frame rate of the audio content 101.

[0286] Furthermore, as an optional feature, for example, as an alternative to the short-term intensity determiner 110, the audio processor 300 may include, for example, Figure 1 The combination of the short-term intensity determiner and analysis result provider shown receives audio content 101 and provides analysis results 102 to the metadata provider 320 and / or modifier 310. Therefore, the audio processor 300 may include information about... Figure 2 The discussion focuses on neural network-based implementations.

[0287] Therefore, the audio processor 300 may include an audio analyzer according to the first and / or second aspects, and thus include any or all of the corresponding features to provide the metadata provider 320 and / or modifier 310 with the corresponding determined analysis results (e.g., 101, e.g., 202, e.g., such as...). Figures 6 to 10 (As shown in the implementation diagram).

[0288] Figure 4 A schematic diagram of a bitstream provider according to an embodiment of the present invention (e.g., according to the fourth aspect) is shown. Figure 4 A bitstream provider 400 is shown, configured to include an encoded representation 401 of audio content and quality control information 402 into a bitstream 403. The quality control information 402 may be, for example, quality control metadata, such as any type of metadata modified, provided, or determined by the metadata provider 320 and / or modifier 310. Therefore, the bitstream provider may include, according to... Figure 3 Any or all functions, details, and features discussed in the context of the audio processor, and therefore also including those based on... Figure 1 and 2 And therefore also according to Figures 6 to 18 An audio analyzer (e.g., in a corresponding manner with respect to a corresponding decoder). Therefore, in particular, quality control information 402 can be determined based on the analysis results as described above. For example, unlike analysis results in the form of intensity metrics, quality control information 402 can include information about the "location" (e.g., temporal location) of problematic portions of the audio content, as well as information based on which such segments can be improved. Therefore, information 402 can be an interpretative version of the analysis results, such as analysis results with added optional classification information and / or instructions on how to overcome such problems in segments of the audio content (e.g., regarding intelligibility).

[0289] As another example, the encoded representation 401 may include pre-existing metadata, which is added to the quality control information 401 in the bitstream 403.

[0290] Furthermore, the bitstream provider 400 can be configured, for example, to adjust the decoding processing parameters of the audio content based on the interpretation of quality control information.

[0291] Figure 5 A schematic diagram of an audio decoder according to an embodiment of the present invention (e.g., according to the fifth aspect) is shown. Figure 5 An audio decoder 500 is illustrated, including a quality control information provider 510 and a decoded audio representation provider 520. The audio decoder 500 is configured to obtain quality control information 512 from an encoded media representation 501 using the quality control information provider 510, and to provide a decoded audio representation 502 using the decoded audio representation provider 520 based on the quality control information 512 and therefore based on the encoded media representation 501. As an example, the encoded media representation 501 may be included in a bitstream, such as bitstream 403, and the quality control information 512 may correspond to the quality control information 402.

[0292] Therefore, as Figure 5 The illustrated embodiments may include corresponding features, such as decoder-side features, like... Figure 4This is discussed in the context of, for example, particularly regarding the relevant quality control information 512.

[0293] Consistently, the decoder according to the embodiment can be configured to determine information about problematic segments of the audio content based on quality control information 512, and information about how such segments can be improved. Furthermore, the decoder can be configured to improve said segments, for example, to provide a decoded audio representation, such as as modified audio content, for example, conforming to... Figure 3 Explanation.

[0294] Specifically, the decoder 500 can be configured to determine, based on quality control information 512, whether and which enhancements to perform and for which segments of the audio content. Therefore, various types of information can be considered, such as metadata information, obtained from the encoded media representation 501 or as input (e.g., decoder-side input), such as information about the listening environment, information about user input, information about one or more sound reproduction devices, and / or information about system settings.

[0295] For example, this improvement in the decoded audio content can be performed based on filtering, where, for example, the parameterization of the filter is set according to quality control information 512. In particular, the filter can be set according to additional information such as system-level information, information about one or more sound reproduction devices, information about the listening environment, and / or information about user input.

[0296] Problematic audio segments can be defined as critical segments, for example, as discussed later. Figures 6 to 18 As described in [the text].

[0297] It should be noted that, optionally, in the above embodiments, the short-term intensity metric (e.g., 112) can be provided at a defined absolute temporal resolution, such as no more than 3000 milliseconds, no more than 1000 milliseconds, no more than 400 milliseconds, no more than 100 milliseconds, no more than 40 milliseconds, or no more than 20 milliseconds, or, for example, a temporal resolution between 3000 milliseconds and 400 milliseconds. Furthermore, the short-term intensity metric (e.g., 112) can also be provided according to the temporal resolution of audio frames, such as the temporal resolution of one, two, or no more than 10 audio frames.

[0298] In addition, as another optional feature, in the above embodiments, the short-term intensity measure (e.g., 112) can be determined by the corresponding short-term intensity determiner 110 as a short-term loudness measure (e.g., difference, such as instantaneous difference, or loudness, such as instantaneous loudness) or a short-term energy measure (e.g., ratio, such as instantaneous ratio, or energy, such as instantaneous energy).

[0299] Therefore, correspondingly, for example, in the form of loudness or energy measures, short-term intensity can be identified as a low-level feature of the audio content. Optionally, the corresponding analysis results can therefore depend only on such low-level features and thus be independent of the higher-order cognitive features of the audio content.

[0300] The embodiments of the present invention will be discussed further below. In particular, the following sections relate to apparatus and methods for quality control and enhancement in audio scenarios. At least some of these embodiments relate to methods and / or apparatus for audio content creation, post-production, and / or quality control (QC), for example, addressing different characteristics related to metrics of audio quality, mix level, intelligibility, and / or listening effort. Other embodiments relate to methods and / or apparatus for automatically and / or dynamically improving audio quality, mix level, intelligibility, and / or listening effort, for example, based on transmitted metadata and / or user settings and / or other device settings. Furthermore, other embodiments will be defined by the appended claims.

[0301] It should also be noted that this disclosure explicitly or implicitly describes the characteristics of content creation and / or post-production, QC, and / or decoding and / or encoding systems and / or methods.

[0302] Furthermore, it should be noted that the various aspects described herein can be used individually or in combination. Therefore, details can be added to each individual aspect without needing to add details to the other aspect.

[0303] It should be noted that any embodiment defined in the claims can be supplemented by any details (features and functions) described below. Furthermore, the embodiments described below can be used alone or supplemented by any feature in another part or any feature included in the claims.

[0304] Different embodiments and aspects of the invention will be described in the sections “Introduction to Embodiments,” “Terms and Definitions According to Embodiments,” “Problem Statement and Current Solution,” “Embodiments According to the Invention” (particularly the corresponding subsections), and “Other Embodiments.” Furthermore, other embodiments will be defined by the appended claims.

[0305] Furthermore, the features and functions disclosed herein related to a method can also be used in a device (e.g., configured to perform such functions). Additionally, any features and functions disclosed herein regarding a device (e.g., an audio analyzer, an audio processor, a bitstream provider, a decoder) can also be used in the corresponding method. In other words, the method disclosed herein can be supplemented by any features and functions described regarding a device.

[0306] Furthermore, any features and functions described herein may be implemented in hardware or software, or using a combination of hardware and software, as described in the "Implementation Alternatives" section.

[0307] In addition, the following features, functions, and details can be implemented as optional features in Figures 1 to 5 Any of the embodiments discussed in the context of this document.

[0308] Implement alternative solutions: ‌

[0309] Although some aspects are described herein in the context of a device, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block, item, or feature of the corresponding device. Some or all of the method steps may be performed by (or using) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.

[0310] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be performed using digital storage media, such as floppy disks, DVDs, Blu-ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, storing electronically readable control signals that cooperate (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.

[0311] Some embodiments of the present invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0312] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, performs one of the methods described herein. The program code may, for example, be stored on a machine-readable medium.

[0313] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0314] In other words, therefore, one embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is run on a computer.

[0315] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. Data carriers, digital storage media, or recording media are typically tangible and / or non-transient.

[0316] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0317] Another embodiment includes a processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0318] Another embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.

[0319] Another embodiment of the invention includes a device or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The device or system may include, for example, a file server for transmitting the computer program to the receiver.

[0320] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0321] The device described in this article can be implemented using hardware devices, or using a computer, or using a combination of hardware devices and a computer.

[0322] The device described herein, or any component thereof, may be implemented, at least in part, in hardware and / or software.

[0323] The methods described in this article can be performed using hardware devices, computers, or a combination of hardware devices and computers.

[0324] The methods described herein, or any component of the device described herein, may be performed at least in part by hardware and / or software.

[0325] The embodiments described herein are for illustrative purposes only. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the intent is limited only by the scope of the forthcoming patent claims and not by the specific details presented herein through the description and explanation of the embodiments.

[0326] Next, an introduction to the following embodiments is provided: As previously mentioned, audio mixing for television (broadcast or streaming) is often reported to be difficult to follow due to excessive background noise. Excessive background music and effects in audio mixing can obscure foreground dialogue, leading to viewer fatigue and frustration [1]. This is a problem known for decades [2]. Modern audio coding systems, such as Next Generation Audio (NAG) systems, offer the ability to provide technical solutions at the user end. These solutions include the possibility of dynamic range compression (DRC) and personalized speech levels, also known as dialogue enhancement [3].

[0327] Embodiments of the invention include and / or propose a system that optionally automatically detects key segments of an audio scene, for example, an audio scene may be a complete audio mix, a combination of audio channels and / or audio objects commonly used in NGA systems, and / or any other audio format used for the production and transmission of audio data. According to embodiments, key segments may be defined as those portions of the audio scene in which a particular characteristic of the dialogue does not meet desired criteria. One of the most important characteristics may be, for example, a measure of the intensity of the foreground speech (e.g., also referred to as dialogue). Key segments may be defined, for example, as segments in which the intensity of the dialogue is locally too low in an absolute sense and / or relative to the background sound, for example, defined by a specified threshold. As shown in [4], the relative level between the dialogue and the background may be, for example, an important factor in determining listening effort and intelligibility, although, for example, it is not the only factor. Other factors may include or may include unfamiliar vocabulary and / or accents, fluency of speech, sentence complexity, speech rate, slurred and / or ambiguous dialogue [5].

[0328] Various embodiments and aspects of the invention are described below. At least some of these embodiments relate to methods and / or apparatus for content creation, post-production, and / or quality control (QC), such as including or incorporating a novel method for providing supporting information, which may, for example, provide statistics on the number, severity, and / or temporal location of key segments. The supporting information may, for example, include other ways of describing key segments of an audio scene.

[0329] Other embodiments relate to methods and / or apparatuses for efficient audio delivery (e.g., broadcasting, streaming, file playback), optionally in conjunction with metadata, which may include, for example, statistical information regarding the number, severity, and / or temporal location of key segments. Based on the settings of the receiving or playing device, user interaction with the content, and / or supporting metadata, the receiving and / or playing device may, for example, optionally automatically and / or dynamically improve the intelligibility of the received content and / or reduce listening effort.

[0330] Other embodiments relate to methods and / or apparatus for encoding audio bitstreams, such as including supporting information, for example, regarding the timing and / or severity of key segments. The supporting information may be embedded in the audio bitstream and / or otherwise conveyed, for example, but not limited to, a transport layer descriptor, such as part of an MPEG-2 transport stream and / or the ISO Basic Media File Format (ISO BMFF).

[0331] The methods and / or apparatus described in the context of this application may, for example, use an MPEG-H audio system as an example of an audio system that can be enhanced with additional metadata and / or decoder processing to dynamically improve the levels of elements constituting an audio scene (e.g., when they are detected as critical). It should be noted that the methods and / or apparatus described according to embodiments are not limited to the MPEG-H audio system and may be used, for example, with other audio systems such as MPEG-D USAC, HE-AAC, E-AC-3, AC-4, etc.

[0332] In the following discussion, terms and definitions according to embodiments will be used: The following terms are used in the technical field:

[0333] ■ Audio Scene: For example, the entirety of all audio components that constitute a complete audio program. An audio component includes or even consists of an audio signal (e.g., often referred to as the audio essence) and associated metadata. A component can be, for example, an object with associated location metadata, and / or a channel signal, and / or a complete or partial mix in a particular channel layout (e.g., stereo), and / or a mixture of all these. For example, one (e.g., fundamental) part of an audio scene can be that it has associated metadata that can, for example, describe the audio component (e.g., including signal-related information such as location, and / or content-related information such as genre), define interactivity and / or personalization options, and / or define the relationships between components within the scene.

[0334] ■ Dialogue: For example, speech in the foreground of an audio scene can be referred to as dialogue, even if it may include or even consist of speakers who are not engaged in dialogue. Common examples are a speaker's narration or multiple overlapping speakers.

[0335] ■ NGA systems: For example, next-generation audio, such as MPEG-H audio.

[0336] ■ Audio Master: For example, a production format file that encapsulates, optionally, all audio essence (e.g., typically in an uncompressed format, such as PCM) and optionally all associated metadata.

[0337] It can be used during audio production and / or as an input format for audio encoding.

[0338] Examples of formats used include BWF / ADM, S-ADM in MXF, S-ADM in IMF, and MPEG-H control tracks.

[0339] Audio masters can be files in a file-based post-production workflow or linear data streams in a linear real-time production workflow.

[0340] The description of the methods according to the embodiments in this document may, for example, revolve around information carried in the final mix (typically but not necessarily uncompressed) and / or audio master file and / or audio bitstream. For content delivery, these methods are not limited to audio bitstreams and may be used, for example, with other delivery environments such as MMT, MPEG-2 transport streams, DASH-ROUTE, file formats for file playback, etc.

[0341] The following discussion will address the problem statement and current solutions: As stated in [1], creating sound that meets audience expectations and needs as well as national regulations is one of the major challenges in the production of television programs and films. In addition to dialogue, creating atmosphere and mood with music and other background sounds is a complex creative task. At the same time, audiences expect to fully understand the story and dialogue in a comfortable way, i.e., without requiring high listening effort.

[0342] A significant portion of viewers experience difficulty following audio on television. It is estimated that 90% of people over the age of 60 frequently or very often have problems understanding television audio [1]. This suggests that state-of-the-art solutions and regulations are insufficient to deliver audio scenarios that meet the needs and preferences of viewers. To address this issue, two main paths have been identified. The first path deals with support tools for the production, post-production, and QC stages. The second path deals with audio encoding / decoding and delivery to receiving devices.

[0343] Modern QC requirements and recommendations (see, for example [6]) include loudness specifications, such as:

[0344] • The overall average loudness of the program measured using a version of ITU-R BS.1770.

[0345] • Overall average dialogue gating loudness.

[0346] Peak value and true peak value.

[0347] Loudness range (LRA) refers to the entire program, and sometimes only to dialogue.

[0348] • Maximum short-term loudness (see, for example, EBU R128 s1) to avoid a strong deviation from the average loudness level.

[0349] These values ​​may, for example, fail to capture key local segments. In fact, even with very good dialogue LRA and dialogue gated loudness, background elements of the scene may obscure the dialogue.

[0350] As summarized in [7], some efforts have been made to develop recommendations for loudness difference (LD) between dialogue and background elements. However, these have not become common practice for two main reasons, as identified by the inventors: 1) they do not use short-term intensity measures, but rather loudness values ​​integrated throughout the program; 2) calculating LD is not straightforward when only the final mix is ​​available. Embodiments of the invention address these two problems by: 1) considering short-term measures, such as short-term and / or transient loudness difference between dialogue and background; 2) proposing scalable methods applicable to all production formats, from the final mix (e.g., at least one of mono, stereo, surround, immersive) to combinations of audio channels and / or audio objects commonly used in NGA systems, and / or any other audio formats used for the production and / or delivery of audio data.

[0351] Furthermore, tools exist for measuring local intelligibility. Embodiments of the invention (e.g., at least some) differ significantly from these tools, which can be used, for example, in a complementary manner to embodiments of the invention. One, or even a major, difference and / or novelty may be that, for example, embodiments of the invention may not measure intelligibility and may not necessarily consider, for example, the various higher-order cognitive factors that determine intelligibility, such as unfamiliar vocabulary and / or accents, fluency, sentence complexity, speech rate, phoneme pronunciation, ambiguous and / or unclear dialogue. Instead, it is suggested that low-level characteristics of audio signals be considered, such as their short-term energy ratio or short-term LD or transient LD (therefore, as an example, embodiments may be configured to consider low-level characteristics of audio signals, such as their short-term energy ratio or short-term LD or transient LD). These can, for example, be proxies of intelligibility and / or listening effort, but more importantly, they may optionally provide the possibility of direct improvement, such as locally altering the energy ratio or LD. If intelligibility is measured, the situation may be different, as the causes of low intelligibility can be very diverse and cognitive in nature.

[0352] Therefore, using low-level features of the audio signal (such as short-term and / or transient LDs) rather than abstract high-order, cognitively relevant intelligibility can, for example, have the following advantages: the audio signal can be modified, for example, based on measured audio signal features. Such modification can be accomplished, for example, directly by changing the audio mix, and / or indirectly by capturing modifications that can be applied as metadata in the receiving and playback devices.

[0353] At the same time, these low-level features can be a good estimate or proxy for the intelligibility of dialogue in audio mixes, for example, as shown in [4], for audio mixes made for television broadcasts and / or streaming content.

[0354] Modern audio codec systems offer features such as DRC (Dialogue Controlled Conversation). Furthermore, NGA (Natural Gauge Control) systems, for example, allow users to personalize speech levels for better intelligibility in various listening environments, but are limited by production constraints. While these options can help users better understand dialogue, they are insufficient to guarantee optimal experience quality in every situation and for every user. Using an NGA system, users can manually increase and / or decrease the level of dialogue during a program (e.g., a movie). Another option available in various audio codecs allows users to select a DRC profile. These two options—increasing the level of dialogue in an NGA system and selecting a specific DRC profile—can, for example, improve the intelligibility of problematic segments, but may also, for example, alter or even change the rest of the mix (where the content was originally perfectly intelligible) and / or even alter segments without speech. This is because they are typically applied statically to the entire program. For example, a key segment might appear, for example, in a movie's intro trailer. Users can optionally reduce the relative level of the background and / or select a DRC profile that contributes to intelligibility. In this way, users may, for example, or even better understand the dialogue during this segment. However, the background might be kept at a reduced and / or compressed level, at least partially, throughout the film, even though the content might be fine after the problematic intro. Users might, for example, or even will, experience the film without all the music, effects, and other elements designed in during its creation. In this case, a better understanding of the key scenes comes at the cost of sacrificing a complete and authentic enjoyment of the rest of the content.

[0355] Some NGA systems, such as MPEG-H audio, support passing dynamic metadata along with the audio data. However, how to set this metadata, i.e., how to derive the information written into it, is not yet defined. For example, one approach could be to set it manually during audio production, which could be time-consuming for sound producers, for instance.

[0356] Therefore, the inventors recognized that using information about the precise location of problematic segments—for example, for QC during production, followed by special attention from the sound producer and / or potential remixing of key locations, and / or during playback on the receiving device—could ensure that content is processed only during the problematic segments, rather than applied to the entire content. This may or even will preserve artistic intent while optionally reducing listening effort and / or improving end-user intelligibility.

[0357] Furthermore, creating multiple DRC sequences for various enhancement levels can, for example, lead to a significant increase in bit rate and / or a reduction in functionality based on the number of transmission sequences. Additionally, DRC is a global tool, meaning the selected DRC profile or sequence is applied to the entire program, so the entire program can be processed, for example. As mentioned above, it may not be suitable, for example, for situations where intelligibility issues only exist in certain problematic segments of the program and / or when it is desirable to avoid DRC processing for segments without dialogue.

[0358] Therefore, embodiments of the present invention propose a more efficient way to carry information about problem fragments, which optionally highly optimizes the transmission of information about problem fragments.

[0359] The following sections will introduce and discuss different implementations, covering various use cases:

[0360] First, optional details regarding the generation of QC reports are discussed: In one possible embodiment, a method and / or apparatus for generating QC reports for audio scenes are proposed, for example, as a support tool during production, post-production, and / or QC, and optionally providing, for example, automated improvement and / or support for human operators to improve the audio quality of the audio scenes. The method and / or apparatus may, for example, accept different production audio formats as input, such as final mixes of mono, stereo, surround, and / or immersive formats, and combinations of audio channels and / or audio objects commonly used in NGA systems. The generated QC report may, for example, include or contain information about critical segments, such as their location and / or severity. Possible formats for the QC report may include or contain, for example, human-readable text, and / or machine-readable formats (e.g., CSV, XML), and / or as visual information, and / or as control tracks, and / or combinations thereof. Optionally, for example, automated fixes may be proposed to enhance critical segments. For example, such a system of the present invention can be implemented as a standalone tool, and / or as part of an audio production suite, and / or integrated into a DAW, or as a VST plugin. For example, such a system of the present invention can be implemented in different ways, optionally as described in the embodiments below.

[0361] Next, we will discuss optional details regarding the generation of QC reports using a step-by-step approach (see, for example...). Figure 6 ). Figure 6 This diagram illustrates an example of generating a QC report using a step-by-step method according to one embodiment. It should be noted that... Figure 6 All elements of the analyzer 600 shown are optional.

[0362] therefore, Figure 6 A schematic diagram of an audio analyzer with optional features according to an embodiment of the present invention (e.g., according to a first aspect) is shown. Figure 6 An audio analyzer 600 is shown, including a measurement tool 610 and a quality control processor 620. Furthermore, the audio analyzer 600 includes an optional source separation unit 630. The source separation unit 630 may be provided with a final mix (e.g., audio content mixed with a speech portion and a background portion, optionally with attached metadata) to extract different portions of the final mix, such as speech (e.g., dialogue), background, and metadata, see 631. Alternatively, the analyzer 600 may be provided directly with different portions of the audio content. As another optional feature, the measurement tool 610 may also be configured to combine portions of the audio content, such as portions of the same type, such as speech or background, for subsequent determination of intensity measures (e.g., forming component groups). Optionally, the measurement tool 610 may be provided with parameters 611, such as integration time, to determine short-term intensity measures, such as short-term intensity information 612 for the speech portion and short-term intensity information 613 for the background portion, and / or to determine additional QC information 614. Optionally, as described above, the difference or ratio between the speech and background portions may be determined. As another optional feature, additional quality control information 614, such as information on overall loudness and / or true peak value, can be determined by measurement tool 610. Short-term intensity measures 612, 613, and the additional QC information 614, as exemplified herein, are provided to quality control processor 620 to determine critical segments 622. As another optional feature, QC processor 620 may be provided with thresholds or other criteria, see 621 (e.g., for determining and optionally classifying critical segments).

[0363] It should be noted that optional combinations of preprocessing (e.g., using source separation 630) and / or audio components can be implemented accordingly in neural network-based embodiments, such as... Figure 2 As shown, and according to Figure 3 In some embodiments, for example, it is used as a preprocessing step for the content 101 provided to the determiner 110.

[0364] In other words, in possible embodiments (e.g., such as...) Figure 6As shown, a method and / or apparatus for generating audio scene QC reports is proposed, for example, as a support tool during production, post-production, and / or QC, and / or to provide automated improvement and / or support for human operators to improve the audio quality of audio scenes. For example, this method and / or apparatus may embody or include, but is not limited to, [the following]. Figure 6 The described system (e.g., 600) includes the following parts:

[0365] ■ Optional measurement tool 610, for example, configured to analyze the input audio signal (e.g., 601, e.g., 631) and / or determine the short-term intensity of different audio components associated with, for example, dialogue type (e.g., speech, narration, audio description, etc.) and / or background type (e.g., music and effects, stadium atmosphere, etc.).

[0366] ○ The short-term intensity of the audio signal (e.g., 612, 613) can be calculated, for example, using a measuring tool (e.g., 610), by means of:

[0367] ■ Local power of the audio signal; and / or

[0368] ■ The power of the filtered signal, wherein the filter may optionally simulate the frequency-selective sensitivity of the human ear; and / or

[0369] ■ Short-term and / or instantaneous loudness according to ITU-R BS.1770[8] and / or EBU Recommendation R 128 or its variants, for example, using different time window sizes; and / or

[0370] ■ A model for calculating loudness; and / or

[0371] ■ AI-based intensity estimation.

[0372] In different embodiments, for example, multiple audio signals of the same or similar type may optionally be combined before measuring or determining an approximation of their short-term intensity. The process of combining the signals may be accomplished based on:

[0373] ■ The importance of multiple audio signals, wherein the importance of the audio signals can be set manually and / or determined based on the speech portions included or contained in the signals; and / or

[0374] ■ The contribution of each audio signal to the final mix, taking into account audio masking effects and the characteristics of the human auditory system.

[0375] ■ Optional QC processor module 620, for example, is configured to receive information about the short-term intensity of the audio signal (e.g., 612, 613) and / or optionally receive decision criteria (e.g., 621) for detecting key segments in the audio signal as input.

[0376] In a particular embodiment, the decision criteria may include, for example, at least a threshold, see, for example, 621, and optionally, local intensity differences (e.g., between dialogue and background, e.g., 612, 613) may be compared to this threshold. For example, all portions of the audio signal with local intensity differences less than the threshold may be marked as key segments in the audio signal, e.g., 622.

[0377] ■ In different embodiments, multiple thresholds may be used, for example, for different parts of the audio signal and / or different types of signals.

[0378] ■ In different embodiments, for example, a frequency-based threshold may be optionally set according to the frequency-selective sensitivity of the human ear and / or other psychoacoustic models.

[0379] ■ In different embodiments, for example, the standard may be based on an AI-based module (e.g., based on a DNN) that will or may detect key segments.

[0380] For example, the detection results may or will return information about key segments, which may include, but are not limited to:

[0381] ■ The beginning of each key segment in the audio signal; and / or

[0382] ■ The end and / or duration of each key segment in the audio signal; and / or

[0383] ■ The level of criticality, which may be associated, for example, with the level of comprehension and / or listening effort required to understand key passages; and / or

[0384] ■ It can be used, for example, to provide additional information to support the production, post-production, and / or QC stages.

[0385] ■ Optional source separation module 630 (e.g., may be AI-based, as in [9]), for example, is configured to estimate dialogue and / or background elements given a mix of dialogue and / or background elements to be used if the separated dialogue and / or background (e.g., the corresponding part of the audio content) cannot be obtained from the production (see, for example, 631).

[0386] For example, the following threshold can be set regarding short-term loudness difference.

[0387] ■ For example, short-term loudness differences (ST-LD) below 4 LU (e.g., loudness units) can be marked as very critical (and / or, for example, marked in red), short-term loudness differences between 4 LU and 10 LU can be marked as slightly critical (and / or, for example, marked in yellow), and those above 10 LU can be considered non-critical (and / or, for example, marked in blue). These example values ​​follow [7], but it is important to note that in [7] and related work, loudness values ​​are integrated over the entire audio program. In contrast, embodiments according to the invention include or focus on local (i.e., short-term) intensity differences.

[0388] ■Example output visualizations can, for example, display ST-LD changing over time, such as... Figure 7 As shown. Therefore, Figure 7 A schematic example visualization of Short-Term Loudness Difference (ST-LD) for Quality Control (QC) according to an embodiment of the present invention is shown, wherein segments are labeled as Very Critical 710, Slightly Critical 720, and Not Critical 730. For example, an audio analyzer can provide, as Figure 7 The visualization shown.

[0389] ■ In addition, summaries can be generated, for example, reporting the percentage of key segments. For example, the location of key segments can be displayed or stored in an output file that can be imported by another DAW or production tool.

[0390] As another example, QC processor 620 can, for instance, be configured to detect a critical segment 622 if the absolute level of the speech (whose information may be contained in 612) deviates locally from the integrated or average speech loudness over the entire program by more than a threshold (e.g., 10 LU). Specifically, a combination of absolute speech loudness deviation and short-term speech loudness relative to the background can be used. Thus, measurement tool 610 can be configured to provide average speech information to the QC processor, and threshold 621 can contain corresponding information about the criteria used for evaluation.

[0391] Next, reference is made to aspects of the invention according to embodiments concerning the generation of QC reports using a step-by-step method, wherein the measuring tool directly outputs the short-term strength difference instead of the short-term strength (see, for example...). Figure 8 ). Figure 8 A schematic diagram of an audio analyzer for optionally determining short-term intensity differences is shown according to one embodiment. In other words, Figure 8 This diagram illustrates an audio analyzer that generates a QC report using a step-by-step method according to one embodiment, where the measurement tool directly outputs the short-term intensity difference instead of the short-term intensity. It should be noted that... Figure 8 The elements of the audio analyzer 800 shown are optional. Analyzer 800 includes... Figure 6The elements discussed in the context of the present invention, but with a measurement tool 810 configured to provide short-term intensity differences (e.g., short-term intensity differences between speech and background) to the QC processor 820 to obtain key segments 622.

[0392] Next, referring to aspects of the invention regarding the one-step generation of QC reports according to embodiments: In such embodiments, key segments can optionally be estimated directly from the audio input, for example, via an end-to-end detector module. For example, the detector module can be an artificial neural network (ANN), optionally trained using a previous embodiment as a teacher, such as... Figure 9 and Figure 10 As shown.

[0393] Figure 9 A schematic diagram of an audio analyzer with a detector according to an embodiment of the present invention is shown. Specifically, Figure 9 This diagram illustrates an example of generating a QC report in one step according to one embodiment. It should be noted that all elements of the analyzer 900 are optional. Figure 9 An audio analyzer 900 is shown, including a detector 910 and a measurement tool 920. As shown, the detector 910 can be configured to acquire key segments 622 based on audio content in the form of speech portions (e.g., dialogue) and background portions (e.g., optionally with additional metadata information, see 631). Based on the audio content 631, the measurement tool 910 can be configured to provide optional additional QC information 614. Similarly, the analyzer 900 may optionally include a source separation unit 630 for extracting information 631 from the final mix 601.

[0394] Figure 10 A schematic diagram of an audio analyzer with two detectors according to an embodiment of the present invention is shown. Specifically, Figure 10 This illustrates an example of generating a QC report in one step without explicit source separation, according to one embodiment. It should be noted that all elements of analyzer 1000 are optional. This method can be used in conjunction with measurement tools for providing additional QC information, such as 614. Figure 10 The audio analyzer 1000 shown includes a first detector 1010 and a second detector 1020. For example, the appropriate detector can be selected based on the format or structure of the input, such as whether it provides the final mix 601 or different parts of the audio content, such as speech and background parts, as shown in 631.

[0395] Next, referring to aspects of the invention according to embodiments regarding deriving QC metadata from QC reports and encapsulating (e.g., for encapsulation) it into an audio master: embodiments include or incorporate generating QC metadata structures and data packets, the metadata attributes of which can be derived (e.g.) from QC reports.

[0396] refer to Figure 11 . Figure 11 A schematic diagram of an audio processor with optional features according to an embodiment of the present invention is shown. In particular, Figure 11 This diagram illustrates an example of deriving QC metadata from a QC report and encapsulating it (e.g., for encapsulation) into an audio master, according to an embodiment. It should be noted that all elements of the processor 1100 are optional.

[0397] The audio processor 1100 includes a measurement tool 1110, a key segment detector 1120, a metadata processor and embedder 1130, and an audio master 1140. Audio content 1001 is... Figure 11 The signal is indicated as an audio signal and provided to the measurement tool 1110. As an optional feature, the measurement tool 1110 may be provided with a set of parameters 611, such as integration time. The measurement tool 1110 is configured to determine the short-term intensity 1112 of the speech portion and the short-term intensity 1113 of the background portion, and provide them to the key segment detector 1120 (optionally, as shown in the image). Figure 8 The difference can be determined, or the processing can be performed based on AI.

[0398] Furthermore, as an optional feature, measurement tool 1110 is configured to provide additional QC information 1114, such as information about the overall loudness or true peak (e.g., corresponding to the processors 620, 820 discussed previously), to the key segment detector 1120 and the metadata processor and embedder 1130, respectively. According to the above embodiments, the key segment detector 1120 may optionally be provided with a threshold or other criterion 621 for determining the key segment 622, which is provided to the metadata processor and embedder 1130. The metadata processor and embedder 1130 is configured to use the key segment 622 and the optional additional QC information 1114 to determine metadata information 1131, such as QC / signal descriptive metadata. The metadata information 1131 may, for example, be encapsulated in a data structure and may optionally be embedded in the metadata of the audio master 1140 to provide an audio master file or stream 1002.

[0399] Therefore, in one possible embodiment (e.g., as...) Figure 11As shown, a method and / or apparatus are proposed for creating metadata structures and / or metadata data packets, which may, for example, encapsulate quality control and / or signal descriptive metadata, for example, to drive enhancement of an audio scene. These metadata structures may, for example, be embedded in an original file containing or including an audio signal and / or a newly created audio master file. For example, in all cases, the metadata structure may, for example, accompany the audio signal, and the audio signal itself may, for example, remain unchanged.

[0400] refer to Figure 12 . Figure 12 A schematic diagram of a second audio processor with optional features according to an embodiment of the present invention is shown. Specifically, Figure 12 This diagram illustrates an embodiment of deriving QC metadata from a QC report and encapsulating it (e.g., for encapsulation) into an audio master, according to an embodiment. It should be noted that all elements of the audio processor 1200 are optional. The audio processor 1200 is configured to receive an audio master file or stream 1201, and is configured to (e.g., with...) Figure 11 (Compared) Extract the audio signal 1241 and provide it to the measuring tool 1110.

[0401] Therefore, in another possible embodiment (e.g., as Figure 12 As shown, a method and / or apparatus are proposed for creating metadata structures and / or metadata data packets, which may, for example, encapsulate quality control and / or signal descriptive metadata, for example, to drive audio scene enhancement. These metadata structures (e.g., 1131) may, for example, be embedded in an audio master file (e.g., 1202) that contains or includes audio signals and existing audio scene metadata. The quality control and / or signal descriptive metadata structures are added to the audio scene metadata (which may be added to the audio scene metadata), while the audio signal itself may, for example, remain unchanged.

[0402] refer to Figure 13 . Figure 13 A schematic diagram of a third audio processor with optional features according to an embodiment of the present invention is shown. Specifically, Figure 13 This diagram illustrates an embodiment of deriving audio elements and metadata and encapsulating (e.g., for encapsulation) them into an audio master. It should be noted that all elements of the audio processor 1300 are optional. Figure 11 Compared to the audio processor shown, audio processor 1300 includes an optional source separation module 1310 (e.g., corresponding to element 630), which is configured to provide separated signals (e.g., dialogue, background, and metadata) to measurement tool 1110 and audio master 1140 based on audio signal 1001, in order to provide an audio master file or stream 1302.

[0403] Therefore, in another possible embodiment (e.g., as Figure 13 As shown, a method and / or apparatus for creating metadata structures and / or metadata data packets are proposed. These data packets may, for example, encapsulate quality control and / or signal descriptive metadata, for example, to drive enhancement of an audio scene. These metadata structures (e.g., 1131) may, for example, be embedded in a newly created audio master file (e.g., 1302). Furthermore, the method and / or apparatus may, for example, include or contain a source separation module (e.g., 1310), which may, for example, be configured to estimate dialogue and / or background elements, given a mixture of dialogue and / or background elements to be used, if the separated dialogue and / or background (e.g., corresponding portions of audio content) is, for example, not available from the production. These dialogue and / or background elements may, for example, be extracted from an input audio signal (e.g., 1001), and new audio signals may, for example, be created for these dialogue and / or background elements. These newly created audio signals may also be used, for example, by a measurement tool (e.g., 1110), and may, for example, be optionally embedded in an audio master file, optionally along with, for example, a metadata structure that may accompany the audio signal.

[0404] Such methods and / or devices may, for example, embody or include, but are not limited to, [the following]. Figure 11 , 12 The system described in section 13 may optionally include:

[0405] ■ Optional measurement tools, such as 1110, are configured to analyze the input audio signal, such as 1001, such as 1241, and / or determine the short-term intensity (e.g., short-term intensity) of different audio components associated with dialogue type (e.g., speech, narration, audio description, etc.) and / or background type (e.g., music and effects, stadium atmosphere, etc.).

[0406] ○ The short-term intensity of an audio signal can be calculated, for example, using measurement tools, such as:

[0407] ■ Local power of the audio signal; and / or

[0408] ■ The power of the filtered signal, wherein the filter may optionally simulate the frequency-selective sensitivity of the human ear; and / or

[0409] ■ Short-term and / or transient loudness, for example, according to ITU-R BS.1770 [8] and EBU Recommendation R 128 or its variants, for example, using different time window sizes; and / or

[0410] ■ A model for calculating loudness; and / or

[0411] ■ AI-based intensity estimation.

[0412] In different embodiments, multiple audio signals of the same or similar type may be optionally combined, for example, before short-term intensity measurements. The process of combining signals may be accomplished, for example, based on the following:

[0413] ■ The importance of multiple audio signals, wherein the importance of the audio signals can be, for example, manually set and / or determined based on the speech portions contained in and / or included in the signals; and / or

[0414] ■ For example, the contribution of each audio signal to the final mix, for instance, considering audio masking effects and / or the characteristics of the human auditory system.

[0415] ■ An optional key segment detection module, such as 1120, may be configured, for example, to receive information about the short-term intensity of the audio signal (e.g., 1112, 1113, or its ratio or difference) and / or decision criteria for detecting key segments in the audio signal (see, for example, 621) as input.

[0416] In a particular embodiment, the decision criteria may include, for example, at least a threshold, and the local intensity difference may be compared, for example, to that threshold. Optionally, all portions of the audio signal with local intensity differences less than the threshold may, for example, be marked as key segments in the audio signal.

[0417] ■ In different embodiments, multiple thresholds may be used, for example, for different parts of the audio signal and / or different types of signals.

[0418] ■ In different embodiments, for example, a frequency-based threshold may be optionally set according to the frequency-selective sensitivity of the human ear and / or other psychoacoustic models.

[0419] ■ In different embodiments, the standard may be based, for example, on an AI-based module (e.g., based on a DNN), which may detect key segments.

[0420] For example, the detection results may or will return information about key segments, which may include, but are not limited to:

[0421] ■ The beginning of each key segment in the audio signal; and / or

[0422] ■ The end and / or duration of each key segment in the audio signal; and / or

[0423] ■ The level of importance, which may be associated, for example, with the level of intelligibility and / or listening effort required to understand key passages; and / or

[0424] ■ For example, additional information that can be used to trigger different processes in receiving and / or playback devices may include, for example, filter coefficients, gain values, etc.

[0425] ■ An optional QC metadata processor, such as 1130, may be configured, for example, to format information received from the key segment detection module into QC metadata packets optionally aligned with the audio data frame rate, and to provide them, for example, to the QC metadata embedder.

[0426] ■ An optional QC metadata embedder, such as 1130, can be configured, for example, to encapsulate quality control and signal descriptive metadata in packets and / or data structures, and optionally insert them into audio master files and / or streams. These structures may, for example, accompany the audio signals they refer to, and the QC metadata embedder may, for example, not modify these audio signals. Depending on the audio master file / stream format, different encapsulation methods may be used, for example, ADM data structures and / or control track data structures. These data structures and packets may, for example, be encapsulated in various audio master file and / or stream formats as static, file-level data (e.g., in the case of a complete QC report for a complete program), and / or dynamic, time-varying data (e.g., in the case of time-varying metadata), and / or real-time streaming audio master formats.

[0427] It should be noted that embodiments may include, for example, only a QC metadata processor to provide QC metadata as output, for example, without embedding it into a file or stream. Furthermore, the QC metadata processor and the QC metadata embedder may be implemented, for example, as separate processing units.

[0428] Next, referring to aspects of the invention relating to converting QC metadata from an audio master to a bitstream in an audio encoder: a device (e.g., a bitstream provider) according to an embodiment may be configured, for example, to read metadata from an audio master and optionally convert the metadata into a bitstream metadata structure, and optionally, for example, to embed the bitstream metadata structure into the audio bitstream during audio encoding.

[0429] Furthermore, referring to aspects of the invention concerning the enhancement of audio scenes based on additional QC metadata: for example, the following embodiments describe alternative methods for altering audio elements in an audio scene at the decoder and / or system level using supporting metadata and / or metadata manipulation mechanisms in the receiving device, for example, to automatically and / or dynamically improve the understanding of received content, and / or, for example, to reduce listening effort for received content.

[0430] refer to Figure 14 . Figure 14 A schematic diagram of a bitstream provider with optional features according to an embodiment of the present invention is shown. Specifically, Figure 14This diagram illustrates an example system architecture for using metadata to enhance an audio scene (encoder side) according to an embodiment. It should be noted that all elements of the bitstream provider 1400 are optional. Figure 14 A bitstream provider 1400 is shown, including an audio encoder 1410 (e.g., an encoding unit, such as an encoder core or encoding core). As optional features, the bitstream provider 1400 includes a measurement tool 1110, a key segment detector 1120, and a metadata processor 1430. The bitstream provider 1400 is configured to include an encoded representation of audio content 1001 and quality control information 1431 (e.g., in the form of QC metadata) into a bitstream 1402.

[0431] Next, refer to Figure 15 . Figure 15 A schematic diagram of an audio decoder with optional features according to an embodiment of the present invention is shown. Specifically, Figure 15 This diagram illustrates a system architecture for enhancing an audio scene using metadata (decoder side) according to an embodiment. It should be noted that all elements of the decoder 1500 are optional. Figure 15 An audio decoder 1500 is shown, including an audio bitstream parser 1510 and a decoding unit 1520 (which, as an optional feature, is configured to render received audio data, such as a decoder core). Furthermore, the audio decoder 1500 includes an optional quality control processor 1530. Thus, the parser 1510 may be provided with a bitstream 1501 to extract audio data 1513 for the decoder 1520, and metadata 1511 (e.g., additional metadata, such as audio, loudness, and DRC metadata) and 1512 (e.g., QC metadata) for the processor 1530. As additional optional inputs, the processor 1530 may be provided with setup information 1503 (e.g., system and user settings) and environmental information 1504 to provide quality control information 1531 to the decoding unit 1520. The decoding unit 1520 may be configured to enhance the audio data 1513 based on the QC information 1531 to obtain improved audio 1502.

[0432] Therefore, in other words, the example system architecture using the system of the present invention according to the embodiments is as follows: Figure 14 (Encoder / Transmitter Side) and Figure 15 (Decoder / Receiver Side) is shown.

[0433] Referring to aspects of the invention concerning a method and / or apparatus for creating an audio bitstream containing QC metadata according to embodiments: In one possible embodiment, a method and / or apparatus for creating an audio bitstream containing, for example, quality control and / or signal descriptive metadata, is proposed, for example, to drive enhancement of an audio scene. Such a method and / or apparatus may embody or include, but is not limited to, these aspects. Figure 14 The system described herein includes the following parts:

[0434] ■ Optional measurement tools, such as 1110, are configured to analyze the input audio signal and optionally determine the short-term intensity (e.g., short-term intensity 1112, 1113, e.g., their difference or ratio) of different audio components associated with dialogue type (e.g., speech, narration, audio description, etc.) and / or background type (e.g., music and effects, stadium atmosphere, etc.).

[0435] ○ The short-term intensity of an audio signal can be calculated, for example, using measurement tools, such as:

[0436] ■ Local power of the audio signal; and / or

[0437] ■ The power of the filtered signal, wherein the filter may optionally simulate the frequency-selective sensitivity of the human ear; and / or

[0438] ■ Short-term and / or transient loudness, for example, according to ITU-R BS.1770 [8] and / or EBU Recommendation R 128 or its variants, for example, using different time window sizes; and / or

[0439] ■ A model for calculating loudness; and / or

[0440] ■Intensity estimation based on AI.

[0441] In different embodiments, multiple audio signals of the same or similar type can be combined, for example, before measuring short-term intensity. The process of combining signals can be accomplished, for example, based on the following:

[0442] ■ The importance of multiple audio signals, wherein the importance of an audio signal can be determined, for example, manually set and / or based on the speech portion contained in or included in the signal; and / or

[0443] ■ The contribution of each audio signal to the final mix, for example, considering audio masking effects and / or the characteristics of the human auditory system.

[0444] ■ An optional key segment detection module, such as 1120, may be configured, for example, to receive information about the short-term intensity of the audio signal and / or decision criteria (e.g., 621) for detecting key segments in the audio signal (e.g., 622) as input.

[0445] In a particular embodiment, the decision criteria may include, for example, at least a threshold, and the local intensity difference may be compared, for example, to the threshold. For example, all portions of an audio signal with a local intensity difference less than the threshold may be, for example, marked as key segments in the audio signal.

[0446] ■ In different embodiments, multiple thresholds may be used, for example, for different parts of the audio signal and / or different types of signals.

[0447] ■ In different embodiments, for example, a frequency-based threshold may be optionally set according to the frequency-selective sensitivity of the human ear and / or other psychoacoustic models.

[0448] ■ In different embodiments, the standard may be based, for example, on an AI-based module (e.g., based on a DNN), which may, for example, detect key segments.

[0449] ○ The detection results may, or will, return information about key segments, including but not limited to:

[0450] ■ The beginning of each key segment in the audio signal; and / or

[0451] ■ The end and / or duration of each key segment in the audio signal; and / or

[0452] ■ The level of criticality can be, for example, associated with the level of intelligibility and / or listening effort required to understand key passages; and / or

[0453] ■ Additional information, such as filter coefficients and gain values, can be used to trigger different processes in receiving and / or playback devices.

[0454] ■ An optional QC metadata processor, such as 1430, may be configured, for example, to format information received from the key segment detection module into QC metadata packets, for example, aligned with the audio data frame rate, and / or to provide the QC metadata packets to the audio encoder.

[0455] QC metadata may not be used directly during audio production, but may be encapsulated in data packets and optionally inserted into the audio bitstream during encoding. Depending on the audio codec used, different encapsulation methods may be used, optionally based on the codec's capabilities; for example, MHAS packets may be used in the case of MPEG-H audio, or extended payloads may be used in the case of MPEG AAC, XHE-AAC, or USAC.

[0456] In different embodiments, the measurement tool, key segment detection module (e.g., 1120), and QC metadata processor component (e.g., 1430) may, for example, be combined into a single intelligibility processor. In a further embodiment, the intelligibility processor may, for example, be configured to derive QC metadata (e.g., 1431) to provide to the audio encoder, for example, using an AI-based solution, optionally trained using an audio dataset, which may include, but is not limited to:

[0457] ■ Audio scenes marked as difficult to understand / easy to understand; and / or

[0458] ■ Marked as audio scenarios requiring low / medium / high listening effort to understand.

[0459] Next, referring to aspects of the invention concerning methods and / or apparatus for receiving audio bitstreams containing quality control metadata: In one possible embodiment, a method and / or apparatus is proposed for receiving audio bitstreams containing, for example, QC metadata and optionally enhancing (e.g., for enhancing) the audio scene (therefore, embodiments include such apparatus or such methods). Such methods and / or apparatus may, for example, embody or include, but are not limited to, [examples of methods]. Figure 15 The system described herein includes the following parts:

[0460] ■ An optional bitstream parser, such as 1510, may be configured, for example, to extract metadata embedded in the audio bitstream (e.g., 1511, 1512), optionally decode and / or dequantize the metadata when necessary, and optionally provide the results to a quality control processor, such as 1530.

[0461] ○ The metadata provided to the quality control processor may include, for example, at least one of the following:

[0462] ■QC metadata,

[0463] ■Audio metadata

[0464] ■ Loudness and / or DRC metadata,

[0465] ■ Other available metadata

[0466] ■ Information about audio frames and / or frame groups and / or audio samples that may optionally correspond to available metadata.

[0467] In one specific embodiment, the QC metadata, such as 1512, may optionally include or contain information about key segments, which may include, for example, but is not limited to:

[0468] ■ The beginning of one or alternatively each key segment in the audio signal; and / or

[0469] ■ The end and / or duration of one or even every key segment in the audio signal; and / or

[0470] ■ The level of criticality can be, for example, associated with the level of intelligibility and / or listening effort required to understand key passages; and / or

[0471] ■ Additional information, which may be used, for example, to trigger different processes in the receiving and / or playback devices, may optionally include filter coefficients, gain values, etc.

[0472] ■ An optional quality control processor, such as 1530, may be configured, for example, to receive information from the system level (e.g., 1503, 1504) in addition to information received from the audio bitstream parser (e.g., 1510).

[0473] Information from the system level may include at least one of the following:

[0474] ■ Information about system settings, such as permanent settings for the receiving device (e.g., preferred language, dialogue enhancement options, options for viewers with hearing or visual impairments, etc.).

[0475] ■ User settings during the current program, such as interacting with the audio scene by increasing the level of a specific audio object and / or changing the position of the audio object.

[0476] ■ Information about the environment received from different sensor interfaces of the receiving device (e.g., microphone, lavalier microphone, camera, GPS location information). For example, if the content is consumed on a mobile device in a noisy environment (such as a bus or home).

[0477] ■ Information about additional devices connected to the receiving device, such as if the television uses external sound devices or internal television speakers to reproduce sound.

[0478] Based on metadata received from an audio bitstream parser (e.g., 1510) and information received from the system level (e.g., 1503, 1504), the quality control processor (e.g., 1530) may perform at least one of the following operations:

[0479] ■ Determine if any audio signal contains key segments that need improvement.

[0480] ■Determine the level and / or intensity of the improvement to be applied.

[0481] ■ To obtain the quality control information needed for the audio decoder to enhance the audio quality of key segments (e.g., to improve intelligibility and / or reduce listening effort).

[0482] ○ Quality control information, such as 1531, may include, for example, at least one of the following:

[0483] ■ One or more gain sequences, which may be applied to, for example, one or more portions of an audio signal in an audio scene.

[0484] ■ Information about which audio signals might need to be processed, for example, to improve key segments.

[0485] ■ Information regarding the duration of key segments.

[0486] ■ An optional audio decoder (and optional renderer), such as 1520, may be configured, for example, to receive quality control information (e.g., 1531) from a quality control processor (e.g., 1530) and optionally apply it to audio signals that may, for example, need to be improved to obtain better intelligibility (e.g., 1513).

[0487] ■ For example, this might be, for instance, a typical case of audio delivered as a full mix, or in the case of NGA, if the audio stream contains only static metadata.

[0488] In another embodiment, a quality control processor (e.g., 1530) may be configured, for example, to receive information from the system level regarding current preferences, device settings, and / or listening environment (e.g., noisy or quiet) (e.g., 1503), and optionally may receive additional information about the current playback situation, such as the user's personal circumstances (e.g., known to the device by preferences and / or by the device's sensors and / or from other devices such as hearing aids) and / or the time of day (e.g., night or day). Based on this information, the quality control processor may, for example, determine whether information about key segments in the audio content should be evaluated, and apply improvements based on quality control metadata (e.g., 1512), which may optionally be adapted to the current situation, for example, as described by additional information optionally from the system level. For example, in the following cases:

[0489] - The device has the following activity settings for: enabling conversation enhancement, improving conversation intelligibility, hearing impairment and / or other settings that can be used to improve intelligibility and / or reduce listening effort, and / or

[0490] - The device's sensor interface indicates a noisy environment, and / or

[0491] - If the user has actively chosen the option to improve intelligibility and / or reduce listening effort, the quality control processor may, for example, or even will trigger the improvement. Otherwise, the improvement may, for example, or even will not be triggered.

[0492] In another embodiment, the quality control processor may, for example, or even use information from quality control metadata to indicate whether an improvement should be triggered.

[0493] In another embodiment, the quality control processor may be configured, for example, to derive quality control information based on the information described in the foregoing embodiments, and (or wherein) the quality control information includes or contains one or more filter coefficients.

[0494] refer to Figure 16 . Figure 16 A schematic diagram of an audio decoder with a filter according to an embodiment of the present invention is shown. Figure 15 Compared to the embodiment shown, the decoder 1600 includes separate decoding unit 1620 and rendering unit 1640, with an enhancement filter 1630 implemented between them.

[0495] like Figure 16 As shown, the enhancement filter 1630 can, for example, use quality control information 1531 to process one or more decoded audio signals, for example, to improve one or more audio signals, optionally before one or more audio signals are rendered together into the final audio output.

[0496] therefore, Figure 16 An example system architecture for using metadata to enhance the audio scene (decoder side) according to an embodiment can be shown. It should be noted that all elements of the audio decoder 1600 are optional.

[0497] In different embodiments, the enhancement filter (e.g., 1630) may be applied directly to the final rendered output, for example. In different embodiments, the enhancement filter (e.g., 1630) may be used, for example, before and after the final rendered output, optionally based on available quality control information.

[0498] Next, referring to aspects of the invention concerning the enhancement of audio scenarios based on additional metadata on different channels according to embodiments: In another embodiment, the above-described method and / or device may, for example, use different channels to transmit additional metadata. For example, the additional metadata may be embedded in new descriptors (e.g., for MPEG-2 transport streams) and / or file format frames (for ISOBMFF). Figure 17 and 18 This describes a workflow according to an embodiment, in which QC metadata is embedded in a dedicated bearer mechanism of the transport layer.

[0499] Figure 17 A schematic diagram of a bitstream provider with an optional multiplexer according to an embodiment of the present invention is shown. Figure 14Compared to the illustrated embodiment, the bitstream provider 1700 includes an additional, optional multiplexer 1710, and QC metadata 1431 is provided to the multiplexer 1710 instead of the encoding unit 1410. The multiplexer 1710 is configured to provide a transport stream 1711 based on the audio bitstream 1702 and the QC metadata 1431. Therefore, in other words, Figure 17 This diagram illustrates an example system architecture that uses metadata at the transport layer to enhance audio scenarios (encoder side) according to an embodiment. It should be noted that all elements of the bitstream provider 1700 are optional.

[0500] Figure 18 A schematic diagram of an audio decoder with an optional demultiplexer according to an embodiment of the present invention is shown. Figure 16 Compared to the illustrated embodiment, the audio decoder 1800 includes an additional, optional demultiplexer 1810, configured to extract audio bitstream 1701 and QC metadata 1131 from transport stream 1702. Therefore, in other words, Figure 18 This diagram illustrates an example system architecture that uses metadata at the transport layer to enhance the audio scene (decoder side) according to an embodiment. It should be noted that all elements of the audio decoder 1800 are optional.

[0501] As an illustrative example, regarding Figures 6 to 18 It should be noted that measuring tools 610, 810, 920, and 1110 may include the same, similar, or corresponding features and functions. Specifically, measuring tools 610, 810, 920, and 1110 may correspond to or be similar to... Figure 1 as well as Figure 3 The example shown is of a short-term intensity determiner 110 (or at least a portion thereof). Quality control processors 620, 820 and key fragment detector 1120 may correspond to or be examples of analysis result provider 120. Therefore, the analysis results may include information about key fragments.

[0502] Therefore, in Figures 6 to 18 Any details, features, and functions discussed in the context may be incorporated (e.g., directly, similarly, or correspondingly) into such Figure 1 and Figure 3 The embodiments discussed herein.

[0503] Accordingly, in Figures 6 to 18 Any details, features, and functions discussed in the context may be incorporated (e.g., directly, similarly, or correspondingly) into such Figure 2 , Figure 3 and Figure 4 In the embodiments discussed herein, for example, Figure 9 and Figure 10 Show Figure 2Possible AI-based implementations of the embodiments; Figure 11 , Figure 12 and Figure 13 Show Figure 3 Possible implementations of the embodiments; Figure 14 and Figure 17 Show Figure 4 Possible implementations of the embodiments; Figure 15 , Figure 16 and Figure 18 Show Figure 5 Possible implementations of the embodiments.

[0504] Therefore, as an example, the audio master 1140, the key segment detector 1120, and the metadata processor and embedder 1130 can correspond to or be examples of the metadata provider 320 and the file or stream provider 330.

[0505] As another example, audio bitstream 1501 may correspond to an example of encoded media representation 501, quality control processor 1530 may correspond to an example of quality control information provider 510, and audio decoder 1520 may correspond to an example of decoded audio representation provider 520.

[0506] Therefore, for the sake of brevity, elements with the same or similar names or the same or similar reference numerals may include the same, similar, or corresponding features and functions. For the sake of brevity, embodiments are explained by way of example. Therefore, it should be noted that any combination of the respective features of the above embodiments can be performed.

[0507] Next, reference will be made to aspects of the invention used in the examples of the embodiments:

[0508] Alternatives or examples of the above-mentioned "quality control metadata" may include one or more of the following: clarity information metadata, accessibility enhancement metadata, speech transparency metadata, speech enhancement metadata, intelligibility metadata, content description metadata, local enhancement metadata, and signal description metadata.

[0509] In addition, MPEG-H 3D audio systems can be used, for example, to carry QC metadata and enhance the decoded and rendered audio based on the QC metadata to achieve better intelligibility and reduced listening effort.

[0510] the following Figures 19 to 21 Examples of the corresponding syntax that may be implemented according to the embodiments are shown.

[0511] For example, a new MHAS package can be defined to carry QC metadata: Reference Figure 19 (For example, also known as the syntax of Table 1 — MHASPacketPayload()) and Figure 20(For example, also known as the value of Table 2 — MHASPacketType).

[0512] The following new definitions may be added, for example, to the MPEG-H 3D audio standard or any other audio standard (where the names of packets, packet types, and information items may be optionally chosen as appropriate, for example, in the terminology of their respective standards).

[0513] Next, we will discuss an example of PACTYP_QUALITYCONTROL according to an embodiment: MHASPacketTypePACTYP_QUALITYCONTROL can, for example, be used to embed information about audio quality control metadata available in the audioQualityControlInfo() structure, and provide quality control information data to the decoder in the form of an audioQualityControlInfo() structure.

[0514] For example, if present, for each random access point and streaming access point, MHASPacketType PACTYP_QUALITYCONTROL should follow PACTYP_MPEGH3DACFG.

[0515] Updated audio quality control information can be available, for example, between two random access points, in which case the quality control information is associated with the next MHAS packet of type PACTYP_MPEGH3DAFRAME. MHASPacketType PACTYP_QUALITYCONTROL can be used, for example, to pass the updated audio quality information to the decoder without reconfiguring the audio decoder.

[0516] Next, we will discuss an example of audio quality control according to the embodiments:

[0517] In general, it is important to note that audio quality control metadata, for example, is used to signal key parts (or key segments, or key sections) of the audio signal to improve audio quality, thereby achieving better intelligibility and reducing listening effort.

[0518] Next, refer to an example of the syntax for this type of audio quality control: Figure 21 (For example, also known as Table 334) lists the syntax of audio quality metadata. Figure 21 Specifically, it can be referred to as the syntax of Table 334 — audioQualityControlInfo().

[0519] The qcInfoCount field, for example, indicates the number of available structures carrying audio quality control information in the stream.

[0520] The `qcInfoActive` flag, for example, signals when audio quality control information should be applied. Based on the value of the `qcInfoActive` flag, audio quality control information can be decoded and applied to the audio scene according to receiver settings.

[0521] The qcInfoType field indicates, for example, whether the subsequent qcInfo() block points to a specific audio element (mae_groupID) or to an audio scene defined by, for example, a combination of audio elements (mae_groupPresetID).

[0522] Alternatively, or additionally, a new extension element can be defined, for example, in mpegh3daConfigExtension() or usacConfigExtension() or an extension element (usacExtElement).

[0523] Therefore, optionally and generally, the bitstream provider according to the embodiments can be configured to include quality control information in the extended payload of the bitstream, such as... Figure 19 As indicated in [the document]. Furthermore, as [the document states]. Figure 19 As shown, the bitstream provider according to the embodiment can be configured to include quality control information into the MHAS packet.

[0524] Therefore, the implementation can be carried out as a standalone QC tool or as an extension of existing audio codec standards.

[0525] Further embodiments: According to embodiments of the present invention, the following examples are provided:

[0526] Example 1: A method for decoding a bitstream that includes or contains an audio scene and controlling the improved level of the audio scene, comprising:

[0527] Receive a bitstream that includes or contains encoded audio data, the encoded audio data including one or more audio signals that contain or contain at least two different audio types that can be characterized as, for example, dialogue and / or narration and / or background and / or music and effects;

[0528] Specifically, at least two different audio types can be included or contained in at least two different audio signals, for example, a stereo channel signal with music and effects and a mono dialogue audio object; or they can be included or contained in the same one or more audio signals, for example, a stereo full tone containing a mixture of music and effects with dialogue.

[0529] Detect key segments present in at least one audio signal included or contained in the audio stream, which may, for example, require improvement under the current system settings;

[0530] A segment is considered a key segment if the reproduction of at least two different audio types (e.g., dialogue and background) leads to an increase in the user's listening effort.

[0531] The current system settings include information about user choices (e.g., enabled conversation enhancement options, hearing impairment options, or preferred language) and / or information about the environment (e.g., whether the content is consumed on a mobile device in a noisy environment (such as a bus) or at home using a dedicated audio system).

[0532] Decode audio data and process detected key segments to improve audio quality of the entire audio scene and reduce the user's listening effort.

[0533] Example 2: Based on the method of Example 1, it further includes:

[0534] Receive metadata associated with an audio scene, including or containing information about at least one key segment present in an audio signal contained in an audio stream;

[0535] Processing information about a critical segment present in at least one audio signal included or contained in the audio stream, along with at least one additional piece of information from the system level or from other metadata in the audio stream, to determine whether the critical segment present in at least one audio signal included or contained in the audio stream can be improved.

[0536] Decode the encoded audio stream, and when it is determined that a key segment present in at least one audio signal included or contained in the audio stream can be improved, use information about the key segment present in at least one audio signal to improve the audio quality of the entire audio scene.

[0537] Example 3: According to the method of either Example 1 or 2, wherein the information about the key segment includes or contains at least one parameter associated with the short-term intensity of the audio signal in the audio scene, or at least one parameter associated with the short-term intensity difference between two or more audio types included or contained in the audio scene.

[0538] Example 4: According to the method of any of Examples 1 to 3, the information about the key fragment includes or contains at least one of the following parameters:

[0539] ■ Information about which audio signals include or contain key segments

[0540] ■ Information on which audio signals need to be processed to improve key segments, which may not exist in all audio signals.

[0541] ■ One or more gain sequences are required to be applied to one or more audio signals.

[0542] ■ Information about the start, end, and / or duration of at least one key segment.

[0543] ■ Short-term intensity value associated with at least one audio signal

[0544] ■ Short-term intensity difference associated with at least two audio types that can be characterized, for example, dialogue and / or narration and / or background and / or music and effects;

[0545] Example 21: A system supporting audio production, post-production, or quality control (QC) stages is configured to receive an audio scene as input and generate a QC report for the audio scene, wherein...

[0546] The input audio scene can be provided in different formats commonly used in audio production, such as final mixes in mono, stereo, surround, or immersive formats (e.g., compressed or uncompressed), as well as combinations of audio channels and audio objects commonly used in the NGA system, such as those encapsulated in audio master files.

[0547] QC reports include information about critical segments, namely segments in an audio scene where a specific signal characteristic of at least one audio component does not conform to expected standards, and

[0548] Short-term strength is used as a signal characteristic, where the short-term strength of the signal can be calculated in one of the following ways:

[0549] ■ Local power of audio signals;

[0550] ■ The power of the filtered signal, where the filtering simulates the frequency selectivity sensitivity of the human ear;

[0551] ■ Short-term or instantaneous loudness according to ITU-R BS.1770 and EBU R 128 or their variants (e.g., using different time window sizes);

[0552] ■ A model for calculating loudness;

[0553] ■ AI-based intensity estimation.

[0554] Example 22: The system according to Example 21 is configured to use at least one absolute or relative threshold with respect to an audio component or a group of audio components as a desired criterion, wherein

[0555] The absolute threshold is related to the short-term intensity of a selected audio component or group of audio components, and

[0556] The relative threshold is related to the short-term intensity difference between audio components or groups of audio components.

[0557] Example 23: A system according to any one of Examples 21-22 is configured to combine multiple audio signals of the same or similar type to form a component group. The process of combining the signals can be accomplished, for example, based on the following:

[0558] ■ The importance of multiple audio signals, wherein the importance of the audio signals is set manually or determined based on the speech portion contained in the signals; and / or

[0559] ■ The contribution of each audio signal to the final mix, taking into account audio masking effects and the characteristics of the human auditory system.

[0560] Example 24: A system according to any of Examples 21-23 is configured to analyze audio components associated with dialogue type (e.g., speech, narration, audio description, etc.) and audio components associated with background type (e.g., music and effects, stadium atmosphere, etc.).

[0561] Example 25: In the system of Example 21, the QC report may include, but is not limited to:

[0562] ■ The beginning of each key segment in the audio signal; and / or

[0563] ■ The end or duration of each key segment in the audio signal; and / or

[0564] ■ The level of criticality, which may be related to the level of intelligibility or listening effort required to understand key passages; and / or

[0565] ■ Additional information that can be used to support the production, post-production, or QC stages.

[0566] Example 26: Based on the system of Example 21, it is configured to use a source separation module (possibly AI-based) to estimate the dialogue and background component groups given a mixture of dialogue and background component groups to be used if separate dialogue and background component groups cannot be obtained from the input audio scene.

[0567] Example 27: Based on the system in Example 21, it is configured to output an enhanced audio scene in which key segments have been automatically enhanced.

[0568] Example 28: A system supporting audio production, post-production, or quality control (QC) stages is configured to receive an audio scene as input and generate a QC report for the audio scene, wherein...

[0569] The input audio scene is provided in various formats commonly used in audio production, such as final mixes in mono, stereo, surround, or immersive formats (compressed or uncompressed), and combinations of audio channels and audio objects encapsulated in audio master files, as commonly used in the NGA system.

[0570] QC reports include information about critical segments, namely segments in an audio scene where a specific signal characteristic of at least one audio component does not conform to expected standards, and

[0571] An end-to-end detector module (possibly AI-based) is used to detect key segments directly from the input, where the detector can switch to different sub-modules (detector 1, detector 2, etc.) depending on the input format type.

[0572] References ‌

[0573] [1] M. Torcoli, C. Simon, J. Paulus, D. Straninger, A. Riedel, V.Koch, S. Wirts, D. Rieger, H. Fuchs, C. Uhle, S. Meltzer and A. Murtaza, "Dialog+ in Broadcasting: First Field Tests Using Deep-Learning-Based DialogueEnhancement," in IBC (International Broadcasting Convention), 2021.

[0574] [2] CD Mathers, "A Study of Sound Balances for the Hard of Hearing," BBC White Paper, 1991.

[0575] [3] C. Simon, M. Torcoli and J. Paulus, "MPEG-H Audio for ImprovingAccessibility in Broadcasting and Streaming," arXiv:1909.11549, 2019.

[0576] [4] M. Torcoli, T. Robotham and E. Habets, "Dialogue Enhancement andListening Effort in Broadcast Audio: A Multimodal Evaluation," in 14th IEEEInternational Conference on Quality of Multimedia Experience (QoMEX), 2022.

[0577] [5] M. Armstrong, "From Clean Audio to Object Based Broadcasting,"BBC R&D White Paper WHP 324, 2016.

[0578] [6] Netflix, "Netflix Sound Mix Specifications & Best Practicesv1.4," 2021. [Online]. Available: https: / / partnerhelp.netflixstudios.com / hc / en-us / articles / 360001794307-Netflix-Sound-Mix-Specifications-Best-Practices-v1-4.

[0579] [7] M. Torcoli, A. Freke-Morin, J. Paulus, C. Simon and B. Shirley, "Preferred Levels for Background Ducking to Produce Esthetically PleasingAudio for TV with Clear Speech," J. Audio Eng. Soc., vol. 67, no. 12, pp.1003-1011, 2019.

[0580] [8] Recommendation, "ITU-RBS.1770-4, Algorithms to measure audioprogramme loudness and true-peak audio level," 2015.

[0581] [9] J. Paulus and M. Torcoli, "Sampling Frequency IndependentDialogue Separation," in 30th European Signal Processing Conference(EUSIPCO), 2022.

[0582]

[10] B. C. J. Moore, B. R. Glasberg and T. Baer, "A model for theprediction of thresholds, loudness, and partial loudness," J. Audio Eng.Soc., vol. 45, no. 4, p. 224–240, 1997.

Claims

1. An audio analyzer (100, 600, 800). The audio analyzer is configured to acquire audio content including a speech portion and a background portion (101, 601, 631, 1001, 1201, 1241). The audio analyzer is configured to determine the short-term intensity difference (112, 812) between the speech portion of the audio content and the background portion of the audio content, and / or The audio analyzer is configured to determine short-term intensity information (112, 612, 1112) of the speech portion of the audio content, and The audio analyzer is configured to provide a representation of the short-term intensity difference and / or a representation of the short-term intensity information of the speech portion as an analysis result (102), or The audio analyzer is configured to derive analysis results (102, 622) from the short-term intensity difference and / or from the short-term intensity information of the speech portion.

2. The audio analyzer (100, 600, 800) according to claim 1. The short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) includes a time resolution of no more than 3000 milliseconds, or includes a time resolution of no more than 1000 milliseconds, or includes a time resolution of no more than 400 milliseconds, or includes a time resolution of no more than 100 milliseconds, or includes a time resolution of no more than 40 milliseconds, or includes a time resolution of no more than 20 milliseconds; or The short-term intensity difference and / or the short-term intensity information include a time resolution between 3000 milliseconds and 400 milliseconds.

3. The audio analyzer (100, 600, 800) according to any one of claims 1 or 2. The short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion include the temporal resolution of an audio frame, or The short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion include the temporal resolution of two audio frames, or The short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion include a temporal resolution of no more than 10 audio frames.

4. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The short-term intensity difference (112, 812) mentioned above is a short-term loudness difference or an instantaneous loudness difference; and / or The short-term intensity information (112, 612, 1112) of the speech portion is short-term loudness or instantaneous loudness.

5. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The short-term intensity difference (112, 812) mentioned above is the short-term energy ratio or instantaneous energy ratio; and / or The short-term intensity information (112, 612, 1112) of the speech portion is short-term energy or instantaneous energy.

6. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) are low-level characteristics of the audio content (101, 601, 631, 1001, 1201, 1241).

7. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to provide analysis results (102, 622) that are independent of the features of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241) that exceed the intensity features.

8. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to provide the analysis results (102, 622) based solely on one or more features of the audio content (101, 601, 631, 1001, 1201, 1241), wherein one or more features of the audio content can be modified by scaling one or more portions of the audio content.

9. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to separate the acquired audio content (101, 601, 631, 1001, 1201, 1241) into a speech portion of the audio content and a background portion of the audio content.

10. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to determine or estimate the intensity of the speech portion and the intensity of the background portion of the acquired audio content (101, 601, 631, 1001, 1201, 1241).

11. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to provide metadata of audio content (101, 601, 631, 1001, 1201, 1241) and / or encoded audio content as analysis results (102, 622).

12. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to provide the analysis results in character-encoded form and / or in binary form.

13. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to provide visualizations of the analysis results (102, 622) (710, 720, 730).

14. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to acquire the short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion based on the local power of one or more audio signals, or based on the local power of multiple portions of the audio content (101, 601, 631, 1001, 1201, 1241); or The audio analyzer is configured to acquire the short-term intensity difference and / or the short-term intensity information of the speech portion based on short-term or instantaneous loudness according to ITU-R BS.1770 and EBU R 128 or variants thereof.

15. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to obtain the short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion based on one or more filtered portions of the audio content (101, 601, 631, 1001, 1201, 1241).

16. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to use a loudness calculation model to obtain the short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion.

17. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to use one or more AI-based short-term intensity estimates to obtain the short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion.

18. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to combine multiple portions of the audio content (101, 601, 631, 1001, 1201, 1241) to obtain the short-term intensity difference (112, 812) and / or to obtain the short-term intensity information (112, 612, 1112) of the speech portion, and / or The audio analyzer is configured to combine multiple audio signals of the audio content (101, 601, 631, 1001, 1201, 1241) to obtain the short-term intensity difference (112, 812) and / or to obtain the short-term intensity information (112, 612, 1112) of the speech portion.

19. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to determine one or more key segments (622) of the audio content (101, 601, 631, 1001, 1201, 1241) based on the short-term intensity difference (112, 812) and / or based on the short-term intensity information (112, 612, 1112) of the speech portion. The key segment is a portion of the audio content in which the intensity of the speech portion is absolutely and / or relative to the background portion; and The analysis results (102, 622) include information about the one or more key segments.

20. The audio analyzer (100, 600, 800) according to claim 19. The audio analyzer is configured to determine the one or more key segments (622) based on a comparison of the short-term intensity difference (112, 812) with a single threshold (621) and / or multiple thresholds (621) and / or based on a comparison of the short-term intensity information (112, 612, 1112) of the speech portion with a single threshold (621) and / or multiple thresholds (621).

21. The audio analyzer (100, 600, 800) according to claim 20. The audio analyzer is configured to use different thresholds (621) for different portions of the audio content (101, 601, 631, 1001, 1201, 1241); and / or The audio analyzer is configured to use different thresholds (621) for different types of audio signals of the audio content.

22. The audio analyzer (100, 600, 800) according to any one of claims 20 to 21. The audio analyzer is configured to use one or more frequency-related thresholds (621).

23. The audio analyzer (100, 600, 800) according to any one of claims 20 to 22. The audio analyzer is configured to use artificial intelligence to adapt one or more thresholds (621).

24. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to perform neural network inference to determine one or more key segments (622) of the audio content (101, 601, 631, 1001, 1201, 1241). The key segment is a part of the audio content in which the intensity of the speech portion is absolutely and / or locally lower relative to the background portion; and The analysis results (102, 622) include information about the one or more key segments.

25. The audio analyzer (100, 600, 800) according to any one of claims 19 to 24. The audio analyzer is configured to determine information about at least one of the following for the one or more key segments (622): Beginning; Finish Duration; quantity Severity and / or criticality; and / or Critical level; Time and location; And the analysis results (102, 622) mentioned therein include the information.

26. The audio analyzer (100, 600, 800) according to any one of claims 19 to 25. The audio analyzer is configured to use two or more states to determine information about the severity and / or criticality of the one or more critical segments (622).

27. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to provide the analysis results (102, 622) in the following form: The form of the binary analysis result; or The form of the ternary analysis result; or The form of the quaternary analysis result.

28. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to determine the short-term intensity difference (112, 812) and / or the short-term intensity information (112, 612, 1112) of the speech portion as an approximation of the intelligibility or listening effort of the audio content (101, 601, 631, 1001, 1201, 1241).

29. The audio analyzer (100, 600, 800) according to any one of the preceding claims. The audio analyzer is configured to determine additional quality control information (614, 1114) based on the audio content (101, 601, 631, 1001, 1201, 1241).

30. The audio analyzer (100, 600, 800) according to any one of claims 1 to 29. The audio analyzer is configured to use at least one absolute or relative threshold (621) with respect to one or more audio components or groups of one or more audio components as one or more desired criteria; The absolute threshold is related to the short-term intensity (112, 612, 613, 1112, 1113) of one or more selected audio components or groups of audio components, and The relative threshold is related to the short-term intensity difference (112, 812) between audio components or groups of audio components.

31. The audio analyzer (100, 600, 800) according to any one of claims 1 to 30. The audio analyzer is configured to combine multiple audio signals of the same or similar type to form component groups.

32. The audio analyzer (100, 600, 800) according to claim 31. The audio analyzer is configured to combine multiple audio signals of the same or similar type to form component groups based on the following: The importance of the plurality of audio signals is based on the fact that the speech portion contained in the signal is manually set or determined, and / or Based on the contribution of each audio signal to the final mixture, the audio masking effect and the characteristics of the human auditory system are taken into account.

33. An audio analyzer (200, 900, 1000). The audio analyzer is configured to acquire audio content including a speech portion and a background portion (101, 601, 631, 1001, 1201, 1241). The audio analyzer includes a neural network (910, 1010, 1020) configured to derive quality control information (202, 622) based on the audio content.

34. The audio analyzer (200, 900, 1000) according to claim 33. The neural network (910, 1010, 1020) is configured to acquire a representation of the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content (101, 601, 631, 1001, 1201, 1241) and / or a representation of the short-term intensity information (112, 612, 1112) of the speech portion of the audio content as an analysis result (102, 622). The audio analyzer is configured to provide a representation of the short-term intensity difference and / or a representation of the short-term intensity information (112, 612, 1112) of the speech portion of the audio content as the quality control information (202, 622).

35. The audio analyzer (200, 900, 1000) according to any one of claims 33 to 34. The neural network (910, 1010, 1020) is trained using the audio analyzer (100, 600, 800) according to claim 1.

36. The audio analyzer (200, 900, 1000) according to any one of claims 33 to 35. The audio analyzer is configured to separate the speech portion and background portion of the audio content (101, 601, 631, 1001, 1201, 1241) so as to provide the separated speech portion and / or background portion to the neural network (910, 1010, 1020).

37. The audio analyzer (200, 900, 1000) according to any one of claims 33 to 36. The audio content (101, 601, 631, 1001, 1201, 1241) comprises, in combination (601), the speech portion and the background portion; and The audio analyzer is configured to provide the speech portion and the background portion to the neural network (910, 1010, 1020) in a combined manner (601).

38. The audio analyzer (200, 900, 1000) according to any one of claims 33 to 37. The audio analyzer is configured to provide the speech portion and the background portion to the neural network (910, 1010, 1020) as separate signals (631).

39. The audio analyzer (200, 900, 1000) according to any one of claims 33 to 38. The audio analyzer includes a first neural network (1010) for deriving the quality control information (202, 622) based on an audio mixture of the audio content (101, 601, 631, 1001, 1201, 1241); and The audio analyzer includes a second neural network (1020) for deriving the quality control information based on the speech and background portions of the audio content provided as separate information entries.

40. The audio analyzer (200, 900, 1000) according to any one of claims 33 to 39. The audio analyzer includes end-to-end detectors (910, 1010, 1020) to detect key segments (622) directly from one or more input signals (601, 631).

41. The audio analyzer (200, 900, 1000) according to claim 40. The end-to-end detector is configured to switch between two sub-modules (1010, 1020) based on the input format type.

42. An audio processor (300, 1100, 1200, 1300) for processing audio content (101, 601, 631, 1001, 1201, 1241). The audio processor is configured to acquire audio content (101, 601, 631, 1001, 1201, 1241) including a speech portion and a background portion. The audio processor is configured to determine a short-term intensity difference (112, 812) between the speech portion of the audio content and the background portion of the audio content, and / or The audio processor is configured to determine the short-term intensity (112, 612, 1112) of the speech portion. and The audio processor is configured to modify (310) the audio content based on the short-term intensity difference and / or based on the short-term intensity of the speech portion, or The audio processor is configured to determine metadata information (321) about the audio content based on the short-term intensity difference and / or based on the short-term intensity of the speech portion, and to provide a file or stream (332) including the audio content and the metadata information.

43. The audio processor (300, 1100, 1200, 1300) according to claim 42. The audio processor is configured to modify (310) the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content (101, 601, 631, 1001, 1201, 1241) to obtain a processed version of the audio content; and / or The audio processor is configured to modify the short-term intensity of the speech portion of the audio content in order to obtain a processed version of the audio content.

44. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 43. The audio processor is configured to modify the acquired audio content (101, 601, 631, 1001, 1201, 1241) at a time resolution not exceeding 3000 milliseconds, or not exceeding 1000 milliseconds, or not exceeding 400 milliseconds, or not exceeding 100 milliseconds, or not exceeding 40 milliseconds, or not exceeding 20 milliseconds, in order to acquire a processed version (312) of the audio content; or The audio processor is configured to modify the acquired audio content (310) at a time resolution between 3000 milliseconds and 400 milliseconds.

45. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 44. The audio processor is configured to modify (310) the acquired audio content (101, 601, 631, 1001, 1201, 1241) at a time resolution of one audio frame, or at a time resolution of two audio frames, or at a time resolution of no more than 10 audio frames, in order to acquire a processed version (312) of the audio content.

46. ​​The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 45. The audio processor is configured to scale the speech portion and / or the background portion of the acquired audio content (101, 601, 631, 1001, 1201, 1241) to obtain a processed version (312) of the audio content.

47. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 46. The audio processor is configured to provide or modify metadata in order to obtain a processed version (312) of the audio content (101, 601, 631, 1001, 1201, 1241).

48. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 47. The audio processor is configured to determine metadata information (321, 1131) about the audio content (101, 601, 631, 1001, 1201, 1241) based on the short-term intensity difference (112, 812) and / or based on the short-term intensity of the speech portion, such that the relationship between the speech portion and the background portion of the audio content can be modified based on the metadata information, and / or the speech portion can be modified based on the metadata information; and The audio processor is configured to provide modified audio content that includes the audio content and the metadata information.

49. The audio processor (300, 1100, 1200, 1300) according to claim 48. The audio processor is configured to format the metadata information according to the audio data frame rate of the audio content (101, 601, 631, 1001, 1201, 1241).

50. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 49. The audio processor is configured (1310) to separate the speech portion and the background portion of the audio content (101, 601, 631, 1001, 1201, 1241).

51. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 50. Includes an audio analyzer according to any one of claims 1 to 29 or 100 to 106; and The audio processor is configured to modify (310) the audio content (101, 601, 631, 1001, 1201, 1241) according to the analysis results (102, 112).

52. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 51. Includes an audio analyzer (100, 200, 600, 800, 900, 1000) according to any one of claims 1 to 32 or 33 to 41. and The audio processor is configured to modify or generate metadata information based on the results of the audio analyzer.

53. The audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 52. Includes an audio analyzer (100, 200, 600, 800, 900, 1000) according to any one of claims 1 to 32 or 33 to 41. and The audio processor is configured to store metadata information aligned with the audio data in a file or stream (332).

54. A bitstream provider (400, 1400, 1700). The bitstream provider is configured to include an encoded representation (401) of the audio content (101, 601, 631, 1001, 1201, 1241) and quality control information (402, 512, 1131, 1512) into the bitstream (403, 501, 1402, 1702).

55. The bitstream provider (400, 1400, 1700) according to claim 54. The quality control information (402, 512, 1131, 1512) implements or supports decoder-side modification of the relationship between the intensity of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241) and the background portion of the audio content; and / or The quality control information described therein enables or supports decoder-side modifications to the speech portion.

56. The bitstream provider (400, 1400, 1700) according to any one of claims 54 or 55. The quality control information (402, 512, 1131, 1512) achieves or supports decoder-side improvement of the speech intelligibility of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241).

57. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 56. The quality control information (402, 512, 1131, 1512) selectively enables and disables decoder-side improvement of the speech intelligibility of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241), or The quality control information indicates which segments of the audio content are permissible for improvements in the speech intelligibility of the speech portion.

58. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 57. The quality control information (402, 512, 1131, 1512) includes information about key time portions (622) of the audio content (101, 601, 631, 1001, 1201, 1241).

59. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 58. The quality control information (402, 512, 1131, 1512) includes decoder-side improvements in the speech intelligibility of the speech portion of the audio content under blocked listening conditions, indicating which segments of the audio content (101, 601, 631, 1001, 1201, 1241) are considered recommendable.

60. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 59. The quality control information (402, 512, 1131, 1512) includes information indicating whether a portion of the audio content (101, 601, 631, 1001, 1201, 1241) includes a speech intelligibility metric or speech intelligibility-related characteristic that is predetermined to one or more thresholds (621).

61. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 60. The quality control information (402, 512, 1131, 1512) includes information that quantitatively describes the speech intelligibility-related characteristics of a portion of the audio content (101, 601, 631, 1001, 1201, 1241).

62. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 61. The quality control information (402, 512, 1131, 1512) includes information indicating whether an audio scene is considered difficult or easy to understand.

63. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 62. The quality control information (402, 512, 1131, 1512) includes information indicating whether an audio scene is considered comprehensible with low listening effort, medium listening effort, or high listening effort.

64. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 63. The quality control information (402, 512, 1131, 1512) includes information indicating segments in the audio content (101, 601, 631, 1001, 1201, 1241) where the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content is less than or equal to a threshold (621); and / or The quality control information (402, 512, 1131, 1512) includes information indicating segments in the audio content whose short-term intensity (112, 612, 1112) of the speech portion of the audio content is less than or equal to a threshold (621).

65. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 64. The quality control information (402, 512, 1131, 1512) describes the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content (101, 601, 631, 1001, 1201, 1241); and / or The quality control information described therein describes the short-term intensity (112, 612, 1112) of the speech portion of the audio content.

66. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 65. The bitstream provider is configured to add the quality control information (402, 512, 1131, 1512) to pre-existing metadata (1511, 1503, 1504).

67. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 66. The bitstream provider is configured to adapt processing parameters for decoding the audio content (101, 601, 631, 1001, 1201, 1241) based on the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content and / or based on the short-term intensity (112, 612, 1112) of the speech portion.

68. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 67. The bitstream provider is configured to include the quality control information (402, 512, 1131, 1512) into an extended payload of the bitstream (403, 501, 1402, 1702).

69. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 68. The bitstream provider includes an audio analyzer (100, 200, 600, 800, 900, 1000) according to any one of claims 1 to 32 or 33 to 41, and the bitstream provider is configured to determine the quality control information (402, 512, 1131, 1512) based on the analysis results (102, 622); and / or The bitstream provider described herein includes an audio processor (300, 1100, 1200, 1300) according to any one of claims 42 to 53.

70. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 69. The bitstream provider is configured to format the quality control information (402, 512, 1131, 1512) into quality control metadata packets.

71. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 69. The bitstream provider is configured to encapsulate quality control metadata in a data packet and insert the data packet into the audio bitstream (403, 501, 1402, 1702) during encoding.

72. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 71. The bitstream provider described above is implemented using a neural network. The neural network is configured to receive a representation of audio content (101, 601, 631, 1001, 1201, 1241) and provide the quality control information (402, 512, 1131, 1512) based on the representation of the audio content (101, 601, 631, 1001, 1201, 1241). The neural network is trained using training audio scenarios, which are labeled according to speech intelligibility.

73. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 72. The bitstream provider described above is implemented using a neural network. The neural network is configured to receive a representation of audio content (101, 601, 631, 1001, 1201, 1241) and provide the quality control information (402, 512, 1131, 1512) based on the representation of the audio content (101, 601, 631, 1001, 1201, 1241). The neural network is trained using an audio analyzer (100, 600, 800) according to any one of claims 1 to 32, and The audio analyzer is configured to provide reference quality control information for the training of the neural network based on multiple audio scenarios.

74. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 73. The quality control information (402, 512, 1131, 1512) includes sharpness information metadata, or The quality control information (402, 512, 1131, 1512) includes accessibility enhancement metadata, or The quality control information (402, 512, 1131, 1512) includes voice transparency metadata, or The quality control information (402, 512, 1131, 1512) includes voice enhancement metadata, or The quality control information (402, 512, 1131, 1512) includes understandable metadata, or The quality control information (402, 512, 1131, 1512) includes content description metadata, or The quality control information (402, 512, 1131, 1512) includes locally enhanced metadata, or The quality control information (402, 512, 1131, 1512) includes signal description metadata.

75. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 74. The bitstream provider is configured to include the quality control information (402, 512, 1131, 1512) into the MHAS packet.

76. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 75. The bitstream provider is configured to provide audio quality control metadata (e.g., qcInfo()) and information about the audio quality control metadata; The information regarding the audio quality control metadata describes how many data structures qcInfo() carrying the audio quality control metadata exist in the MHAS packet, and / or The information describing the audio quality control metadata or indicating when the audio quality control metadata should be applied (e.g., qcInfoActive), and / or The information describing the audio quality control metadata or indicating in what circumstances the audio quality control metadata should be applied, and / or The information regarding the audio quality control metadata describes or indicates under what circumstances the corresponding decoder or renderer can choose to apply the audio quality control metadata, and / or The information regarding the audio quality control metadata indicates whether the audio quality control metadata is associated with a specific audio element or with an audio scene defined by a combination of audio elements; and / or The information regarding the audio quality control metadata indicates the type of the audio content (101, 601, 631, 1001, 1201, 1241) associated with the audio quality control metadata; and / or The information regarding the audio quality control metadata indicates which type of audio content the audio quality control metadata can be applied to in order to manipulate audio elements of the corresponding type. and / or The information regarding the audio quality control metadata includes an identifier that indicates which audio element or group of audio elements the corresponding audio quality control metadata is associated with.

77. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 76. The bitstream provider is configured to provide audio quality control information (402, 512, 1131, 1512) with a first granularity and audio quality control information with a second granularity.

78. The bitstream provider (400, 1400, 1700) according to any one of claims 54 to 77. The bitstream provider is configured to provide multiple different audio quality control metadata associated with different audio elements and / or different combinations of audio elements.

79. An audio decoder (500, 1500, 1600, 1800) for providing decoded audio representations (502, 1502, 1602) based on encoded media representations (403, 501, 1402, 1702). The audio decoder is configured to obtain quality control information (402, 512, 1131, 1512) from the encoded media representation; and The audio decoder is configured to provide the decoded audio representation based on the quality control information.

80. The audio decoder (500, 1500, 1600, 1800) according to claim 79. The encoded media representations (403, 501, 1402, 1702) include representations of audio content (101, 601, 631, 1001, 1201, 1241), which includes a speech portion and a background portion. The audio decoder is configured to receive quality control information (402, 512, 1131, 1512) and provide the decoded audio information based on the quality control information (402, 512, 1131, 1512), wherein the quality control information includes at least one of the following: Information used to modify the relationship between the intensity of the speech portion of the audio content and the background portion of the audio content; Information used to modify the speech portion of the audio content; Information used to improve the speech intelligibility of the speech portion of the audio content; Information for selectively enabling and disabling improvements in speech intelligibility of the speech portion of the audio content; Information indicating which segments of the audio content are permissible for improvements in the speech intelligibility of the speech portion; Information indicating key time segments (622) of the audio content; This indicates which segments of the audio content are considered recommendable information due to improvements in the speech intelligibility of the speech portion under blocked listening conditions; Information indicating whether a portion of the audio content includes a speech intelligibility metric or speech intelligibility-related characteristic that is predetermined to one or more thresholds; Indicate whether the audio scene is considered difficult or easy to understand. The information indicates whether the audio scene is considered comprehensible with low listening effort, medium listening effort, or high listening effort. Information indicating segments in the audio content whose short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content is less than or equal to a threshold. Information indicating segments in the audio content whose short-term intensity is less than or equal to a threshold.

81. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 or 80. The audio decoder is configured to perform speech enhancement based on the quality control information (402, 512, 1131, 1512).

82. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 81. The audio decoder is configured to selectively perform speech enhancement on segments of the audio content (101, 601, 631, 1001, 1201, 1241) indicated by the quality control information (402, 512, 1131, 1512).

83. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 82. The audio decoder is configured to selectively perform speech enhancement on segments of audio content (101, 601, 631, 1001, 1201, 1241) that are difficult to understand according to the quality control information (402, 512, 1131, 1512).

84. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 83. The audio decoder is configured to receive control information defining whether speech enhancement should be performed, and The audio decoder is configured to activate and deactivate speech enhancement based on control information that defines whether speech enhancement should be performed.

85. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 84. The audio decoder is configured to receive control information defining interaction with the audio scene; and The audio decoder is configured to activate and deactivate speech enhancement based on control information defined for interaction with the audio scene, and / or The audio decoder is configured to adjust one or more parameters of the speech enhancement based on control information that defines the interaction with the audio scene.

86. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 85. The audio decoder is configured to adjust one or more parameters of the speech enhancement based on the quality control information (402, 512, 1131, 1512).

87. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 86. The audio decoder is configured to acquire information about the listening environment (1504), and The audio decoder is configured to determine whether to perform speech enhancement based on the information about the listening environment and the quality control information (402, 512, 1131, 1512).

88. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 87. The audio decoder is configured to acquire information about the listening environment (1504), and The audio decoder is configured to adjust one or more parameters of the speech enhancement based on the quality control information (402, 512, 1131, 1512) and the information about the listening environment.

89. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 88. The audio decoder is configured to acquire user input (1503), and The audio decoder is configured to determine whether to perform speech enhancement based on the user input and the quality control information (402, 512, 1131, 1512).

90. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 89. The audio decoder is configured to acquire system-level information (1503), and The audio decoder is configured to determine whether to perform speech enhancement based on the information about the listening environment and the quality control information (402, 512, 1131, 1512).

91. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 90. The audio decoder is configured to acquire information about one or more sound reproduction devices, and The audio decoder is configured to adjust one or more parameters of speech enhancement based on the quality control information (402, 512, 1131, 1512) and the information regarding one or more sound reproduction devices.

92. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 91. The audio decoder is configured to acquire one or more of the following system-level information (1503): o Information about system settings, Information about user settings. Information about the environment, and Information about one or more additional devices, and The audio decoder is configured to perform one or more of the following functions based on the quality control information (402, 512, 1131, 1512) and the system-level information: o It was determined that the audio content (101, 601, 631, 1001, 1201, 1241) contained key segments that needed improvement; Determine the level and / or intensity of the quality improvement to be applied; The goal is to obtain the quality control information required by the audio decoder to enhance the audio quality of one or more key segments in order to improve intelligibility and / or reduce listening effort.

93. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 92. The quality control information (402, 512, 1131, 1512) includes one or more of the following: o requires one or more gain sequences to be applied to one or more audio signals that are part of an audio scene, or to one or more portions of the audio content (101, 601, 631, 1001, 1201, 1241); Information regarding which signals or portions of the audio content should be processed to improve one or more key segments (622); Information regarding the duration of the one or more key segments.

94. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 93. The audio decoder is configured to apply the quality control information (402, 512, 1131, 1512) to obtain a quality-enhanced version of the audio content.

95. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 94. The audio decoder is configured to apply the quality control information (402, 512, 1131, 1512) to the audio content (101, 601, 631, 1001, 1201, 1241) to obtain a quality-enhanced version of the audio content.

96. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 95. The audio decoder is configured to perform filtering (1630) to obtain a quality-enhanced version of the audio content (101, 601, 631, 1001, 1201, 1241). The audio decoder is configured to determine one or more filter coefficients of the filter based on the quality control information (402, 512, 1131, 1512).

97. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 96. The audio decoder is configured to perform filtering (1630) to obtain a quality-enhanced version of the audio content (101, 601, 631, 1001, 1201, 1241). The audio decoder is configured to determine one or more filter coefficients of the filter based on the quality control information (402, 512, 1131, 1512).

98. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 97. The audio decoder is configured as follows: Based on system-level information (1503); and / or Based on information regarding one or more sound reproduction devices; and / or Based on information about the listening environment (1504); and / or Based on the information about user input (1503) Determine one or more filter coefficients for the filtering.

99. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 98. The audio decoder is configured to apply the filter (1630) to one or more output signals of the decoder core (1520, 1620), or the audio decoder is configured to apply the filter to one or more rendered audio signals (1502) obtained by rendering the output signals of the decoder core.

100. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 99. The audio decoder is configured to trigger time-frequency modification based on the quality control information (402, 512, 1131, 1512).

101. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 100. The audio decoder is configured to detect one or more key segments present in at least one audio signal contained in the encoded media representation (403, 501, 1402, 1702).

102. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 101. The audio decoder is configured to decode the encoded audio data and process one or more detected key segments.

103. The audio decoder (500, 1500, 1600, 1800) according to claim 102. The audio decoder is configured to process one or more detected key segments to improve the audio quality of the audio scene and / or reduce the user's listening effort.

104. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 103. The quality control information (402, 512, 1131, 1512) includes metadata associated with the audio scene, which contains information about key segments present in at least one audio signal contained in the audio stream; The audio decoder is configured to process information about key segments present in at least one audio signal contained in the encoded media representation (403, 501, 1402, 1702) and at least one additional piece of information from the system level or from other metadata in the encoded media representation, to determine whether the key segments present in at least one audio signal contained in the encoded media representation can be improved. The audio decoder is configured to decode an encoded audio stream included in or constituting the encoded media representation, and to improve the audio quality of the entire audio scene by using information about the key segments present in at least one audio signal when it is determined that the key segments present in at least one audio signal contained in the audio stream can be improved.

105. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 104. The information about the key segment includes at least one parameter associated with the short-term intensity (112, 612, 1112) of the audio signal in the audio scene or with the short-term intensity difference (112, 812) between two or more audio types contained in the audio scene.

106. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 105. The information about the key segment includes at least one of the following parameters: - Information about which audio signals contain key segments; - Information regarding which audio signals need to be processed to improve key segments, which may not be present in all audio signals; - One or more gain sequences to be applied to one or more audio signals, wherein the time resolution of the gain sequences is no more than 1000 milliseconds, or no more than 400 milliseconds, or no more than 100 milliseconds, or no more than 40 milliseconds, or no more than 20 milliseconds, or between 3000 milliseconds and 400 milliseconds, or the time resolution of one audio frame, or the time resolution of two audio frames, or the time resolution of no more than ten audio frames; - Information regarding the start and / or end and / or duration of at least one key segment; - Short-term intensity value associated with at least one audio signal; - Short-term intensity difference (112, 812) associated with at least two audio types, which may be characterized, for example, as dialogue and / or narration and / or background and / or music and sound effects.

107. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 106. The audio decoder is configured to evaluate the quality control information (402, 512, 1131, 1512) contained in the MHAS package.

108. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 107. The audio decoder is configured to evaluate audio quality control metadata (e.g., qcInfo()) and information about the audio quality control metadata; The information regarding audio quality control metadata describes how many data structures qcInfo() carrying audio quality control metadata exist in the MHAS package, and / or The information described therein regarding audio quality control metadata describes or indicates when the audio quality control metadata should be applied (e.g., qcInfoActive), and / or The information regarding the audio quality control metadata described therein or indicates under what circumstances the audio quality control metadata should be applied, and / or The information regarding audio quality control metadata describes or provides indications of under what conditions the corresponding decoder or renderer can choose to apply the audio quality control metadata, and / or The information regarding audio quality control metadata indicates whether the audio quality control metadata is associated with a specific audio element or with an audio scene defined by a combination of audio elements; and / or The information regarding audio quality control metadata indicates the type of audio content (101, 601, 631, 1001, 1201, 1241) associated with the audio quality control metadata; and / or The information regarding audio quality control metadata indicates which type of audio content the audio quality control metadata can be applied to in order to manipulate audio elements of that type. and / or The information regarding audio quality control metadata includes an identifier that indicates which audio element or group of audio elements the corresponding audio quality control metadata is associated with.

109. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 108. The audio decoder is configured to evaluate audio quality control information (402, 512, 1131, 1512) with a first granularity and audio quality control information with a second granularity.

110. The audio decoder (500, 1500, 1600, 1800) according to any one of claims 79 to 109. The audio decoder is configured to evaluate multiple different audio quality control metadata associated with different audio elements and / or different combinations of audio elements.

111. A method for analyzing audio content (101, 601, 631, 1001, 1201, 1241), comprising: The audio content is obtained, which includes a speech portion and a background portion; Determine the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content, and / or Determine the short-term intensity information (112, 612, 1112) of the speech portion of the audio content, and Provide a representation of the short-term intensity difference (112, 812) and / or a representation of the short-term intensity information of the speech portion as an analysis result, or The analysis results are derived from the short-term intensity difference (112, 812) and / or from the short-term intensity information of the speech portion.

112. A method for analyzing audio content (101, 601, 631, 1001, 1201, 1241), comprising: Obtain the audio content, including the speech portion and the background portion; The quality control information (402, 512, 1131, 1512) is derived from the audio content using a neural network.

113. A method for processing audio content (101, 601, 631, 1001, 1201, 1241), comprising: The audio content is obtained, wherein the audio content includes a speech portion and a background portion; Determine the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content; and / or Determine the short-term intensity of the speech component, and The audio content is modified based on the short-term intensity difference (112, 812) and / or based on the short-term intensity of the speech portion.

114. A method for providing a bitstream (403, 501, 1402, 1702), comprising: The encoded representation of the audio content (101, 601, 631, 1001, 1201, 1241) and the quality control information (402, 512, 1131, 1512) are included in the bitstream.

115. A method for providing a decoded audio representation based on an encoded media representation, comprising: Quality control information (402, 512, 1131, 1512) is obtained from the encoded media representation. and The decoded audio representation is provided based on the quality control information.

116. A computer program for performing the method according to claim 111, 112, 113, 114 or 115 when the computer program is run on a computer.

117. A bitstream (403, 501, 1402, 1702), comprising: Encoded representation of audio content (101, 601, 631, 1001, 1201, 1241); And quality control information (402, 512, 1131, 1512).

118. The bitstream (403, 501, 1402, 1702) according to claim 117. The quality control information (402, 512, 1131, 1512) implements or supports decoder-side modification of the relationship between the intensity of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241) and the background portion of the audio content; and / or The quality control information (402, 512, 1131, 1512) enables or supports decoder-side modifications to the speech portion.

119. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 or 118. The quality control information (402, 512, 1131, 1512) achieves or supports decoder-side improvement of the speech intelligibility of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241).

120. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 119. The quality control information (402, 512, 1131, 1512) selectively enables and disables decoder-side improvement of the speech intelligibility of the speech portion of the audio content (101, 601, 631, 1001, 1201, 1241), or The quality control information indicates which segments of the audio content are permissible for improvements in the speech intelligibility of the speech portion.

121. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 120. The quality control information (402, 512, 1131, 1512) includes information about key time portions of the audio content (101, 601, 631, 1001, 1201, 1241).

122. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 121. The quality control information (402, 512, 1131, 1512) includes decoder-side improvements in the speech intelligibility of the speech portion of the audio content under blocked listening conditions, indicating which segments of the audio content (101, 601, 631, 1001, 1201, 1241) are considered recommendable.

123. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 122. The quality control information (402, 512, 1131, 1512) includes information indicating whether a portion of the audio content (101, 601, 631, 1001, 1201, 1241) includes a speech intelligibility metric or speech intelligibility-related characteristic that is predetermined to one or more thresholds.

124. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 123. The quality control information (402, 512, 1131, 1512) includes information that quantitatively describes the speech intelligibility-related characteristics of a portion of the audio content (101, 601, 631, 1001, 1201, 1241).

125. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 124. The quality control information (402, 512, 1131, 1512) includes information indicating whether an audio scene is considered difficult or easy to understand.

126. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 125. The quality control information (402, 512, 1131, 1512) includes information indicating whether an audio scene is considered comprehensible with low listening effort, medium listening effort, or high listening effort.

127. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 126. The quality control information (402, 512, 1131, 1512) includes information indicating segments in the audio content (101, 601, 631, 1001, 1201, 1241) where the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content is less than or equal to a threshold; and / or The quality control information includes information indicating segments in the audio content whose short-term intensity is less than or equal to a threshold.

128. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 127. The quality control information (402, 512, 1131, 1512) describes the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content (101, 601, 631, 1001, 1201, 1241); and / or The quality control information described therein describes the short-term intensity of the speech portion of the audio content.

129. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 128. The bitstream includes information for adapting processing parameters for decoding the audio content (101, 601, 631, 1001, 1201, 1241) based on the short-term intensity difference (112, 812) between the speech portion and the background portion of the audio content and / or based on the short-term intensity of the speech portion.

130. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 129. The bitstream includes an extended payload, and the extended payload includes the quality control information (402, 512, 1131, 1512).

131. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 130. The bitstream includes the quality control information (402, 512, 1131, 1512) formatted into a quality control metadata packet.

132. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 131. The bitstream includes data packets, and the quality control metadata is encapsulated within the data packets.

133. The bitstream (403, 501, 1402, 1702) according to any one of claims 117 to 132. The quality control information (402, 512, 1131, 1512) includes sharpness information metadata, or The quality control information (402, 512, 1131, 1512) includes accessibility enhancement metadata, or The quality control information (402, 512, 1131, 1512) includes voice transparency metadata, or The quality control information (402, 512, 1131, 1512) includes voice enhancement metadata, or The quality control information (402, 512, 1131, 1512) includes understandable metadata, or The quality control information (402, 512, 1131, 1512) includes content description metadata, or The quality control information (402, 512, 1131, 1512) includes locally enhanced metadata, or The quality control information (402, 512, 1131, 1512) includes signal description metadata.

134. An audio analyzer (100, 600, 800), as claimed in any one of claims 1 to 41, configured to Obtain the average speech information of the audio content; The short-term intensity information is compared with the average speech information to obtain a comparison result; Based on the comparison results, an analysis is derived, which includes information about the deviation between the short-term intensity information and the average speech information.

135. The audio analyzer according to claim 134, The average speech information includes at least one of the following: average speech level, average speech intensity, or average speech loudness of the audio content.

136. The audio analyzer according to claim 134 or 135, configured to The average speech information is determined based on the average of the audio content over a predetermined time interval.

137. The audio analyzer according to any one of claims 134 to 136, configured to The analysis results are provided based on a combined evaluation of the short-term intensity difference and the deviation between the short-term intensity information and the average speech information.

138. The audio analyzer according to any one of claims 134 to 137, configured to Information about local speech levels is determined as the short-term intensity information; The average speech information is determined based on the average speech loudness over all audio content, or the average speech loudness over a time period of at least ten times the duration that determines the short-term intensity difference or the short-term intensity information. The local speech level is compared with the average speech information to obtain the comparison result; The analysis results are derived based on the evaluation of the comparison results relative to the threshold.

139. The audio analyzer according to any one of claims 134 to 138, configured to Information on the evolution of the short-term intensity difference over time between the speech portion and the background portion of the audio content, and / or Information on the evolution of short-term intensity information (112, 612, 1112) over time regarding the speech portion of the audio content; and The analysis results (102, 622) are derived from the information regarding the evolution of short-term intensity differences over time and / or from the evolution of short-term intensity information of the speech portion over time.

140. An audio analyzer (100, 600, 800). The audio analyzer is configured to acquire audio content (101, 601, 631, 1001, 1201, 1241) of an audio scene including a speech portion and a background portion. The audio analyzer is configured to determine the short-term intensity difference (112, 812) between the speech portion of the audio content and the background portion of the audio content, and / or The audio analyzer is configured to determine short-term intensity information (112, 612, 1112) of the speech portion of the audio content, and The audio analyzer is configured to derive analysis results (102, 622) from the short-term intensity difference and / or from the short-term intensity information of the speech portion. In order to provide information about key segments of an audio scene, for which a particular signal characteristic of at least one audio component in the audio scene does not meet one or more predefined criteria.