Audio system and method

Through deep learning DL classification algorithm to estimate the reverb category in the audio system and add matching reverbs, the problem that reverb simulation in the prior art is difficult to achieve a satisfactory listening experience, and high-quality audio signal generation is achieved.

CN120075696APending Publication Date: 2025-05-30HARMAN BECKER AUTOMOTIVE SYST GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411626001.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art may produce unnecessary artifacts when acoustic simulation of concert halls or other listening spaces by appropriate addition and reproduction of matching reverbs, and the calculation amount is large, making it difficult to achieve a satisfactory listening experience.

Method used

Using a deep learning DL classification algorithm, the audio input signal is received through the reverb classification unit, the appropriate reverb category is estimated, and a prediction is output to the processing unit, and a matching reverb is added to the audio output signal based on the prediction.

Benefits of technology

The addition of appropriate reverb to the audio signal with relatively less computational amount is achieved, which significantly improves the quality of the listening experience, reduces artifacts, and the generated audio signal is satisfactory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075696A_ABST
    Figure CN120075696A_ABST
Patent Text Reader

Abstract

An audio system includes a processing unit and a reverberation classification unit configured to receive a first plurality of audio input signals, estimate a reverberation category suitable for the first plurality of audio input signals by a deep learning (DL) classification algorithm, and output a prediction to the processing unit, the prediction comprises information about an estimated reverberation category, and the processing unit is configured to receive the first plurality of audio input signals, generate a second plurality of audio output signals based on the first plurality of audio input signals, and output the second plurality of audio output signals, wherein generating the second plurality of audio output signals includes adding reverberation to at least one of the second plurality of audio output signals based on the prediction received from the reverberation classification unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an audio system and related methods, and more particularly to an audio system and method for adding reverberation to an audio signal. Background Art

[0002] By expanding an audio signal with surround or 3D information, for example, adding a reverberation effect that matches the reverberation already present in the audio signal, thereby simulating a certain listening environment, the listening experience of the user to whom the audio signal is presented can be significantly improved. For example, in the context of an upmixing process, the audio signal can be expanded by adding reverberation that matches the original audio signal or by creating additional reverberation channels or signals that match the existing audio signal. However, it can be challenging to acoustically simulate a concert hall or any other kind of listening space by appropriately adding and reproducing matching reverberation. The resulting audio signal may include unwanted artifacts, may not sound satisfactory, and may require a relatively high computational load to generate the expanded audio signal.

[0003] There is a need for an audio system and related methods that allow for simulating a listening environment by expanding an audio signal with surround or 3D information, thereby providing a very satisfactory listening experience for the listener while only requiring a relatively small amount of computational power. Summary of the Invention

[0004] An audio system includes a processing unit and a reverberation classification unit, wherein the reverberation classification unit is configured to receive a first plurality of audio input signals, estimate a reverberation class suitable for the first plurality of audio input signals by a deep learning DL classification algorithm, and output a prediction to the processing unit, the prediction including information about the estimated reverberation class, and the processing unit is configured to receive the first plurality of audio input signals, generate a second plurality of audio output signals based on the first plurality of audio input signals, and output the second plurality of audio output signals, wherein generating the second plurality of audio output signals includes adding reverberation to at least one of the second plurality of audio output signals based on the prediction received from the reverberation classification unit.

[0005] A method includes: estimating a reverberation class suitable for a first plurality of audio input signals by a deep learning DL classification algorithm and making a prediction including information about the estimated reverberation class; generating a second plurality of audio output signals based on the first plurality of audio input signals, wherein generating the second plurality of audio output signals includes adding reverberation to at least one of the second plurality of audio output signals based on the prediction; and outputting the second plurality of audio output signals.

[0006] In studying the following specific embodiments and the drawings, other systems, features, and advantages of the present disclosure will be or will become apparent to those skilled in the art. It is intended that all such additional systems, methods, features, and advantages be included in this specification, be within the scope of the invention, and be protected by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The apparatus and method can be better understood with reference to the following description and the drawings. The components in the drawings are not necessarily drawn to scale, but rather emphasis is placed on illustrating the principles of the invention. Additionally, in the drawings, the same reference numerals indicate corresponding parts in all different views.

[0008] Figure 1 An audio system according to an embodiment of the present disclosure is schematically illustrated.

[0009] Figure 2 A reverberation classification unit according to an embodiment of the present disclosure is schematically illustrated.

[0010] Figure 3 Different steps performed in the reverberation classification unit according to an embodiment of the present disclosure are schematically illustrated.

[0011] Figure 4 An audio system according to a further embodiment of the present disclosure is schematically illustrated.

[0012] Figure 5 An audio system according to a further embodiment of the present disclosure is schematically illustrated.

[0013] Figure 6 A method according to an embodiment of the present disclosure is schematically illustrated. DETAILED DESCRIPTION

[0014] The audio systems and related methods according to the various embodiments described herein allow for simulating different listening environments by adding reverberation to an audio signal. Only a relatively small amount of computation is required, and the resulting audio signal sounds very satisfactory because it includes few or even no artifacts. The audio systems and methods disclosed herein do not process the audio signal "blindly" as in conventional audio systems, but rather perform "informed" processing on the audio signal.

[0015] Generally, by adding 3D or surround information to an audio signal, the listening experience of a user listening to the audio signal can be significantly improved. In particular, multi-channel playback provides the possibility for the user to feel as if they are in a certain location or event when listening to a musical work. Additional ambient playback can create an atmosphere comparable to the experience of a live event. For example, if the center signal of a multi-channel audio signal is extracted and added to a speaker located at the center, the optimal listening area (so-called sweet spot) can be expanded, and the stability of the front image can be significantly improved. Using 3D speakers, the true immersive feeling of the audio event can be further enhanced. It is even possible to lift the stage and create an overhead playback effect.

[0016] To improve the quality of the reproduced sound scene, the perception of the sound scene is generally modeled as a combination of foreground sound and background sound, which are also commonly referred to as the primary (or direct) component and the ambient (or diffuse) component, respectively. The primary component consists of point-like directional sound sources, while the ambient component generally consists of diffuse ambient sounds (reverberation). Due to the perceptual differences between the primary component and the ambient component, different rendering schemes are usually applied to the primary component and the ambient component to achieve the best spatial audio reproduction of the sound scene. However, channel-based audio only provides a mixed signal. Therefore, some methods focus on extracting the primary component and the ambient component from the mixed signal. Known methods (which may include, for example, ambience estimation in the frequency domain) usually require a large amount of computational effort, and the resulting audio signal is usually not satisfactory because it includes a large number of artifacts.

[0017] Compared with conventional methods, the audio systems and related methods disclosed herein use artificial intelligence (deep learning DL) algorithms to classify the spatial components of an existing audio signal, and further use this information to create an artificial reverberation that matches the original reverberation. That is, instead of extracting the ambient component from the mixed signal, the environment in which a musical work may have been recorded is estimated by artificial intelligence. Musical works usually include specific reverberation categories. The reverberation included in a musical work usually has certain characteristics typical of a specific reverberation category, for example, a long reverberation tail, specific room modes, early reflections, etc. As a result, reverberation can be added to the audio signal that matches the reverberation included in the musical work. That is, for example, based on the estimated reverberation category, a matching artificial reverberation is added to the audio signal.

[0018] Now referring to Figure 1 , an audio system according to an embodiment of the present disclosure is schematically shown. The audio system includes a processing unit 100 and a reverberation classification unit 200. The reverberation classification unit 200 is configured to receive a first plurality of audio input signals IN 1 、...、IN N , estimate, by a deep learning DL classification algorithm, a reverberation suitable for the first plurality of audio input signals IN 1 、...、INN the reverberation category and outputs a prediction P1 to the processing unit 100, the prediction P1 including information about the estimated reverberation category. The processing unit 100 is configured to receive a first plurality of audio input signals IN 1 、...、IN N and, based on the first plurality of audio input signals IN 1 、...、IN N generate a second plurality of audio output signals OUT 1 、...、OUT M and output the second plurality of audio output signals OUT 1 、...、OUT M where generating the second plurality of audio output signals OUT 1 、...、OUT M includes adding reverberation to at least one of the second plurality of audio output signals OUT 1 、...、OUT M based on the prediction P1 received from the reverberation classification unit 200.

[0019] In Figure 1 , a single input signal IN and a single output signal OUT are schematically shown. However, there may be more than one input signal IN and more than one output signal OUT, as schematically shown in Figure 4 . According to one example, the first plurality of audio input signals IN 1 、...、IN N can be two channels of a stereo audio signal (e.g., left (L) channel and right (R) channel). The second plurality of audio output signals OUT 1 、...、OUT M can be different channels of an upmixed 5.1 surround signal (e.g., front left (FL) channel, front right (FR) channel, center (C) channel, left surround (LS) channel, right surround (RS) channel). However, any other numbers N, M of the input signals IN 1 、...、IN N and output signals OUT 1 、...、OUT M are also possible, where M ≥ N. Reverberation can be added to each of the plurality of audio output signals OUT 1 、...、OUT M , or only to some of the plurality of audio output signals OUT 1 、...、OUT M . In a 5.1 surround signal, reverberation may only be added to the left surround (LS) channel and the right surround (RS) channel of the upmixed 5.1 surround signal, which is just one example.

[0020] The reverberation classification unit 200 is configured to estimate a reverberation class suitable for the first plurality of audio input signals IN 1 、...、IN N (a reverberation class matching the first plurality of audio input signals IN 1 、...、IN N ). According to some embodiments of the present disclosure, estimating a reverberation class suitable for the first plurality of audio input signals IN 1 、...、IN N includes separating the first plurality of audio input signals IN 1 、...、IN N into a plurality of consecutive individual frames, extracting one or more features from each of the individual frames, each of the one or more features being a characteristic of one of a plurality of types of listening environments, identifying a specific pattern in each of the individual frames through the extracted features, and estimating a reverberation class suitable for each of the individual frames based on the identified specific pattern. The length of each frame may affect the accuracy of the deep learning DL classification algorithm and can be selected in any suitable manner.

[0021] Reference Figure 2 , schematically shows a reverberation classification unit 200 according to an embodiment of the present disclosure. The reverberation classification unit 200 may include a receiving unit 202 configured to receive the first plurality of audio input signals IN 1 、...、IN N . The reverberation classification unit 200 may further include a preprocessing unit 204. The preprocessing unit 204 may be configured to preprocess the input audio signals according to the requirements of the deep learning DL classification algorithm in a first step. The preprocessing unit 204 may resample the audio input signals to a target sampling rate and slice the audio input signals into frames with a defined length. The length of each frame may correspond to a desired time window for the output prediction. The reverberation classification unit 200 may further include a feature extraction unit 206 configured to transform the individual signal frames into a signal representation that allows the classification algorithm to more easily identify patterns in the frames that are related to the amount of reverberation present in the corresponding frames of the audio input signals.

[0022] For example, a suitable signal representation may include a time-frequency representation, such as, for example, a log-frequency spectrogram. This is schematically shown in Figure 3 . Figure 3Schematically shows that a logarithmic frequency spectrogram can be generated for each of the different frames into which an audio input signal is sliced. A (logarithmic frequency) spectrogram is a standard sound visualization tool that allows the visualization of the energy distribution in both time and frequency. A logarithmic frequency spectrogram is simply an image formed by the magnitude of the short-time Fourier transform on a logarithmic intensity axis (e.g., dB).

[0023] Then, the transformed audio input signal frames (i.e., input samples) can be processed in batches in the classification unit 208 by a classification algorithm. This results in predicting the appropriate reverberation class for each of the audio input signal segments (frames). The reverberation classification unit 200 may also include a prediction unit 210 that is configured to make a prediction P1 and output it to the processing unit 100.

[0024] The reverberation classification unit 200 can be configured to make a global prediction P1 based on the estimated reverberation classes in a defined plurality of individual consecutive frames. The defined plurality of individual consecutive frames can constitute a musical work, and the global prediction P1 can include information about the estimated reverberation class suitable for the entire musical work. That is, a musical work can be separated into a plurality of individual frames. An appropriate reverberation can be determined for each of the plurality of individual frames. For example, the appropriate reverberation can be "high reverberation" or "low reverberation". If it is determined that within the plurality of individual frames of a musical work, the result "high reverberation" dominates compared to "low reverberation", then high reverberation can be added to the entire musical work, and vice versa. However, "high reverberation" and "low reverberation" are merely examples. Other reverberation classes can include but are not limited to, for example, "large / medium / small jazz hall", "large / medium / small living room", "wooden large / medium / small concert hall", etc. Other more general reverberation classes can include "hall 1", "hall 2", "hall 3", etc.

[0025] Alternatively, a sub-prediction P1 can also be made based on the estimated reverberation classes in a subset of the defined plurality of individual consecutive frames. The defined plurality of individual consecutive frames can constitute a musical work, and the sub-prediction P1 can include information about the estimated reverberation class suitable for a part of the musical work. That is, a musical work can be separated into a plurality of individual frames. An appropriate reverberation can be determined for each of the plurality of individual frames. For example, the appropriate reverberation can be "high reverberation" or "low reverberation". Different reverberations can be added to each of the different frames of the audio input signal based on the corresponding predictions. However, a combined prediction can also be made for several but not all of the individual frames.

[0026] Reference Figure 3 , an appropriate reverberation can be determined for each of the plurality of frames, resulting in Figure 3 the prediction shown on the right. Then several frames (e.g., frames 1 to 5 (or any other number of frames)) can be combined into a group of frames, and a sub-prediction P1 can be determined and output for that group of frames. If inFigure 3 In the example shown, if frames 1 to 5 are combined, the result "high reverberation" dominates. That is, the sub-prediction P1 "high reverberation" can be output to the processing unit 100, and then this processing unit can add "high reverberation" to the corresponding frames 1 to 5. However, a musical work may include more than just frames 1 to 5. That is, different reverberations can be added to different segments of the musical work, with each segment including more than one frame. This allows reverberation to be added to the musical work in a very accurate manner, thus bringing a very satisfactory listening experience.

[0027] The deep learning DL classification algorithm can be based on a deep learning DL model, where this DL model is trained using annotated data consisting of audio signals with different known reverberation levels. For example, the DL model can learn hierarchical representations from input samples. To be able to predict the reverberation category with high accuracy, it can be trained using annotated data consisting of audio signals with different known reverberation levels. For example, the reverberation level can be measured perceptually. Usually, one or more different databases can be used for this purpose. The one or more databases can be obtained in any suitable way.

[0028] The audio system described herein is capable of directly classifying the amount of reverberation present in an audio input signal, which is directly consistent with the perceptual measurement of reverberation. The audio signal is highly flexible in terms of the amount of estimable reverberation categories. According to one example, two reverberation categories can be estimated, for example, "high reverberation" and "low reverberation". According to another example, three reverberation categories can be estimated, for example, "high reverberation", "medium reverberation", and "low reverberation", where "medium reverberation" is a reverberation that is less than "high reverberation" and greater than "low reverberation". Usually, any other intermediate reverberation categories between "high reverberation" and "low reverberation" can also be estimated. As mentioned above, other additional or alternative reverberation categories can include, but are not limited to, for example, "large / medium / small jazz hall", "large / medium / small living room", "wooden large / medium / small concert hall", "hall 1", "hall 2", "hall 3", etc.

[0029] For example, the audio system described above can be a surround sound system, or any kind of 3D audio system (e.g., VR / AR applications). That is, the first plurality of audio input signals IN 1 、...、IN N The number of audio input signals included in may be equal to the number of audio signals included in the second plurality of audio output signals OUT 1 、...、OUT M as shown in Figure 1 However, the number of audio input signals included in the first plurality of audio input signals IN 1 、...、IN N may also be less than the number of the second plurality of audio output signals OUT1 、...、OUT M the number of audio signals included in, such as Figure 4 and 5 schematically shown (N < M). That is, the processing unit 100 can be or can include an upmix processor 102, as Figure 4 exemplarily shown in

[0030] Figure 5 A surround sound system is schematically shown. The surround sound system includes a stereo source 30. The processing unit 100 can receive a first plurality of audio input signals IN from the stereo source 30 (for example, two channels of a stereo audio signal (left (L) channel and right (R) channel)) 1 、...、IN N . A second plurality of audio output signals OUT 1 、...、OUT M can be different channels of an upmixed 5.1 surround signal (for example, front left (L') channel, front right (R') channel, center (C) channel, surround left (LS) channel, surround right (RS) channel). A plurality of speakers can be arranged in the listening environment 50. The speakers can be arranged at appropriate positions relative to the listener 40 present in the listening environment 50. Different channels of the audio output signal OUT can be fed to the corresponding speakers. As Figure 4 and Figure 5 exemplarily shown, the surround sound system can generate an artificial sense of space (ambience) in order to achieve acoustic enclosure of the listener 40 in the listening environment 50. The output signals generated by the surround sound system can be routed to the surround channels and height channels of a multi-channel speaker system

[0031] Reverb added to one or more audio output signals OUT 1 、...、OUT M can be generated, for example, by a reverb engine 104 based on a prediction P1, as Figure 4 exemplarily shown in. For example, the spatial information (prediction P1) obtained and provided by the reverb classification unit 200 can be used to adjust the parameters of an artificial reverb generator algorithm (reverb engine 104). That is, for different predictions P1 received from the reverb classification unit 200, the reverb engine 104 can apply different parameters when generating artificial reverb. Then, the artificially generated reverb can be fed to the distribution block 106 of the processing unit 100. The distribution block 106 can route the reverb ("ambience" signal component) to the desired speakers (for example, the surround speakers and / or height speakers of the surround sound system) according to the desired upmix settings

[0032] However, the processing unit 100 may also include or be coupled to a memory 110, in which different types of reverberations are stored. The processing unit 100 (i.e., the reverberation engine 104 of the processing unit 100) may retrieve a suitable reverberation from the memory 110 based on the prediction P1 and add it to one or more audio output signals OUT accordingly 1 、...、OUT M 。

[0033] Reference Figure 6 According to an embodiment of the present disclosure, a method includes: estimating a reverberation category suitable for a first plurality of audio input signals IN 1 、...、IN N using a deep learning DL classification algorithm, and making a prediction P1 including information about the estimated reverberation category (step 601); and generating a second plurality of audio output signals OUT 1 、...、OUT N based on the first plurality of audio input signals IN 1 、...、OUT M wherein generating the second plurality of audio output signals OUT 1 、...、OUT M includes adding a reverberation to at least one of the second plurality of audio output signals OUT 1 、...、OUT M based on the prediction P1 (step 602). The method further includes outputting the second plurality of audio output signals OUT 1 、...、OUT M (step 603).

[0034] According to some embodiments of the present disclosure, estimating a reverberation category suitable for a first plurality of audio input signals IN 1 、...、IN N may include separating the first plurality of audio input signals IN 1 、...、IN N into a plurality of consecutive individual frames, extracting one or more features from each of the individual frames, each of the one or more features being a characteristic of one of a plurality of types of listening environments, identifying a specific pattern in each of the individual frames through the extracted features, and estimating a reverberation category suitable for each of the individual frames based on the identified specific pattern.

[0035] According to some embodiments, the method may further include transforming the individual signal frames into logarithmic frequency spectrograms after separating the first plurality of audio input signals IN 1 、...、IN N into a plurality of consecutive individual frames and before extracting one or more features from each of the individual frames.

[0036] The artificially generated sense of space matches the (possible) sense of space present in the original audio signal (the first plurality of audio input signals IN 1 、...、IN N ). In an ideal situation, it has the same room acoustic properties. Classical ambience extraction methods (which can be used to extract the original ambience signal from the input signal) are algorithmically complex and result in perception-related artifacts. Multiplying the extracted environmental signal and distributing it to the speaker channels of a multichannel system further increases the perceived artifacts. The audio system described above overcomes these drawbacks. The reverberation category information of the ambience part of the original signal is used to control or configure the algorithm for the artificially generated output ambience signal. The way the artificial ambience component is created generally does not matter. The artificial ambience component can generally be created in any suitable way. Any type of reverberation method can generally be used, for example, a feedback delay network (FDN), convolutional reverberation, etc.

[0037] Based on the requirements of a specific application (e.g., upmixing techniques), the number and semantic properties of the reverberation categories are adjusted, and a deep learning DL classification network can be trained accordingly. Additionally, categories can be defined based on perception-related characteristics (e.g., reverberation length, reflection density, early reflection pattern, dry-wet ratio, decay rate, spectral behavior, etc.), and the artificial reverberation algorithm can be configured accordingly. For example, if a DL network is trained to estimate the reverberation length in an input signal, the output prediction can be used to set the same parameter in the artificial reverberation engine.

[0038] It is understood that the system shown is merely an example. Although various embodiments of the present invention have been described, it will be apparent to those of ordinary skill in the art that there can be more embodiments and implementations within the scope of the present invention. Specifically, those skilled in the art will recognize the interchangeability of various features from different embodiments. Although these technologies and systems have been disclosed in the context of certain embodiments and examples, it should be understood that these technologies and systems can extend beyond the specifically disclosed embodiments to other embodiments and / or their uses and obvious modifications. Therefore, the present invention is not limited except in accordance with the appended claims and their equivalents.

[0039] The description of the embodiments has been presented for purposes of illustration and description. Suitable modifications and variations of the embodiments can be effected in accordance with the above description or can be obtained through practice of the methods. The arrangements are exemplary in nature and may include additional elements and / or omit elements. As used in this application, an element recited in the singular and preceded by the word "a" or "an" should be understood as not excluding a plurality of such elements, unless such exclusion is stated. Furthermore, reference to "an embodiment" or "an example" of the present disclosure is not intended to be construed as excluding the existence of additional embodiments that also incorporate the recited features. The terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements or a particular positional order on their objects. The described systems are exemplary in nature and may include additional elements and / or omit elements. The subject matter of the present disclosure includes all novel and non-obvious combinations and subcombinations of the various systems and configurations and other features, functions, and / or properties disclosed. The appended claims particularly point out the subject matter that is regarded as novel and non-obvious from the foregoing disclosure.

Claims

1. An audio system comprising a processing unit (100); and A reverberation classification unit (200), wherein The reverberation classification unit (200) is configured to receive a first plurality of audio input signals (IN1, . . . , IN N ), estimating suitable for the first plurality of audio input signals (IN1, ..., IN2) through a deep learning DL classification algorithm N) and outputting a prediction (P1) to the processing unit (100), the prediction (P1) comprising information about the estimated reverberation class, and The processing unit (100) is configured to receive the first plurality of audio input signals (IN1, . . . , IN N ), based on the first plurality of audio input signals (IN1, ..., IN N ) generates a second plurality of audio output signals (OUT1, ..., OUT M ), and output the second plurality of audio output signals (OUT1, . . . , OUT M ), wherein the second plurality of audio output signals (OUT1, . . . , OUT M ) comprises providing, based on the prediction (P1) received from the reverberation classification unit (200), to the second plurality of audio output signals (OUT1, . . . , OUT M ) Add reverb.

2. The audio system of claim 1, wherein the estimation is suitable for the first plurality of audio input signals (IN1, ..., IN N )'s reverb categories include: The first plurality of audio input signals (IN1, . . . , IN N ) is separated into multiple consecutive individual frames, extracting one or more features from each of the individual frames, each of the one or more features being characteristic of one of a plurality of types of listening environments, identifying specific patterns in each of the individual frames through the extracted features, and A reverberation class appropriate for each of the individual frames is estimated based on the identified specific pattern.

3. The audio system of claim 2, further comprising: After the first plurality of audio input signals (IN1, . . . , IN N ) into a plurality of consecutive individual frames and before extracting one or more features from each of the individual frames, each of the plurality of consecutive individual frames is transformed into a logarithmic frequency spectrogram.

4. An audio system as claimed in claim 2 or 3, further comprising making a global prediction (P1) based on the estimated reverberation class in a defined plurality of individual consecutive frames.

5. An audio system as claimed in claim 4, wherein the defined plurality of individual consecutive frames constitutes a musical piece and the global prediction (P1) comprises information about the estimated reverberation class suitable for the musical piece.

6. An audio system as claimed in claim 2 or 3, further comprising making a sub-prediction (P1) based on the estimated reverberation class in a defined subset of a plurality of individual consecutive frames.

7. An audio system as claimed in claim 6, wherein the defined plurality of individual consecutive frames constitutes a musical piece and the sub-prediction (P1) comprises information about the estimated reverberation class that is suitable for a part of the musical piece.

8. An audio system as claimed in any one of the preceding claims, wherein the first plurality of audio input signals (IN1, ..., IN N ) is equal to the number of the second plurality of audio output signals (OUT1, . . . , OUT2, . . . , OUT3, . . . ) M ) includes the number of audio signals.

9. An audio system as claimed in any one of the preceding claims, wherein the first plurality of audio input signals (IN1, ..., IN N ) is smaller than the number of the second plurality of audio output signals (OUT1, . . . , OUT M ) includes the number of audio signals.

10. The audio system of claim 9, wherein the first plurality of audio input signals (IN1, . . . , IN2 N ) is composed of two channels (L, R) of a stereo audio signal, and wherein the second plurality of audio output signals (OUT1, . . . , OUT M ) consists of five channels (FL, FR, C, LS, RS) of upmixed 5.1 surround signals.

11. An audio system as claimed in any one of the preceding claims, wherein the deep learning (DL) classification algorithm is based on a deep learning (DL) model, wherein the DL model is trained using annotated data consisting of audio signals with different known reverberation levels.

12. An audio system as claimed in any one of the preceding claims, wherein suitable for the first plurality of audio input signals (IN1, ..., IN N )'s estimated reverberation category is one of low reverberation, medium reverberation, and high reverberation.

13. A method comprising: The deep learning (DL) classification algorithm is used to estimate the first plurality of audio input signals (IN1, . . . , IN2, IN3, IN4, IN5, IN6, IN7, IN8, IN9, IN10, IN111, IN122, IN13, IN143, IN154, IN165, IN176, IN187, IN190, IN201, IN212, IN230, IN248, IN310 N ) and making a prediction (P1) including information about the estimated reverberation class, Based on the first plurality of audio input signals (IN1, . . . , IN N ) generates a second plurality of audio output signals (OUT1, ..., OUT M ), wherein the second plurality of audio output signals (OUT1, . . . , OUT M ) comprises providing, based on the prediction (P1), to the second plurality of audio output signals (OUT1, . . . , OUT M ) adds reverberation, and Output the second plurality of audio output signals (OUT1, . . . , OUT M ).

14. The method of claim 13, wherein estimating a first plurality of audio input signals (IN1, ..., IN N )'s reverb categories include: The first plurality of audio input signals (IN1, . . . , IN N ) is separated into multiple consecutive individual frames, extracting one or more features from each of the individual frames, each of the one or more features being characteristic of one of a plurality of types of listening environments, identifying specific patterns in each of the individual frames through the extracted features, and A reverberation class appropriate for each of the individual frames is estimated based on the identified specific pattern.

15. The method of claim 14, further comprising: N ) into a plurality of consecutive individual frames and before extracting one or more features from each of the individual frames, the individual signal frames are transformed into logarithmic frequency spectrograms.