Acoustic signal processing method

JP2024528735A5Pending Publication Date: 2025-07-01アリール·ベスローテン·フェンノートシャップ
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024505561
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-30
Filing Date
2022-07-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing methods for upmixing stereo audio signals to create three-dimensional soundscapes often compromise sound quality due to non-linear phase responses, uneven frequency distribution, and reliance on stereo channel relationships, leading to poor user experiences.

Method used

A computer-implemented method using Mid-Side decoding, filtering banks with linear phase response, and high-pass filtering to generate spatially distributed pseudo-surround channels, maintaining original sound characteristics without adding reverb, and applying amplitude compensation and polarity inversion to achieve high-quality 3D upmixing.

Benefits of technology

The method preserves the clarity and spatial presence of the original sound, providing an immersive and focused 3D audio experience by accurately positioning sound objects and maintaining tonal characteristics, enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented audio signal processing method for upmixing an input audio stereo signal (S) into a set of multi-channel output signals (O).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a computer-implemented audio signal processing method for upmixing an input audio stereo signal (S) into a set of multi-channel output signals (O). [Background technology]

[0002] In many applications, it is desirable to generate a three-dimensional soundscape that can embed multiple directional sound sources to simulate reality in order to enhance the user's perception. However, most prior art methods simply rely on regular stereo feeds to attempt to create an articulated, multi-dimensional soundscape. Such attempts tend to suffer from poor sound quality and a poor user experience due to the inherent presence of artifacts scattered throughout the soundscape.

[0003] Since the advent of stereo, music and media productions have tended to allocate side information in the feed to promote enhanced spatial characteristics of the song / recording. However, most prior art methods rely on regular stereo feeds to create an articulated soundscape. Existing upmixing methods generally compromise the quality of the user experience. To this extent, the user experiences significant spatial orientation differences resulting from the excess information contained in the height layer, leading to a poor user experience.

[0004] Most of the methods presented in the prior art rely on processing filters that do not achieve a linear phase response in the frequency domain, resulting in directional sound artifacts in the soundscape, ultimately degrading the user experience and compromising the multidimensionality of the feed. In this range, the non-linear phase response results in a soundscape that sounds over-processed to the final user. In addition, most of the methods presented in the prior art rely on processing filters that operate in a wide frequency range that spans outside the human audible frequency support. This results in parts of the frequency spectrum being uneven in terms of amplitude in a given direction, resulting in a poor user experience as the sound is perceived as being unevenly distributed in space.

[0005] Most of the methods presented in the prior art rely on the relationship between the Left and Right stereo channels to achieve the representation of the Center channel, which results in the user's perception being split towards the Left and Right sides. More specifically, this approach may not accurately position important sound items in the soundscape and may not preserve their timbral characteristics, which may result in an over-processed feeling for the user. Ultimately, this results in a poor user experience for important sound objects such as vocals.

[0006] Existing upmixing methods usually use reverb to create a sense of space. This is achieved by adding artificial information to the original signal, but in return the original sound loses its clarity and the spatial presence of the sound changes unnaturally, hindering the user experience. In addition, reverb can only be used in spaces with limited size, making it unsuitable for large live settings. Summary of the Invention [Problem to be solved by the invention]

[0007] Therefore, what is needed is a method for rendering an input acoustic stereo signal into multiple spatially distributed pseudo surround channels to enhance the user experience. [Means for solving the problem]

[0008] Surprisingly, the inventors have found that one or more of these problems can be solved by the present invention and its embodiments. The method allows detailed presets of new insights and features that meet the high surround sound standards required. The present invention respects the creative process of sound design. The present invention can preserve the original characteristics of the music, since it does not need reverb to create a sense of space. There is no need to add effects to the music, it is just distributed spatially (evenly).

[0009] The present invention provides a computer-implemented audio signal processing method for upmixing an input audio stereo signal (S) into multiple spatially distributed pseudo-surround channels to define a height layer, said method preferably comprising the following steps: - receiving at least one input acoustic stereo signal (S); - performing a pre-processing step on an input acoustic stereo signal (S), said pre-processing step comprising: performing mid-side decoding to generate at least one Sum (SUM) signal and at least one Difference (DIFF) signal; performing a polarity inversion on at least one Difference (DIFF) signal; performing filtering of at least one Difference (DIFF) signal with at least two, preferably at least four, filtering banks (PFs); carrying out a pre-treatment process, comprising: - reconstructing at least two, preferably at least four, signals from a filtering bank (PF) to obtain an upmixed output signal (O); - performing a high-pass filtering on at least one upmixed output signal (O), preferably on all upmixed output signals (O); - performing a level adjustment for at least one upmixed output signal (O), preferably for all upmixed output signals (O); - Routing the upmixed reconstructed acoustic signal (O) to the audio speaker channels (C) to define a matrix of spatially distributed channels forming a height layer by feeding them into a Top Channel, for example at least a Top Front Left channel (TFL), a Top Front Right channel (TFR), a Top Rear Left channel (TRL) and a Top Rear Right channel (TRR).

[0010] In some embodiments, the filtering bank (PF) is configured to have a linear phase response in the frequency domain.

[0011] In some embodiments, the filtering bank (PF) is configured to operate around filter sub-bands (PSB), each of which has a center frequency FSB-C and is configured to operate around a range of low-frequency sounds above a low cut-off frequency FSB-L, and a range of high-frequency sounds below a high cut-off frequency FSB-U.

[0012] In some embodiments, each of the filter sub-bands (PSB) is configured to have an amplitude about a sub-band center frequency FSB-C selected in the range of -3 dB to -15 dB, preferably -6 dB to -12 dB, more preferably -9 dB.

[0013] In some embodiments, the filtering bank (PF) is configured to have a width between 1 / 9 octave and 1 octave.

[0014] In some embodiments, the operating frequency range of the filter sub-band (PSB) is F L ~F U and configured to operate within a frequency range of F L ~F U The frequency is 350 Hz to 20 kHz, preferably 400 Hz to 10 kHz, and more preferably 500 Hz to 9 kHz.

[0015] In some embodiments, the amplitude compensation is performed outside the frequency support of a filtering bank (PF), said amplitude compensation being F L and F U and the amplitude compensation is performed on the amplitude level at the resulting F L and F U With amplitude levels around these are selected in the range of -3dB to -12dB, preferably -6dB to -9dB, more preferably within -6dB.

[0016] In some embodiments, high-pass filtering is performed on each of the height channels using a high-pass filter (HPF), such high-pass filter (HPF) configured to operate with a center frequency FHFC=500 Hz.

[0017] In some embodiments, the high-pass filter is a high-shelf filter and has a linear phase response in the frequency domain.

[0018] In some embodiments, a level adjustment is performed on each of the upmixed output audio signals.

[0019] In some embodiments, the processing time is on the order of milliseconds, preferably less than 5 ms, more preferably less than 3 ms, and even more preferably less than 1 ms.

[0020] In some embodiments, synchronization is performed during the Mid-Side decoding step, and latency compensation is performed for input channels that are not subject to the Mid-Side decoding step.

[0021] In some embodiments, the method further comprises the steps of: - performing a delay adjustment (D-ADJ) on at least one (SUM) signal; - defining a matrix of spatially distributed channels forming a center layer by routing the upmixed reconstructed acoustic signal (O) to audio speaker channels (C) and feeding at least a Center channel (CE) and a Low Frequency Effect (LFE) channel; - performing low pass filtering (LPF) on the LFE channel.

[0022] In some embodiments, the method further comprises performing compensation filtering on the obtained upmixed output signal (O).

[0023] In some embodiments, the compensation filters are low and / or high shelf filters and have a linear phase response in the frequency domain. [Brief description of the drawings]

[0024] [Figure 1] 4 is a graph illustrating a filtering bank according to a preferred embodiment of the present invention. [Diagram 2] FIG. 2 is a block scheme diagram according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] The present invention will be described with respect to particular embodiments but the invention is not limited thereto but only by the claims, and any reference signs in the claims are not to be construed as limiting the scope thereof.

[0026] As used herein, the singular forms "a," "an," and "the" include both singular and plural referents unless the context clearly dictates otherwise.

[0027] As used herein, the terms "comprising," "comprises," and "comprised of" are synonymous with "including," "includes," or "containing," and are inclusive or open-ended and do not exclude additional, unrecited elements, elements, or method steps. The terms "comprising," "comprises," and "comprised of" when referring to recited elements, elements, or method steps also include embodiments that "consist of" the recited elements, elements, or method steps.

[0028] Moreover, the terms first, second, third, etc. in this specification and claims are used to distinguish between similar elements, unless expressly stated otherwise, and are not necessarily used to describe a sequential or chronological order. The terms so used are interchangeable under appropriate circumstances, with it being understood that the embodiments of the invention described herein are capable of operating in sequences other than those described or illustrated herein.

[0029] As used herein, the term "about" when referring to a measurable value such as a parameter, amount, time duration, and the like, is intended to encompass variations from the stated value of no more than ±10%, preferably no more than ±5%, more preferably no more than ±1%, and even more preferably no more than ±0.1%, to the extent that such variations are appropriate for practice in the disclosed invention. It is to be understood that the value to which the modifier "about" refers is itself specifically and preferably disclosed.

[0030] The recitation of numerical ranges by endpoints includes not only the recited endpoints but also all values ​​and fractions subsumed within each range. All documents cited herein are incorporated by reference in their entirety. Unless otherwise defined, all terms used in disclosing the present invention, including technical and scientific terms, have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs.

[0031] As further guidance, definitions of terms used herein are included to better understand the teachings of the present invention. Terms or definitions used herein are provided only to aid in understanding the present invention. References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the present invention.

[0032] Thus, although the phrases "in one embodiment" or "in an embodiment" appear in various places throughout this specification, they do not necessarily all refer to the same embodiment. Furthermore, particular features, structures or characteristics may be combined in any suitable manner in one or more embodiments, as would be apparent to one of ordinary skill in the art from this disclosure. Furthermore, some embodiments described herein include some features included in other embodiments but not other features included in other embodiments, as would be understood by one of ordinary skill in the art, while combinations of features of different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the following claims and description, any of the claimed or described embodiments may be used in any combination.

[0033] When focusing on height layers, there are several important steps that are taken: Changing one or more of these steps may result in a significant difference in the results for the user.

[0034] For the creation of the height layer, the present invention uses a technique known as MS encoding. By MS encoding a stereo signal, there is no longer a left and right signal, but a sum and a difference (side), known as mono. In music production, the width or spatial character of a piece of music is defined by the amount of side information. For example, artificial stereo reverbs create a lot of side information to create a sense of space. Therefore, tracks with a lot of side information are made to feel more spatial, and therefore they are at their best in 3D upmixes. Therefore, the information of the MS encoder in the difference channel is perfectly utilized in the height layer.

[0035] In this invention, we use an MS matrix to generate a difference signal from the left and right stereo signals. This difference signal is used to create the height layer, achieving a true 3D upmix. If the MS matrix were skipped and the height layer were instead created using regular stereo feeds, there would be too much stereo center information in the height speakers. Typically, the most important items in music are in the center of the stereo feed. By not including this information in the height layer, this invention allows these items to maintain the correct focus.

[0036] An MS matrix can be applied to the original stereo signals to produce a mono (sum) of the original left and right signals and a difference (side) signal. The mono sum can be made by simply adding the left and right signals together. The difference signal is made by subtracting the right signal from the left signal. The subtraction is preferably done by inverting the phase of the right signal to create a negative right signal, which is then added to the left signal. The result is a difference signal that contains signals where the left and right signals are not identical. These signals usually contain the "spatial" information of the stereo track. Sounds such as reverb information or extreme panoramic sounds are contained in this difference signal.

[0037] In sound design, the most important features of a music track are placed in the center of the stereo image. This center is called "mono" when the stereo feed is processed with an MS matrix. The difference contains all information except the stereo center. So, using stereo to create a height layer means that important sounds, such as lead vocals, are split into 3D instead of just 2D. When this happens, the user experiences too much spread out from important features of the song, making it difficult to focus on specific sounds.

[0038] In our invention, the combination of a 2D split and mono feed of the center channel is unobtrusive in the low frequency layer of the 3D upmixed signal, and the use of a differential signal for the height speakers does not further degrade important sounds.

[0039] Additionally, in music production, if a song is made to sound big, there is a lot of information in the difference signal. If a song is made to sound small, there is very little detail in the difference signal. As a result, height layering becomes an extension of the creative process designed for each song. A big, epic song will have a sense of height more than a small, compact song, exactly how the song was intended to sound.

[0040] The difference (side) signal is the signal used to generate the height layer. There is preferably one major difference between the left to Left A and Left B processing and the processing used to generate Height A and Height B. The processing used to generate Height A and Height B preferably uses a compensation filter only for the area above the processing range (e.g. above 9 kHz). The area below the processing range is preferably removed from the height speakers by introducing a high-pass filter at the lower boundary frequency of the processing range. Most spatial sounds do not contain low frequencies, so there is no need to generate low frequencies in the height speakers.

[0041] The level of the height speakers is preferably attenuated to match the front and rear speakers. If the level is too high, there will be too much focus on the height speakers and they will end up being distracting. Therefore, the height speakers are preferably attenuated, for example by 5 dB. The exact amount may depend on the speakers used in the surround setup.

[0042] To avoid too much interaction between the front and front height speakers, the Height B signal is preferably routed to the front height speakers, which can be perfectly combined with the Left A and Right A signals generated by the front speakers.

[0043] Since the difference signal is also present in the original stereo signal, undesirable interactions may appear between the height and front speakers. To avoid this, it is preferable to apply B processing to the height speakers and A processing to the front speakers. Similar processing is applied to the rear speakers and the rear height speakers. Since only the rear speakers have Left B and Right B signals, the Height A signal is introduced into the height rear channel.

[0044] In this situation, the rear height speakers generate a partially identical signal diagonally towards the front speakers. Since these speakers are usually facing each other, unpleasant interactions between the speakers can occur in the center of the setup. To solve this problem, it is preferable to invert the polarity of all signals routed to the height speakers. When these signals are combined with the front speaker signals, they add instead of subtract. It is especially preferable to perform polarity inversion for the height channels, since the low-frequency layer front speakers and the high-frequency layer rear speakers generate a (small amount of) identical signal, and since these speakers are facing each other. As a result, no degradation of sound quality occurs.

[0045] Therefore, in some embodiments, the signal used in the height layer is polarity inverted towards the low-frequency layer. In a 3D upmix, the processing filters used in the low-frequency level front speakers are preferably identical to the processing filters used in the high-frequency speakers. The signal coming from the difference channel of the MS matrix contains a small amount of stereo signal. Therefore, the rear height speakers and the front low-frequency layer speakers generate signals identical to each other. Since these speakers face each other, the signal can be subtracted in the center of the surround setup. The user will notice a big difference in sound and quality depending on whether they listen while sitting or standing.

[0046] However, by inverting the polarity of the height layer, the signals that meet in the surround center are added together, giving the user a more consistent sound experience along the Z axis.

[0047] In the present invention, the same series of filters designed to create LA, LB, RA, and RB are preferably used for the height speakers, with the difference that the low shelf filter (starting at 502 Hz for example) is preferably changed to a high pass filter (at 502 Hz for example) for the height channel, since there is no need for low frequencies in the height channel. The series of filters used in the height channel (either A or B processing) is preferably the opposite series to that used in the low frequency layer. For example, the height front left and right channels generate side information processed by the B series of filters. Thus, the low frequency layer and the high frequency layer cooperate perfectly. This also applies to the high rear left and right channels, which generate side information processed by the A series of filters.

[0048] The height layer in a 3D upmix is ​​additive to the low-frequency layer, creating a more immersive feeling in the music and a more robust stereo mix than a 2D upmix. However, it is preferable that the height layer remains additive and does not become the main source of sound. Therefore, the signal going to the height layer is preferably attenuated. If the signal is not attenuated, the user may experience a disruption in the signal of the low-frequency layer, which contains all the main sounds in a typical musical production. This sounds unpleasant and does not meet expectations. However, if the height layer is attenuated according to the preferred embodiment (for example in the range of -3dB to -12dB), the main sources of sound remain focused in the low-frequency layer, and the height layer feels like a more natural addition to the experience.

[0049] In some embodiments, the filtering bank (PF) is configured to have a linear phase response in the frequency domain. Lack of linear phase leads to over-processed sound and sound artifacts, degrading the spatial perception of the final user. The use of linear phase filters is necessary to reach the quality level desired in the present invention. The use of linear phase filters can prevent the phase shift that occurs with conventional filters. Since the phase response of the signal does not change, the interaction becomes natural instead of sounding over-processed.

[0050] The basic principle of the invention consists of splitting one signal into two signals by dividing the frequency. The left signal is split into Left A and Left B. The division from left to Left A and Left B is done by a series of frequency filters, more specifically, the invention employs linear phase filters to eliminate the phase shift between the speakers caused by conventional filters. The filters used have three specific characteristics: amplitude, frequency, and width. Each of these characteristics is related to the other characteristics. The right channel is split into Right A and Right B in the same way that the left signal is split into Left A and Left B. The Left A signal is routed to the front left speaker, the Right A signal is routed to the front right speaker, the Left B signal is routed to the left rear speaker, and the Right B signal is routed to the right rear speaker.

[0051] In a preferred example, a magnitude of -9db is used, a width of 1 / 3 octave is used, and the following exemplary frequencies are used: 502Hz, 652.6Hz, 848.8Hz, 1102.9Hz, 1433.7Hz, 1863.8Hz, 2423Hz, 3149.9H, 4094.9Hz, 5323.3Hz, 6920.3Hz, 8999.4Hz. As exemplified herein, the frequency spacing is preferably 1 / 3 octave. This is directly related to the filter width used. If the filter width becomes narrower, the frequency spacing needs to be adjusted to match the filter width used. However, it is preferable to avoid filters that are too small or too wide. According to some preferred embodiments, each of the filtering banks (PFs) is configured to have a width between 1 / 9 of an octave and 1 octave.

[0052] In some embodiments, the filtering bank (PF) is configured to operate around a filter subband (PSB), F L ~F U and configured to operate within an operating frequency range spanning F L ~FU is preferably between 350 Hz and 20 kHz, preferably between 400 Hz and 10 kHz, more preferably between 500 Hz and 9 kHz. Furthermore, said filtering subbands may be configured to extract at least 4 (PSB) subband signals per filtering bank, preferably 8 (PSB) subband signals, more preferably 16 (PSB) subband signals. Surprisingly, the inventors have found that extracting a sufficient number of subband signals used to reconstruct the audio feed results in an output feed exhibiting an improved dynamic range and improved spatial resolution.

[0053] The average person cannot locate areas below 500Hz in space and above 9kHz. Therefore, the processing used to split frequencies between Left A and Left B is preferably only active in this region. If the filtering goes higher above 9kHz, the open sounds associated with these frequencies will be unevenly split in the listening area, resulting in an unfulfilled "open sound" coverage. If the filtering goes lower below 500Hz, the summation of the low frequencies will be almost nonexistent, resulting in a poor sounding upmixed signal with a soft feel due to this low frequency summation.

[0054] In some preferred embodiments, the amplitude compensation is performed outside the operating band of the sub-band filtering, and such compensation preferably has an amplitude of -9 dB to -3 dB, more preferably about -6 dB. In some preferred embodiments, the compensation filters are low-shelf and high-shelf filters. In the case of a high-layer, the low-shelf filter may be a high-pass filter.

[0055] Since the filters may overlap, there may be an overall amplitude reduction within the processing frequency range. It is therefore desirable to introduce compensation filters above and below the processing range. These filters are preferably low- and high-shelf filters, and are also preferably of linear phase type. When using a magnitude of -9 dB at 1 / 3 width and 1 / 3 spacing, a compensation filter of -6 dB is preferred.

[0056] The frequency of the compensation filter is the boundary frequency of the processing range. As can be seen from the above list, the exemplary boundary frequencies are 502Hz and 8999.4Hz. The compensation filter is set at this frequency, so it will reduce the magnitude at that frequency by 3dB. Therefore, the magnitude of the boundary frequency filter used for processing is preferably set to -6dB so that the sum result of the two is -9dB.

[0057] In some embodiments, a (slight) delay is introduced to the center signal. If the speaker setup is done properly according to the ITU-R BS.775 standard, there is no need for a delay for the center signal. However, in most practical setups, the front-left, front-right and center speakers are physically placed in the same line. In this case, a slight delay (e.g. in the range of 1ms to 5ms) can prevent the focus from being directed towards the center channel instead of evenly across all speakers.

[0058] In some embodiments, the Low Frequency Effect channel (LFE) can also receive the mono sum from the MS matrix. The same principle of the center channel can be applied to the LFE. In addition to the level and delay, a low pass filter can also be introduced. This prevents the LFE from generating frequencies that are too high for the application. The frequency of this filter depends on the frequency response of the speakers used, for example in a 5.1.4 setup. The frequency can vary from 60Hz to 200Hz. A level use for the LFE signal is preferably -9dB, but this may also vary depending on the surround setup.

[0059] In some preferred embodiments, the present invention uses dynamic EQ filters. These filters have fixed frequencies and bandwidths, similar to those described in the previous section. The dynamic filters may interact with the signal fed into them in the magnitude domain. The filters are preferably set up to reduce the magnitude as the input signal rises. In preferred embodiments, a series of filters with two layers is used. The first layer preferably includes static filters with frequencies and bandwidths as described herein, with the magnitude set to -6 dB instead of -9 dB, for example. The second layer preferably includes a series of dynamic filters with a maximum magnitude range of -6 dB, for example.

[0060] The advantage of this technique is that the method can isolate a particular sound that pops out of the piece of music and place it in space. Once the sound fades away and the level drops below the dynamic filter threshold, the method returns to its static position. The result is an organic upmix that interacts with the music, leading to more creative upmix techniques.

[0061] A variation of this technique can be realized using multiband compression. A multiband compressor should be introduced in the second layer of the dynamic upmix and replace the dynamic filter. With a multiband compressor, a specific frequency region can be compressed (attenuated) or expanded (boosted). Most multiband compressors have a wider frequency range to operate in than dynamic filters. A dynamic upmix of the present invention can use a multiband compressor to attenuate a frequency region in the rear speakers and boost the same region in the front speakers, for example. This results in a dynamic interaction between the processing and the music. When a sound pops out of the music, it is projected in front, and when the sound stops, the upmix returns to its static position.

[0062] The method can also be used in a live setting. The method even has the possibility to live interact with the music to contribute on a creative level, whereby the basic principles (e.g. linear phase, filtering from 500Hz to 9khz) remain the same. In some embodiments, the frequency range is separated from the algorithm and the method includes the mono sum of the separated range without applying any processing. The result is a specific frequency range of the song (e.g. 800Hz to 3kHz) that can be separated from the static upmix and moved in the sound field, which can be described hereafter as an "object". This movement can be performed by Vector / Intensity / Layer Based Amplitude Panning type. However, it is preferable not to use a pan delay based system, as this creates a time difference between the object and the upmix. The method can enable or disable the creation of the object. The method can adjust the width of the frequency range of the object. The method can move the object through the sound field. The use of the method goes beyond the previously known stereo experience and adds another creative layer to the upmix algorithm.

[0063] In some embodiments, the method is used on a digital sound processor unit (DSP), or alternatively on a field-programmable gate array (FPGA).

[0064] In some embodiments, the method uses delay compensation for audio channels that are not subject to Mid / Side encoding. In the present invention, the non-mid-side encoded audio channels are subject to an adjustment signal path that introduces delays resulting from processing / adjusting the signal path. The method introduces delay compensation for the non-mid-side encoded signal path to establish time synchronization with the mid-side encoded signal path. The compensation delay is selected in the range of 0.1 ms to 2.0 ms. [Example] The current method and its embodiments also allow for the creation of height layers from 2D surround formats, for example making a 5.1.4 mix from a 5.1 mix.

[0065] The height layer is achieved through the use of an MS matrix. For example, the front height speakers receive the signal generated by the MS matrix placed on the Front Left and Front Right signals of the 5.1 mix. The difference signal is used for the height. A and B filtering is applied to the difference signal to create Front Height A and Front Height B. The Front Height A signal is fed to the Front Height Left speaker and the Front Height B signal is fed to the Front Height Right speaker.

[0066] For the center channel, it is preferable to attenuate the level so that it does not become too present in a 5.1.4 setup, but enough to close the gap between the front left and right speakers. In the example setup, the signal is attenuated by 10 dB. This level may vary depending on the music / content used.

[0067] When upmixing 2D surround formats to 3D, the problem of phase inversion does not necessarily arise. The height layer does not produce a small amount of the same information as the low-level front speakers, so it may not be necessary to phase invert the height layer. However, since front height left and front height right produce exactly the same signal, it is preferable to apply the above processing to the front height signal. This results in two separate channels for the front height speakers.

[0068] In cinema sound design, designers often use front left and right speakers to widen a voice or important sound, for example. That sound is then present equally in both the front left and right signals. Because our method uses MS technology for the height layer, this important sound remains untouched and is not upmixed to 3D. On the other hand, if a cinema sound designer wants to create a sense of space, they use artificial reverb to create that sense. Reverb is primarily present in the difference signal when it is processed with the MS matrix. Therefore, the reverb is upmixed to 3D, which further enhances the intended sense of space. With these basic principles in mind, a variety of 2D to 3D upmixes are possible. For example, 7.1 to 7.1.2 can be created by upmixing the front left / right speakers.

Claims

1. A computer-implemented audio signal processing method for upmixing an input audio stereo signal (S) into a plurality of spatially distributed pseudo surround channels to define a height layer, the method comprising: receiving at least one input audio stereo signal (S); performing a preprocessing step on the input audio stereo signal (S), the preprocessing step comprising: performing Mid-Side decoding to generate at least one Sum (SUM) signal and at least one Difference (DIFF) signal; performing a polarity inversion on the at least one Difference (DIFF) signal; performing filtering on the at least one Difference (DIFF) signal by at least two, preferably at least four, (PF) filter banks; performing a preprocessing step including; reconstructing at least two, preferably at least four, signals from the filter bank (PF) to obtain an upmixed output signal (O); performing high-pass filtering on at least one upmixed output signal (O), preferably on all upmixed output signals (O); performing level adjustment on at least one upmixed output signal (O), preferably on all upmixed output signals (O); routing the upmixed reconstructed audio signal (O) to audio speaker channels (C) to define a matrix of spatially distributed channels forming the height layer by feeding to Top Channels, such as at least Top Front Left channel (TFL), Top Front Right channel (TFR), Top Rear Left channel (TRL) and Top Rear Right channel (TRR); An audio signal processing method including.

2. The method according to claim 1, wherein the filter bank (PF) is configured to have a linear phase response in the frequency domain.

3. The filtering bank (PF) is configured to operate in the vicinity of the filter sub-bands (PSB), and each of these sub-bands (PSB) has a center frequency F SB-C and is configured to operate in the range of low-frequency sound waves higher than the low-pass cut-off frequency F SB-L and in the vicinity of the range of high-frequency sound waves lower than the high-pass cut-off frequency F SB-U The method according to claim 1, wherein the method is configured to operate.

4. Each of the filter sub-bands (PSBs) is configured to have an amplitude in the range of -3 dB to -15 dB, preferably -6 dB to -12 dB, more preferably around the sub-band center frequency F selected from -9 dB. SB-C The method according to claim 3, wherein the method is configured to have an amplitude in the vicinity thereof.

5. The method according to claim 1, wherein each of the filtering banks (PF) is configured to have a width between 1 / 9 octave and 1 octave.

6. The operating frequency range of the filter sub-band (PSB) is configured to operate within the frequency range of F L to F U where F L to F U is from 350 Hz to 20 kHz, preferably from 400 Hz to 10 kHz, more preferably from 500 Hz to 9 kHz, the method according to claim 3.

7. Amplitude compensation is performed outside the frequency support of the filtering bank (PF), and the amplitude compensation is F L and F U is performed on the amplitude levels in, and the amplitude compensation results in the resulting F L and F U with amplitude levels in the vicinity, the amplitude levels being in the range of -3 dB to -12 dB, preferably -6 dB to -9 dB, more preferably selected from -6 dB, according to the method of claim 1.

8. Using a high-pass filter (HPF), high-pass filtering (HPF) is performed on each of the high channels, and such a high-pass filter (HPF) has a center frequency F HFC = 500 Hz, the method according to claim 1, configured to operate at

9. The method according to claim 8, wherein the high-pass filter is a high-shelf filter and has a linear phase response in the frequency domain.

10. The method according to claim 1, wherein the level adjustment is performed for each of the upmixed output audio signals.

11. The method according to claim 1, wherein the processing time is on the order of milliseconds, preferably shorter than 5 ms, more preferably shorter than 3 ms, and even more preferably shorter than 1 ms.

12. The method according to claim 1, wherein time synchronization is performed in the step of Mid-Side decoding, and latency compensation is performed for input channels not subject to the step of Mid-Side decoding.

13. The method performing delay adjustment (D-ADJ) on at least one (SUM) signal; routing the upmixed reconstructed audio signal (O) to an audio speaker channel (C) and feeding it to at least a Center channel (CE) and a Low Frequency Effect (LFE) channel to define a matrix of spatially distributed channels that form a center layer; performing low-pass filtering (LPF) on the LFE channel and further comprising the method according to claim 1.

14. The method according to claim 1, further configured to perform compensation filtering on the obtained upmixed output signal (O).

15. The method according to claim 14, wherein the compensation filter is a low and / or high-shelf filter and has a linear phase response in the frequency domain.