Apparatus and method for reducing spectral distortion in a system for reproducing virtual sounds through loudspeakers

The apparatus and method address spectral distortions in virtual sound reproduction by applying adaptive and time-dynamic equalization techniques, ensuring high-quality audio output with preserved spatial characteristics.

JP7820540B2Active Publication Date: 2026-02-25FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024548661
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-18
Filing Date
2023-02-08
Publication Date
2026-02-25
Estimated Expiration
2043-02-08

AI Technical Summary

Technical Problem

Existing systems for reproducing virtual sounds through loudspeakers introduce spectral distortions due to crosstalk cancellation, which affect the timbre and spatial characteristics of the audio, particularly when using headphones or earphones.

Method used

An apparatus and method that perform adaptive equalization and time-dynamic equalization to reduce spectral distortions by applying correction filters based on crosstalk cancellation filters and similarity information, adjusting equalizers for different signal components to preserve the intended virtual spatial image.

Benefits of technology

The solution effectively reduces spectral and tonal distortions while maintaining the intended spatial effect, enhancing the overall audio quality and timbre of virtual sound reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007820540000088
    Figure 0007820540000088
  • Figure 0007820540000089
    Figure 0007820540000089
  • Figure 0007820540000090
    Figure 0007820540000090
Patent Text Reader

Abstract

An apparatus (100) for reducing spectral distortion in a system (200) for reproducing virtual sounds through loudspeakers is provided, the apparatus (100) being configured to reduce said spectral distortion by performing adaptive equalization and / or by performing time dynamic equalization.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to audio signal coding, processing and decoding, and in particular to an apparatus and method for reducing spectral distortion in a system for reproducing virtual sounds. [Background technology]

[0002] When sound waves are emitted from a speaker to a listener's ear, the sound is modified multiple times, for example by the reflection of the sound waves on walls, so that the sound that reaches the pinna of the ear contains, in addition to e.g. music and speech, information about the listening environment.

[0003] Additionally, sounds arriving from multiple directions are shaped differently in the listener's head and pinnae, and using this information the listener's brain can determine the approximate direction and distance of the sound source.

[0004] However, when headphones are used, all such information is typically missing, as the audio is emitted almost directly to the listener's eardrum, creating the impression that the sound is generated inside the listener's head, which can be perceived as inconvenient and can result in, for example, spectral coloration, especially when earphones are used for longer periods of time.

[0005] It has been determined that the above-mentioned modifications of sound waves directed towards the listener's pinna and eardrum can be measured and replicated by digital filters, for example, by using head-related impulse responses, head-related transfer functions, binaural room impulse responses, and binaural room transfer functions. When such filters are applied to audio signals to be reproduced by headphones or earphones, a spatial sound is created that creates a realistic sound impression.

[0006] Virtual sound, also called virtual acoustics (see [7]) or virtual auditory space, is an audio technology in which sounds presented through headphones appear to originate from any desired spatial direction, creating the illusion of one or more virtual sound sources outside the listener's head.

[0007] Head-related transfer functions (HRTFs) are acoustic transfer functions from a sound source to the two ears. HRTFs contain the location information of the corresponding sound source. Virtual sounds from a specific direction can be generated by convolving the corresponding HRTF with the audio signal when heard through headphones.

[0008] To render spatial sound binaurally, the HRTFs of relevant locations around the listener are measured and stored. HRTFs are frequency dependent and provide the essential psychoacoustic cues for a plausible binaural effect.

[0009] For example, when using speaker boxes instead of headphones to play binaural audio signals, the signal played by one of the speaker boxes reaches both ears, thus causing crosstalk. To correctly play back the binaural signal through a pair of speakers, this signal should be pre-filtered to compensate for the crosstalk effect, which significantly impairs the spatial characteristics of the binaural signal, for example, crosstalk cancellation (CTC) applied before playback, to avoid or at least reduce crosstalk.

[0010] To achieve crosstalk cancellation, the applied filter matrices introduce spectral distortions. This can be due, for example, to extreme dynamics in the filter's phase / magnitude response. For example, the spectral dynamics of a crosstalk cancellation filter matrix can reach extreme values ​​in certain frequency bands. This affects the overall timbre, especially the intelligibility, the presence of center source timbre, and the perceived quality of a crosstalk cancellation-based playback system.

[0011] The concept of crosstalk cancellation is described in [1] and [2]. TIFF0007820540000001.tif1032 is a two speaker signal TIFF0007820540000002.tif1222 shows the transfer function when replayed by two speaker boxes.

[0012] Listener's left ear TIFF0007820540000003.tif44 and right ear The two signals in TIFF0007820540000004.tif45 can be expressed as follows: TIFF0007820540000005.tif1056 signal TIFF0007820540000006.tif45 is the first speaker TIFF0007820540000007.tif43 is fed to the first speaker box (e.g., the left speaker). TIFF0007820540000008.tif45 is the second speaker TIFF0007820540000009.tif43 is fed into a second speaker box (e.g., right speaker).

[0013] signal TIFF0007820540000010.tif44 is the first signal received at the listener's first ear (e.g., the listener's left ear). TIFF0007820540000011.tif45 is a second signal received at the listener's second ear (eg, the listener's right ear).

[0014] First speaker TIFF0007820540000012.tif43 (e.g., left speaker), crosstalk coefficient TIFF0007820540000013.tif69 is the speaker mentioned above TIFF0007820540000014.tif43 shows the direct path and crosstalk coefficients TIFF0007820540000015.tif69 is the speaker mentioned above Shows the cross-pathway of TIFF0007820540000016.tif43.

[0015] Second Speaker TIFF0007820540000017.tif43 (e.g., right speaker), crosstalk coefficient TIFF0007820540000018.tif610 is the speaker mentioned above TIFF0007820540000019.tif43 shows the direct path and crosstalk coefficients TIFF0007820540000020.tif69 is the speaker mentioned above Shows the cross-pathway of TIFF0007820540000021.tif43.

[0016] therefore, TIFF0007820540000022.tif44 describes the correction of the speaker signal to the ipsilateral ear and the crosstalk to the contralateral ear (see [1], [2]). In TIFF0007820540000023.tif44, the coefficient TIFF0007820540000024.tif58 and TIFF0007820540000025.tif58 shows the crosstalk component that should be canceled or at least reduced.

[0017] Two crosstalk-canceled speaker signals are generated before the audio signal is output by the two speakers. To obtain TIFF0007820540000026.tif1368, the filter matrix C is used to filter the audio signals for the two speakers. TIFF0007820540000027.tif45, When applied to TIFF0007820540000028.tif45, perfect reconstruction of the signal at the listener's ears, e.g., perfect crosstalk cancellation, is achieved. is obtained by inverting the HRTF matrix H according to TIFF0007820540000029.tif1060, where TIFF0007820540000030.tif44 is TIFF0007820540000031.tif449. Figure 2 shows a schema for such a two-channel crosstalk cancellation system.

[0018] A perfect crosstalk cancellation system would introduce no additional coloration into the binauralized source signals, i.e., perfect separation of the ear signals when the listener is positioned in the sweet spot. However, in real-world crosstalk cancellation systems, unwanted coloration is generally unavoidable.

[0019] One important factor affecting the spectral distortion is the CTC coefficients in C. Inversion of the matrix H is likely to be an undesirable problem. To achieve sufficient crosstalk cancellation performance, the CTC filter matrix may exhibit extreme spectral dynamics in certain frequency bands.

[0020] In prior art approaches, the dynamics of the filter matrix C is reduced by (frequency-dependent) regularization of the inverse problem (see [3]).

[0021] FIG. 3 shows an example transfer function matrix H, assuming symmetric HRTFs and a total loudspeaker opening angle of 30°.

[0022] FIG. 4 shows an example transfer function matrix C(H) with low regularization applied, where TIFF0007820540000032.tif418.

[0023] FIG. 5 shows an example transfer function matrix C(H) with increased regularization applied, where The file is TIFF0007820540000033.tif417.

[0024] Consider the example of a virtual center component: the summation of the direct and crosstalk signals on a single system loudspeaker can cause coloration and reduce their presence (see [4]). Because the input signals to both system channels are correlated, the coloration expected in this case differs from other cases, such as surround components, where the crosstalk canceling filters are orthogonal to each other.

[0025] Some approaches apply dynamic adaptation of the crosstalk cancellation signal.

[0026] Some prior art approaches appear to propose pre-processing of the input signal. U.S. Patent No. 9,532,156 (see [5]) provides an apparatus and method for acoustic stage enhancement. A spatial ratio is determined from a center component and a side component. The digital audio input signal is adjusted based on the spatial ratio to form a pre-processed signal. The center component of the crosstalk cancellation signal is re-adjusted to generate the final digital audio output.

[0027] US Patent No. 10 063 984 (see [6]) provides a method for creating a virtual acoustic stereo system with a distortion-free acoustic center. Medial / lateral separation of the CTC input signal is performed in order to apply crosstalk cancellation only to the lateral components, leaving the medial component undistorted. [Prior art documents] [Patent documents]

[0028] [Patent Document 1] U.S. Patent No. 9,532,156 [Patent Document 2] U.S. Patent No. 10,063,984 Summary of the Invention [Problem to be solved by the invention]

[0029] The object of the present invention is to provide an improved concept for reducing spectral distortions in systems for reproducing virtual sounds. [Means for solving the problem]

[0030] The object of the present invention is solved by an apparatus according to claim 1, a system according to claim 29, a method according to claim 34 and a computer program according to claim 35.

[0031] An apparatus is provided for reducing spectral distortion in a system for reproducing virtual sound through loudspeakers, the apparatus being configured to reduce spectral distortion by performing adaptive equalization and / or by performing time-dynamic equalization.

[0032] Further provided is a system for reproducing virtual sound via speakers, the system comprising a speaker signal generator for generating two or more audio output signals from one or more audio input signals, the system further comprising an apparatus for reducing spectral distortion as described above, the apparatus being configured to reduce the spectral distortion by performing adaptive equalization and / or time-dynamic equalization on at least one of the one or more audio input signals and / or on at least one of the two or more audio output signals and / or on filter information used by the speaker signal generator on the one or more audio input signals or on one or more processed signals dependent on the one or more audio input signals.

[0033] Further provided is a method for reducing spectral distortion in a system for reproducing virtual sounds through loudspeakers, the method comprising reducing spectral distortion by performing adaptive equalization and / or by performing time-dynamic equalization.

[0034] Furthermore, a computer program is provided for carrying out the above-described method when run on a computer or signal processor.

[0035] Some embodiments aimed at counteracting spectral distortions may apply signal component specific equalizers to components of, for example, the input signal or the crosstalk-canceled speaker signal in order to reduce signal coloration while preserving the acquired virtual spatial image at the designated listening position.

[0036] According to some embodiments, it is intended to reduce the spectral distortion through a virtual sound stereo system by jointly equalizing two speaker signals depending on the applied crosstalk cancellation filters and similarity information of the crosstalk cancellation signals.

[0037] To reduce expected spectral distortions while preserving the intended virtual spatial image, according to one embodiment, a correction filter is applied for each output signal frame, which can be derived in advance, for example, from a sum of crosstalk correlation filter matrices. In some embodiments, it can be assumed that different signal components, such as center, surrounding, and side components, require different correction filters. In some embodiments, a combination of correction filter sets can be determined and applied, for example, depending on the output signal. An advantage is that the applied correction equalizer can be adjusted to improve the timbre of specific components of the input signal.

[0038] Some embodiments aim to reduce tonal distortion to an acceptable level while preserving the CTC performance, e.g., "spatial effect," as well as possible. In some embodiments, a dynamic equalizer (CTC-DynEQ) is used to adjust the overall tonal distortion of the two-channel processed signal, which can operate, e.g., in the QMF domain, e.g., on a per-buffer basis, and can be, e.g., user-adjustable.

[0039] In one embodiment, the dynamic equalizer may act on the output signal, for example, by compensating to a variable degree for the expected level of coloration, which may be approximated, for example, by simulating a sum of active CTC filters in the output speaker path.

[0040] According to one embodiment, a set of compensation equalizer magnitude responses can be generated, for example, depending on the expected coloration of three basic cases: center equalizer (EQ), side EQ, and surround EQ.

[0041] In one embodiment, the compensation filter applied may be user adjustable, for example.

[0042] According to one embodiment, the compensation filter applied may be obtained, for example, from a combination of these equalizers, for example, as a function of an inter-channel similarity metric that may be derived from the processed signal. In one embodiment, for example, weighting of the equalizer components may be performed, for example, before and / or during the combination of these equalizers.

[0043] Some embodiments provide dynamic equalization for crosstalk cancellation.

[0044] According to some embodiments, the input signal is taken into consideration.

[0045] In some embodiments, timbre correction is applied during runtime depending on the input signal components, with timbre emphasis tailored specifically to the virtual center signal component, while ambient signals may, for example, be corrected differently.

[0046] According to an embodiment, application of equalization to the input signal and / or the output signal and / or the crosstalk cancellation filter matrix may for example be performed.

[0047] In one embodiment, the equalization may be determined by performing a calculation based on, for example, crosstalk cancellation coefficients.

[0048] According to one embodiment, the calculation of the equalizer components may be performed, for example, based on a combination of complex crosstalk cancellation coefficients in the frequency domain.

[0049] In one embodiment, for example, a linear combination of complex crosstalk cancellation coefficients may be used to calculate the equalizer components.

[0050] According to one embodiment, multiple equalizer components for a single correction equalizer may be combined, for example.

[0051] In one embodiment, the equalizer components may be weighted, for example, before being combined into a single correction equalizer.

[0052] According to one embodiment, the combination of equalizer components may be updated at a particular time based on, for example, one or more particular characteristics of the signal.

[0053] In one embodiment, the equalizer may be updated in response to, for example, information regarding signal similarity.

[0054] According to one embodiment, the equalizer may be updated, for example, depending on signal similarity in one or more frequency bands.

[0055] In one embodiment, the equalizer may be updated based on, for example, the average similarity of multiple frequency bands.

[0056] According to one embodiment, for example, additional weighting may be used before calculating the average.

[0057] In one embodiment, for example, magnitude-based weighting may be used.

[0058] According to one embodiment, the factors derived from the signal similarity may be weighted, for example, with a particular weighting function.

[0059] In one embodiment, for example, a sigmoid function may be used as the weighting function.

[0060] According to one embodiment, the magnitude of the frequency bands may be used, for example, to detect the frequency bands that are used to calculate similarity information.

[0061] In one embodiment, a certain number of frequency bands with the largest magnitude may be used, for example, to calculate similarity information.

[0062] Some embodiments relate to head-related transfer functions and / or crosstalk cancellation filter matrices, for example for two speakers, for example for mobile devices.

[0063] In some embodiments, the aim is to achieve reduced spectral distortion and / or reduced tonal distortion. According to some embodiments, for example, post-processing of the crosstalk-canceled signal can be performed.

[0064] The signal similarity and / or filter size of the crosstalk-canceled signal based on the addition of crosstalk cancellation coefficients can be determined, for example. Equalization of the middle, side, and surrounding signals can be provided, for example, to achieve a distortion-free center for crosstalk cancellation and / or a center enhancement for crosstalk cancellation, for example, by using dynamic equalization.

[0065] In the following, embodiments of the invention will be explained in more detail with reference to the drawings. [Brief explanation of the drawings]

[0066] [Figure 1a] FIG. 1 illustrates an apparatus for reducing spectral distortion according to one embodiment. [Figure 1b] 1a shows an embodiment in which the device of FIG. 1a interacts with a system for reproducing virtual sounds through speakers, but the device of FIG. 1a is not part of the system. [Figure 1c] 1 shows a system for reproducing virtual sounds through speakers according to one embodiment, the system including the device of FIG. 1a; [Figure 2] FIG. 1 illustrates a schema for a two-channel crosstalk cancellation system. [Figure 3] FIG. 10 is a diagram of an exemplary transfer function matrix assuming symmetric head-related transfer functions. [Figure 4] FIG. 10 illustrates an example transfer function matrix with low regularization applied. [Figure 5] FIG. 10 illustrates an example transfer function matrix with increased regularization applied. [Figure 6] FIG. 1 illustrates the generation of a central equalizer according to one embodiment. [Figure 7] FIG. 1 illustrates the generation of a side equalizer according to one embodiment. [Figure 8] FIG. 1 illustrates the generation of an ambient equalizer according to one embodiment. [Figure 9] FIG. 1 illustrates a sigmoid activation function according to one embodiment. [Figure 10] FIG. 10 illustrates an example of a resulting equalizer according to one embodiment. [Figure 11] FIG. 10 illustrates the average gain across all subbands introduced by an exemplary dynamic equalizer according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0067] FIG. 1a illustrates an apparatus 100 for reducing spectral distortion in a system 200 for reproducing virtual sounds through speakers according to one embodiment.

[0068] The apparatus 100 is configured to reduce spectral distortion by performing adaptive equalization and / or by performing time-dynamic equalization.

[0069] Figure 1b shows an embodiment in which the apparatus for reducing spectral distortion 100 of Figure 1a and a system for reproducing virtual sounds via speakers 200 interact with each other, but the apparatus 100 of Figure 1a is not part of the system 200. In other words, in the embodiment of Figure 1b, the system 200 does not include the apparatus 100.

[0070] Fig. 1c shows a system 200 for reproducing virtual sounds through speakers, according to one embodiment. In contrast to the embodiment of Fig. 1b, in the embodiment of Fig. 1c, the device 100 of Fig. 1a is part of the system 200. In other words, in the embodiment of Fig. 1c, the system 200 includes the device 100.

[0071] The following specific embodiments relate to the embodiment of FIG. 1a, the embodiment of FIG. 1b, and the embodiment of FIG. 1c.

[0072] According to an embodiment, the apparatus 100 may be configured to reduce spectral distortions, for example by performing adaptive equalization and / or time-dynamic equalization on at least one of the one or more audio input signals of the system 200 for reproducing virtual sounds and / or on at least one of the two or more audio output signals of the system 200 and / or on filter information applied by the system 200 to the one or more audio input signals or to one or more processed signals that depend on the one or more audio input signals.

[0073] In one embodiment, apparatus 100 may be configured to determine equalization information dependent on at least two of the audio input signals, and / or dependent on at least two of the audio output signals, and / or dependent on at least two of the processed signals. Apparatus 100 may be configured, for example, to perform adaptive equalization and / or by performing time-dynamic equalization by using the equalization information.

[0074] According to one embodiment, the system 200 for reproducing virtual sound comprises a crosstalk cancellation system 200 for performing crosstalk cancellation to remove and / or reduce and / or avoid crosstalk generated by the system 200 when reproducing the virtual sound through speakers. The apparatus 100 is configured, for example, to reduce spectral distortion resulting from performing crosstalk cancellation.

[0075] In one embodiment, the device 100 includes an equalizer. The device 100 may be configured to update the equalizer at specific times, for example.

[0076] According to one embodiment, apparatus 100 may be configured to determine similarity information, for example, by determining information regarding the similarity of at least two audio signals. Apparatus 100 may be configured to perform adaptive equalization and / or time-dynamic equalization, for example, using the similarity information. Furthermore, one or more audio input signals of system 200 may include at least two audio signals, two or more audio output signals of system 200 may include at least two audio signals, or one or more processed signals may include at least two audio signals.

[0077] In one embodiment, to determine the similarity information, apparatus 100 may be configured to, for example, determine information regarding the similarity of at least two audio signals in each of one or more frequency bands. Apparatus 100 may be configured to perform adaptive equalization and / or time-dynamic equalization, for example, by using the information regarding the similarity of the signals in each of one or more frequency bands.

[0078] According to one embodiment, to determine the similarity information, the apparatus 100 may be configured to, for example, determine an average of the similarities of at least two audio signals in each of a plurality of frequency bands. The apparatus 100 may be configured to perform adaptive equalization and / or time-dynamic equalization, for example, by using the average of the similarities of the signals in each of the plurality of frequency bands.

[0079] In one embodiment, to determine the similarity information, apparatus 100 may be configured to determine a magnitude-based weighted similarity by performing a magnitude-based weighting of the similarities of the at least two audio signals in each of a plurality of frequency bands. Apparatus 100 may be configured to perform adaptive equalization and / or time-dynamic equalization, for example, by using the magnitude-based weighted similarity.

[0080] According to one embodiment, the apparatus 100 may be configured to perform magnitude-based weighting, for example, by using a weighting function.

[0081] In one embodiment, the weighting function may be, for example, a sigmoid function.

[0082] According to one embodiment, apparatus 100 may be configured to determine a suitable subset of one or more frequency bands from the plurality of frequency bands, for example, by using a magnitude of each of the plurality of frequency bands of the at least two audio signals to determine the suitable subset. Apparatus 100 may be configured to determine the similarity information by determining similarity information for each of the one or more frequency bands of the suitable subset, for example, without determining similarity information for each of the one or more frequency bands of the plurality of frequency bands that are not included in the suitable subset.

[0083] In one embodiment, each frequency of the plurality of frequency bands may be associated with a magnitude that depends on, for example, the magnitude of one or more frequency bands of the at least two audio signals. Apparatus 100 may be configured to determine a proper subset of the one or more frequency bands such that, for example, a magnitude associated with each frequency band of the one or more frequency bands of the proper subset may be greater than or equal to, for example, a magnitude associated with each of the one or more frequency bands of the plurality of frequency bands not included in the proper subset.

[0084] According to an embodiment, the system 200 for reproducing virtual sound may be configured to perform crosstalk cancellation, for example, by using a plurality of crosstalk cancellation coefficients. The apparatus 100 may be configured to reduce spectral distortion, for example, by performing adaptive equalization using a plurality of equalizer components and / or by performing time-dynamic equalization. The apparatus 100 may be configured, for example, to determine the plurality of equalizer components depending on one or more of the plurality of crosstalk cancellation coefficients.

[0085] In one embodiment, the apparatus 100 may be configured to determine the plurality of equalizer components, for example, by selecting a set of pre-calculated equalizer components from two or more sets of pre-calculated equalizer components depending on the crosstalk cancellation coefficients.

[0086] In one embodiment, the apparatus 100 may be configured to determine a plurality of equalizer components at runtime, for example, in response to similarity information indicative of information regarding the similarity of at least two audio signals and / or in response to a plurality of crosstalk cancellation coefficients.

[0087] According to one embodiment, the apparatus 100 may be configured to determine the plurality of equalizer components by determining one or more combinations of the plurality of crosstalk cancellation coefficients, for example, the plurality of complex crosstalk cancellation coefficients in the frequency domain.

[0088] In one embodiment, to determine one or more combinations of multiple crosstalk cancellation coefficients to determine the multiple equalizer components, the apparatus 100 may be configured to determine, for example, one or more linear combinations of complex crosstalk cancellation coefficients in the frequency domain.

[0089] According to one embodiment, the apparatus 100 may be configured to determine a single correction equalizer from multiple equalizer components, for example.

[0090] In one embodiment, the apparatus 100 may be configured to determine a single correction equalizer from multiple equalizer components, for example, by weighting the multiple equalizer components before combining them to obtain a single correction equalizer.

[0091] According to one embodiment, the apparatus 100 may be configured to weight a plurality of equalizer components depending on, for example, a similarity value, the similarity value being dependent on the similarity information.

[0092] In one embodiment, the device 100 may be configured to perform adaptive equalization and / or time-dynamic equalization on one or more audio input signals of the system 200, for example, to reproduce virtual sounds.

[0093] According to one embodiment, the device 100 may be configured to perform adaptive equalization and / or time-dynamic equalization on two or more audio output signals of the system 200, for example, to reproduce virtual sounds.

[0094] In one embodiment, the apparatus 100 may be configured, for example, to perform adaptive equalization and / or time-dynamic equalization on crosstalk cancellation filter matrices used for crosstalk cancellation by the system for reproducing virtual sound (200).

[0095] FIG. 1c illustrates a system 200 for reproducing virtual sounds through speakers, according to one embodiment.

[0096] The system 200 of FIG. 1c includes a speaker signal generator 150 that generates two or more audio output signals from one or more audio input signals.

[0097] Furthermore, the system 200 of FIG. 1c includes the apparatus 100 of FIG. 1a for reducing spectral distortion.

[0098] In system 200 of FIG. 1c, apparatus 100 is configured to reduce spectral distortion by performing adaptive equalization and / or time-dynamic equalization on at least one of the one or more audio input signals and / or on at least one of the two or more audio output signals and / or on filter information used by speaker signal generator 150 on the one or more audio input signals or on one or more processed signals that are dependent on the one or more audio input signals.

[0099] According to one embodiment, the system 200 for reproducing virtual sound may comprise a crosstalk cancellation system (not shown) for performing crosstalk cancellation, e.g., to remove and / or reduce and / or avoid crosstalk generated by the system for performing crosstalk cancellation when reproducing the virtual sound via speakers. The apparatus 100 is configured, e.g., to reduce spectral distortion resulting from performing crosstalk cancellation.

[0100] In one embodiment, the one or more audio input signals may include, for example, two binaural audio signals.

[0101] According to one embodiment, the system 200 of Figure 1c includes, for example, a speaker. In another embodiment, the system of Figure 1c does not include, for example, a speaker.

[0102] In one embodiment, apparatus 100 may be configured to perform adaptive equalization and / or time-dynamic equalization, e.g., by applying an average gain across two or more subbands, e.g., to achieve loudness preservation. In particular embodiments, apparatus 100 may be configured to apply an average gain across all subbands, e.g.

[0103] Specific embodiments of the present invention are described below. In an embodiment, a correction equalizer can be applied to, for example, the crosstalk-canceled speaker signals, e.g., in the QMF domain, to combat expected coloration. According to one embodiment, the correction filter can be determined / estimated, for example, for the expected magnitude response for, e.g., three fundamental signal components, by a summation of complex CTC filter matrices, e.g., three dependent ones.

[0104] For example, for the middle component, the middle / center equalizer ( TIFF0007820540000034.tif513 / TIFF0007820540000035.tif517 r ) may be determined, for example, and / or, for example, in the case of the side components, a side equalizer ( TIFF0007820540000036.tif513) may be determined, for example, and / or, for example, in the case of the ambient component, an ambient equalizer ( TIFF0007820540000037.tif514) may be determined, for example. For example, the terms median equalizer and center equalizer may be used interchangeably, for example.

[0105] According to one embodiment, a combination of three component equalizers can determine, for example, the equalizer to be applied, which may depend, for example, on signal similarity information about the output signal and / or may depend, for example, on manual adjustment of component weights.

[0106] The following describes determining the component equalizers according to some embodiments. According to some embodiments, two or more, e.g., three, component equalizers (also referred to as equalizer components) can be determined, for example. The determination of the three component equalizers may be performed, for example, before runtime. For example, the component equalizers may be determined, for example, as described below.

[0107] Some embodiments may include TIFF0007820540000038.tif48( Phase difference between channels for each TIFF0007820540000039.tif43 ( TIFF0007820540000040.tif48) the expected coloration of a two-channel input signal s with, for example, speaker index QMF bandwidth in TIFF0007820540000041.tif45 Resulting amplitude spectrum per TIFF0007820540000042.tif43 This is based on the knowledge that the direct path for each speaker can be estimated based on TIFF0007820540000043.tif613. TIFF0007820540000044.tif510 and cross-path Depends on the sum of the CTC filters in TIFF0007820540000045.tif510. TIFF0007820540000046.tif676

[0108] Center Equalizer The magnitude response of TIFF0007820540000047.tif517 is, for example, the IPD across all bands (e.g., phantom center image) averaged across both loudspeakers. center (b) Two-channel input signal s with s = 0° center This allows compensation for the expected coloration of the

[0109] Side Equalizer The magnitude response of TIFF0007820540000048.tif513 is averaged across both speakers, for example: IPD side (b) Two-channel input signal s with = 180° side This allows compensation for the expected coloration of the

[0110] Ambient Equalizer In the case of TIFF0007820540000049.tif514, the left and right input signals may be uncorrelated, for example. The average expected coloration per speaker may assume, for example, unit power in the input spectrum. TIFF0007820540000050.tif1579TIFF0007820540000051.tif1575TIFF0007820540000052.tif2287In the above formula, i can represent, for example, the speaker index: i=sp.

[0111] In the following, taking into account the similarity of speaker signals according to some embodiments will be described.

[0112] In some embodiments, the speaker signal similarity may be obtained / determined, for example, for each input buffer. For example, the speaker signal similarity may be obtained, for example, during runtime.

[0113] To modulate the frequency response of the resulting compensation equalizer, according to one embodiment, the similarity vector TIFF0007820540000053.tif512, for example, can be derived for TIFF0007820540000054.tif32. TIFF0007820540000055.tif615 and TIFF0007820540000056.tif615 can, for example, show a two-channel complex-valued signal for each buffer and frequency band.

[0114] TIFF0007820540000057.tif511 is, for example, Similarity metric for bands 0 and 1 weighted by TIFF0007820540000058.tif517 The combination TIFF0007820540000059.tif513 can be shown. The sigmoid function in TIFF0007820540000060.tif512 is center ,s side ) to the advantage of The intention is to tilt the values ​​of TIFF0007820540000061.tif511.

[0115] In one embodiment, the first two frequency bands ( TIFF0007820540000062.tif43=0,1) can be considered, for example. However, in another preferred embodiment, the two QMF bands with the largest magnitudes can be determined first, for example, and selected for signal similarity estimation and weighting instead, for example. This improves stability.

[0116] Similarity Vector To stabilize the TIFF0007820540000063.tif511, for example, a weighting factor is used to introduce relative weights between the inter-channel similarity values. TIFF0007820540000064.tif517 can be used. This may depend, for example, on the distribution of input levels between the two QMF bands across the input channels. A low signal amplitude in one frequency band may, for example, have a disproportionate effect on the resulting similarity vector. For example, if useful signals are present only in QMF band 1 and band 2 consists only of a low-amplitude noise floor, the resulting similarity values ​​may exhibit unpredictable behavior between adjacent input buffers.

[0117] TIFF0007820540000065.tif412 has a slightly reduced range of possible values ​​and can be adjusted between 0 and 1. TIFF0007820540000066.tif1247TIFF0007820540000067.tif869TIFF0007820540000068.tif32170Or, for example, TIFF0007820540000069.tif13114

[0118] The following describes how the resulting equalization is obtained by combining similarity vectors and / or manual adjustment coefficients according to some embodiments.

[0119] According to certain embodiments, the magnitude of the applied equalizer, e.g., a dynamic equalizer, TIFF0007820540000070.tif519 or TIFF0007820540000071.tif523 is, for example, the ambient equalizer TIFF0007820540000072.tif519 Side Equalizer TIFF0007820540000073.tif519 or Center Equalizer It can be combined with either TIFF0007820540000074.tif523. Coefficient TIFF0007820540000075.tif413, TIFF0007820540000076.tif49, and TIFF0007820540000077.tif49 may be calculated, for example, once per input buffer. They may be used, for example, as similarity vectors TIFF0007820540000078.tif56 and / or adjustment parameters TIFF0007820540000079.tif524 and / or TIFF0007820540000080.tif521 and / or TIFF0007820540000081.tif521, which may be, for example, a user-adjustable tuning parameter, for example ranging from 0 to 1. The relative weighting of the equalizer components may be adjustable, for example, to balance the spatial and temporal impression of the system. TIFF0007820540000082.tif44120TIFF0007820540000083.tif6114TIFF0007820540000084.tif6122

[0120] for example, TIFF0007820540000085.tif513, for example TIFF0007820540000086.tif533, which would result in the following: TIFF0007820540000087.tif6124

[0121] According to an embodiment, the resulting equalizer may be applied to the output speaker signal, for example.

[0122] 6 through 11 provide visual examples of the embodiments provided. 6 is a diagram illustrating the generation of a center equalizer (EQ) according to one embodiment, where the x-axis labels indicate the center frequencies of the QMF bands. FIG. 7 is a diagram illustrating the generation of a side equalizer according to one embodiment. FIG. 8 is a diagram illustrating the generation of an ambient equalizer according to one embodiment.

[0123] FIG. 9 illustrates a sigmoid activation function with a signal similarity value of −0.5 and a weight of 0.4 according to one embodiment. 10 shows an example of the resulting equalizer, according to one embodiment. The ambient EQ is labeled "EQ 90." FIG. 11 is a diagram illustrating the average gain across all subbands introduced by an exemplary dynamic equalizer (DynEQ) according to one embodiment.

[0124] The inventive concepts may, for example, be used in another domain, for example another frequency domain, for example the FFT domain instead of the QMF domain. Some embodiments may, for example, be implemented in the Fast Fourier Transform (FFT) domain.

[0125] In one embodiment, the selection of the QMF bands may be used, for example, for signal similarity estimation. In one implementation (e.g., headphone library headphonelib), speaker signal similarity may be based, for example, on the two bands with the largest magnitude. This allows for stability when, for example, the signal energy in bands 0 and 1 is low.

[0126] According to one embodiment, device 100 may be configured to reduce spectral distortions, for example, by performing adaptive equalization in a loudness-preserving manner and / or by performing time-dynamic equalization in a loudness-preserving manner and / or by adjusting one or more audio input signals to ensure loudness preservation. For example, loudness preservation may be ensured by, for example, an applied equalizer. This can be countered by applying a component or applied equalizer and / or a make-up gain factor to the signal, so that the average or root-mean-square (RMS) volume of the output signal is not affected.

[0127] In one embodiment, for example, a different configuration for the component equalizer magnitude response may be used. The above-described approach for estimating the magnitude of the component correction filter may be modified, for example. The variation may be related to the sum of complex CTC filters, for example, by applying variable weighting to specific frequency regions. According to one embodiment, weighting between direct and crosstalk components may be introduced, for example, to specifically address coloration by one component. In another embodiment, the equalizer components may be calculated, for example, at run time.

[0128] According to one embodiment, for example, a (e.g., frequency-selective) compression or expansion of the spectral dynamics for a particular frequency region of the component or applied filter can be used. According to another embodiment, the equalizer components may be calculated, for example, at run time.

[0129] In one embodiment, for example, different combinations of component equalizers can be applied, for example, a center equalizer (EQ center ) and side equalizer (EQ side ) is used, for example, to adjust the size of the ambient equalizer (EQ amb ) may be summed to create a similarity (corrLR) between the left and right signals. For example, in the intermediate case where the similarity (corrLR) between the left and right signals is +-0.5, it is not guaranteed that the applied equalizer will match well with the model of the assumed coloration. Variations and / or combinations of the component equalizers can be implemented in any suitable manner. For example, a complex addition of correction EQs can be performed.

[0130] According to one embodiment, for example, a constrained optimization approach can be used: the filter to be applied can be generated for each frequency band with respect to its signal similarity, for example, taking into account the expected crosstalk cancellation in the sweet spot within this band. Further embodiments are provided below. Embodiment 1: An apparatus (100) for reducing spectral distortion in a system (200) for reproducing virtual sounds through speakers, comprising: An apparatus (100), wherein the apparatus (100) is configured to reduce spectral distortion by performing adaptive equalization and / or by performing time-dynamic equalization. Embodiment 2: The apparatus (100) of embodiment 1, wherein the apparatus (100) is configured to reduce spectral distortion by performing adaptive equalization and / or time-dynamic equalization on at least one of one or more audio input signals of the system (200) for reproducing virtual sound and / or on at least one of two or more audio output signals of the system (200) and / or on filter information applied by the system (200) to one or more audio input signals or one or more processed signals dependent on the one or more audio input signals. Embodiment 3: The apparatus (100) is configured to determine equalization information depending on at least two of the audio input signals, and / or depending on at least two of the audio output signals, and / or depending on at least two of the processed signals; 3. The apparatus (100) of embodiment 2, wherein the apparatus (100) is configured to perform adaptive equalization and / or to perform time-dynamic equalization by using the equalization information. Embodiment 4: A system (200) for reproducing virtual sounds comprises a crosstalk cancellation system (200) for performing crosstalk cancellation to remove and / or reduce and / or avoid crosstalk generated by the system (200) when reproducing the virtual sounds through speakers, 4. The apparatus (100) of embodiment 2 or 3, wherein the apparatus (100) is configured to reduce spectral distortion resulting from performing crosstalk cancellation. Embodiment 5: The device (100) comprises an equalizer; 5. The apparatus (100) of any one of embodiments 2 to 4, wherein the apparatus (100) is configured to update the equalizer at specific times. Embodiment 6: An apparatus (100) configured to determine similarity information by determining information regarding the similarity of at least two audio signals, the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization using the similarity information; An apparatus (100) according to any one of embodiments 2 to 5, wherein one or more audio input signals of the system (200) comprise at least two audio signals, or two or more audio output signals of the system (200) comprise at least two audio signals, or one or more processed signals comprise at least two audio signals. Embodiment 7: For determining similarity information, an apparatus (100) is configured to determine information relating to the similarity of at least two audio signals in each of one or more frequency bands, An apparatus (100) as described in embodiment 6, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization by using information regarding the similarity of signals in each of one or more frequency bands. Embodiment 8: To determine the similarity information, the apparatus (100) is configured to determine an average of similarities of at least two audio signals in each of a plurality of frequency bands; The apparatus (100) of embodiment 6 or 7, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization by using an average of the signal similarity in each of a plurality of frequency bands. Embodiment 9: To determine the similarity information, the apparatus (100) is configured to determine a magnitude-based weighted similarity by performing a magnitude-based weighting of the similarity of at least two audio signals in each of a plurality of frequency bands; 9. The apparatus (100) of any one of embodiments 6 to 8, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization by using magnitude-based weighted similarity. Embodiment 10: The apparatus (100) of embodiment 9, wherein the apparatus (100) is configured to perform magnitude-based weighting by using a weighting function. Embodiment 11: The apparatus (100) of embodiment 10, wherein the weighting function is a sigmoid function. Embodiment 12: An apparatus (100) described in any one of embodiments 9 to 11, wherein the apparatus (100) is configured to perform magnitude-based weighting by using the magnitudes of multiple frequency bands to detect which of the multiple frequency bands to use to calculate similarity information. Embodiment 13: The device (100) of any one of embodiments 9 to 12, wherein the device (100) is configured to perform magnitude-based weighting by using a specific number of frequency bands with the largest magnitudes to calculate the similarity information. Embodiment 14: An apparatus (100) configured to determine a proper subset of one or more frequency bands from a plurality of frequency bands by using a magnitude of each of the plurality of frequency bands of at least two audio signals to determine the proper subset; An apparatus (100) as described in any one of embodiments 6 to 13, wherein the apparatus (100) is configured to determine similarity information by determining similarity information for each of one or more frequency bands of a plurality of frequency bands that are not included in the appropriate subset, without determining similarity information for each of one or more frequency bands of a plurality of frequency bands that are not included in the appropriate subset. Embodiment 15: Each frequency of the plurality of frequency bands is associated with a magnitude that depends on the magnitude of one or more frequency bands of the at least two audio signals; An apparatus (100) as described in embodiment 14, wherein the apparatus (100) is configured to determine a suitable subset of one or more frequency bands such that a magnitude associated with each frequency band of the suitable subset is greater than or equal to a magnitude associated with each of one or more frequency bands of a plurality of frequency bands not included in the suitable subset. Embodiment 16: A system (200) for reproducing virtual sounds is configured to perform crosstalk cancellation by using a plurality of crosstalk cancellation coefficients, The apparatus (100) is configured to reduce spectral distortion by performing adaptive equalization using a plurality of equalizer components and / or by performing time-dynamic equalization; 16. The apparatus (100) of any one of embodiments 1 to 15, wherein the apparatus (100) is configured to determine a plurality of equalizer components according to one or more of a plurality of crosstalk cancellation coefficients. Embodiment 17: The apparatus (100) described in embodiment 16, wherein the apparatus (100) is configured to determine a plurality of equalizer components by selecting a set of pre-calculated equalizer components from a set of two or more pre-calculated equalizer components according to a crosstalk cancellation coefficient. Embodiment 18: An apparatus (100) according to embodiment 16, further dependent on any one of embodiments 6 to 15, wherein the apparatus (100) is configured to determine multiple equalizer components at runtime in response to similarity information indicating information regarding the similarity of at least two audio signals and / or in response to multiple crosstalk cancellation coefficients. Embodiment 19: The apparatus (100) of embodiment 16 or 18, wherein the apparatus (100) is configured to determine the plurality of equalizer components by determining one or more combinations of the plurality of crosstalk cancellation coefficients, the combinations being a plurality of complex crosstalk cancellation coefficients in the frequency domain. Embodiment 20: The apparatus (100) of embodiment 19, wherein, in order to determine one or more combinations of multiple crosstalk cancellation coefficients to determine multiple equalizer components, the apparatus (100) is configured to determine one or more linear combinations of complex crosstalk cancellation coefficients in the frequency domain. Embodiment 21: The apparatus (100) of embodiment 20, wherein the apparatus (100) is configured to determine a single correction equalizer from a plurality of equalizer components. Embodiment 22: The apparatus (100) of embodiment 21, wherein the apparatus (100) is configured to determine a single correction equalizer from multiple equalizer components by weighting the multiple equalizer components before combining the multiple equalizer components to obtain a single correction equalizer. Embodiment 23: The device (100) of embodiment 20, further dependent on any one of embodiments 6 to 15, wherein the device (100) is configured to weight multiple equalizer components according to a similarity value, the similarity value depending on the similarity information. Embodiment 24: An apparatus (100) described in any one of embodiments 1 to 23, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization on one or more audio input signals of the system (200) to reproduce virtual sound. Embodiment 25: An apparatus (100) described in any one of embodiments 3 to 24, further dependent on embodiment 2, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization on two or more audio output signals of the system (200) to reproduce virtual sound. Embodiment 26: An apparatus (100) described in any one of embodiments 1 to 25, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization on a crosstalk cancellation filter matrix used for crosstalk cancellation by the system (200) for reproducing virtual sound. Embodiment 27: An apparatus (100) described in any one of embodiments 3 to 26, further dependent on embodiment 2, wherein the apparatus (100) is configured to reduce spectral distortion by performing adaptive equalization in a manner that maintains volume, and / or by performing time-dynamic equalization in a manner that maintains volume, and / or by adjusting one or more audio input signals to ensure that volume is maintained. Embodiment 28: The apparatus (100) of any one of embodiments 1 to 27, wherein the apparatus (100) is configured to perform adaptive equalization and / or time-dynamic equalization by applying an average gain across two or more subbands. Embodiment 29: A system (200) for reproducing virtual sounds through a speaker, the system (200) comprising: a speaker signal generator (150) for generating two or more audio output signals from one or more audio input signals; 29. A system (200) comprising the apparatus (100) of any one of embodiments 1-28 for reducing spectral distortion; A system (200) in which the apparatus (100) is configured to reduce spectral distortion by performing adaptive equalization and / or time-dynamic equalization on at least one of one or more audio input signals and / or on at least one of two or more audio output signals and / or on filter information used by a speaker signal generator on one or more audio input signals or one or more processed signals dependent on the one or more audio input signals. Embodiment 30: A system (200) for reproducing virtual sounds comprises a crosstalk cancellation system for performing crosstalk cancellation to remove and / or reduce and / or avoid crosstalk generated by the crosstalk cancellation system when reproducing the virtual sounds through speakers, 30. The system (200) of embodiment 29, wherein the device (100) is configured to reduce spectral distortion resulting from crosstalk cancellation. Embodiment 31: The system (200) described in embodiment 30, wherein the one or more audio input signals include two binaural audio signals. Embodiment 32: A system (200) described in any one of embodiments 29 to 31, wherein the system (200) does not include a speaker. Embodiment 33: A system (200) described in any one of embodiments 29 to 31, wherein the system (200) comprises a speaker. Embodiment 34: A method for reducing spectral distortion in a system (200) for reproducing virtual sounds through speakers, comprising: A method, wherein the method includes reducing spectral distortion by performing adaptive equalization and / or by performing time dynamic equalization. Embodiment 35: A computer program for performing the method according to embodiment 34 when the computer program is run on a computer or signal processor.

[0131] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0132] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partly in hardware, or at least partly in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.

[0133] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0134] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and the program code can be stored on, for example, a machine-readable carrier.

[0135] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0136] In other words, therefore, an embodiment of the inventive methods is a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0137] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.

[0138] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can for example be arranged to be transmitted via a data communication connection, for example via the Internet.

[0139] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0140] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0141] Further embodiments according to the invention comprise an apparatus or system configured to transfer (e.g. electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may for example be a computer, a mobile device, a memory device, etc. The apparatus or system may for example comprise a file server for transferring the computer program to the receiver.

[0142] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0143] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0144] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0145] The above-described embodiments are merely illustrative of the principles of the present disclosure. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0146] literature [1] Masiero, B., Fels, J., & Vorlaender, M. (2011).Review of the crosstalk cancellation filter technique.Proc. of ICSA, 112. [2] Kaiser, F.(2011).Transaural Audio-The reproduction of binaural signals over loudspeakers (Doctoral dissertation, Diploma Thesis, Universitaet fuer Musik und darstellende Kunst Graz / Institut fuer Elekronische Musik und Akustik / IRCAM, March 2011). [3] Choueiri, E.Y.(2008).Optimal crosstalk cancellation for binaural audio with two loudspeakers.Princeton University, 28. [4] Canfield, G.H., & Kuo, S.M.(1997, September).Dual-Channel Audio Equalization and Cross-Talk Cancellation for Correlated Stereo Signals.In Audio Engineering Society Convention 103.Audio Engineering Society. [5] US 9 532 156 B2, Apparatus and Method for Sound Stage Enhancement. [6] US 10 063 984 B2, Method for creating a virtual acoustic stereo system (200) with an undistorted acoustic center. [7] https: / / en.wikipedia.org / wiki / Acoustic_space .

Claims

1. An apparatus (100) for reducing spectral distortion in a system (200) for reproducing virtual sounds through speakers, comprising: the apparatus (100) is configured to reduce the spectral distortion by performing adaptive equalization and / or by performing time-dynamic equalization, the device (100) is configured to apply a filter matrix for performing crosstalk cancellation to obtain a crosstalk-canceled speaker signal; the device (100) is configured to reduce the spectral distortion by performing the adaptive equalization and / or the time-dynamic equalization on at least one of one or more audio input signals of the system (200) for reproducing virtual sounds and / or on at least one of two or more audio output signals of the system (200) and / or on filter information applied by the system (200) to the one or more audio input signals or to one or more processed signals dependent on the one or more audio input signals, the apparatus (100) is configured to determine similarity information by determining information regarding the similarity of at least two audio signals; the apparatus (100) is configured to use the similarity information to perform the adaptive equalization and / or the time-dynamic equalization; The apparatus (100) wherein the one or more audio input signals of the system (200) include the at least two audio signals, the two or more audio output signals of the system (200) include the at least two audio signals, or the one or more processed signals include the at least two audio signals.

2. To determine the similarity information, the device (100) is configured to determine information relating to the similarity of at least two audio signals in each of one or more frequency bands; the device (100) is configured to perform the adaptive equalization and / or the time-dynamic equalization by using the information regarding the similarity of the signals in each of the one or more frequency bands. The apparatus (100) of claim 1.

3. To determine the similarity information, the apparatus (100) is configured to determine an average similarity of at least two audio signals in each of a plurality of frequency bands; the device (100) is configured to perform the adaptive equalization and / or the time-dynamic equalization by using an average of the similarities of the signals in each of the plurality of frequency bands. The apparatus (100) of claim 1.

4. and to determine the similarity information, the apparatus (100) is configured to determine a magnitude-based weighted similarity by performing a magnitude-based weighting of similarities of at least two audio signals in each of a plurality of frequency bands; the apparatus (100) is configured to perform the adaptive equalization and / or the time-dynamic equalization by using the magnitude-based weighted similarity. The apparatus (100) of claim 1.

5. the apparatus (100) is configured to perform the magnitude-based weighting by using a weighting function; and / or the weighting function is a sigmoid function, and / or the apparatus (100) is configured to perform the magnitude-based weighting by using the magnitudes of the plurality of frequency bands to determine which of the plurality of frequency bands to use to calculate similarity information; and / or the apparatus (100) is configured to perform the magnitude-based weighting by using a certain number of frequency bands with the largest magnitudes for calculating the similarity information. The apparatus (100) of claim 4.

6. the apparatus (100) is configured to determine the appropriate subset of one or more frequency bands from the plurality of frequency bands by using a magnitude of each of the plurality of frequency bands of the at least two audio signals to determine the appropriate subset; the apparatus (100) is configured to determine the similarity information by determining similarity information for each of the one or more frequency bands of the suitable subset without determining similarity information for each of the one or more frequency bands of the plurality of frequency bands not included in the suitable subset. The apparatus (100) of claim 1.

7. each frequency of the plurality of frequency bands is associated with a magnitude that depends on the magnitude of the frequency band(s) of one or more of the at least two audio signals; the apparatus (100) is configured to determine the suitable subset of the one or more frequency bands such that the magnitude associated with each frequency band of the one or more frequency bands of the suitable subset is greater than or equal to the magnitude associated with each of the one or more frequency bands of the plurality of frequency bands not included in the suitable subset.

7. The apparatus (100) of claim 6.

8. the system (200) for reproducing virtual sound is configured to perform crosstalk cancellation by using a plurality of crosstalk cancellation coefficients; the apparatus (100) is configured to reduce the spectral distortion by performing adaptive equalization using a plurality of equalizer components and / or by performing time-dynamic equalization; The apparatus (100) of claim 1, wherein the apparatus (100) is configured to determine the plurality of equalizer components in response to one or more of the plurality of crosstalk cancellation coefficients.

9. the apparatus (100) is configured to determine the plurality of equalizer components by selecting a set of pre-calculated equalizer components from two or more sets of pre-calculated equalizer components in response to the crosstalk cancellation coefficients; and / or the apparatus (100) is configured to determine the plurality of equalizer components by determining one or more combinations of the plurality of crosstalk cancellation coefficients, the combinations being a plurality of complex crosstalk cancellation coefficients in the frequency domain.

9. The apparatus (100) of claim 8.

10. A system (200) for reproducing virtual sounds through speakers, said system (200) comprising: a speaker signal generator (150) for generating two or more audio output signals from one or more audio input signals; The system (200) comprises the apparatus (100) of claim 1 for reducing spectral distortion; the device (100) is configured to reduce the spectral distortion by performing adaptive equalization and / or time-dynamic equalization on at least one of the one or more audio input signals and / or on at least one of the two or more audio output signals and / or on filter information used by the loudspeaker signal generator on the one or more audio input signals or on one or more processed signals dependent on the one or more audio input signals, System (200).

11. 1. A method for reducing spectral distortion in a system (200) for reproducing virtual sounds through speakers, comprising: the method comprising reducing the spectral distortion by performing adaptive equalization and / or by performing time-dynamic equalization; the method includes applying a filter matrix for performing crosstalk cancellation to obtain a crosstalk-canceled speaker signal; the method comprising reducing the spectral distortion by performing the adaptive equalization and / or the time-dynamic equalization on at least one of one or more audio input signals of the system (200) for reproducing virtual sounds and / or on at least one of two or more audio output signals of the system (200) and / or on filter information applied by the system (200) to the one or more audio input signals or to one or more processed signals dependent on the one or more audio input signals, the method comprising determining similarity information by determining information regarding the similarity of at least two audio signals; the method comprising using the similarity information to perform the adaptive equalization and / or the time-dynamic equalization; the one or more audio input signals of the system (200) include the at least two audio signals, the two or more audio output signals of the system (200) include the at least two audio signals, or the one or more processed signals include the at least two audio signals; method.

12. 12. A computer program for performing the method according to claim 11 when the computer program is run on a computer or signal processor.

Citation Information

Patent Citations

  • Audio system for providing separate listening zones

    JP2018528681A

  • Method for creating a virtual acoustic stereo system with an undistorted acoustic center

    US10063984B2

  • Methods, apparatus and systems for dynamic equalization for cross-talk cancellation

    US20190373398A1

  • Apparatus and method for sound stage enhancement

    US9532156B2

  • Apparatus and method for providing individual sound zones

    WO2017178454A1