Audio reproduction system and method
The audio control unit determines the delay based on the distance weighted audio signal and the use of auxiliary microphone or pilot signals, and solves the problem that the virtual position of the karaoke participant does not correspond to the actual position in the noisy environment, achieving accurate reproduction of sound and good experience.
Patent Information
- Application Number
- CN202480006719.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-04
- Filing Date
- 2024-01-02
- Publication Date
- 2025-08-08
AI Technical Summary
In noisy environments, such as bars or clubs, the prior art is difficult to correspond to the virtual location of the karaoke participant with the actual location, resulting in inaccurate sound reproduction.
The audio control unit weights the audio signal according to the distance between the audio input device and the reproduction device, and determines the delay using the auxiliary microphone or pilot signal, and adjusts the audio output to achieve the correspondence between the virtual position and the actual position.
Accurately reproduce participants' voices in noisy environments, so that their perceived location corresponds to the actual location, reduce the impact of environmental noise, and provide a good user experience.
Smart Images

Figure CN120457480A_ABST
Abstract
Description
Background Art
[0001] The present invention relates to audio reproduction systems.
[0002] The invention also relates to an audio reproduction method.
[0003] Karaoke is a popular form of interactive entertainment in which participants use microphones to sing along to recorded music, with the original voice partially or completely removed. The recorded music can originate from a local data store or can be streamed from an external source. Similarly, participants can be instrumentalists, such as guitarists, who play along with the recorded music, with the original instrumental sound partially or completely removed. The participant's vocal or instrumental sound is then combined with the modified recorded music and reproduced by an audio reproduction system.
[0004] For a realistic experience, it is desirable that the virtual positions of the participants perceived by the listener based on the output of the audio reproduction system correspond to the actual positions of the participants. Summary of the Invention
[0005] According to a first aspect of the present invention, there is provided an improved audio reproduction system which effectively adjusts the reproduction of the voices of participants to achieve a correspondence between the virtual positions and the actual positions of the participants.
[0006] According to a second aspect of the present invention, there is provided an audio reproduction method that effectively adjusts the reproduction of the voices of participants to achieve correspondence between the virtual positions of the participants and the actual positions of the participants.
[0007] The improved audio reproduction system according to the first aspect comprises: an audio input device, a pre-recorded audio source, an audio control unit and a plurality of audio reproduction devices. The audio input device is configured to convert an acoustic signal into an audio input signal. Usually, the audio input device comprises a microphone that converts the voice of a singer, but alternatively or additionally, the audio input device can comprise an element that converts the acoustic output of a musical instrument into an audio input signal. The pre-recorded audio source provides the other audio input of the pre-recorded music that the singer or instrument player wishes to play accompaniment. The audio control unit is configured to combine the audio input signal from the audio input device with one or more other audio input signals, and provide corresponding audio output signals to each audio reproduction device. One or more other audio input signals can be streamed from an external source and / or can be derived from a local source, for example, from a local data storage with pre-recorded audio and / or from an electronic input such as a keyboard.
[0008] For a good user experience, it is desirable that the virtual positions of the participants perceived by the audience based on the output of the audio reproduction system correspond to their actual positions. This can be achieved in a recording studio by placing a pair of microphones at a sufficient distance from the singer. However, this is impossible in a noisy environment, which is typical of bars or clubs where karaoke is performed. In this case, the voice of the singer or instrumentalist can only be properly recorded at a receiving position close to the sound source. Therefore, it is impossible to properly reproduce the voices of the participants using the audio reproduction device so that the virtual positions of the participants correspond to their actual locations.
[0009] In the improved audio reproduction system, the audio control unit is configured to provide respective audio output signals with respective amplitudes weighted according to respective distances of the audio input devices relative to the respective audio reproduction devices.
[0010] In one embodiment, the improved audio reproduction system SYS is configured to determine the respective distances using a first correlation unit and a second correlation unit. The correlation units correlate the respective signal components with reference signal components to determine the respective acoustic signal delays. By measuring the acoustic delays in each path between the audio input device and the respective one of the audio reproduction devices, the position of the participant can be readily determined, thereby optimizing the use of the components of the audio reproduction system that are anyway required to reproduce the audio input signal.
[0011] In one example of this embodiment of the improved audio reproduction system, a first correlation unit correlates an audio input signal obtained from the microphone with an auxiliary signal from a first auxiliary microphone disposed near a first audio reproduction device in the audio reproduction apparatus to determine a delay of an acoustic signal originating from the vicinity of the microphone as perceived by the first auxiliary microphone. A second correlation unit correlates the audio input signal obtained from the microphone with an auxiliary signal from a second auxiliary microphone disposed near a second audio reproduction device in the audio reproduction apparatus to determine a delay of an acoustic signal originating from the vicinity of the microphone as perceived by the second auxiliary microphone.
[0012] By determining the delay of the auxiliary signal from the auxiliary microphone, the participant's position can be reliably determined, and the audio input signal from the participant can be appropriately weighted based on the participant's actual position, so that the participant's virtual position corresponds to the participant's actual position. The auxiliary audio input signal itself is not used to reproduce the participant's voice, but only to ensure that the participant's voice is perceived as originating from a virtual position corresponding to the participant's actual position. Therefore, the auxiliary microphone does not introduce ambient noise into the participant's voice reproduced by the audio reproduction system.
[0013] In another example of an embodiment of the improved audio reproduction system, the audio control unit is configured to include a corresponding pilot signal in each of the audio output signals for the audio reproduction devices. In this example, a first correlation unit determines a delay for a pilot signal reproduced by a first one of the audio reproduction devices to be received by the audio input device, and a second correlation unit determines a delay for a pilot signal reproduced by a second one of the audio reproduction devices to be received by the audio input device.
[0014] As in the previous example, the audio reproduction system thereby determines the distance of the audio input device indicating the position of the participant (receiving position) to each of the audio reproduction positions where the audio reproduction device is located. An advantage of this embodiment is that no additional microphone is required to measure the transmission delay.
[0015] Due to the fact that the pilot signal has predefined properties, the corresponding component can be identified in the audio input signal so that the corresponding component can be separated from the component due to the participant's voice. Thus, the participant's voice can be correctly reproduced.
[0016] The pilot signal is preferably inaudible to a human audience. In one example, the pilot signal has a frequency within a frequency range inaudible to humans. For example, the pilot signal is a wavelet having a center frequency selected within the range between 25 kHz and 30 kHz. Signals within this range are inaudible to humans but can be processed by standard audio equipment.
[0017] In another example, the pilot signal is a low-power broadband signal or a low-power frequency modulated signal with a characteristic phase-frequency relationship. In this case, the signal may have frequency components that are within the audible range, but are inaudible due to the low power level. If the phase-frequency relationship or frequency modulation is in a pattern sufficient to distinguish, the components corresponding to the pilot signal in the audio input signal can be identified in the audio input signal despite their low power level, so that they can be separated from the components generated by the participant's voice. In yet another example, the pilot signal includes a corresponding Barker sequence. Thus, the pilot signal has a strong peak autocorrelation function.
[0018] In some examples, the audio reproduction device placed at the audio reproduction position is actually a multi-channel audio reproduction device. In this case, it is advantageous if the multi-channel audio reproduction device is configured to reproduce the corresponding audio output signal received via each audio channel.
[0019] The improved audio reproduction method according to the second aspect of the present invention comprises:
[0020] converting the acoustic signal into an audio input signal at an audio receiving position in the space;
[0021] receiving one or more additional audio input signals;
[0022] combining audio input signals to provide a plurality of audio output signals;
[0023] Each of the audio output signals is reproduced at a respective audio reproduction position in the space, each of the audio output signals having a respective amplitude weighted according to a respective distance of each of the audio reproduction positions to the audio receiving position.
[0024] An embodiment of the improved audio reproduction method includes correlating the respective signal components with a reference signal component to determine a respective acoustic signal delay occurring at each of the respective distances.
[0025] In one example, the associating includes associating the audio input signal with an auxiliary signal indicative of the acoustic signal perceived at a first audio reproduction position in the audio reproduction positions to determine a delay of the acoustic signal perceived at the first audio reproduction position in the audio reproduction positions; and associating the audio input signal with an auxiliary signal indicative of the acoustic signal perceived at a second audio reproduction position in the audio reproduction positions to determine a delay of the acoustic signal perceived at the second audio reproduction position in the audio reproduction positions.
[0026] In another example, the improved audio reproduction method further includes:
[0027] including a corresponding pilot signal in each of the audio output signals;
[0028] reproducing a first pilot signal of the corresponding pilot signals as a first acoustic signal component at a first audio reproduction position of the audio reproduction positions;
[0029] converting, at an audio receiving location, a version of the first acoustic signal component perceived at the location into a first delay-indicative signal component;
[0030] correlating a first delay-indicative signal component with a first one of the pilot signals to determine a delay at which the first one of the pilot signals reproduced at a first one of the audio reproduction locations was received at the audio reception location;
[0031] reproducing a second pilot signal among corresponding pilot signals at a second one of the audio reproduction positions as a second acoustic signal component;
[0032] converting, at an audio receiving location, a version of the second acoustic signal component perceived at the location into a second delay-indicative signal component;
[0033] A second delay-indicative signal component is correlated with a second one of the pilot signals to determine a delay at which the second one of the pilot signals reproduced at a second one of the audio reproduction locations was received at the audio reception location. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematically illustrating an embodiment of an improved audio reproduction system;
[0035] Figure 2 Illustrative components of an embodiment of an improved audio reproduction system are shown;
[0036] Figure 3 Another exemplary component is shown in greater detail;
[0037] Figure 4 Another embodiment of the improved audio reproduction system is schematically shown. DETAILED DESCRIPTION
[0038] Figure 1 Schematically shows an audio input device M L , an audio reproduction system SYS of a first audio reproduction device ARD1 and a second audio reproduction device ARD2. Audio input device M L Coupled to the audio control unit (ACU, see Figure 2 ), audio input device M L The audio control unit is provided with an audio input signal S ML .
[0039] Figure 2 An exemplary audio control unit ACU is shown, comprising a karaoke processor KP and a mixer Mix. The karaoke processor KP is configured to receive audio signals L, R from a first channel and a second channel, and to provide modified outputs L', R' at its output channels. In an embodiment, the modified output signals L', R' are obtained from the audio signals L, R by completely or partially removing the original singing voice from the audio signals L, R. Cano et al. describe an exemplary method by which this can be achieved in the karaoke processor KP in "Musical source separation: An introduction", IEEE Signal Processing Magazine, 36(1): 31-40, 2019.
[0040] Alternatively or additionally, one or more other components are fully or partially removed from the received audio signal L, R. For example, the contribution from accompanying instruments such as guitar, piano etc. may be fully or partially removed.
[0041] The sound mixer Mix will output L', R' and the audio input device M from singer P1 L The audio input signal S ML The sound mixer mixes the audio input signals S and combines them to generate the respective audio output signals L", R" for the audio reproduction devices ARD1, ARD2. ML are added to the respective audio output signals L", R" to create the impression that the singer P1 is at an apparent position in the room in which the audio reproduction system SYS is arranged. In an embodiment, the apparent position corresponds to the actual position of the singer P1.
[0042] exist Figure 1 In the embodiment shown, each of the audio reproduction devices ARD1, ARD2 is provided with a corresponding auxiliary microphone M S1 、M S2 The auxiliary microphone presents the corresponding auxiliary signal S MS1 、S MS2 , auxiliary signal S MS1 、S MS2 Indicates the voice of the singer audible at the position of the corresponding audio reproduction device ARD1, ARD2. The sound mixer Mix receives these auxiliary signals and is configured to determine the voice of the singer audible at the position of the corresponding auxiliary signal S MS1 、S MS2 Each of the above represents the singing voice relative to the microphone M held by the singer P1. L The obtained audio input signal S ML The time delay of the singing voice is represented by ML It is essentially determined by the singing voice, so it can be associated with each auxiliary signal relatively easily.
[0043] Figure 3 An exemplary embodiment is shown in which the first correlation unit CR1 causes the microphone M held by the singer P1 to L The obtained audio input signal S ML The singing voice represented by the corresponding auxiliary signal S MS1 The singing voice represented by the first auxiliary microphone M is associated and determined S1 Similarly, the second correlation unit CR2 makes the microphone M held by the singer P1 L The obtained audio input signal S ML The singing voice represented by the second auxiliary microphone M S2 Auxiliary signal S MS2 The singing voice represented by the second auxiliary microphone M is associated and determined S2The received singing voice has a delay Δt2. The weight calculation unit CW calculates corresponding weights w1 and w2 for each output channel, and uses the weights w1 and w2 to weight the signal from the microphone held by the singer.
[0044] exist Figure 3 In the embodiment, the singer's microphone M L The signal S ML The delays Δt1, Δt2 measured by the correlation units are further delayed by the corresponding delay units PH1, PH2 of each channel. For example, the delays imposed on the microphone signals by the delay units may be proportional to the delays Δt1, Δt2 measured by the correlation units. By giving one of the channels a greater delay than the other, the perception of position will shift to the channel with the smaller delay. Even if the signal of the singer's microphone is reproduced with the same intensity in each of the channels, the listener will perceive the signal of the channel with the smaller delay with a higher intensity than the other channel. This phenomenon is called "time intensity trading" and is discussed in more detail by RMAarts in: Time / intensity trading stereophony for (HD) TV and audio applications. 14th International Conference on Acoustics (Beijing, China), September 1992. (Conference L8-3). Therefore, in the first delay unit PH1, the signal S ML is delayed by a delay depending on Δt1 to obtain a first intermediate signal S ML1 The first intermediate signal S ML1 The multiplier M1 is weighted with weight w1 to obtain another first intermediate signal S ML11 , the other first intermediate signal S ML11 The modified output signal L' is added by the adder A1 to obtain the audio output signal L' for the first audio reproduction device ARD1. Similarly, in the second delay unit PH2, the signal S ML is delayed by a delay depending on Δt2 to obtain a second intermediate signal S ML2 The second intermediate signal S ML2 The second intermediate signal S is weighted by weight w2 in the multiplier M2 to obtain another second intermediate signal S ML22 , the second intermediate signal S ML22 The modified output signal R' is added by the adder A2, thereby obtaining an audio output signal R" for the second audio reproduction device ARD2.
[0045] It should be noted that the auxiliary signal S MS1 、S MS2It is only used to determine the position of the singer and thereby control the weights w1, w2 and / or the delay of the signal of the singer's microphone being reproduced by the audio reproduction device. Thus, the signal components in the auxiliary signal contributed by other audio sources do not affect the quality of the reproduced singer's voice. The weights and / or delays can be adjusted at a relatively low frequency (for example, a frequency range of about 1HZ to 10Hz) so that the process is also basically insensitive to noise. However, if necessary, noise can be suppressed by auxiliary microphones that have sufficient directional sensitivity and are far away from the audio reproduction device so that they hardly receive the reproduced audio signal. Additionally or alternatively, the relevant unit is configured to consider that the received auxiliary signal includes a signal component originating from the corresponding audio reproduction device. This component can be easily identified because it has basically no delay except for the delay optionally explicitly introduced by the delay units PH1 and PH2.
[0046] In the example shown, the first audio reproduction device ARD1 is a stereo device that reproduces a first audio output signal L" at its two channels. Similarly, the second audio reproduction device ARD2 is a stereo device that reproduces a second audio output signal R" at its two channels. In an alternative embodiment, the sound mixer Mix is configured to present a pair of corresponding output signals S to be reproduced by the left channel and the right channel, respectively, of the first audio reproduction device ARD1. 1L and S 1R , and / or is configured to present a pair of corresponding output signals S to be reproduced by the left channel and the right channel, respectively, of the second audio reproduction device ARD2 2L and S 2R .
[0047] The audio reproduction system SYS proposed here can be easily extended to more channels. For each additional channel, the sound from the singer's microphone M is detected according to the same method as described above for the left and right channels. L The signal S ML The appropriately weighted and optionally delayed signal from the singer's microphone M L The signal S ML .
[0048] Alternatively or additionally, the audio reproduction system SYS can be easily extended for one or more additional singers. In this case, the microphone signal of each singer is also correlated with each auxiliary signal to determine the corresponding signal delay and thus the corresponding position of the singer, and the weights and / or delays used to reproduce the additional microphone signals via the respective channels are appropriately adjusted.
[0049] exist Figure 4 In the embodiment, the audio control unit ACU converts the corresponding pilot signal S A1 、SA2 Included in each of the audio output signals L", R" for the audio reproduction devices ARD1, ARD2. Pilot signal S A1 、S A2 Reproduced by audio reproduction devices ARD1, ARD2 as components A in their respective reproduced audio signals A1, A2 A1 、A A2 . Singer's Microphone M L Output microphone signal S ML , the microphone signal S ML is supplied to a signal separator SPL which separates the microphone signal S ML Divided into basic components S ML0 , first delay indication component S MLD1 and the second delay indication component S MLD2 . Basic component S ML0 Corresponding to the singer's voice, the first delay indication component S MLD1 Corresponding to the reproduction by the first audio reproduction device ARD1 as component A A1 The first pilot signal S A1 , and the second delay indication component S MLD2 Corresponding to the reproduction by the second audio reproduction device ARD2 as component A A2 The second pilot signal S A2 .
[0050] The first correlation unit CR1 determines the component A reproduced by the first audio reproduction device ARD1. A1 Relative to the pilot signal S A1 The second correlation unit CR2 determines the component A reproduced by the second audio reproduction device ARD2. A2 Relative to the pilot signal S A2 is received with a delay Δt2.
[0051] The sound mixer Mix combines the modified output signals L', R' with the basic component S corresponding to the singer's voice. ML0 The corresponding weighted versions of are mixed to obtain the corresponding output signals L", R". Figure 2 In the embodiment of FIG. 5 , the respective weights for obtaining the respective output signals L″, R″ are determined based on the calculated delays Δt1, Δt2, respectively.
[0052] In an embodiment, the pilot signal S used to calculate the delays Δt1 and Δt2 is A1 、S A2 The audio control unit ACU may, for example, periodically generate a wavelet having a first center frequency as the pilot signal S A1, and generate a wavelet with a second center frequency as a pilot signal S A2 The signal separator SPL can easily separate the first delay indication component S using a bandpass filter. MLD1 and the second delay indication component S MLD2 With the basic component S ML0 Distinguish.
[0053] Alternatively, the pilot signal S A1 、S A2 is a low-power broadband signal with a characteristic phase-frequency relationship or a low-power frequency modulated signal that can be distinguished from the microphone signal. A1 、S A2 are not necessarily outside the range of audible frequencies, since they are inaudible anyway due to their low power. Although these pilot signals are inaudible, they can be detected in the microphone signal S due to their characteristic phase-frequency relationship or their characteristic frequency modulation pattern. ML These pilot signals are distinguished in .
[0054] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. In this regard, it should be noted that in the examples presented in this application, the multiple audio reproduction devices include a first audio reproduction device, an audio reproduction device, and a second audio reproduction device. However, the present invention is also applicable to audio reproduction systems with more audio reproduction devices. A single component or other unit may fulfil the functions of multiple items recited in the claims. For example, Figure 3 The embodiments are described with respect to individual functional units CR1, CR2, CW, etc. In practice, it is conceivable to construct an audio reproduction system using corresponding components that implement these functional units. Alternatively, two or more functional units can be implemented using a single component. For example, a single component can implement the related functions CR1, CR2 on a time-sharing basis, and another component can implement delay units PH1, PH2, etc. on a time-sharing basis. It is also possible for a single component to implement mutually different functions on a time-sharing basis. Any component can be provided as dedicated hardware specifically having a specified function, or can be provided as a programmable processor that is appropriately programmed. In other embodiments, the functions of the audio reproduction system are implemented as an integrated circuit or by a trained neural network.
[0055] The mere fact that certain measures are recited in mutually different claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.
[0056] It should be noted that the present invention is similarly applicable to other input sources. For example, the audio input device AID may respond to a musical instrument, such as a guitar or flute, rather than voice input. A combination of audio input devices may also be provided.
Claims
1. An audio reproduction system (SYS), comprising: an audio input device (AID) for converting an acoustic signal into an audio input signal; Audio Control Unit (ACU); as well as Multiple audio reproduction devices (ARD1, ARD2) The audio control unit is configured to convert the audio input signal (S M , L, R) with one or more further audio input signals and providing a respective audio output signal (L", R") to each of the audio reproduction devices, wherein the audio control unit (ACU) is configured to provide the respective audio output signal (L", R") with a respective amplitude weighted according to the respective distance of the audio input device (AID) relative to the respective audio reproduction device (ARD1).
2. The audio reproduction system (SYS) according to claim 1, comprising a first correlation unit (CR1) and a second correlation unit (CR2), said correlation units (CR1, CR2) correlating respective signal components with reference signal components to determine respective acoustic signal delays.
3. The audio reproduction system (SYS) according to claim 2, wherein The first correlation unit (CR1) makes the microphone (M L ) obtained by the audio input signal (S ML ) and a first auxiliary microphone (M) arranged near a first audio reproduction device (ARD1) in the audio reproduction device S1 ) of the auxiliary signal (S MS1 ) is associated to determine the source from the microphone (M L ) is detected by the first auxiliary microphone (M S1 ) perceived delay (Δt1), and wherein the second correlation unit (CR2) causes the second correlation unit (CR2) to receive the signal from the microphone (M L ) obtained by the audio input signal (S ML ) and a second auxiliary microphone (M) arranged near a second audio reproduction device (ARD2) in the audio reproduction device. S2 ) of the auxiliary signal (S MS2 ) is associated to determine the source from the microphone (M L ) is detected by the second auxiliary microphone (M S2 ) perceived delay (Δt2).
4. The audio reproduction system (SYS) according to claim 2, wherein The audio control unit (ACU) is configured to convert the corresponding pilot signal (S A1 , S A2 ) is included in each of the audio output signals (L", R") for the audio reproduction devices (ARD1, ARD2), wherein the first correlation unit (CR1) determines a delay (Δt1) for the pilot signal reproduced by a first audio reproduction device (ARD1) of the audio reproduction devices to be received by the audio input device (AID), and wherein the second correlation unit (CR2) determines a delay (Δt2) for the pilot signal reproduced by a second audio reproduction device (ARD2) of the audio reproduction devices to be received by the audio input device (AID).
5. An audio reproduction system (SYS) according to claim 4, wherein The pilot signal is inaudible to a listener.
6. Audio reproduction system (SYS) according to claim 5, wherein The pilot signal (S A1 , S A2 ) has a frequency in the frequency range inaudible to humans.
7. An audio reproduction system (SYS) according to claim 5, wherein The pilot signal (S A1 , S A2 ) is a low-power broadband signal or a low-power frequency modulated signal with a characteristic phase-frequency relationship.
8. An audio reproduction system (SYS) according to claim 5, wherein The pilot signal (S A1 , S A2 ) including the Barker sequence.
9. An audio reproduction system according to any one of the preceding claims, wherein The audio reproduction device is a multi-channel audio reproduction device configured to reproduce a received respective audio output signal via each audio channel.
10. An audio reproduction method, comprising: converting the acoustic signal into an audio input signal at an audio receiving position in the space; receiving one or more additional audio input signals; combining the audio input signals to provide a plurality of audio output signals (L", R"); Each of the audio output signals (L", R") is reproduced at a corresponding audio reproduction position in the space, and each of the audio output signals (L", R") has a corresponding amplitude weighted according to the corresponding distance of each of the audio reproduction positions to the audio receiving position.
11. The audio reproduction method according to claim 10, comprising: The respective signal components are correlated with the reference signal components to determine a respective acoustic signal delay occurring at each of the respective distances.
12. The audio reproduction method according to claim 11, wherein: The association includes: making the audio input signal (S ML ) and an auxiliary signal (S) indicative of the acoustic signal perceived at a first one of the audio reproduction positions MS1 ) to determine a delay (Δt1) of the acoustic signal perceived at the first of the audio reproduction positions; and causing the audio input signal (S ML ) and an auxiliary signal (S) indicative of the acoustic signal perceived at a second one of the audio reproduction positions MS2 ) to determine a delay (Δt2) of the acoustic signal perceived at the second one of the audio reproduction positions.
13. The audio reproduction method according to claim 11, comprising: The corresponding pilot signal (S A1 , S A2 ) is included in each of said audio output signals (L", R"); A first pilot signal (S) among the corresponding pilot signals is played at a first audio reproduction position among the audio reproduction positions. A1 ) is reproduced as the first acoustic signal component (A A1 ); At the audio receiving location, a version of the first acoustic signal component perceived at that location is converted into a first delay-indicating signal component (S MLD1 ); The first delay indication signal component (S MLD1 ) and the first pilot signal (S A1 ) to determine a delay (Δt1) at which the first pilot signal, of the pilot signals reproduced at the first audio reproduction position in the audio reproduction positions, is received at the audio reception position; A second pilot signal (S) in the corresponding pilot signal is played at a second audio reproduction position in the audio reproduction position. A2 ) is reproduced as the second acoustic signal component (A A2 ); At the audio receiving position, a version of the second acoustic signal component perceived at that position is converted into a second delay-indicating signal component (S MLD2 ); The second delay indication signal component (S MLD2 ) and the second pilot signal (S A2 ) to determine a delay (Δt2) at which the second pilot signal, among the pilot signals reproduced at the second audio reproduction position among the audio reproduction positions, is received at the audio reception position.
14. The audio reproduction method according to claim 13, wherein: The reproduced pilot signal is inaudible to humans.