Acoustic scene playback method and apparatus

By configuring and assigning virtual speaker objects (VLOs) to the microphone, generating encoded data streams, and decoding and rendering them, the dependence on sound source location and room attributes in existing technologies is resolved, enabling high-quality audio playback in real acoustic scenarios and real-time virtual listening position changes.

WO2025218310A9PCT designated stage Publication Date: 2025-12-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/075216
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-01-26
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing technologies require explicit knowledge of room properties and source location when reproducing recorded audio scenes. Furthermore, object-based solutions are only suitable for synthesized scenes and not for high-quality rehearsals of real acoustic scenes, and require multi-microphone recording or source separation technology.

Method used

By providing recording data and microphone signal and metadata of microphone configuration, specifying virtual listening position, and assigning virtual speaker objects (VLOs) to each microphone configuration, generating encoded data streams for decoding and rendering, the virtual listening position can be changed in real time.

Benefits of technology

It enables real-time adjustment of audio playback position within a realistically recorded acoustic environment, avoiding dependence on sound source location and room properties, improving computational efficiency and audio playback realism, and supporting real-time streaming and interactive experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075216_11122025_PF_FP_ABST
    Figure CN2025075216_11122025_PF_FP_ABST
Patent Text Reader

Abstract

An acoustic scene playback method, comprising: providing recording data, comprising a microphone signal of one or more microphone configurations in an acoustic scene and microphone metadata of the one or more microphone configurations (200); specifying a virtual listening position, wherein the virtual listening position is a position within the acoustic scene (210); allocating one or more virtual loudspeaker objects (VLOs) to each microphone configuration among the one or more microphone configurations (220); generating an encoded data stream on the basis of the recording data, the virtual listening position, and VLO parameters allocated to the one or more microphone configurations (230); decoding the encoded data stream on the basis of a playback configuration to generate a decoded data stream (240); and inputting the decoded data stream to a rendering device (250). Further provided are a playback apparatus for implementing the acoustic scene playback method and a data carrier.
Need to check novelty before this filing date? Find Prior Art

Description

Acoustic scene playback method and apparatus

[0001] This application claims priority to the Chinese Patent Application No. 202410452474.3, filed on April 15, 2024, and entitled "Acoustic scene playback method and apparatus", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present invention relates to an acoustic scene playback method and apparatus. BACKGROUND

[0003] In classical recording techniques, a surround image of a spatial audio scene, also referred to as acoustic scene or sound scene, is captured and reproduced in the original sound scene from a single listener's perspective. Single perspective recording is usually achieved by stereo (channel-based) recording and reproduction techniques or ambisonic (scene-based) recording and reproduction techniques. The advent of interactive audio displays and the shift of audio transmission media from tapes or CDs to more flexible media has made the use of audio more dynamic, e.g. interactive client-side audio rendering of multi-channel data or server-side rendering and transmission of pre-rendered audio streams for clients. While the above mentioned techniques are already common in gaming, they are rarely used for reproducing recorded audio scenes.

[0004] So far, it has been possible to traverse a sound scene in the reproduction process only by audio rendering based on separate isolated recordings of the sound involved and additional recordings or renderings of the ambience (object-based). By changing the arrangement of the recorded sound sources, the playback angle at the reproduction side can be adjusted.

[0005] Furthermore, another possibility is to infer a kind of parallax adjustment to create the impression of a perspective change from a single perspective recording by remapping directional audio coding. This is done by assuming the source position after projecting the direction of the source position onto a convex hull. This arrangement relies on time-varying signal filtering using the spectral separation assumption of direct / early sound. However, this leads to signal attenuation. Furthermore, the assumption that the source lies on the convex hull only applies to small position changes.

[0006] Therefore, the limitation of the prior art is that when using object-based audio rendering for rendering a rehearsal, the room properties, the source positions and the properties of the sources themselves need to be displayed. Furthermore, obtaining an object-based representation from a real scene is a difficult task and requires either many microphones close to all the required sources or source separation techniques to extract the individual sources from a mixed source. Therefore, the object-based approach is only suitable for synthetic scenes and cannot be used to achieve high quality rehearsals in real acoustic scenes.

[0007] The present invention solves the problems of the prior art and allows for continuously changing the virtual listening position at which the sound in the recorded acoustic scene is played back, when playing back the sound in the acoustic scene at a virtual listening position. Thus, the present invention solves the problem of playing back an acoustic scene using an improved method and apparatus. SUMMARY

[0008] In a first aspect, a method of acoustic scene playback is provided, wherein the method comprises:

[0009] providing recording data comprising microphone signals of one or more microphone configurations located within an acoustic scene and microphone metadata of the one or more microphone configurations, wherein each of the one or more microphone configurations comprises one or more microphones and has a recording point as a center position of the respective microphone configuration;

[0010] specifying a virtual listening position, wherein the virtual listening position is a position within the acoustic scene;

[0011] assigning one or more virtual loudspeaker objects (VLOs) to each of the one or more microphone configurations, wherein each VLO is an abstract sound output object within a virtual free field;

[0012] generating an encoded data stream based on the recording data, the virtual listening position and VLO parameters assigned to the one or more microphone configurations;

[0013] decoding the encoded data stream based on a playback configuration, thereby generating a decoded data stream; and

[0014] inputting the decoded data stream to a rendering device, thereby driving the loudspeaker device to reproduce the sound in the acoustic scene at the virtual listening position.

[0015] A virtual free field is an abstract (i.e. virtual) sound field comprising direct sound but not reverberation. Virtual means modelled or represented on a machine, such as a computer, or on a system of interacting computers. An acoustic scene comprises a spatial region and sounds within that spatial region. In addition to an acoustic scene, it can alternatively be referred to as a sound field or a spatial audio scene. Furthermore, a rendering device can be one or more loudspeakers and / or one or more headphones. Thus, a listener listening to a reproduced sound of the acoustic scene at a virtual listening position is able to change the desired virtual listening position and virtually walk through the acoustic scene. In this way, the listener is able to experience or re-experience the whole original sound stage, e.g. a concert. The user can walk through the whole acoustic scene and listen at any point within the scene. Thus, the user can explore the whole acoustic scene in an interactive way by determining and inputting a desired position within the acoustic scene and is able to listen to the sounds within the acoustic scene at the chosen position. For example, at a concert, the user can choose to listen at the back, in the crowd, right in front of the stage or even on the stage surrounding the musicians. Furthermore, applications in virtual reality (VR) are conceivable which extend from rotations to also being able to translate. In the present invention, only the recording positions and the virtual listening positions are required to be known. Thus, in the present invention, no information about the original sound sources, e.g. the musicians, such as their number, position or orientation, is required. In particular, since virtual loudspeaker objects (VLOs) are used, the spatial distribution of the sound sources is inherently encoded and no estimation of the actual positions is required. Furthermore, room properties such as reverberation are inherently encoded and drive signals for driving the VLOs are used which do not correspond to source signals, thus no actual sound source signals need to be recorded or estimated. These drive signals are derived from the microphone signals by a data-independent linear processing. Furthermore, the present invention is computationally efficient and enables real-time encoding and rendering. Thus, a listener is able to interactively change the desired virtual listening position and virtually walk through the (recorded) acoustic scene, e.g. a concert. Due to the computational efficiency of the present invention, the acoustic scene can be streamed to a remote, e.g. playback device, in real-time. The present invention does not rely on prior information about the number or position of sound sources. Similar to classical mono or surround sound recording techniques, all sound source parameters are inherently encoded and no estimation is required. In contrast to object-based audio approaches, the sound source signals do not need to be isolated and thus no close microphones are required and audible artefacts due to source signal separation are avoided.

[0016] A Virtual Loudspeaker Object (VLO) can be implemented on a computer, e.g. as an object in an object-based spatial audio layer. Each VLO can represent a mixture of a sound source, early reflections and diffuse sound. In this context, a sound source is a local primary source, e.g. a person speaking or singing, an instrument or a physical loudspeaker. Generally, a joint of several, i.e. two or more, VLOs is needed to reproduce an acoustic scene.

[0017] According to the first aspect, in a first implementation form of the method, after assigning one or more VLOs to each microphone configuration, for each microphone configuration, the one or more VLOs are placed at positions within the virtual sound field corresponding to the recording points of the respective microphone configuration within the acoustic scene.

[0018] This facilitates virtually configuring a virtual reproduction system comprising VLOs for each recording point in a common virtual free field. Thus, these features of the first implementation form facilitate implementing an arrangement in which a user is able to change a virtual listening position of an audio playback within a truly recorded acoustic scene when playing back a signal corresponding to the selected virtual listening position.

[0019] According to the first aspect, in a second implementation form of the method, the VLO parameters comprise one or more static VLO parameters of the one or more VLOs being independent of the virtual listening position and describing properties, the static VLO parameters being fixed for the acoustic scene playback.

[0020] Thus, the VLO parameters within the virtual free field describe properties of the VLOs, the properties being fixed for a certain playback configuration arrangement, which facilitates adequately configuring a reproduction system in the virtual free field and describing properties of the VLOs within the virtual free field. For example, if the playback is provided through loudspeakers in a room or in-ear, the playback configuration arrangement refers to properties of the playback device itself, etc.

[0021] According to the first aspect, in a third implementation form of the method, the method further comprises calculating the one or more static VLO parameters based on the microphone metadata and / or a critical distance, wherein the critical distance is a distance at which a sound pressure level of a direct sound and a sound pressure level of a reverberant sound are equal for a directional sound source, or receiving the one or more static VLO parameters from a transmission device prior to generating the encoded data stream.

[0022] Thus, the static VLO parameters can be calculated within the playback device or can be received from elsewhere, e.g. from a transmission device. Moreover, since the static VLO parameters take into account the microphone metadata and / or the critical distance, the static VLO parameters take into account parameters at the time of recording the acoustic scene, such that the playback device plays back a certain sound corresponding to a certain virtual listening position as realistically as possible.

[0023] According to the first aspect, in a fourth implementation form of the method, for each of the one or more microphone configurations, the one or more static VLO parameters comprise: a plurality of VLOs, and / or a distance of each VLO to the recording point of the respective microphone configuration, and / or an angular layout of the one or more VLOs assigned to the respective microphone configuration (e.g. with respect to the direction of the one or more microphones of the respective microphone configuration), and / or a mixing matrix B defining a mixing of the microphone signals of the respective microphone configuration i .

[0024] Thus, the static VLO parameters are parameters that are fixed for a certain acoustic scene playback and do not change during the acoustic scene playback and are independent of the selected virtual listening position.

[0025] According to the first aspect, in a fifth implementation form of the method, the VLO parameters comprise one or more dynamic VLO parameters that are dependent on the virtual listening position, the method comprising: calculating the one or more dynamic VLO parameters based on the virtual listening position prior to generating the encoded stream, or receiving the one or more dynamic VLO parameters from a transmission device.

[0026] Thus, both the static VLO parameters and the dynamic VLO parameters can easily be generated within the playback device or can both be received from a separate (e.g. remote) transmission device. Moreover, the dynamic VLO parameters are dependent on the selected virtual listening position, such that the sound playback will be determined according to the selected virtual listening position and the dynamic VLO parameters.

[0027] According to the first aspect, in a sixth implementation form of the method, for each of the one or more microphone configurations, the one or more dynamic VLO parameters comprise: one or more VLO gains, wherein each VLO gain is a gain of a control signal of a corresponding VLO, and / or one or more VLO delays, wherein each VLO delay is a time delay of a sound wave propagating from the corresponding VLO to the virtual listening position, and / or one or more VLO incidence angles, wherein each VLO incidence angle is an angle between a line connecting the recording point and the corresponding VLO and a line connecting the corresponding VLO and the virtual listening position, and / or one or more parameters indicative of a radiation directivity of the corresponding VLO.

[0028] By providing the VLO gain, proximity regularization can be performed by regularizing the gain depending on the distance between the corresponding VLO corresponding to the VLO gain and the virtual listening position. Furthermore, the direction dependency can be ensured, since the VLO gain can be determined depending on the virtual listening position relative to the VLO position within the virtual free field. Thus, a more realistic sound impression can be conveyed to the listener. Furthermore, the VLO delay, the VLO incidence angle and the parameter indicating the directivity of the radiation contribute to the realistic sound impression as well.

[0029] According to the first aspect, in a seventh implementation form of the method, the method further comprises, prior to generating the encoded data stream, calculating an interactive VLO format comprising, for each recording point and each VLO assigned to the recording point, a resulting signal and an incidence angle φ ij , wherein wherein g ij is a gain factor of a control signal x ij of the j-th VLO at the i-th recording point, τ ij is a time delay of a sound wave propagating from the j-th VLO of the i-th recording point to the virtual listening position, t denotes time, the incidence angle φ ij is an angle between a line connecting the i-th recording point and the j-th VLO at the i-th recording point and a line connecting the j-th VLO of the i-th recording point and the virtual listening position.

[0030] Thus, an interactive VLO format of a certain kind can be effectively used as input for encoding, such that this interactive VLO format contributes to effectively performing the encoding.

[0031] According to the first aspect, in an eighth implementation form of the method, the gain factor g ij is determined depending on the incidence angle φ ij and a distance d ij between the j-th VLO at the i-th recording point and the virtual listening position.

[0032] Thus, if the virtual listening position is close to the corresponding VLO, proximity regularization is possible. Furthermore, the direction dependency can be ensured, such that the gain factor acknowledges proximity regularization and direction dependency.

[0033] According to the first aspect, in a ninth implementation form of the method, for generating the encoded data stream, each resulting signal and incidence angle φ ij is input to an encoder, in particular a binaural reverberation sound encoder.

[0034] Thus, a prior art ambisonics encoder can be used, wherein the specific signals are input into the ambisonics encoder for encoding, i.e. each resulting signal and the angle of incidence φ ij The above mentioned effects are achieved in connection with the first aspect. Thus, the application according to the first aspect or any of the implementation forms provides a very simple and cost-effective arrangement, wherein the application can be implemented using a prior art ambisonics encoder.

[0035] According to the first aspect, in a tenth implementation form of the method, for each of the one or more microphone configurations, the one or more VLOs assigned to the respective microphone configuration are provided on a circle line having the recording point of the respective microphone configuration as center of the circle line within the virtual free field, the circle line having a radius R i According to the order of directivity of the microphone configuration, the reverberation of the acoustic scene and the average distance d i is determined.

[0036] Thus, the VLOs can be effectively arranged within the virtual free field, thereby providing a very simple arrangement for achieving the effects of the application.

[0037] According to the first aspect, in an eleventh implementation form of the method, the number of VLOs on the circle line and / or the angular position of each VLO on the circle line and / or the directivity of the acoustic radiation of each VLO on the circle line is determined according to the order of directivity of the microphone of the respective microphone configuration and / or the recording principle of the respective microphone configuration and / or the radius R i and / or the distance d ij between the j-th VLO of the i-th microphone configuration and the virtual listening position.

[0038] These features help to generate a realistic sound impression for the listener and help to achieve all the advantages already mentioned above in connection with the first aspect.

[0039] According to the first aspect, in a twelfth implementation form of the method, the recording data is received from outside, i.e. from outside of the apparatus implementing the VLOs, in particular by applying streaming media, in order to provide the recording data.

[0040] This makes it possible that the recorded data is not generated in any playback device, but can be received directly from a corresponding transmission device or the like, for example, which is recording an acoustic scene, for example a concert, and provides the recorded data in a live stream to the playback device. The playback device can then perform the acoustic scene playback method provided herein. Thus, in the present application, a live stream of a concert or the like acoustic scene can be realized. The VLO parameters in the present application can be adjusted in real time depending on the selected virtual listening position. Thus, the present application is computationally efficient and enables real-time encoding and rendering. Therefore, a listener is able to interactively change the desired virtual listening position and virtually walk through the recorded acoustic scene. Due to the computational efficiency of the present application, the acoustic scene can be streamed to the playback device in real time.

[0041] According to the first aspect, in a thirteenth implementation form of the method, for providing the recorded data, the recorded data is extracted from a recording medium, in particular from a CD-ROM.

[0042] This is a further possibility to provide the recorded data to the playback device, namely by inserting a CD-ROM into the playback device, wherein the recorded data is extracted from the CD-ROM, thereby providing the recorded data for acoustic scene playback.

[0043] According to a second aspect, a playback device or a computer program or both are provided. The playback device is configured to perform the method according to the first aspect, in particular according to any of its implementation forms. The computer program can be provided on a data carrier, which, when executed on a computer, can instruct the playback device to perform the method according to the first aspect, in particular according to any of its implementation forms. BRIEF DESCRIPTION OF DRAWINGS

[0044] Fig. 1 shows a representative acoustic scene with several virtual listening positions within the acoustic scene;

[0045] Fig. 2a shows an acoustic scene playback method according to an embodiment of the present application;

[0046] Fig. 2b shows an acoustic scene playback method according to a further embodiment of the present application;

[0047] Fig. 2c shows an acoustic scene playback method according to a further embodiment of the present application;

[0048] Fig. 2d shows an acoustic scene playback method according to a further embodiment of the present application;

[0049] Fig. 2e shows an acoustic scene playback method according to a further embodiment of the present application;

[0050] Fig. 3 shows a block diagram of an acoustic scene playback method according to an embodiment of the application;

[0051] Fig. 4 shows exemplary microphone and sound source distribution within an acoustic scene;

[0052] Fig. 5 shows exemplary reproduction configurations for different microphone configurations;

[0053] Fig. 6 shows VLOs in a virtual free field and corresponding virtual listening positions;

[0054] Fig. 7 shows a block diagram of an interactive VLO format for computing microphone signals according to an embodiment of the application;

[0055] Fig. 8 shows a block diagram of encoding / decoding of an interactive VLO format according to an embodiment of the application;

[0056] Fig. 9 shows an arrangement and configuration of VLOs assigned to corresponding microphone configurations according to an embodiment of the application;

[0057] Fig. 10 shows a directional pattern of a VLO according to an embodiment of the application;

[0058] Fig. 11 shows a relationship between VLOs and virtual listening positions in a virtual free field according to an embodiment of the application;

[0059] Fig. 12a shows another relationship between VLOs and virtual listening positions in a virtual free field according to another embodiment of the application;

[0060] Fig. 12b shows another relationship between VLOs and virtual listening positions in a virtual free field according to another embodiment of the application;

[0061] Fig. 13 shows a relationship between a function f indicating a gain of a corresponding VLO and a distance of the VLO to a virtual listening position according to an embodiment of the application.

[0062] Generally, it is noted that all apparatuses, devices, elements, units and means described in the present application can be implemented by software or hardware elements or any combination thereof. All steps which are performed by the various entities described in the present application as well as the functionalities described to be performed by the various entities are intended to mean respective specific apparatuses or devices adapted to perform the respective steps and functionalities, or programmed to perform the respective steps and functionalities. Even if in the following description or in specific embodiments a specific functionality or step to be performed by a general entity is not reflected in the description of a specific detailed apparatus or device performing that specific step or functionality, it should be clear for the skilled person that these elements can be implemented in respective software or hardware elements, or any combination thereof. Further, the methods of the present application and its various steps are embodied in the functionalities of the described apparatus elements. DETAILED DESCRIPTION

[0063] Fig. 1 shows an acoustic scene (e.g. a concert hall) and sounds in the acoustic scene. Here, some people in the crowd are enjoying the music played by a band. The person near the lower left corner represents a certain virtual listening position. In general, and not only in the present example, a virtual listening position can be selected, e.g. by a user of a playback device for acoustic scene playback according to embodiments of the present application. Fig. 1 shows several virtual listening positions within the acoustic scene, which can be selected at will by a user of the playback device or by an automated process without any manual input by the user of the playback device. For example, Fig. 1 shows virtual listening positions behind the crowd, within the crowd, in front of the crowd and next to the musicians on stage or on stage.

[0064] Fig. 2a shows a method of acoustic scene playback according to an embodiment of the present application. In step 200, recording data is provided, comprising microphone signals of one or more microphone configurations placed within an acoustic scene and microphone metadata of the one or more microphone configurations. The one or more microphone configurations each comprise one or more microphones. In this context, the microphone metadata can be microphone positions, microphone directions and microphone characteristics within the acoustic scene of Fig. 1, etc. According to step 200, only the recording data needs to be provided. The recording data can be calculated within any playback device performing the method of acoustic scene playback or can be received from elsewhere; the method step 200 of providing the recording data (to the playback device) is a method step covering both alternatives.

[0065] Subsequently, in step 210, a virtual listening position can be specified. The virtual listening position is a position within the acoustic scene. The virtual position can be specified, e.g. by a user using the playback device. For example, the user can be enabled to specify a certain virtual listening position by entering the virtual listening position into the playback device. However, specifying the virtual listening position is not limited to this example and can also be specified in an automated manner without manual input by a listener. For example, it is conceivable to read the virtual listening position from a CD-ROM or to extract the virtual listening position from a storage unit, thus without any manual determination by a listener.

[0066] Furthermore, in a subsequent step 220, one or more virtual loudspeaker objects (VLOs) can be assigned to each of the one or more microphone configurations. Each microphone configuration comprises (or defines) a recording point, which is a central position of the microphone configuration. Each VLO is an abstract sound output object within a virtual free field. A virtual sound field is an abstract sound field comprising direct sound but not reverberation. This method step 220 contributes to the advantage of embodiments of the present application, namely to virtually establish a reproduction system comprising VLOs for each recording point in a virtual free field. In embodiments of the present application, the desired effect is achieved by virtual loudspeaker objects (VLOs), namely to reproduce sound in an acoustic scene at a desired virtual listening position. These VLOs are abstract sound objects placed in a virtual free field.

[0067] In step 230, an encoded data stream is generated (e.g. in a playback phase after the recording phase) based on the recording data, the virtual listening position and the VLO parameters assigned to the one or more microphone configurations. For each of the one or more microphone configurations, the encoded data stream can be generated by virtually driving the one or more VLOs assigned to the respective microphone configuration such that the one or more VLOs virtually reproduce the sound recorded by the respective microphone configuration. Then, the virtual sound at the virtual listening position can be obtained by superimposing (i.e. by forming a linear combination of) the virtual sounds from all VLOs (i.e. from the VLOs of all microphone configurations) in the method at the virtual listening position.

[0068] In step 240, the encoded data stream is decoded based on a playback configuration, thereby generating a decoded data stream. In this context, the playback configuration can be a configuration corresponding to a loudspeaker array arranged in a room of a home, for example, in which a listener wants to listen to sound corresponding to the virtual listening position, or to headphones worn by the listener when listening to sound in the acoustic scene at the virtual listening position.

[0069] Furthermore, in step 250, the decoded data stream can be input to a rendering device, thereby driving the rendering device to reproduce sound in the acoustic scene at the virtual listening position. The rendering device can be one or more loudspeakers and / or headphones.

[0070] Thus, it is possible to change the desired virtual listening position for (3D) audio playback within a real recorded acoustic scene by a user of a playback device. For example, this enables the user to walk through the acoustic scene and listen at any point in the scene. Thus, the user can explore the acoustic scene in an interactive way by entering the desired virtual listening position into the playback device. In the present application, according to the embodiment of Fig. 2a, the VLO parameters are adjusted in real-time when the virtual listening position is changed. Thus, the embodiment according to Fig. 2a corresponds to a computationally efficient approach and enables real-time encoding and rendering. According to the embodiment of Fig. 2a, only the recorded data and the virtual listening position need to be provided. The current embodiment of Fig. 2a does not rely on prior information about the number or position of sound sources. Furthermore, all sound source parameters are inherently encoded and do not need to be estimated. In contrast to object-based audio approaches, the sound source signals do not need to be isolated and thus closed microphones are not required and audio artifacts due to sound source signal separation are avoided.

[0071] Fig. 2b shows a further embodiment of the present application of a method for playback of an acoustic scene. In contrast to the embodiment of Fig. 2a, the embodiment of Fig. 2b further comprises a step 225 of placing one or more VLOs within the virtual sound field at positions corresponding to the recording points of the microphone configuration in the acoustic scene for each microphone configuration. For example, the VLOs corresponding to each recording point in the virtual free field can be placed as shown in Fig. 9. In Fig. 9, a set of microphones 2 at the ith recording point can be considered as one quasi-coincident microphone array, provided that the distance between the microphones 2 in the set is less than, for example, 20 cm, and the ith recording point is the center position of the set of microphones 2. For each (quasi-coincident) microphone array at the ith recording point, the average distance of the microphone array to its neighboring (quasi-coincident) microphone arrays can be estimated based on the Delaunay triangle of the sum of all microphone positions, i.e. all coordinate points of the microphones. For one (quasi-coincident) microphone array with the ith recording point, the average distance d i is the average distance of the ith recording point to the center position of the set of microphones 2. Furthermore, the signal of the microphone array located at the ith recording point is played back by the VLOs provided on a circle with radius R i around the position r i , where r i is the vector from the coordinate origin to the center position of the ith recording point. The circle contains L i virtual loudspeaker objects, and the radius R i can be calculated according to the following equation:

[0072] R i = c0max(d i , 3m)

[0073] Here, c0is a design parameter depending on the directivity order of the microphones and the reverberation of the recording room (especially the critical distance r H , i.e. the distance at which the sound pressure level of the direct sound and the reverberant sound are equal for a directional sound source). Thus, for a microphone directivity order N = 0, c0is 0; for a microphone directivity order N > 1, for a reverberation room (low r H < 1 m), c0is 0.4; for an "average room" (r H < 2 m), c0is 0.5; and for a dry room (r H > 3 m), c0is 0.6. The number L i of virtual loudspeaker objects recording the signals of the microphone array at the i-th recording point, the angular position of each virtual loudspeaker object, and the virtual loudspeaker directivity control are determined depending on the microphone directivity order N i , the channel- or scene-based recording principle of the microphone array, the radius R i of the arrangement of the virtual loudspeakers around the end point of the vector r i , and the distance d ij between the j-th VLO at the i-th recording point and the virtual listening position.

[0074] Furthermore, for a directivity order N i = 0 and a single microphone, L i = 1 at the i-th recording point, the virtual loudspeaker directivity control is not provided (omnidirectional mode) without virtual sound wave directivity. In this case, a virtual loudspeaker object is provided at the recording position of the single microphone.

[0075] Moreover, for the case N i > 1, a decision has to be made between a channel-based microphone array and a scene-based microphone array:

[0076] • For a channel-based microphone array with a directivity order N i > 1 of K i channels (e.g. single-channel cardioid, single-channel shotgun, dual-channel XY recording, dual-channel ORTF recording, small front-facing three-channel arrangement), as a default adjustment, each of the L i VLOs at the i-th recording point is placed on a coaxial line with respect to the microphone it is assigned to, using R i as the distance from the center position of the recording point i to the corresponding VLO. Coaxial means that the VLO of a microphone in the microphone array is provided on the same line connecting the microphone and the i-th recording point.

[0077] Otherwise, no default adjustment is used as long as there is a standard loudspeaker layout for the channel-based microphone array configuration that is used to place the VLOs at R i ORTF is the same, where one playback loudspeaker pair is used specifically for a directional order N

[0078] • For a scene-based microphone array with a directional order N i ≥ 1 (e.g. B-format), the VLOs are generated according to the following parameters:

[0079] R i ≤ 2.5 m: L i = 4N i , an angular spacing of 90° / N i and a controlled directivity determined according to the virtual listening position, wherein the angular spacing represents the angular spacing of two adjacent VLOs assigned to the same i-th recording point;

[0080] 2.5 m < R i ≤ 3.5 m: L i = 5N i , an angular spacing of 72° / N i and a controlled directivity determined according to the virtual listening position;

[0081] R i > 3.5 m: L i = 6N i , an angular spacing of 60° / N i and a controlled directivity determined according to the virtual listening position.

[0082] Furthermore, for a scene-based microphone array (stereo reverb microphone array), the arrangement of the VLOs can potentially overlap in the virtual free field. To avoid this, each arrangement of VLOs assigned to a corresponding recording point is rotated with respect to the other VLO arrangements in the virtual free field such that the minimum distance of adjacent VLO arrangements becomes maximum.

[0083] Thus, the positions of the VLOs corresponding to the corresponding recording points can be determined within the virtual free field. As mentioned above, Fig. 9 merely represents one example, wherein e.g. a microphone configuration 1 comprising five microphones 2 is provided. Furthermore, the corresponding VLOs 3 corresponding to the microphones 2 are also shown together with the construction lines supporting the correct determination of the positions of the corresponding VLOs 3.

[0084] Furthermore, all other method steps as shown in Fig. 2b are the same as in Fig. 2a.

[0085] Fig. 2c shows another embodiment which additionally provides a method step 227 of calculating one or more static VLO parameters based on microphone metadata and / or a critical distance, which is the distance at which the sound pressure level of the direct sound and the reverberant sound are equal for a directional sound source, or receiving one or more static VLO parameters from a transmission device. In this context, it is noted that in principle, the method step 227 can also be provided before any of the method steps 200, 210, 220 and 225 are performed or between two of the method steps 200, 210, 220 or 225. Thus, the position of the step 227 in Fig. 2c is merely an example position. In this context, the static VLO parameters do not depend on any desired virtual listening position and are determined only once for a certain recording configuration and acoustic scene playback and do not change for the acoustic scene playback. In this context, the recording configuration refers to all microphone positions, microphone directions, microphone characteristics and other characteristics of the acoustic scene recording site. For example, the static VLO parameters can be the number of VLOs per recording point, the distance of a VLO to an assigned recording point, the angular layout of VLOs and a mixing matrix B i of the i-th recording point. The term "angular layout" can refer to the angle between a line connecting a recording point and a VLO assigned to the recording point and a line starting from the microphone and pointing in the main extraction direction of the microphone. However, the term "angular layout" can also refer to the angular distance between adjacent VLOs assigned to the same recording point. These static VLO parameters are determined depending on the microphone positions, the microphone characteristics, the microphone directions and the estimated or assumed critical distance. In a room, the critical distance is the distance to a sound source at which its direct sound equals the reverberant sound in the room. The shorter the distance, the louder the direct sound and the longer the distance, the louder the reverberant sound.

[0086] Fig. 2d shows yet another embodiment of the present application. In comparison to the embodiment of Fig. 2c, Fig. 2d also involves a method step 228: calculating one or more dynamic VLO parameters based on the virtual listening position, or receiving one or more dynamic VLO parameters from a transmission device. In this context, it should be noted that step 228 in Fig. 2d is disclosed after step 227 and before step 230, however, the position of step 228 within the method flowchart of Fig. 2d is merely an example, in principle, step 228 can be moved at any position within Fig. 2d, as long as the method step is executed before the encoded data stream is generated and after the virtual listening position is specified. Thus, method step 228 involves two possibilities, namely, calculating the dynamic VLO parameters within the playback device, or alternatively, receiving the dynamic VLO parameters from an external source, e.g. from a transmission device. In this context, the dynamic parameters are determined according to the desired virtual listening position and are recalculated whenever the virtual listening position changes. Examples of dynamic VLO parameters include: VLO gains, where each VLO gain is the gain of the control signal for the corresponding VLO; VLO directivity, i.e. the directivity of the virtual sound waves radiated by the corresponding VLO; VLO delays, where each VLO delay is the time delay for a sound wave to propagate from the corresponding VLO to the virtual listening position; and VLO angles of incidence, where each VLO angle of incidence is the angle between the line connecting the recording point and the corresponding VLO and the line connecting the corresponding VLO and the virtual listening position. For example, as can be seen in Fig. 11, Fig. 11 or Fig. 12b provide a schematic view in which the angles of incidence φ 12 , φ 22 and φ 31 are indicated, the three angles φ ij are angles of incidence, each being the angle between the line connecting the corresponding i-th recording point and the corresponding j-th VLO and the line connecting the corresponding j-th VLO and the virtual listening position. Furthermore, Fig. 11 also shows the distances d ij , i.e. the distances d 12 , d 22 and d 31 , indicating the distance between the corresponding j-th VLO at the corresponding i-th recording point and the virtual listening position. Thus, as can be seen in Fig. 12a, the distance vector d ij may be calculated as d ij = r ij - r, where r is the vector connecting the position of the virtual listening position and the origin of the coordinate system as seen in Fig. 12a, and r ij is the vector indicating the position of the corresponding j-th VLO at the i-th recording point within the coordinate system. Furthermore, the VLO delay τ ij indicates the time it takes for a virtual sound wave to propagate out of the j-th VLO at the i-th recording point, and can be defined as τ ij = d ij / c, where c is the velocity of the sound wave. Additionally, the VLO gain g... ij It can be calculated as: g ij =f(φ ij ,d ij ) / d ij In this context, the function f(φ) ij ,d ij ) is a kind of dependency on d ij And provides near-regularization and due to φ ij And functions that provide direction dependency.

[0087] In this context, the function f(φ) ij ,d ij As exemplarily shown in Figure 13, Figure 13 illustrates a VLO f(d) on the y-axis. ij (180°), the x-axis indicates the distance d from this VLO. ij Therefore, from the gain g mentioned above... ij The definition clearly shows that it implements the classical free field 1 / d corresponding to the virtual loudspeaker object. ij Attenuation, and due to the function f(φ) ij ,d ij This provides additional distance-dependent attenuation, which avoids an unrealistically loud signal whenever the virtual listening position is very close to the virtual speaker object. This can be seen in Figure 13, which indicates this additional distance-dependent attenuation. As seen in Figure 13, for example, if the distance d from the virtual listening position to the corresponding VLO is... ij For distances greater than 0.5m, classical free-field 1 / r attenuation is provided. However, if the distance d... ij =0, then it provides, for example, 15dB of attenuation. Furthermore, it can be clearly seen from Figure 13 that when 0... <d ij For values ​​<0.5m, linear interpolation is provided. Furthermore, f(φ) ij ,d ij Therefore, it can be calculated using the following formula:

[0088] in, d min d ij The linear interpolation begins when φ = 0, because φ ij =0°, d min2 This represents the limit of linear interpolation, which is in φ ij =180° from d min2 to d min Provided in the interval, as shown in d min2 As shown in Figure 13.

[0089] Here, the first item denotes distance regularization, the second term denotes the direction dependence of the virtual sound wave radiated by the corresponding VLO.

[0090] The radiation characteristics of each VLO can be adjusted such that the interactive directivity (determined according to the virtual listening position) differentiates between an "inner" and an "outer" part within the arrangement of VLOs corresponding to the corresponding microphone configuration, in a way that the signal amplitude dominating the "outer" part is reduced to some extent to avoid misplacement at the diffuse far field. Furthermore, the directivity is formulated as a mixture of an omni- and a cardioid directivity pattern, with controllable order where a and b indicate parameters used when calculating the direction dependence of the virtual sound wave radiated by the corresponding VLO. Here, a determines the weight of the omni-radiation, b determines the weight of the cardioid directivity pattern in the above expression. Furthermore, it is also conceivable to have a directivity in the shape of a hemispherical slepian function. Furthermore, in particular for larger distances d ij between the virtual loudspeaker object and the virtual listening position, the backward amplitude of each VLO can be reduced by controlling a. One implementation example is: for d ij < 1 m, the backward amplitude of the corresponding VLO is a = 1, while for d ij > 3 m, the backward amplitude of the VLO is a = 0, with linear interpolation between both. Furthermore, the exponent b controls the selectivity between the inner and outer part of the larger distance d ij between the virtual listening position and the j-th VLO at the i-th recording point, such that misplacement of distant sound sources or unnecessary diffuse appearance is minimized. One implementation example is: for distances d ij < 3 m, such that b = 1, while for distances d ij > 6 m, then b = 2, with linear interpolation between both. This way, the recording positions are restricted from being part of the common acoustic convex hull in distant or diffuse audio scenes due to their direction. In this context, Fig. 10 shows a cardioid plot of a virtual loudspeaker object. Here, the omni-directivity is shown with a circle, where d ij > 1 m, and the other directivity patterns are generated by superposition of the omni- and cardioid directivity patterns for d ij < 3 m and d ij < 6 m, respectively, with

[0091] Furthermore, all other steps in the implementation according to 2d are identical to the previous implementation according to Fig. 2c.

[0092] Fig. 2e shows another embodiment, wherein the embodiment in Fig. 2e requires, in addition to the embodiment shown in Fig. 2d, a method step 229 of calculating an interactive VLO format comprising, for each recording point and for each VLO assigned to the recording point, a resulting signal and an angle of incidence φ ij , wherein g ij is a gain factor of the control signal x ij of the j-th VLO at the i-th recording point, τ ij is a time delay of the sound wave propagating from the j-th VLO at the i-th recording point to the virtual listening position, t denotes time, and an angle of incidence φ ij is the angle between the line connecting the i-th recording point and the j-th VLO at the i-th recording point and the line connecting the j-th VLO at the i-th recording point and the virtual listening position.

[0093] An example for performing the method step 229 of generating an interactive VLO format can also be seen in Fig. 7, which shows a block diagram for calculating an interactive VLO format in microphone signals. For each of the P recording points in the acoustic scene, i.e. recording positions, the control signal of the corresponding VLO is obtained from the microphone (array) signal assigned thereto. The control signal at the i-th recording point is obtained as follows:

[0094] x i (t) = B i s i (t),

[0095] wherein is the control signal vector (VLO signal vector) of all VLOs assigned to the i-th recording point (dimension L i x 1, i.e. a column vector of length L i , is the microphone signal vector (dimension K i x 1), B i is the L i x K i mixing matrix, L i is the number of VLOs, K i is the number of microphones, and t is time.

[0096] This can also be seen clearly in Fig. 7, which shows the corresponding microphone signals as input to the mixing matrix B i . For each VLO, the VLO format stores one resulting signal and the corresponding angle of incidence φ ij .

[0097] In Fig. 7, an overall block diagram for computing an interactive VLO format is illustrated based on the corresponding microphone signals, wherein in the present example it is assumed that a total of P recording positions, i.e. P microphone points, are given. The resulting signals are accordingly schematically plotted in Fig. 7.

[0098] In some embodiments of the application, the method further comprises, after assigning one or more VLOs to each microphone configuration, for each microphone configuration, placing the one or more VLOs within the virtual sound field at positions corresponding to the recording point of the respective microphone configuration within the acoustic scene;

[0099] wherein,

[0100] for each of the one or more microphone configurations, the one or more VLOs assigned to the respective microphone configuration are provided on a circle line having the recording point of the respective microphone configuration as center within the virtual free field, the circle line having a radius r i in accordance with a directivity order of the microphone configuration, a reverberation of the acoustic scene, and an average distance d between the recording point of the respective microphone configuration and a recording point of an adjacent microphone configuration i is determined;

[0101] the radius r of the circle line i in accordance with a directivity order of the microphone configuration, a reverberation of the acoustic scene, and an average distance d between the recording point of the respective microphone configuration and a recording point of an adjacent microphone configuration i satisfies the following equation:

[0102] wherein the r i corresponds to a radius of the corresponding circle line;

[0103] the d i avg corresponds to an average distance d between the recording point of the respective microphone configuration and a recording point of an adjacent microphone configuration i;

[0104] the dir i corresponds to a directivity order of the microphone configuration;

[0105] the rev i corresponds to a reverberation of the acoustic scene;

[0106] the i denotes an i-th microphone of the plurality of microphones;

[0107] The distanceFactor is a quantized decimal value, which is obtained from the bitstream or determined by the decoder.

[0108] In some embodiments of the present application, the distanceFactor is in the range of [0.25, 4].

[0109] In some embodiments of the present application, the directivity order of the microphone configuration is determined according to the HOA order.

[0110] In some embodiments of the present application, when the HOA order is 1, the directivity order of the microphone configuration is 1.5,

[0111] when the HOA order is 2, the directivity order of the microphone configuration is 1.2,

[0112] when the HOA order is 3 or larger, the directivity order of the microphone configuration is 1.

[0113] In some embodiments of the present application, the reverberation rev of the acoustic scene is determined according to the distance between the sound source and the decoder. i The distance between the sound source and the decoder is determined according to the distance between the sound source and the decoder.

[0114] The reverberation factor is based on the HOA source-specific scaling factor, which should be adjusted according to the room's reverberation level (e.g., defined as the distance between the sound source and the decoder is the critical distance, which can also be referred to as the receiver. At this distance, the sound pressure level of the direct sound and the reverberant sound is equal). The acoustic scene reverberation factor can vary according to the position of the HOA source.

[0115] In some embodiments of the present application, the reverberation rev of the acoustic scene is determined according to the distance between the sound source and the decoder. i In some embodiments of the present application, the reverberation rev of the acoustic scene is determined according to the distance between the sound source and the decoder.

[0116] wherein the distance between the sound source and the decoder is determined according to the distance between the sound source and the decoder. denotes the average distance between all sound sources and the decoder;

[0117] The distance between the sound source and the decoder is determined according to the distance between the sound source and the decoder. denotes the distance between the i-th sound source and the decoder.

[0118] For example, the radius (r i ) of the circle line of the given VLO set should be calculated as the distance between the position of the corresponding HOA source and the position of each of its neighboring HOA sources a statistical average (e.g., median) of the directional order factor (diridir), the acoustic scene reverberation factor (rever), and a scene range scaling factor defined by the bitstream (distance factor). For example, ri is calculated by:

[0119] In the above equation, · denotes multiplication.

[0120] The directional order factor is a HOA source-specific scaling factor that should be adjusted according to the directivity order of the corresponding microphone setup (e.g., mapped to the HOA order), as shown in the correspondence in Table 1. In scenarios where the HOA order (HOA Source Order) is not available, a default directional factor of 1.0 is used.

[0121] Table 1

[0122] The acoustic scene reverberation factor is based on a HOA source-specific scaling factor that should be adjusted according to the room’s reverberation level (e.g., defined as the distance between the sound source and the decoder, which can also be referred to as the receiver. At this distance, the sound pressure level of the direct sound and the reverberant sound are equal). The acoustic scene reverberation factor can vary depending on the location of the HOA source (e.g., in a scene containing several rooms with different acoustic characteristics, the acoustic scene reverberation factor is also different). When the acoustic characteristics of the room are uniform throughout the space, a default acoustic scene reverberation factor of 1.0 is used.

[0123] Using the VLO approach to reproduce an acoustic scene recorded by multiple HOA sources allows the listener to move with 6DoF. The VLO approach represents each HOA source with a set of virtual loudspeaker objects (VLOs) arranged in a sphere. The result is accurate when the listener is at the center of the VLO sphere. However, when the listener is far from the center of the sphere, the output audio signal is strongly affected by the sphere’s radius. For example, if the radius is very large, the output signal is hardly affected as the listener moves. On the other hand, if the radius is small, the output signal will change rapidly as the listener moves. Therefore, careful adjustment of the sphere’s radius for each HOA source is needed to ensure that the output signal provides accurate spatial audio cues (e.g., correct positioning of sound sources) to the listener. The VLO approach has provided some automatic methods to define the radius based on the neighbor distance, but some correction factors are needed to fine-tune it.

[0124] HOA source signals have an order (N >= 1). For smaller N, HOA signals provide a coarser spatial resolution, while higher N allows to draw sharper directivity patterns and thus provides a finer spatial resolution. For a coarser spatial resolution, it is less meaningful to quickly adapt the output signal as the listener moves. Therefore, it is recommended to use a larger VLO sphere radius for smaller orders, as indicated by diri in Table 1 as mentioned before.

[0125] Reverberation coefficient (rev):

[0126] HOA signals usually contain reverberation, depending on the characteristics of the room in which they were recorded. The reverberant part of a recorded signal is usually more diffuse and less dependent on the receiver position (it is more influenced by the distance between source and receiver, the directivity of the source, etc.) than the direct part of the signal. Therefore, if the HOA signals correspond to a more reverberant space, it is less meaningful to quickly adapt the output signal as the listener moves. Therefore, it is recommended to use a larger VLO sphere radius in more reverberant scenarios. To this end, a critical distance is associated with each HOA source, the critical distance being defined as the distance between the sound source and the receiver at which the sound pressure levels of the direct and reverberant sound are equal (a smaller critical distance equals a more dominant reverberation). By default, all HOA sources have the same associated critical distance (e.g. all HOA signals were recorded in the same room, where the acoustics are uniform throughout the space), and the VLO sphere radius is not modified. However, if the HOA sources are assigned different critical distances, their VLO sphere radius is corrected accordingly (e.g. when the HOA signals were recorded in two or more different rooms).

[0127] Fig. 3 shows an overall block diagram of an acoustic scene playback method according to an embodiment of the present application. Here, on the left side, the recorded data is provided, wherein the recorded data comprises microphone signals and microphone metadata. In this context, the present application is not limited to any recording hardware, e.g. a specific microphone array. The only requirement is that the microphones are distributed within the acoustic scene to be captured and that the positions, characteristics (omnidirectional cardioid curve etc.) and directions are known. However, the best results are obtained if a distributed microphone array is used. These arrays can be (first or higher order) spherical microphone arrays or any compact classic stereo or surround recording configuration (e.g. XY, ORFT, MS, OCT surround, Fukada tree). Furthermore, as seen in Fig. 3, the microphone metadata is used to calculate static VLO parameters. Furthermore, the microphone signals and the static VLO parameters can be used to calculate control signals, i.e. VLO signals, which are used to control each VLO in a virtual free field, wherein each control signal is used to control one corresponding VLO within the virtual free field. Furthermore, as seen in Fig. 3, dynamic VLO parameters can be calculated based on a selected virtual listening position and based on the static VLO parameters. Furthermore, the dynamic VLO parameters and the control signals are used as input to be encoded, preferably to be high order ambisonic encoded. Then, the resulting encoded data stream is decoded as a function of a certain playback configuration. An example of a certain playback configuration can be a configuration corresponding to a loudspeaker arrangement of a room or the playback configuration can reflect the use of headphones. Depending on this playback configuration, the corresponding decoding is performed, as also seen in Fig. 3. Then, the resulting decoded data stream is input to a rendering device, which can be a loudspeaker or headphones, as also seen in Fig. 8.

[0128] The block diagram of Fig. 3 can be performed by a playback device. In this context, it should be mentioned that, in principle, the method steps shown in Fig. 3: providing recorded data, calculating static VLO parameters, calculating control signals, i.e. VLO signals, can be performed at a place outside the playback device, e.g. at a location remote from the playback device, but can also be performed within the playback device. Since the virtual listening position has to be provided to the playback device, the only thing that has to be performed within the playback device is the dynamic VLO parameter calculation together with the encoding and decoding steps. However, all other method steps shown in Fig. 3 do not need to be performed within the playback device, but can also be performed outside the playback device. Thus, for example, the recorded data can be provided to the playback device by any conceivable means, i.e. for example by using live streaming or similar means of recording data via an internet connection. Another option is that the recorded data is generated within the playback device itself by extracting the recorded data from a recording medium provided within the playback device. Furthermore, the block diagram of Fig. 3 only shows one example and the method steps of Fig. 3 do not have to be performed in the way described in Fig. 3.

[0129] Figure 4 shows an example of microphone and sound source distribution in an acoustic scene, where an acoustic scene has been recorded with three distributed compact microphone configurations. Configuration 1 is a 2D B-format microphone, configuration 2 is a standard surround configuration, and configuration 3 is a unidirectional microphone.

[0130] Figure 5 shows each of the three microphone configurations 1, 2 and 3 (see upper row of Figure 5) and the corresponding loudspeaker configuration (see lower row of Figure 5) that can be used to reproduce the acoustic scene (soundfield) captured by the respective microphone configuration. That is, each of these loudspeaker configurations contains one or more virtual loudspeaker objects that will accurately reproduce the spatial soundfield at the recording point, i.e. at the center position of the corresponding microphone configuration associated with the respective loudspeaker configuration. Thus, the present invention aims at virtually establishing a reproduction system for each microphone configuration in a virtual free field comprising the loudspeaker configuration. The VLOs assigned to the corresponding microphone configuration are placed at positions within the corresponding virtual free field that correspond to the positions of the corresponding microphone configuration.

[0131] Figure 6 shows possible configurations of VLOs within a virtual free field. If the virtual listening position substantially coincides with one of the center positions of the microphone configurations, i.e. one of the recording points, and assuming that the control signals of all VLOs corresponding to the other recording points are sufficiently attenuated, it is clear that the spatial image delivered to the listener when the VLOs are encoded and rendered accordingly is accurate. In this context, it should be noted that for these virtual listening positions only the angular layout of the VLOs is important, while the radius of the reproduction system (indicated as the gray circle in Figure 6) is not important. Figure 6 shows the arrangement of VLOs corresponding to the microphone configurations 1, 2, 3 as shown in Figure 4. However, if the virtual listening position does not coincide with a recording point, the spatial image of the acoustic scene is likely to be corrupted and the listener will likely mislocate the sound sources. In addition, mixing time-shifted related signals can produce phase artifacts. Therefore, in embodiments of the present invention these difficulties are overcome by automatic parameterization of the VLOs (e.g. VLO positions, gains, directivity, etc.) to minimize mislocation and deliver a reasonable spatial image to the listener at an arbitrary listening position while avoiding phase artifacts.

[0132] If the virtual listening position is the center position of a recording point (recording position), the signal connection of the virtual loudspeaker objects is free of disturbing artifacts: typical acoustic delays are in the range of 10 ms to 50 ms. Mixing of audio technology independent signals and distance dependent attenuation here will not produce any disturbing tonal artifacts. In addition, the precedence effect supports proper localization at all recording positions. Moreover, if there are only a few virtual loudspeaker objects per playback point in the virtual free field, a multitude of it playback points supports localization and room impression.

[0133] However, for the case that the virtual listening position deviates from the center position of any recording point, potential localization confusion can be avoided by adjusting the position, gain and delay of the corresponding virtual loudspeaker object determined from the virtual listening position. Furthermore, by choosing appropriate distances between the virtual loudspeakers, interference is reduced, which controls the phase and delay properties to ensure high sound quality. The arrangement and position of the VLOs assigned to the corresponding recording point can be automatically generated from the metadata of the microphone configuration. This results in an arrangement of VLOs whose superimposed playback is controllable in order to achieve the following properties for an arbitrary virtual listening position: Minimization of perceived interference (phase) by optimally considering the phenomenon of auditory precedence effect. In particular, the localization advantage can be exploited by choosing appropriate distances between the virtual loudspeaker objects. In doing so, the sound propagation delay is adjusted in order to achieve superior sound quality. Furthermore, the angular distances between the virtual loudspeaker objects are chosen in order to produce maximum achievable stability of the phantom sound source, which will depend on the order of the gradient microphone directivity associated with the virtual loudspeaker objects, the critical distance of the room reverberation and the degree of coverage in the acoustic scene recorded by the microphones.

[0134] Fig. 8 shows the Nth order HOA encoding / decoding of the VLO format. Since each VLO is defined by its corresponding resulting signal and angle of incidence, any rendering system capable of rendering sound objects can be used (e.g. wave field synthesis, binaural coding). However, in the present embodiment, the higher order ambisonics (HOA) format can be used to achieve maximum flexibility with respect to the rendering system. First, the interactive VLO format is encoded into the HOA signals, which can be rendered for a certain loudspeaker arrangement or binaural headphone reproduction. The block diagram of the HOA encoding and decoding is shown in Fig. 8, where the corresponding resulting signals and angles of incidence are input to the corresponding encoder. After performing the encoding, the encoded data streams are summed and input to the corresponding ambisonics decoder provided inside the loudspeakers or headphones via the ambisonics bus. Optionally, a head tracker can be provided to fully perform the ambisonics rotation as seen in Fig. 8.

[0135] In Fig. 8, the VLOs generated virtual sound field within the virtual free field are encoded into higher-order ambisonics (HOA) using the VLO parameters (static and dynamic VLO parameters). That is, the signals are input on the ambisonics bus of the Nth order ambisonics signals:

[0136] where y N is the circular or spherical harmonic evaluated at the VLO angle of incidence φ ij at the corresponding virtual listener position. Furthermore, Li P is the number of VLOs at the i-th microphone recording point, P denotes the total number of microphone configurations within the acoustic scene. The recommended order of encoding is greater than 3, typically 5th order yields stable results.

[0137] Furthermore, with respect to decoding, the decoding of the scene-based material uses a headphone- or loudspeaker-based HOA decoding method. Generally speaking, the most flexible, and thus most advantageous, decoding method for loudspeakers or for a set of head-related impulse responses (HRIRs) in case of headphone playback is called ALLRAD. Other methods can be used, e.g. decoding by sampling, energy preserving or regularized pattern matching. All these methods yield similar performance on well-directed loudspeaker or HRIR layouts. The decoder typically uses a frequency-independent matrix to obtain the signals for the known configuration of loudspeakers or for convolution with a given set of HRIRs:

[0138] y(t) = D x N (t)

[0139] In headphone-based playback, the directional signals y(t) are convolved with the left and right HRIRs of the corresponding direction and then summed for each ear:

[0140] To represent a static virtual audio scene, the head rotation β measured by head tracking has to be compensated in headphone-based playback. To keep the set of HRIRs static, it is preferable to modify the ambisonic signal by a rotation matrix before decoding to the set of HRIRs.

[0141] χ' N (t) = R(-β) χ N (t)

[0142] A playback device for performing the method of playback of an acoustic scene can comprise a processor for performing any of the method steps and a storage medium for storing the microphone signals and / or metadata of one or more microphone configurations, the static and / or dynamic VLO parameters, and / or any information required for performing the method in the embodiments of the present application. The storage medium can further store a computer program comprising program code for performing the method in the embodiments, the processor being configured to read the program code and to perform the method steps in the embodiments of the present application according to the program code. In yet another embodiment, the playback device can further comprise units for performing the method steps in the disclosed embodiments, wherein for each method step a corresponding unit can be provided which is dedicated to performing the assigned method step. Optionally, a certain unit within the playback device can be used for performing more than one method step disclosed in the embodiments of the present application.

[0143] The application has been described in connection with the various embodiments disclosed herein. However, those skilled in the art will understand and appreciate that the disclosed embodiments are merely examples and that variations of the disclosed embodiments are possible without departing from the scope of the application. In the claims, the term "including" does not exclude other elements or steps, the term "a" or "an" does not exclude a plurality, and the term "one" does not exclude a plurality unless otherwise specified. A single processor or other unit can fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state storage medium supplied together with or as part of other hardware, but also distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

Claims

1. An acoustic scene playback method, characterized by, The method comprises: providing recording data comprising microphone signals of one or more microphone configurations located within an acoustic scene and microphone metadata of the one or more microphone configurations, wherein each of the one or more microphone configurations comprises one or more microphones and has a recording point as a central position of the respective microphone configuration; specifying a virtual listening position, wherein the virtual listening position is a position within the acoustic scene; assigning one or more virtual loudspeaker objects VLOs to each of the one or more microphone configurations, wherein each VLO is an abstract sound output object within a virtual free field; wherein a virtual free field is a virtual sound field consisting of direct sound of un-reverberant sound; generating an encoded data stream based on the recording data, the virtual listening position and VLO parameters assigned to the one or more microphone configurations; decoding the encoded data stream based on a playback configuration, thereby generating a decoded data stream; and inputting the decoded data stream to a rendering device, thereby driving the rendering device to reproduce sound in the acoustic scene at the virtual listening position; wherein further comprising, after assigning one or more VLOs to each microphone configuration, for each microphone configuration, positioning the one or more VLOs within the virtual sound field at positions corresponding to the recording point of the respective microphone configuration within the acoustic scene; wherein, for each of the one or more microphone configurations, providing the one or more VLOs assigned to the respective microphone configuration on a circle with the recording point of the respective microphone configuration as center of the circle within the virtual free field, the circle having a radius r i in dependence on an order of directivity of the microphone configuration, a reverberation of the acoustic scene, and an average distance d between the recording point of the respective microphone configuration and a recording point of an adjacent microphone configuration i determined; wherein the radius r of the circle line i a directionality order of the microphone configuration, a reverberation of the acoustic scene, and an average distance d between the recording point of the respective microphone configuration and a recording point of an adjacent microphone configuration i satisfies the following equation: wherein the r i corresponding to the circle line; The d i avg The average distance d between the recording point of the corresponding microphone configuration and the recording point of an adjacent microphone configuration i ; the dir i a directionality order corresponding to the microphone configuration; the rev i reverberation corresponding to the acoustic scene; the i denotes an i-th microphone of the plurality of microphones; the distanceFactor is a quantized decimal value, the distanceFactor being obtained from a bitstream or the distanceFactor being determined by a decoder; reverberation rev of the acoustic scene i determined as a function of the distance between the sound source and the decoder.

2. The method of claim 1, wherein, the VLO parameters comprise one or more static VLO parameters of the one or more VLOs being independent of the virtual listening position and describing properties, the static VLO parameters being fixed for the acoustic scene playback.

3. The method of claim 2, wherein, Further comprising: calculating the one or more static VLO parameters based on the microphone metadata and / or a critical distance, wherein the critical distance is a distance at which a sound pressure level of a direct sound is equal to a sound pressure level of a reverberant sound for a directional sound source, prior to generating the encoded data stream, or receiving the one or more static VLO parameters from a transmission device prior to generating the encoded data stream.

4. The method according to any one of claims 1 to 3, wherein: the one or more static VLO parameters comprise, for each of the one or more microphone configurations: a plurality of VLOs, and / or a distance of each VLO to the recording point of the respective microphone configuration, and / or an angular layout of the one or more VLOs already assigned to the respective microphone configuration with respect to a direction of the one or more microphones of the respective microphone configuration, and / or a mixing matrix defining the microphone signals of the respective microphone configuration.

5. The method of claim 1, wherein, the VLO parameters comprise one or more dynamic VLO parameters determined depending on the virtual listening position, the method comprising, prior to generating the encoded data stream, calculating the one or more dynamic VLO parameters based on the virtual listening position, or receiving the one or more dynamic VLO parameters from a transmission device.

6. The method of claim 5, wherein, The one or more dynamic VLO parameters comprise for each of the one or more microphone configurations: one or more VLO gains, wherein each of the one or more VLO gains is a gain of a control signal of a corresponding VLO, and / or one or more VLO delays, wherein each VLO delay is a time delay of a sound wave propagating from the corresponding VLO to the virtual listening position, and / or one or more VLO incidence angles, wherein each VLO incidence angle is an angle between a line connecting the recording point and the corresponding VLO and a line connecting the corresponding VLO and the virtual listening position, and / or one or more parameters indicating a radiation directivity of the corresponding VLO.

7. The method of claim 1, wherein, Further comprising: Before generating the encoded data stream, an interactive VLO format is computed which comprises for each recording point and each VLO assigned to the recording point a resulting signal and an angle of incidence φ ij wherein wherein g ij is a gain factor of a control signal x ij of the jth VLO at the ith recording point, τ ij is a time delay of a sound wave propagating from the jth VLO of the ith recording point to the virtual listening position, t denotes time, the angle of incidence φ ij is an angle between a line connecting the ith recording point and the jth VLO at the ith recording point and a line connecting the jth VLO of the ith recording point and the virtual listening position.

8. The method of claim 7, wherein, The gain factor is determined depending on the angle of incidence φ ij and the distance d between the jth VLO at the ith recording point and the virtual listening position ij is determined.

9. The method of claim 8, wherein, for generating the encoded data stream, inputting each resulting signal and incidence angle into an encoder, in particular a binaural reverberation sound encoder.

10. The method according to claim 1, wherein: the number of VLOs on the circle line and / or the angular position of each VLO on the circle line and / or the directivity of the sound radiation of each VLO on the circle line is dependent on the microphone directivity order of the respective microphone configuration and / or the recording principle of the respective microphone configuration and / or the radius R of the recording point of the i-th microphone configuration i and / or the distance d between the j-th VLO of the i-th microphone configuration and the virtual listening position ij is determined.

11. The method according to any one of claims 1 to 10, characterized in that, the distanceFactor is in the range [0.25, 4].

12. The method according to any one of claims 1 to 11, characterized in that, when the HOA order is 1, the directivity order of the microphone configuration is 1.5, when the HOA order is 2, the directivity order of the microphone configuration is 1.2, when the HOA order is 3 or larger, the directivity order of the microphone configuration is 1.

0.

13. The method according to any one of claims 1 to 12, characterized in that, reverberation rev of the acoustic scene i determined based on a distance between the sound source and the decoder, including: reverberation rev of the acoustic scene i is determined by: Among them, the denotes an average distance between all sound sources and the decoder; The denotes a distance between the i-th sound source and the decoder.

14. An acoustic scene playback method, characterized by, The method comprises all features of any one of claims 1 to 13, wherein: for providing the recording data, the recording data are received from an external source, in particular by applying streaming media.

15. An acoustic scene playback method, characterized by, The method comprises all features of any one of claims 1 to 13, wherein: for providing the recording data, the recording data are extracted from a recording medium, in particular from a CD-ROM.

16. A playback device for carrying out the method according to any one of claims 1 to 15.

17. A data carrier having stored thereon a computer program, characterized in that a computer program for instructing a playback device to carry out the method according to any one of claims 1 to 15, when the computer program is run on a computer.