Device for generating an immersive stereo signal for playback via headphones and data format for transmitting audio data
By integrating metadata-driven mixing of multiple audio channels based on head angle and distance, the method addresses issues in immersive binaural audio, enhancing sound localization and proximity simulation for headphones.
Patent Information
- Application Number
- DE102024100053
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2044-01-02
AI Technical Summary
Existing methods for generating immersive binaural audio using headphones fail to provide an optimal binaural sound experience due to issues such as level attenuation in the low-tone range, inaccurate localization of sound sources, and inability to simulate sound sources outside the head, leading to a diminished sense of proximity and distance.
A method that combines M audio channels from a microphone array with additional mono- or stereo audio channels, using metadata to adjust mixing ratios based on the listener's head angle and distance from the sound source, to generate a stereo signal that accurately localizes sound sources relative to the head position.
Enhances the binaural hearing experience by providing accurate distance and direction localization, improving the naturalness of sound perception and reducing the perceived distance of sound sources, while maintaining a realistic sound image.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field of the invention
[0001] The invention relates to the generation of stereo signals, in particular, immersive and / or binaural stereo signals for playback via headphones. In some embodiments of the invention, the stereo signal is generated depending on a detected head angle of the listener wearing the headphones. Furthermore, the invention relates to an audio container format for transmitting audio information for generating the stereo signal, as well as to computer-readable media, devices, and systems for generating and / or reproducing the stereo signal. Technical background
[0002] The stereo audio format with two channels (often referred to as left and right channels) has long been standard in the audio industry. Multi-channel audio formats such as 5.1 and 7.1 have since become established in certain applications, such as cinemas or home theaters. Immersive multi-channel audio formats such as Dolby Atmos, Auro 3D, MPEG-H, DTS:X, NHK 22.2, etc., expand the horizontal dimension of the audio format to include the reproduction of sound sources from above. They require greater technical complexity and additional speakers or soundbars. The term "immersive audio" refers to a variety of formats in which sound emanates from all directions in the room, thus immersing the listener deeper into the "action."
[0003] The creation of an immersive listening experience using headphones dates back to the 1980s with the introduction of artificial head stereophony, which, however, failed to gain market acceptance. Due to the ubiquity of headphones and in-ears, there are renewed efforts to make immersive music experience possible using headphones. Common solutions typically use head-related transfer functions (HRTFs). HRTFs describe the complex filtering effect of the head, outer ear (pinna), and torso on sound arriving at the inner ear from different directions. In these processes, the loudspeaker signals are convolved with the corresponding directional HRTFs and can be localized more or less accurately by the listener from that direction. In addition to generic HRTFs, individual HRTFs can also be generated, which can usually better match the listener's own listening experience.
[0004] Some methods allow headphones to track rotation and tilt around the listener's head axis. With Apple Spatial Audio, the Dolby Atmos stream is converted to 7.1.2 format and the HRTFs are then dynamically adjusted to the current speaker positions.
[0005] All HRTF-based methods for creating an immersive listening experience using headphones involve a transformation from a playback for loudspeakers to a playback for headphones: Instead of a performance space, such as a concert hall, one hears the renderer's virtual listening room through the headphones, and instead of the instruments, the loudspeaker signals. Instead of sound source positions, the loudspeaker positions and their phantom sound sources are transmitted. The convolutions with the generic or individual HRTFs alter the original sound, which leads to a attenuation of the level, especially in the low-frequency range; the sound is often perceived as thin. The distance of the sound sources from the listener is determined during production but cannot be less than the specified loudspeaker plane, and satisfactory localization of sources in the frontal plane is impossible for most people.
[0006] Another approach to generating a motion-dynamic and binaural audio signal that records binaurally at the recording location, i.e. actually head-related, is described in US Patent US 7,333,622 B2.
[0007] US 7,333,622 B2 describes a method for capturing and playing back live or recorded 3D audio exclusively for headphone playback. This method is called “motion-tracked binaural” (MTB). It uses a microphone array with multiple microphones, head-tracking headphones, and special signal processing to combine the signals recorded by the microphones into a binaural audio signal that follows the rotation of the listener’s head and creates a counter-rotating sound field. The microphone array is shaped like a sphere roughly the size of a head, in which 8, 16, or 32 individual microphones (microphone capsules) are flush-mounted and evenly arranged in a plane. Alexander Lindau and Sebastian Roos, “Perceptual evaluation of discretization and interpolation for motion-tracked binaural (MTB) recordings,” Report of the 26thTonmeistertagung, pages 660-701, were able to show that when using a High Frequency Spectrum Interpolation (HF-SP), doubling or quadrupling the number of 8 microphones does not result in any measurable improvement in localization and timbre.
[0008] US 2011 / 0 091 056 A1 describes a hearing aid comprising a microphone and an external input jack, a hearing aid processor to which audio signals from the microphone and the external input jack are fed, and a receiver to which the audio signals processed by the hearing aid processor are output. The hearing aid processor has a mixer that mixes audio signals from the microphone with audio signals from the external input jack and outputs these audio signals to the receiver, a mixing unit for determining the mixing ratio between the audio signals from the microphone and the audio signals from the external input jack in the mixer, and a face recognition detector connected to this mixing unit.
[0009] US 2019 / 0 313 200 A1 describes systems and methods that can be configured to identify, manipulate, and play back various audio source components from encoded 3D audio mixes, such as those that may include content mixed for azimuth, elevation, and / or depth relative to a listener. The systems and methods can be designed to decouple depth encoding and decoding to adapt spatial playback to a particular playback environment or platform. In one example, the systems and methods enhance playback in applications with listener tracking, including tracking across six degrees of freedom (e.g., yaw, pitch, roll orientation, and x, y, z position). Summary of the invention
[0010] Even though a fairly high degree of realism is achieved, the solution provided in US Pat. No. 7,333,622 B2 still does not provide the listener with an optimal binaural sound result. In the solution provided in US Pat. No. 7,333,622 B2 (as well as in some embodiments of the invention), two opposing microphones form a binaural audio pair. The spherical separator of the microphone array, on which the eight microphones are arranged in a horizontal plane, generates characteristic ILD (Interaural Level Differences) and ITD (Interaural Time Differences) similar to those found in natural hearing.
[0011] This microphone array is positioned at a suitable distance in front of a sound body. Its function is that of out-of-head localization (AKL) and the location of sound sources during headphone playback. In sound engineering (and in the embodiments), sound sources are referred to as musical sound bodies, e.g., a piano, string quartet, orchestra, or band. As with an artificial head, sound sources are located around the head. Two opposing microphones form an audio pair, which, as with an artificial head, represents the ears. With eight microphones, the Fig. 11b: 0°, +45°, +90°, +135°, 180°, -135°, -90°, and -45°; with 0° as the reference axis, which can correspond, for example, to the central, horizontal line of sight to the center of the stage or a head position facing the sound box. The corresponding microphones of a stereo pair are offset by -90° (left ear) and +90° (right ear) to the head angle, as shown in Fig. 11b is illustrated by way of example for a head angle of 0° and -45°. The "head angle" is understood herein to mean the angle of head rotation in a plane relative to a reference axis. This reference axis of head rotation can coincide with the reference axis relative to which M angular positions are defined, or it can have a (e.g., configurable) offset therefrom.
[0012] When a pair of audio from the microphone array is routed to a pair of headphones—that is, one microphone per channel of the headphones—a sound source produces a realistic out-of-head localization (OHE) for the listener, whose horizontal angle corresponds to natural perception. Out-of-head localization means that the sound source is perceived as being outside the listener's head. In contrast, in-head localization (IHE) localizes sources inside the listener's head. This is often experienced with normal stereo recordings or with monophonic signals heard through headphones.
[0013] Due to the lack of ear cups on the microphone array in the solution described in US patent US 7,333,622 B2, diffuse sound is amplified at the expense of the direct sound component, reducing frontal presence. Sound sources appear farther away than they do in natural hearing. To compensate for this excessive distance impression, the microphone array could be moved closer to the sound source to receive more direct sound. However, this would make the sound source appear wider than the perceived proximity would suggest.
[0014] One object of the invention is to provide solutions that improve the binaural hearing experience to create the most natural hearing impression possible with precise distance localization. Another object of the invention is to provide solutions that improve the binaural hearing experience to create the most natural hearing impression possible with precise directional and distance localization.
[0015] One aspect of the invention is therefore to improve the distance localization of the sound source (e.g. an orchestra, a band, etc. on a stage). In addition to M audio channels (with M ≥ 4) assigned to M angular perspectives, one or more further mono or stereo audio channels (e.g. from spot microphones and / or main microphones used in the conventional recording of the sound source) are received. From the M audio channels (also layer A), a stereo audio signal is generated depending on a detected / determined head angle of the listener, to which a further stereo audio signal is mixed, which is (in each case) also generated depending on the detected / determined head angle and optionally depending on a selectable (virtual) distance from the sound source from the further one or more mono or stereo audio channels.The respective mixing ratio with which a further stereo audio signal generated from a further mono or stereo audio channel is mixed with the stereo signal generated from the M audio channels can depend on the detected / determined head angle of the listener and optionally on a selectable (virtual) distance from the sound source.
[0016] A further aspect of the invention relates to a transmission format or a data structure, e.g., an audio container, by means of which the information (e.g., metadata) and audio signals of the individual audio channels required for binaural headphone playback can be transmitted. To implement the mixing of the additional mono or stereo audio channel(s), metadata is transmitted using a data structure / audio container, which contains a metadata set for each of the at least one additional mono or stereo audio channel. Each metadata set identifies the mixing ratios of the respective mono or stereo audio channel of the second layer (also Layer B) for predefined N angular perspectives.
[0017] The N angular perspectives of the additional mono or stereo audio channel(s) (Layer B) can arise from practical considerations. For example, when panoramas are taken of the mono and / or stereo channels (Layer B), the number N of angular perspectives should not be less than 8 (i.e., N ≥ 8) in order to achieve a continuous, flowing panoramic movement counter to head rotation. For a frontal sound event (i.e., the sound source is largely located at 0° as the reference direction), +90° / -90° should be included as the angular position, as this is where panoramas are most critical. The number N can therefore be chosen, for example, as N = (n + 1) 4, with n ∈ [1, 2, 3, ...]. The N angular perspectives of Layer B can correspond to the M angular perspectives of Layer A (i.e., N = M); however, this is not mandatory; i.e., N ≠ M can also apply. In some embodiments, N = M = 8 may apply.
[0018] If the mixing of the additional mono or stereo audio channel(s) is to be distance-dependent, several such metadata sets can be provided for each additional mono or stereo audio channel, which define the mixing ratio(s) of the respective mono or stereo audio channel of the second layer for each of the specified N angular perspectives for different (in particular three) specified virtual distances from the sound source.
[0019] A further aspect of the invention relates to the implementation of the above aspects and their embodiments described herein in software and / or hardware, and a playback system that generates a stereo audio signal that is played back via headphones and enables a binaural listening experience.
[0020] Some embodiments provide a method for generating an output stereo signal (in particular a binaural stereo signal) with the L and R channels for playback with headphones. The method can be carried out by a device for generating an output stereo signal with the L and R channels, which device comprises a computing unit (e.g., a processor). The method includes receiving M audio channels, with M ≥ 4, which are assigned to M predetermined angular perspectives lying in a plane, and receiving at least one further mono or stereo audio channel. Furthermore, metadata is received which contains a metadata set for each of the at least one further mono or stereo audio channel.The metadata serves to acoustically position the auditory event of the respective additional mono or stereo audio channel relative to an auditory event generated from the M audio channels for the individual M angular positions (similar to a "panoramic view"). For this purpose, for example, each metadata set can define the mixing ratio of the respective mono or stereo audio channel for each of the N angular perspectives.
[0021] The method further comprises determining a head angle ϕ of a listener wearing the headphones. The head angle ϕ can be determined in a plane relative to one of the N angular perspectives and / or the M angular perspectives that define a reference direction. The determination of the head angle ϕ can be based on sensor signals from at least one sensor. In one embodiment, the one or more sensors form a head-tracking sensor system, the sensor signal(s) of which enables the determination of the head angle ϕ.
[0022] The method generates a first binaural stereo audio signal (hereinafter also binaural audio signal) with one channel L and one channel R. The channel L of the first stereo audio signal is generated, for example, by a linear crossfading of two audio channels of the M audio channels (in layer A), which are assigned angular perspectives that are closest to the head angle ϕ - 90°. Accordingly, the channel R of the first stereo audio signal is generated, for example, by a linear crossfading of two audio channels of the M audio channels, which are assigned the M angular perspectives that are closest to the head angle ϕ + 90°. It is assumed that the ears of the head are at an angle of -90° or +90° relative to the head angle ϕ. If the head angle ϕ ± 90° corresponds to one of the M angle positions, the audio channels of the two angle positions (ϕ - 90° and ϕ + 90°) can be used for the L and R channels of the first stereo audio signal (ieIn this case, crossfading of audio channels is not necessary).
[0023] In order to mix the one or more mono or stereo audio channels to the first stereo audio signal, the mixing ratios are first determined for each additional mono or stereo audio channel (e.g. for a stereo audio channel: mixing ratios i 1Ln , i 1Lnn , i 1Rn , i 1Rnn , i 2Ln , i 2Lnn , i 2Rn , i 2Rnn for angle perspectives n and nn from the N angle perspectives) specified in those two metadata sets for the respective mono or stereo audio channel, which are assigned to the two angle perspectives n and nn of the N angle perspectives that are closest to the absolute head angle ϕ, are (linearly) blended to create blended mixing ratios (e.g. for a stereo audio channel: blended mixing ratios: i 1Lϕ =f(i 1Ln , i 1Lnn ), i 1Rϕ =F(i 1Rn, i 1Rnn ), i 2Lϕ =f(i 2Ln , i 2Lnn ), i 2Rϕ =f(i 2Rn , i 2Rnn ), where f(x, y) represents a (linear) function of the mixing ratios x and y). The (linear) crossfading can, for example, combine the respective two mixing ratios according to the respective (angular) distance of the head angle ϕ from the two nearest angular positions n and nn (nearest neighbor n and next-nearest neighbor nn). For each additional mono or stereo audio channel, a second stereo audio signal with an L channel and an R channel is further generated. The second stereo audio signal can be generated based on the respective additional mono or stereo audio channel using the crossfaded mixing ratios. Subsequently, the output stereo signal can be generated by mixing (summing) the first stereo audio signal with the at least one generated second stereo audio signal.
[0024] In a further embodiment, the at least one further mono or stereo audio channel includes or is a stereo audio channel. The metadata set for the stereo audio channel identifies four parameter values for each of the N angular perspectives.
[0025] The four parameter values indicate for each of the N angle perspectives: - a first mixing factor (i 1L ) for the first channel of the stereo audio channel and a second mixing factor (i 2L ) for the second channel of the stereo audio channel for the channel L of the second stereo audio signal, and - a third mixing factor (i 1R ) for the first channel of the stereo audio channel and a fourth mixing factor (i 2R ) for the second channel of the stereo audio channel for the channel R of the second stereo audio signal;
[0026] The above-described (linear) crossfading of the mixing ratios depending on the head angle ϕ for such a stereo audio channel includes a respective linear crossfading of the first two mixing factors i 1Lϕ =f(i 1Ln , i 1Lnn ), the two second mixing factors i 2Lϕ =f(i 2Ln , i 2Lnn ), the two third mixing factors i 1Rϕ =f(i 1Rn , i 1Rnn ), and the two fourth mixing factors i 2Rϕ =f(i 2Rn , i 2Rnn ) specified in the two metadata sets to create a blended first blending factor (i 1LΦ ), a blended second mixing factor (i 2Lϕ ), a blended third mixing factor (i 1Rϕ ) and a blended fourth mixing factor (i 2Rϕ). Accordingly, this may include generating the second stereo audio signal based on the one stereo audio channel: generating the channel L of the second stereo signal by multiplying the first channel of the stereo audio channel by the blended first mixing factor (i 1Lϕ ) and the second channel of the stereo audio channel with the blended second mixing factor (i 2Lϕ ) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel; and generating the channel R of the second stereo signal by multiplying the first channel of the stereo audio channel by the blended third mixing factor (i 1Rϕ ) and the second channel of the stereo audio channel with the blended fourth mixing factor (i 2Rϕ ) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel.
[0027] In a further embodiment, the method further comprises, for each of the N angular perspectives, receiving a parametric equalizer setting for the L and R channels of the second stereo signal. For example, the parametric equalizer setting for a respective one of the two L and R channels can approximate or simulate a head-related or outer ear transfer function (HRTF). The equalizer setting can be a multi-band equalizer setting, in particular an equalizer setting for at least three frequency bands. Based on the equalizer settings, the L channel of the second stereo signal is filtered. During filtering, the two equalizer settings assigned to the two angular perspectives n and nn that are closest to the absolute head angle ϕ can be used.These two equalizer settings are faded in or out linearly according to the distance of the head angle ϕ from the respective angular perspectives n and nn. Channel R of the second stereo signal is filtered accordingly: Here, too, the two equalizer settings assigned to the two angular perspectives n and nn that are closest to the absolute head angle ϕ can be used for filtering. These two equalizer settings are faded in or out linearly according to the distance of the head angle ϕ from the respective angular perspectives n and nn.
[0028] In a further embodiment, for each of at least three predetermined virtual distances from the sound source, an amplification factor (g near , g medium , g far ) is received. The gain factor (g near , g medium , g far) is provided for amplifying the second stereo signal. One or more crossfade curves can also be used to calculate a gain factor (g dist ) for a virtual distance between two of the specified virtual distances from a sound source. If no crossfade curve(s) are included in the received data, a predefined crossfade curve (e.g., a linear crossfade) can be used.
[0029] The second stereo signal can be adjusted accordingly depending on a selected virtual distance from the sound source based on one of the gain factors (g near , g medium , g far ) (if the selected virtual distance corresponds to one of the predefined distances) or a gain factor calculated using one of the blending curves (g dist) (if the selected virtual distance does not correspond to any of the predefined distances).
[0030] In a further embodiment, a respective parametric equalizer setting (EQ near , EQ medium , EQ far ) for the second stereo signal. This can be used to adjust the sound image of the second stereo signal according to a selected virtual distance based on the received parametric equalizer settings (EQ near , EQ medium , EQ far ) for the second stereo signal. For example, using the equalizer settings, the sound of the second stereo signal can be adjusted to simulate the high-frequency attenuation of air at great distances.
[0031] In a further embodiment, a head-locked stereo audio channel is also received. In addition, for each of at least three predetermined virtual distances from the sound source, a gain factor (g HL_near , g HL_medium , g HL_far ) is received. The gain factor (g HL_near , g HL_medium , g HL_far ) can be used to amplify the head-locked stereo audio channel. Furthermore, one or more crossfade curves can optionally be used to calculate a gain factor (g HL_dist ) for a virtual distance between two of the specified virtual distances from a sound source.
[0032] The head-locked stereo audio channel can be adjusted depending on a selected virtual distance from the sound source based on one of the gain factors (g HL_near , g HL_medium , g HL_far) (if the selected virtual distance corresponds to one of the predefined distances) or a gain factor calculated using one of the blending curves (g HL_dist ) (if the selected virtual distance does not correspond to any of the specified distances).
[0033] In a further embodiment, a respective parametric equalizer setting (EQ HL_near , EQ HL_medium , EQ HL_far ) for the head-locked stereo audio channel. These can be used to change the sound image of the head-locked stereo audio channel according to a selected virtual distance. For example, using the equalizer settings, the sound image of the head-locked stereo audio channel can be adjusted to achieve a stronger spatial effect by boosting the Blauert band by 1 kHz.
[0034] Optionally, according to a further embodiment of the invention, an interpolation of the high-frequency spectrum of the first stereo audio signal can be implemented. This can serve to compensate for comb filter effects that occur due to the phase difference in the M audio channels as a result of the spatial offset of the microphones when recording the sound source. The interpolation of the high-frequency spectrum of the first stereo audio signal can include high-pass filtering of the signals, decomposition of the high-frequency signal components into amplitudes and phases, and amplitude addition and introduction of the phase of a microphone. This signal is then added again to the low-frequency signal component.
[0035] In a further embodiment, for each of at least three predetermined virtual distances from the sound source, an amplification factor (g array_near , g array_medium , g array_far ) is received. The gain factor (g array_near, g array_medium , g array_far ) can be used to amplify the first stereo audio signal. Furthermore, one or more crossfade curves can optionally be received for calculating a gain factor for a virtual distance between two of the specified virtual distances from a sound source.
[0036] Furthermore, the first stereo audio signal can be amplified as a function of a selected virtual distance from the sound source based on one of the gain factors (g array_near , g array_medium , g array_far ) (if the selected virtual distance corresponds to one of the predefined distances) or a gain factor calculated using one of the blending curves (if the selected virtual distance does not correspond to any of the predefined distances).
[0037] In a further embodiment, a respective parametric equalizer setting (EQ array_near , EQ array_med i um , EQ array_far ) for the first stereo audio signal. The sound image of the first stereo audio signal can be adjusted according to a selected virtual distance based on the received parametric equalizer settings (EQ array_near, EQ array_medium , EQ array_far ) for the first stereo audio signal.
[0038] Further embodiments relate to one or more computer-readable media storing instructions that, when executed by a computing unit of a device, cause the device to perform the method described herein for generating an output stereo signal.
[0039] Another embodiment relates to a device adapted to carry out the method described herein for generating an output stereo signal. For this purpose, the device can be equipped with a computing unit that carries out the method for generating an output stereo signal according to any of the embodiments of the method described herein. The device can be, for example, a user terminal, in particular a smartphone or tablet computer. The device can be adapted accordingly to feed the generated output stereo signal with the L and R channels to a headphone for playback. Alternatively, the device can also be integrated into a headphone.
[0040] A further embodiment relates to a device adapted to receive and process one or more audio containers according to any of the embodiments described below to generate the output stereo signal. The audio containers can be received as part of a continuous data stream (streaming). In a further embodiment, the computing unit of the device is adapted to execute the method for generating an output stereo signal according to any of the embodiments described herein based on the data in multiple audio containers according to any of the embodiments described below.
[0041] In one embodiment, the device can contain one or more sensors that at least partially provide the sensor signals for determining the head angle ϕ of the listener wearing the headphones in the plane. The head angle ϕ is specified relative to one of the M angular perspectives of the stereo signal of layer A or relative to one of the N angular perspectives of the further mono or stereo signal(s) (layer B). Optionally, it is therefore conceivable that at least one further sensor signal from a sensor that is not part of the device itself is taken into account for determining the head angle ϕ of the listener. The one or more sensors can, for example, comprise one or more of the following sensors: acceleration sensor, gyroscope or magnetic field sensor, in order to detect a change in the orientation of the device.Additionally or alternatively, the one or more sensors may comprise a camera to detect a head movement of the head of the listener wearing the headphones.
[0042] The device's processing unit can, for example, be further adapted to detect a change in the listener's head angle from the sensor signals. Alternatively, the listener's head angle can also be provided to the device directly by a head-tracking sensor.
[0043] Further embodiments relate to a system that may include a device according to any of the embodiments described herein and a headphone for reproducing an output stereo signal generated by the device.
[0044] A further aspect of the invention relates to a data structure, and in particular to an audio container. Embodiments therefore also relate to an audio container for generating an output stereo signal with L and R channels based on M audio channels, with M ≥ 4, which are assigned to M predetermined angular perspectives lying in a plane, and at least one further mono or stereo audio channel.
[0045] The audio container can define the following layers: - a first layer (layer A) for the M audio channels assigned to the M angle perspectives; - a second layer (layer B) for the at least one mono and / or stereo audio channel and possibly further mono and / or stereo audio channels assigned to N angular perspectives (e.g. N ≥ 8); - a third layer with metadata; - a fourth layer (optional) for one or more HL stereo signals (Note: the HL stereo signals can also be considered as part of layer B).
[0046] The metadata serves to acoustically position the auditory event of the respective additional mono or stereo audio channel relative to an auditory event (similar to "panoramicization"). The metadata can contain first metadata, which contains a metadata set for each of the at least one mono or stereo audio channel of the second layer. Each metadata set can define the mixing ratio(s) of the respective mono or stereo audio channel of the second layer for each of the N predefined angular perspectives.
[0047] In a further embodiment of the audio container, the first metadata corresponds to a first specified virtual distance from a sound source. The third layer can further comprise second metadata and / or third metadata corresponding to a second specified virtual distance and a third specified virtual distance from the sound source. The second specified virtual distance can be smaller than the first specified virtual distance from the sound source, and the third specified virtual distance can be greater than the first specified virtual distance from the sound source. The second metadata and / or the third metadata for each of the at least one mono or stereo audio channel of the second layer can further contain a metadata set. Each metadata set can define the mixing ratio of the respective further mono or stereo audio channel of the second layer for the specified N angular perspectives.
[0048] In a further embodiment, each metadata record for a stereo audio channel of the second layer (B) identifies four parameter values for each of the N angular perspectives. The four parameter values can specify for a respective angular perspective: - a first mixing factor (i 1,L ) for the first channel of the stereo audio channel of the second layer and a second mixing factor (i 2,L ) for the second channel of the stereo audio channel of the second layer for the channel L of the output stereo signal, and - a third mixing factor (i 1,R ) for the first channel of the stereo audio channel of the second layer and a fourth mixing factor (i 2,R ) for the second channel of the stereo audio channel of the second layer for the channel R of the output stereo signal.
[0049] In another embodiment, each metadata record for a mono audio channel of the second layer identifies two parameter values for each of the N angular perspectives. The two parameter values for a respective angular perspective specify: - a first mixing factor (i L ) of the mono audio channel of the second layer for the channel L of the output stereo signal, and - a second mixing factor (i R ) of the mono audio channel of the second layer for the R channel of the output stereo signal.
[0050] In a further embodiment, each metadata set may further comprise one or more of the following information for each of the N angular perspectives: parametric equalizer settings. The equalizer settings may, for example, relate to a simulation or approximation of a head-related or outer ear transfer function (HRTF) for the two channels of a stereo signal generated from a mono or stereo audio channel of the second layer. Alternatively or additionally, each metadata set may contain, for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying the stereo signal generated from the mono or stereo audio channel of the second layer. Additionally, one or more blending curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source may be included in each metadata set.
[0051] Furthermore, each metadata set can additionally or alternatively contain, for the at least three predefined virtual distances from a sound source, a parametric equalizer setting for simulating distance-dependent sound absorption by air for the stereo signal generated from the mono or stereo audio channel of the second layer. Applying the equalizer settings depending on a selected virtual distance from the sound source can, for example, result in an attenuation of the high frequencies of the stereo signal generated from the mono or stereo audio channel of the second layer to simulate air absorption of the sound waves.
[0052] In a further embodiment, the third layer may further comprise metadata for the M audio channels of the first layer. The metadata for the M audio channels may further comprise one or more of the following information: - A rotation offset for aligning the M angular perspectives relative to a sound source, which are represented by the M audio channels of the first layer. For example, if the sound source was recorded with a microphone array with M microphones arranged in a single plane, the orientation angle of the microphone array relative to the sound source could be subsequently adjusted or adjusted using the rotation offset. - A divergence control parameter. This parameter can be used, for example, to decrease or increase the binaural angle of the M audio channels of the first layer. - For each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying a stereo signal generated from the M audio channels, and optionally further one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source. - For the at least three specified virtual distances from a sound source, a parametric equalizer setting, e.g. for the simulation of a distance-dependent sound absorption by air for the stereo signal generated from the M audio channels
[0053] In further embodiments, the audio container further comprises information indicating for each audio channel of the second layer whether it is a mono channel or a stereo channel.
[0054] In further embodiments, the audio container further comprises, in a fourth layer or as part of the second layer (B), at least one head-locked stereo audio channel whose channels are assigned to the L and R channels of the output stereo signal. Optionally, the audio container may further comprise, for example as part of the third layer, metadata for the at least one head-locked stereo audio channel, which comprises one or more of the following information: - For each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying or attenuating the head-locked stereo audio channel, and one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source. - A parametric equalizer setting for at least three predefined virtual distances from a sound source. Applying the equalizer settings depending on a selected virtual distance from the sound source can, for example, result in a reduction of the high frequencies of the head-locked stereo audio channel to simulate air absorption of sound waves.
[0055] In the various embodiments described herein, the M angular perspectives can correspond to the respective angles of M microphones in a omnidirectional microphone or microphone array that is or was used to record the sound source. The M angular perspectives can lie in one plane (e.g., horizontal plane). The M angular perspectives can, for example, correspond to the angles i · 360° / M with i ∈ [0,1,2, ..., M - 1]. In the embodiments described herein, it is assumed that: M ≥ 4 . In principle, it is advisable to choose an even number M of angular perspectives, since this means that two angular perspectives are offset by 180° from each other and the associated two audio signals can form a stereo pair / stereo channel.Thus, in an even number of M audio channels (which can be called mono channels), two of the M audio channels, whose angular perspectives are offset by 180° from each other, can form one of M / 2 stereo channels (in particular biaural stereo channels).
[0056] In principle, however, an odd number of M audio channels can also be used, where no two microphones are positioned 180° apart. This results in a microphone on one side (e.g., left) being assigned to the crossfade of two microphones on the other side (here, right).
[0057] The embodiments are not limited to the use of a microphone array with in-plane microphones for recording the M audio channels of a sound source. In principle, other microphone arrangements can also be used for recording, optionally together with appropriate recording-side audio signal processing, to generate the M audio channels of a sound source (first layer in the audio container format). For example, it is also possible to use a microphone array on the recording side in which microphones are arranged in two mutually normal planes (horizontal plane and vertical plane), forming what are essentially two orthogonally aligned microphone arrays.It is also possible to record the sound source using a microphone arrangement of fewer or (significantly) more than M microphones, and to generate the M audio channels of a sound source (first layer (A) in the audio container format) from the respective microphone signals by signal processing.
[0058] Each additional mono or stereo signal that is mixed into the stereo audio signal formed from the M audio channels can, for example, be recorded with a single main or support microphone. However, it is equally conceivable that multiple microphones are used to generate a respective mono or stereo signal, e.g. by means of pre-mixing on the recording side or signal processing of the individual microphone signals. The N angular perspectives can lie in one plane (e.g., horizontal plane). The N angular perspectives can, for example, correspond to the angles i · 360° / N with i ∈ [0,1,2, ..., N - 1]. In the embodiments described herein, it is assumed that N ≥ 6 . In principle, it is advisable to choose the number N as a multiple of 4 (e.g., 8 or 12) as angular perspectives, since this includes the critical angular positions 0°, + / - 90° and 180°. Description of the characters Fig. 1 shows a system 100 according to an embodiment of the invention that exemplifies the recording / live recording, the processing of the recording and the generation of metadata, the streaming up to the playback via headphones; Fig. 2 shows an exemplary numbering of angular positions in a microphone array 112 and the assignment of microphone signals to a headphone signal; Fig. 3 illustrates, by way of example, comb filter effects that can occur when recording a sound source with a microphone array 112 when crossfading two adjacent microphone capsules of the microphone array 112 due to their distance; Fig. Figure 4 illustrates an example of a localization of the sound sources on a line between the ears (in-head localization - ICL) in a pure stereo recording via headphones; Fig. Figure 5 illustrates an example of out-of-head localization (AKL) of sound sources that occurs when microphone array signals and additional spot microphone signals are mixed; Fig. 6 shows an example of a mixing of the mixing factors according to the head angle of the listener for panning the stereo group from layer B, according to an embodiment consisting of four support microphones; Fig. Figure 7 shows an example of the determination of the two most obvious angular positions for a head angle; Fig. Figure 8 shows a schematic representation of the functions of the renderer / player 122 according to an exemplary embodiment, with individual sections of the Fig. 8 in the Fig. 8A-D are shown enlarged; Fig. Figure 9 shows a further schematic representation of the functions of the renderer / player 122 according to another exemplary embodiment, wherein individual sections of the Fig. 9 in the Fig. 9A-E are shown enlarged; Fig. 10 shows an exemplary RF spectrum interpolation and blending 1002 that may be used in the embodiments; Fig. 11a and Fig. 11b shows an exemplary numbering of angular positions in a microphone array 112 and the assignment of pairs of microphone signals to a headphone signal for the head angles ϕ = 0° (reference direction) and ϕ = 45°; Fig. 12a shows the localization of a sound source at a head position of -45° and the corresponding angular position of a microphone array 112 with the transmitted microphone signals; and Fig. Figure 12b shows the same sound source and its localization at head positions between 0° and -45° with blended microphone signals according to the head position. Detailed description
[0059] The invention relates to the generation of immersive and / or binaural stereo signals for reproduction via headphones. Embodiments of the invention are described below with reference to a system that describes the recording, transmission, and reproduction of audio. The various aspects of the recording, transmission (e.g., data format), and reproduction (e.g., generation of stereo signals for the headphones) of audio each form individual aspects of the invention, but also in any combination. In some embodiments of the invention, the stereo signal for the headphones is generated as a function of a detected head angle of the listener wearing the headphones.Furthermore, the invention relates to an audio container format for transmitting audio information for generating the stereo signal in a corresponding device (also "renderer" or "player"), as well as computer-readable media, devices and systems for generating and / or reproducing the stereo signal.
[0060] In audio recordings of sound sources, such as a concert, the sound source is typically recorded by a main microphone system. A main microphone system, for example, consists of several microphones appropriately positioned in a recording room. Additional spot microphones are typically positioned near the instruments and mixed into the main microphone system. By increasing the mixing level of a spot microphone, the acoustic presence and proximity of the respective instrument can be increased. The inventor has discovered that this method, which is also applicable to dummy head stereophony, can also be combined with a microphone array as the main microphone system.A microphone array that can be used in some embodiments of the invention to record the sound source can, for example, be a spherical, head-sized separating body on which, ideally, eight microphones (microphone capsules with pressure transducers) are mounted in a ring on the same plane and at the same distance from each other. The microphones are thus located in a (horizontal) plane at an angle of 45° (= 360° / 8). In particular, the microphone array described in US Pat. No. 7,333,622 B2 can be used, but the invention is not limited thereto.
[0061] The microphone array can be placed at a suitable distance in front of the sound body. The exact positioning of the microphone array depends in particular on the local acoustic conditions of the recording room (e.g. concert hall, open-air stadium, etc.) and is usually determined by the recording manager / sound engineer and / or sound engineer. In a concert hall, the microphone array could, for example, be placed in the center of the second or third row in the auditorium. The invention is not limited to a specific arrangement / positioning of the microphone array in the recording room or to a specific design of the microphone array. However, the recording technology and the main microphone system should be capable of capturing the spatial sound in such a way that individual stereo groups from the mixing console can be dynamically mixed into the array signals. All M (e.g. 8) microphone array signals are fed to a playback unit (e.g.Renderer / Player), where they are blended according to the head position. When the head is turned, a tracking sensor ensures that the playback unit blends neighboring microphones accordingly; the microphones of the array are essentially tracked to the listener's ears. This makes the sound source appear to be stationary in the sound space when listening through headphones. A slight movement of the head immediately establishes a front-to-back orientation.
[0062] Furthermore, in addition to the microphone array, additional main and support microphones are positioned in the recording room. Here, too, the number of main and / or support microphones used and / or the positioning of the microphones is not limited to a specific design; both the number (e.g., string quartet vs. symphony orchestra) and the positioning of the main and / or support microphones and the optional combination of several microphone signals into stereo signals can be determined by the recording manager / sound engineer and / or sound engineer. In practice, stereo signals or stereo groups with similar signals, e.g., main microphone signals, string support microphones, drum microphones, etc., formed at the mixing console and / or a digital audio workstation (DAW), are suitable. These are normally panned accordingly for a stereo presentation via loudspeakers, i.e.The phantom sound source positions of the individual signals between the loudspeakers correspond to the position on the stage or that of the main microphone arrangement. The DAW can be software for recording, editing, and producing audio that runs on a processing unit (e.g., computer, laptop, tablet, etc.). These main and / or support signals are also transmitted to the playback unit (e.g., renderer / player) to be mixed with the audio signals from the microphone arrays. By increasing the mixing level of the support signals, the acoustic presence and proximity or directional localization of the respective sound source (e.g., the instrument or a group of instruments) can be increased. It is also possible to use mono signals as support signals (or mono and stereo signals) as support signals. If necessary, the main and support signals can be delayed by the sound propagation times to the microphone array. In order to provide the playback unit (e.g.,To enable the renderer / player to mix the support signals depending on the respective head rotation of the listener, corresponding metadata is transmitted which determines the mixing levels and thus also the panoramas for all angular positions.
[0063] Fig. 1 shows a system 100 according to an embodiment of the invention. The system 100 consists of a recording-side part 110 that captures a sound source with a microphone array 112 with M microphones and a main microphone system and additional support microphones, and provides the audio data and the associated metadata in a data format for transmission (e.g., via streaming) and / or storage. The system 100 also includes a playback-side part 120 that generates an immersive stereo signal for headphone playback, wherein the stereo signal adaptively follows the listener's head movement.
[0064] The recording-side part 110 of the system 100 can comprise a microphone array 112 with M=8 or more or fewer microphone capsules, by means of which the sound source is detected and corresponding M microphone array signals are generated. The microphone capsules can be arranged in a plane (e.g., on a circular line) and assigned corresponding M angular perspectives. For example, the microphone capsules can be arranged in a circle, wherein the angle between two adjacent microphone capsules is 360° / M. Advantageously, M can be selected such that two microphone capsules opposite one another on the circular line form a 180° angle (i.e., m 360° / M = 180°, with m < M, m ∈ ℕ and M ∈ ℕ, where ℕ denotes the set of natural numbers excluding 0). The audio signals of such a pair of opposing microphone capsules can form a stereo pair.The microphone array 112 is placed at a suitable distance in front of the sound box, along with conventional microphones for a stereo or Dolby Atmos recording. This could be, for example, the second or third row in the audience. The microphone array 112 can, for example, comprise a spherical, approximately head-sized separator, on which the M microphones are mounted in a ring at the same level and at the same distance from each other. It is possible to use a microphone array 112 as in US patent specification US 7,333,622 B2.
[0065] From the signals of the M microphones of the microphone array 112, M / 2 audio or stereo pairs can be formed. For each angle φ i = i · 360° / M with i ∈ [0, 1, 2,..., M - 1], the microphones of the orthogonal angles (φ i - 90° and φ i+ 90°) form an audio or stereo pair that can be fed to the headphones as part of the binaural stereo signal. For example, if you number them as in Fig. 2, respectively Fig. 11a and Fig. 11b shows the microphones in a clockwise direction with No. 0 for the microphone that points in the reference direction ϕ = 0° (e.g. pointing directly to the sound body), for the reference direction ϕ = 0° the microphones No. 6 (left) and 2 (right) form such a stereo pair that can be transmitted to the headphones (cf. Fig. 2). For backward head angles, the microphones are swapped accordingly (e.g., 180°: No. 2: left and No. 6: right). When crossfading two adjacent stereo pairs, e.g., from 0° to -45° for a (tracked) head angle ϕ between these two angles, an acoustic counter-rotation of the transmitted acoustic environment is perceived. This is exemplified in Fig. 11b. The head angle ϕ = 0° defines the reference direction of the head angle, for example, the head position frontally to the sound body. Furthermore, it can be assumed, for example, that the angular perspectives φ i = i · 360° / M have the same reference direction as the head angle ϕ.
[0066] It is also assumed in the embodiments described herein that the M angular perspectives φ i match the angular positions of the M microphones of the microphone array 112 (e.g., with regard to number and / or orientation). However, this is not mandatory. The alignment of the microphone array 112 to the M angular positions φ i could, for example, also be subsequently corrected or adjusted by means of a rotation offset in the metadata. Furthermore, it is also conceivable that the number of microphones of the microphone array 112 is smaller or larger than the number M of angular perspectives φ iIn this case, suitable signal processing (e.g., in a DAW or Producer 118) could convert the smaller or larger number of microphone signals into the corresponding M audio signals for the specified M angular positions.
[0067] It is also conceivable that the microphones in the microphone array 112 are not all provided with the same angle and / or that fewer than 8 microphones are provided in the microphone array 112 (e.g., 4, 5, or 6 microphones). For example, the range (angle) of the head angle to which the headphone signal is adapted could not be 360°. For example, in the microphone array 112 in Fig. 2 Microphones 0 and 4 (0° and 180°) are not provided so that a head angle range of < 180° can be covered.
[0068] For ease of understanding and purely as an example, it is assumed below that M microphone array signals from microphone array 112 are captured and corresponding M audio signals are transmitted or recorded. The M microphone array signals as well as the audio signals from the main and / or support microphones 114 are fed to a mixing console 116. The mixing console 116 can be implemented in hardware or via a DAW on a processing unit, or as a combined software and hardware solution. Using the mixing console 116, the signals from multiple support microphones 114 can optionally be combined into a stereo or mono signal before the resulting support signal is fed to the producer 118.
[0069] The producer 118 can be implemented, for example, using a software solution installed on a processing unit (e.g., a computer). Optionally, the functions of the mixing console 116 and the producer 118 can also be combined in one software solution. The M audio signals of the microphone array 112 as well as the audio signals of the support microphones 114 (which may have been grouped together into individual stereo support signals using the mixing console 116) are fed to the producer 118. The producer 118 can, for example, be a stand-alone application or a plug-in for DAWs. The producer 118 generates metadata and combines it with the audio data in a data format. In order to implement the mixing of one or more audio signals based on the signals of the support microphones 114, the metadata is transmitted using a data structure, which contains a metadata set for each of these support signals or the associated channel.Each metadata set identifies the mixing ratios of the respective support signal for each of N predefined angular perspectives. If the mixing of the support signals is to be distance-dependent, several such metadata sets can be provided for each support signal, which define the mixing ratio(s) of the respective support signal for each of the N predefined angular perspectives for different (in particular, three) predefined virtual distances from the sound source.
[0070] All metadata (e.g., mixing ratios and EQ settings for each additional mono or stereo channel per N angle perspectives) can be entered manually by the sound engineer / sound mixer via a GUI on the producer. If a microphone array is permanently installed in one location (e.g., a concert hall), only a few parameters will need to be changed between different concerts.
[0071] The N angular perspectives of the additional mono or stereo audio channel(s) (Layer B) can arise from practical considerations. The number N can therefore be chosen, for example, as N = (n + 1) · 4, with n ∈ [1, 2, 3, ...]. The N angular perspectives of the additional mono or stereo audio channel(s) can coincide with the M angular perspectives of the microphone array 112 (i.e., N = M); however, this is not mandatory, i.e., N ≠ M can also apply. In the following, it is assumed purely by way of example that M=8 and N=8 apply. The N angular perspectives can lie in one plane (e.g., a horizontal plane). The N angular perspectives can, for example, correspond to the angles i · 360° / N with i ∈ [0, 1, 2,..., N - 1]. In the embodiments described herein, it is assumed that N ≥ 6. In principle, it is advisable to define the number N as a multiple of 4 (e.g.: 8 or 12) as angular perspectives, as this covers the critical angular positions 0°, + / - 90° and 180° for a frontal sound event (0°).
[0072] The audio signals and metadata generated by producer 118 are sent, for example, as a stream 130 to a playback-side part 120 of system 100 and / or recorded on a storage medium. The audio signals contain the M array signals of microphone array 112 and the additional stereo and / or mono signals originating from a mixing console 112 (or the DAW). Producer 118 can be optimized in particular for working on a live production. It may be advisable to process as few stereo groups as possible in producer 118 in order to keep the data rate and additional computing effort as low as possible.The producer 118 can provide a graphical user interface (GUI) by means of a display device (not shown), by means of which the operator can generate the metadata that describes, for example, the behavior of the mono and / or stereo signals based on the main and / or support microphones 114 at different head positions of the listener, as will be explained in more detail below.
[0073] Optionally, it is also possible to integrate the functions of the renderer / player 122 into the producer 118. Using the functions of the renderer / player 122 (or an emulation of its functionality), the settings made by the operator (e.g., the set values in the metadata that describe, for example, the behavior of the mono and / or stereo signals based on the support microphones 114 at different head positions of the listener) can be checked in the recording-side part 110 of the system 100. The producer 118 can enable the operator to freely adjust the head position (e.g., head angle) using the GUI. The immersive audio signal for the headphones 104 (also output stereo signal) corresponding to the set head position, which is generated by the functions of the renderer / player 122 in the producer 118, can thus be "pre-listened to" by the operator with headphones (e.g., during a recording) or "listened to" live (e.g.,during a livestream of the recording).
[0074] The producer 118 processes the supplied M audio signals of the microphone array 112 and the further audio signals that are generated in the mixing console 116 based on the signals of the main and / or support microphones 114 and generates digital audio data in a predetermined data format.
[0075] One aspect of the invention relates to such a data or transmission format or a data structure, e.g. audio container, by means of which the information necessary for binaural headphone playback can be stored and / or transmitted. The data format can be a streaming-capable container format for multimedia, video and / or audio data, by means of which the M audio signals of the microphone array 112, the further audio signals generated in the mixer 116 based on the signals of the main and / or support microphones 114, and the associated metadata can be transmitted. The data format can, for example, be one based on ISO / IEC 14496-12:2022 (ISO base media file format), MPEG:H, EBU Audio Definition Model (ADM), etc., or be implemented as an extension thereof. However, it is also possible to use other data formats orUse and / or extend container formats suitable for transmitting and / or storing audio signals and metadata. The metadata can be specified in an XML-based or XML-like syntax within the data format.
[0076] The data format (or data structure) can define the following layers: - A first layer (Layer A) for M audio channels. These M audio channels transport corresponding M audio signals, which were recorded, for example, with a microphone array 112. The M audio channels or signals are assigned to M angle perspectives. - A second layer (Layer B) for the at least one main and / or support signal channel, hereinafter referred to as the support signal channel. Each support signal channel contained in Layer B can be either a mono audio channel or a stereo audio channel. Each support signal channel contains a mono or stereo signal, recorded, for example, with one or more support microphones 114. - A third layer (Layer C) with metadata containing a metadata set for at least the at least one supporting signal channel in the second layer. Each metadata set identifies the mixing ratios of the respective supporting signal for each of the N predefined angular perspectives of the signals from Layer B. If the mixing of the supporting signals is to be distance-dependent, several such metadata sets can be provided for each supporting signal, which define the mixing ratio(s) of the respective supporting signal for each of the N predefined angular perspectives for different (in particular three) predefined virtual distances from the sound source.
[0077] The metadata for layer B can be used to acoustically position the auditory event of the respective support channel relative to an auditory event generated from the M audio channels for the individual M angular positions (layer A) (similar to a “panoramaization”).
[0078] In a fourth layer (Layer D) or as part of the second layer (B), the data format can also contain headlocked (HL) stereo signals. These HL stereo signals can be transmitted to the headphones without panning, i.e., with a fixed LR assignment. The HL stereo signals can, for example, contain direction-independent signals (e.g., reverberation, presentation, etc.).
[0079] The metadata in Layer C may contain further metadata for Layer A and / or Layer B and / or Layer D. For example, the metadata for different (virtual) distances from the sound source may contain a metadata set for each (virtual) distance. Alternatively or additionally, the metadata may also contain distance levels and / or equalizer settings used by renderer / player 122 to generate an immersive stereo signal for the headphones.
[0080] The renderer / player 122 can be implemented, for example, as software or firmware that runs in a device with a computing unit (e.g., DSP, processor, SoC, SiP, ASIC, etc.). The renderer / player 122 could, for example, be implemented in a smartphone, tablet, laptop, or other device with the ability to integrate tracking data (e.g., camera and gyroscope), which generates the generated headphone audio signal and feeds it to the headphones. The renderer / player 122 can be implemented as a web app or an app on a smart device, or can be used in a gaming environment. The audio and metadata can be transferred to the renderer / player 122 as a stream 130 or a file in a data format described herein. The renderer / player 122 blends the audio signals of layer A according to the head angle (of the listener) between the M angular positions, preferably in “real time” (e.g. in less than 20 ms, preferably less than 10 ms).At the same time, the other microphone signals from Layer B are mixed in depending on the head position (e.g. head angle) and optionally also a (virtual) distance according to the metadata from Layer C.
[0081] A tracking system 124 is used to detect head posture. The tracking system 124 is configured for dynamic detection of head movements or "head detection" (in particular in a (horizontal) plane or for detecting the head angle or a change in the head angle) and can also be referred to as a head tracking system. The sensors of the tracking system 124 can also be integrated into the device implementing the renderer / player 122. Alternatively, this sensor technology of the tracking system 124 can also be implemented in the headphones 104. For example, (in-ear) headphones such as Apple's AirPods Max & AirPod Pro, JBL's Quantum One, HyperX's Cloud Orbit S, etc., contain head tracking functions that implement an exemplary tracking system 124 and that can be used in some embodiments. A headset with a built-in tracking system can transmit full 360° head movements to the renderer / player 122.Combined with a (smartphone) camera, for example, the (virtual) distance to a virtual stage can also be simulated or adjusted. Alternatively, the tracking system 124 can be implemented as a standalone device that can be attached to the headphones, for example (e.g., Waves Nx Head Tracker for headphones from Waves Inc.).
[0082] In a further alternative, the sensors of the tracking system 124 can be implemented both in the device that supports the renderer / player 122 and in the headphones 104. The tracking system 124 is configured to determine the (current) head angle ϕ or a change in the head angle ϕ and to feed it to the renderer / player 122. The angular resolution of the tracking system 124 should be less than or equal to 5° in order to be able to track a head movement of the listener as accurately as possible and to enable the renderer / player 122 to generate an immersive stereo signal for the headphones 104 that follows the head position as accurately as possible. In some embodiments, the tracking system 124 enables an angular resolution of 5° or less, preferably 2° or even less.
[0083] Various systems can be used for head tracking, including camera(s) and image processing, gyroscope(s), magnetometer(s), accelerometer(s), ultrasonic sensor(s) and / or infrared sensor(s), and three-point radio wave measurement systems. Gyroscopes and accelerometers can be used to record the rotational movements and accelerations of a device / headset. By integrating the signals from these sensors, the relative position or orientation of the device / headset in space can be determined. When a gyroscope is combined with a camera, the tracking data is added together, enabling 360° tracking with a moving head.
[0084] Magnetometers can be used to measure the magnetic field around the device / headphones. By analyzing changes in the magnetic field, the movement and orientation of the device / headphones can be tracked. Ultrasonic or infrared sensors can use ultrasonic or infrared signals to determine the position of the head or headphones in space. Typically, multiple sensors are placed in the environment to receive the signals and calculate the position of the headphones.
[0085] Camera(s) and image processing can also be used to track the listener's head rotation based on the images captured by the camera. When using a smartphone camera, the smartphone can be placed or held in front of the listener, as if taking a selfie. Using the camera, software (e.g., facial recognition software) can detect head position within a range of less than + / - 90° from the reference direction (0°). At the same time, the camera can measure the listener's distance from the smartphone and convert this into a (virtual) distance from the virtual stage.
[0086] To improve tracking accuracy, signals from multiple and / or different sensors / sensor types can be used.
[0087] Alternatively or additionally, it is also possible for all movement parameters relating to head rotation / orientation (especially the head angle) and / or the virtual distance from the listener to be manually adjusted via the graphical user interface (GUI) of the renderer / player 122. For example, the frontal view of a (sketched) head can be shown on the display. The listener / user could then turn their head left or right with one or more fingers to adjust the head angle ϕ. Moving the head up or down with one or more fingers could be used to adjust the listener's distance from a virtual stage.
[0088] Furthermore, it is possible for the renderer / player 122, together with the tracking system 124, to enable gesture control of the renderer / player 122. The gestures can be used, for example, to: (1) fix the current head position (distance and angle of rotation) ("freeze"), e.g., by nodding twice, (2) release the "freeze" by nodding three times, (3) set the head angle and / or a virtual distance to a default value (0° and distance to a predefined "sweet spot" - e.g., the distance to the medium), start or pause / stop audio playback, etc. The individual corresponding functions can alternatively or additionally also be implemented using the control elements on the GUI.
[0089] As already mentioned, the renderer / player 122 can crossfade the audio signals of Layer A according to the head angle (of the listener) between the M angular positions (e.g., in real time). At the same time, the additional microphone signals of Layer B are mixed in depending on the head position (e.g., head angle) and optionally also a (virtual) distance according to the metadata from Layer C. By mixing in the additional microphone signals of Layer B, it is possible, from certain levels onwards, to mask comb filter effects that arise when crossfading neighboring microphones of the microphone array 112, without the need for complex signal processing. As in Fig. 3, the spatial offset between adjacent microphones in the microphone array 112 causes the incoming sound to be received with different phase positions. For example, at a distance of 6.8 cm between two microphone capsules on the spherical surface, the passing sound is maximally attenuated at approximately 2.5 kHz when the capsules are mixed, which can be masked by mixing in the other microphone signals from Layer B. In comparison, in US 7,333,622 B2, such comb filter effects are compensated for using more complex signal processing, which includes high-pass filtering of the critical signals, decomposition of the high signal components into amplitudes and phases, amplitude addition, and introduction of the phase of a microphone. In principle, it is also possible in some embodiments to use the signal processing described in US 7,333,622 B2 to compensate for comb filter effects.
[0090] The decisive effect of mixing in the additional microphone signals of Layer B is to increase the presence. When playing a normal stereo recording through headphones alone, the sound sources are localized on a line between the ears (in-head localization - ICL), as shown in Fig. 4. Only in a mixture of the spherical array signals of Layer A and the possibly delayed, panned and possibly regulated microphone signals of Layer B, the sound events migrate out of the head (out-of-head localization - AKL) and appear more three-dimensional to the listener in closer position and presence, as is the case in Fig. 5 is illustrated.
[0091] In the following, purely as an example and for easier understanding, we will consider a string quartet on a stage in which the first violin is positioned on the left and the cello on the right outside of the stage, with the two other musicians in between. Fig. Figure 6 shows a sketch. The reference direction here is ϕ = 0°. When recording the string quartet, for example, a microphone array 112 with 8 microphones can be positioned in the center of the auditorium at the height of the second row of chairs, with microphone number 0 directed toward the stage (in the reference direction ϕ = 0°). Each of the four instruments also has its own spot microphone. In this example, the recording-side part 110 of the system 100 can transmit 8 microphone signals from the microphone array 112 in four stereo pairs (Layer A) in the stream 130; for example, for the head angle 0° (microphones 6 & 2 for left & right), -45° (microphones 5 & 1 for left & right), +45° (microphones 7 & 3 for left & right), and -90° (microphones 4 & 0 for left & right).By swapping the stereo channels in the renderer / player 122 you get the four other head angles +90° (microphones 0 & 4 for left & right), 180° (microphones 2 & 6 for left & right) and +135° (microphones 1 & 5 for left & right) and -135° (microphones 3 & 7 for left & right).
[0092] The audio signals from the support microphones are combined as a stereo group with panned audio signals from the instruments and transmitted as a single stereo audio signal in Layer B: For a head angle of ϕ = 0°, the level of the first violin in the left channel is highest; the cello signal is highest on the right. This corresponds to the stereo group signal in a normal loudspeaker recording. In Producer 118, the stereo group is panned, adjusted, and if necessary delayed (optional) at given angular positions in N head positions, so that the musicians appear in the "correct" positions for each head angle in the immersive stereo signal for headphones 104, which can be generated by Renderer / Player 122. These settings (parameters) made in Producer 118 for each of the N angular positions are transmitted as a metadata record in the metadata (Layer C) of Stream 130 for the stereo group (in Layer B).When the head is rotated, the renderer / player 122 does not blend the stereo signal of Layer B. Instead, the parameter values in the metadata for panoramaizing the corresponding N angular positions are averaged according to the listener's head angle and used to generate a stereo signal from the stereo signal of Layer B. With appropriate parameter settings, the panoramas "move" in the opposite direction to the head movement, and the sound source positions essentially "remain" in place. The stereo groups can be prepared for the producer 116 at the mixing console 116 (or in the DAW). It may be useful for the audio signals from the microphone array 112 to arrive at the listener's ears simultaneously or slightly earlier (law of the first wave front).In the string quartet example, the signals from the support microphones can therefore be delayed by the mixing console 116 (or in the DAW) by at least the propagation time of the distance to the microphone array 112 before being fed to the producer 118. It may also be sufficient to delay a stereo group of support signals as a whole. However, a corresponding temporal adjustment of the propagation times can also be implemented, for example, in the producer 118 and can be adjusted via a user interface.
[0093] In practice, 8 angular positions have proven useful for the additional mono and stereo signals from layer BN in order to obtain stable localizations of the sound sources when the head is rotated. A smaller number of angular positions can lead to lateral and central deflections of the sound source positions when the head is rotated. A larger number of angular positions can also be used. Increasing the number of angular positions (e.g. to 12 angular positions in 30° steps) does not significantly stabilize the localization, but does result in more work in generating the metadata for layer B. Eight angular positions also have the advantage that acoustic events can be easily reproduced mentally in 45° steps in order to make the adjustments to the panorama.
[0094] In the following, N=8 angle positions are assumed for the panorama of the signals of Layer B. The panoramas are calculated in Producer 118 for each stereo channel using four mixing factors, i 1,L , i 1,R , i 2,L and i 2,R (also called mixing ratio) per angular position. With i as the value of the mixing factor, 1 left channel, 2 right channel of the stereo signal, and L and R as the target for the headphone channels L and R of the headphone signal generated in the renderer / player 122, on which the respective channels S1 and S2 of the stereo channel are rendered with the mixing factors i. Thus, each of the N=8 angular positions is assigned a metadata set with the four mixing factors i 1,L , i 1,R , i 2,L and i 2,R assigned and transmitted in the metadata of Layer C in Stream 130. For each additional stereo support signal, the following mixing factors are transmitted as a metadata set in the metadata of Layer C: ij,1,L, ij,1,R, ij,2,L, ij,2,R where j denotes the angular position j (j ∈ 0, 1, .., N - 1). Index 1 denotes channel S1 of the support signal, index 2 denotes channel S2 of the support signal, and indices L and R denote the headphone channels L and R.
[0095] In the renderer / player 122, the mixing factors of the angular positions are read from the metadata of layer C and linearly averaged depending on the head position and the head angle. The mixing factors thus obtained are then rendered with the signals S1 and S2 of the stereo channel from layer B. With appropriate selection of the mixing ratios, this creates the impression of opposing sound sources with a rotating head, albeit still as IKL. In parallel, the M audio signals of the microphone arrays 112 from layer A of stream 130 are blended in the renderer / player 122 to form a binaural stereo signal and summed with the output of the rendered stereo signals from layer B, thus generating the headphone signal for the headphones 104 as the output stereo signal.
[0096] Following the example of the string quartet, Fig. 6 shows an example of a mix of the mixing factors according to the head angle of the listener for panning the stereo group, which is created from the four support microphones 602, 604, 606, 608. As in Fig. As shown in Figure 6 above, the signal level of the first spot microphone 602 of the first violin is approximately 18 dB stronger on the left channel than on the right channel of the stereo group. The second violin (spot microphone 604) is panned with a 6 dB difference in favor of the left channel and appears approximately 50% of the left side (during loudspeaker playback). The level of the cello (spot microphone 608) is approximately 18 dB stronger on the right channel than on the left channel of the stereo group. The viola (spot microphone 606) is panned with a 6 dB difference in favor of the right channel and appears approximately 50% of the right side (the loudspeaker base). This spot microphone stereo mix serves as the starting point and is assigned to a head angle of 0°. In this mix, the instruments appear acoustically to the listener as shown in Figure 610, but still as IKL.
[0097] The mixing factors for the angular position = 0° are, for example: i 1,L (ϕ = 0°) = 1, i1,R (ϕ = 0°) = 0, i 2,L (ϕ = 0°) = 0, i 2,R (ϕ = 0°) = 1
[0098] For a head angle of ϕ = 0°, the spot microphone stereo mix ("original mix") is mixed with these mixing factors to the corresponding stereo signals of microphone array 112 for a head angle of 0° (microphones 6 (left) and 2 (right)). Now, the sound event migrates out of the head and appears three-dimensional in front of the listener.
[0099] For the angular position = -45°, the level of the right stereo channel of the spot microphone stereo mix is raised and that of the left channel is lowered and distributed to both channels L and R, so that the instruments appear to the listener acoustically "shifted" to the right onto channel R, as shown in Figure 612. The mixing factors for the angular position = -45° are, for example: i 1,L (ϕ = -45°) = 0.5, i 1,R (ϕ = -45°) = 0.5, i 2,L (ϕ = -45°) = 0, i 2,R (ϕ = -45°) = 1.2
[0100] For a head angle of ϕ = -45°, the spot microphone stereo mix is mixed with these mixing factors to the corresponding stereo signals of the microphone array 112 for the head angle -45° (microphone 5 (left) and 1 (right)).
[0101] For a head angle that does not correspond to any of the predefined angular positions, for example ϕ = -22.5°, the mixing factors of the two predefined angular positions closest to the head angle - ie here the angular positions 0° and -45° - are blended linearly (see Fig. 6, right side), so that we get: i1, L (ϕ = -22.5°) = 0.75, i 1,R (ϕ = -22.5°) = 0.25, i 2,L (ϕ = -22.5°) = 0, i 2,R (ϕ = -22.5°) = 1.1. The instruments are acoustically "shifted" slightly to the right to channel R, as shown in graph 614, but still with signal components on channel L. The blended metadata is rendered with the stereo signal.
[0102] In Layer A, depending on the head position, it is not the metadata but the M audio signals themselves that are blended. The audio signals of microphone array 112 (here, for example, M=8) in Layer A are blended linearly (microphones 6 and 7 for the left channel, and microphones 2 and 3 for the right channel). For a head angle of ϕ = -22.5°, the spot microphone stereo mix is mixed with the blended stereo signals of microphone array 112 for a head angle of -45° using the averaged mixing factors.
[0103] The linear mixing of the mixing ratios of the spot microphone stereo mix (Layer B) for a head angle ϕ can generally be represented as follows: i1,L(ϕ)=in,1,L⋅(1−b)+in+1,1,L⋅b i2,L(ϕ)=in,2,L⋅(1−b)+in+1,2,L⋅b i1,R(ϕ)=in,1,R⋅(1−b)+in+1,1,R⋅b i1,R(ϕ)=in,2,R⋅(1−b)+in+1,2,R⋅b where the indices n and n + 1 denote the two angular positions closest to the head angle ϕ and with b=ϕ−n⋅360° / N360° / N, where n ∈ 0, 1,.., N - 1 and the value of n is chosen such that: ϕ ≥ n · 360° / N and ϕ ≤ (n + 1) · 360° / N. N denotes the number of angular perspectives (e.g., N=8 or N=12). The two angular positions n and n + 1 can also be referred to as the angular positions n and nn (nearest neighbor n and next-nearest neighbor nn). Which of the angular positions n or n + 1 forms the nearest neighbor n or the next-nearest neighbor depends on the head angle ϕ.
[0104] Accordingly, the stereo support signal of layer B is linearly averaged for a head angle ϕ as follows: L(ϕ)=S1⋅i1,L(ϕ)+S2⋅i2,L(ϕ)=S1(in,1,L⋅(1−b)+in+1,1,L⋅b)+S2(in,2,L⋅(1−b)+in+1,2,L⋅b) R(ϕ)=S1⋅i1,R(ϕ)+S2⋅i2,R(ϕ)=S1(in,1,R⋅(1−b)+in+1,1,R⋅b)+S2(in,2,R⋅(1−b)+in+1,2,R⋅b) where L(ϕ) and R(ϕ) represent the left and right channels of the blended stereo signal and S1 and S2 represent the left and right channels of the stereo support signal of layer B.
[0105] Additionally, further metadata can be transmitted in Layer C to simulate variable distances to the sound body. This allows the listener to choose a virtual distance to the sound body, e.g., from the conductor's position via the "best seat" to the back of the hall. With the distance, the panoramas of the stereo signals in the frontal area also change for the listener, since, for example, the aperture angle of an ensemble decreases with distance. By appropriate panoramas for different distances, these relationships can be credibly reproduced. Accordingly, for example, a respective metadata set can be n,1,L , i n,1,R , i n,2,L , i n,2,R, for (at least) two, advantageously for three (or more) different virtual distances (positions) to the sound source in Layer C. These additional metadata sets can be generated accordingly in the producer 118. For example, three different metadata sets, each containing the mixing ratios i n,1,L , i n,1,R , i n,2,L , i n,2,R for which N angular positions are contained, for three virtual distances (e.g., Close, Medium, Far) for the respective stereo signals of Layer B are provided in the metadata. In the renderer / player 122, virtual distances that lie between two of the defined distances (e.g., distance between Close and Medium or between Medium and Far) can be calculated by linear averaging (analogous to the averaging of the mixing ratios for an angular position ϕ) of the (possibly averaged) mixing ratios for the respective angular position ϕ: i1,L(d, ϕ)=iD1,1,L(ϕ)⋅(1−d)+iD2,1,L(ϕ)⋅d i2,L(d, ϕ)=iD1,2,L(ϕ)⋅(1−d)+iD2,2,L(ϕ)⋅d i1,R(d, ϕ)=iD1,1,R(ϕ)⋅(1−d)+iD2,1,R(ϕ)⋅d i1,R(d, ϕ)=iD1,1,R(ϕ)⋅(1−d)+iD2,1,R(ϕ)⋅d where d denotes the set virtual distance between the two specified distances D1 (e.g., close or medium) and D2 (e.g., medium or far). The values i (ϕ) denote the already averaged mixing ratios for the head angle ϕ, as described above. Accordingly, the stereo support signal of Layer B is linearly averaged for a head angle ϕ and a distance d as follows: L(d, ϕ)=S1⋅i1,L(d, ϕ)+S2⋅i2,L(d, ϕ) R(d, ϕ)=S1⋅i1,R(d, ϕ)+S2⋅i2,R(d, ϕ)
[0106] Alternatively, the mixing ratios can first be averaged depending on the selected / desired virtual distance (analogous to the averaging of the mixing ratios for an angular position ϕ) and then a further averaging of the distance-dependent mixing ratios thus obtained can be carried out for any angular position ϕ (if the angular position ϕ does not correspond to any of the specified angular positions).
[0107] In the embodiments described above, it was assumed that the audio signals from the support microphones are combined into one or more stereo groups (i.e., one or more stereo signals with left and right channels) and that the individual stereo groups are transmitted in Layer B of the data format (e.g., audio container) in stream 130. It is also possible for the signals from the support microphones to be transmitted individually or combined in groups as mono signals in Layer B. In this case, the mixing factors of a metadata record describe how the mono signal in Layer B is distributed between the two left and right channels of the stereo signal generated from the mono signal in the renderer / player 112. If the data format supports both stereo signals and mono signals in Layer B, the metadata can also specify whether a respective channel in Layer B is a mono channel or a stereo channel.
[0108] In further embodiments, the metadata in Layer C can also contain further equalizer settings that, for example, simulate the treble attenuation due to air absorption for a predetermined (virtual) distance. The equalizer settings can - similar to the metadata sets of the mixing factors - also be specified for at least two, for example, for three or more different virtual distances, in order to, for example, simulate the treble attenuation due to air absorption for different virtual distances. The equalizer settings are faded in and out linearly. The metadata for equalizer settings can be specified for each stereo channel of Layer B and / or for each of the predetermined N angular positions. The equalizer settings for the left and right headphone channels can be set individually in Producer 118, in order to, for example, simulate head shadowing at one ear when the head is rotated.Accordingly, a metadata record for the equalizer settings for each stereo channel of Layer B and / or for each of the predefined N angular positions contains equalizer settings for the left and right channels of the stereo signal generated from the stereo channel of Layer B in the renderer / player 122.
[0109] Embodiments of the channel layout or the data structure of the stream 130 take into account, as shown above, the different degrees of freedom of the listener: On the one hand, the additional support signals in layer B and the associated metadata sets of the mixing ratios for the different predetermined N angular positions of layer C allow consideration of a head rotation of the listener (up to 360° rotation) when generating an immersive headphone signal in the renderer / player 122. In addition, the specification of metadata sets of the mixing ratios for different distances to the sound body allows consideration of a (virtual) distance of the listener from the sound body.
[0110] The producer 118 in the recording-side part 110 of the system 100 is configured to provide or generate the audio signals from the microphone array 112 (Layer A) and the support microphones 114 (Layer B), as well as the associated metadata (Layer C). The metadata is generally not time-critical and can be changed during a transmission. For example, the channel layout or data structure of the stream 130 can be structured as follows: (1) Layer A: a. M audio channels
[0111] The M audio channels in M mono tracks. The M audio channels can be generated with a 112-microphone array (e.g., with M=8, with 8 microphones for the angle perspectives 0°, -45°, +45°, +90°, -90°, -135°, +135°, and 180°). (2) Layer B:a. Audio channels B1, ..., Bm
[0112] Any number of audio channels for additional audio signals, which can be mono and / or stereo. For example, mixed signals from 5.1 to Dolby Atmos surround sound can be transmitted in mono format. Possible applications: Transmission of stereo support signals and / or stereo main microphone signals. b. Audio channels C1, ..., Ck (optional):
[0113] Any number of additional ones are available for headlocked (HL) stereo signals. These signals are not further panned in Renderer / Player 122, but are mixed with the headphone signal in Renderer / Player 122. Application: Stereo reverb signals without directional information, presentations, etc. (3) Layer C:a. Metadata set for each audio channel in Layer B
[0114] Metadata set 802-2 contains N mixing factors for panoramas for the N angular positions for the respective audio channel B1, ..., Bm from Layer B. Optionally, several such metadata sets 802-1, 802-2, 802-3 can be included in the metadata, with each metadata set being assigned to a predefined (virtual) distance (e.g., Close, Medium, Far). It is also possible for several audio channels from Layer B to share one or more metadata sets. b. Equalizer settings (optional)
[0115] Dual-band equalizer settings 916-1, 916-2 for the L channel and R channel of a stereo signal generated in the renderer / player 122 from the respective audio channel B1, ..., Bm from layer B. The dual-band equalizer settings 916-1, 916-2 are specified for both channels for each of the N angular positions. The equalizer settings can be parametric or semi-perimetric. In particular, the equalizer settings can affect three or more bands of an equalizer 936-1, 936-2. c. Gain factors (optional)
[0116] Amplification / attenuation factor (gain) 918 for the amplification / attenuation of a stereo signal generated in the renderer / player 122 from the respective audio channel B1,..., Bm from Layer B or for a respective HL signal from Layer B. A amplification / attenuation factor can be specified for each of several predefined (virtual) distances (e.g., close, medium, far). The amplification / attenuation factor (gain) 918 or distance-dependent amplification / attenuation factors (gain) 918 can be defined for each audio channel of Layer B. In addition to the gain / attenuation factor 918 (for a respective distance), one or more blending curves 932 can optionally be defined or specified, which can be used to average two gain / attenuation factors 918 when a (virtual) distance is to be set between specified (virtual) distances. d. Distance-dependent equalizer settings (optional)
[0117] Equalizer settings 920 for a stereo signal generated in renderer / player 122 from the respective audio channel B1, ..., Bm from layer B or for a respective HL signal from layer B, or equalizer settings 966 for a binaural stereo signal generated from the M audio channels of layer A, wherein corresponding equalizer settings 920, 966 are specified for each of several predetermined (virtual) distances (e.g., Close, Medium, Far). The equalizer settings 920, 966 can, in particular, relate to two or more bands of an equalizer 914, 940, 944. e. System-related metadata (optional)
[0118] This metadata 954 includes, for example, settings for stereo width controls 946 generated after summation 820, 822 from the audio signals B1,..., Bm of Layer B (e.g., head size adjustment). Alternatively or additionally, settings 950, 830 for output equalizer 952 and / or limiter 828. f. Metadata for Layer A (optional)
[0119] The metadata for Layer A can contain one or more of the following information: - Contains a rotation offset 960 for the subsequent “alignment” of the microphone array 112. - Divergence control 960 for (subsequent) reducing or enlarging the binaural angle of the microphones of the microphone array 112 (normally 180°). - Gain / attenuation factor 964 for the amplification / attenuation of a stereo signal generated in the renderer / player 122 from the M audio signals of Layer A. A gain / attenuation factor can be specified for each of several predefined (virtual) distances (e.g., Close, Medium, Far). In addition to the gain / attenuation factor 964 (for a respective distance), one or more crossfade curves can optionally be defined or specified, which can be used to crossfade two gain / attenuation factors 964 when a (virtual) distance between predefined (virtual) distances is to be set.
[0120] Fig. 8 ( Fig. 8A-D) shows a schematic and exemplary representation of the functions of the renderer / player 122 (hereinafter referred to as player 122). The player 122 can, for example, receive audio and metadata in the form of a stream 130 or alternatively from a file that is present in a corresponding data format with layers A, B, and C. In the implementation example shown, layer A comprises a number of M audio channels with the signals S(j) with j ∈ 0, 1,..., M - 1 (see (1) a. of the data format), which are assigned to M predetermined angular positions (1 to M). The audio channels of layer A can have been recorded with a microphone array 112. Layer B comprises a stereo group (stereo signal B1 - see (2) a. of the data format) with the audio signals S SB1,L and S SB1,R , as well as an HL stereo signal (HL Stereo HL1 - see (2) b. of the data format) with the audio signals 5 H1,L and S HL1,R. In layer C for the metadata, (optional) system-related metadata (see (3) e. of the data format) and metadata 802-2 for the stereo group of layer B (metadata stereo signal B1 - see (3) a. of the data format) are included.
[0121] The audio blender 810 of the player 122 receives a tracking signal (e.g., continuously, at specified time intervals) from the tracking system 124, which indicates at least the current head angle ϕ of the listener wearing the headphones 104, with which the immersive output stereo signal 808 is reproduced. The M audio signals are rendered in the audio blender 810 according to the head angle ϕ in the tracking signal, so that the currently required adjacent stereo pairs are blended, provided the head angle ϕ does not correspond to any of the predetermined angular positions. For a head angle of, for example, -22.5°, the audio signals of angular positions 6 and 7 for the left headphone channel and 2 and 3 for the right headphone channel are linearly blended 812: L(ϕ)array=S(m−90°)⋅(1−b)+S(m+1−90°)⋅b R(ϕ)array=S(m+90°)⋅(1−b)+S(m+90°)⋅b where the indices m and m + 1 denote the two angular positions closest to the head angle ϕ and with b=ϕ−m⋅360° / M360° / M, where m ∈ 0, 1,.., M - 1 and the value of m is chosen such that ϕ ≥ m · 360° / M and ϕ ≤ (m + 1) · 360° / M. M denotes the number of angular perspectives (e.g., M=6 or M=8). m ± 90° and m + 1 ± 90° denote the two angular positions that are shifted by 90° to the left and right, respectively, with respect to the angular positions m and m + 1. For M=8, the angle between the angular positions would be 45°, so that m ± 90° and m + 1 ± 90° correspond to the audio signals with the angular positions (M + m ± 2) mod M and (M + m + 1 ± 2) mod M, respectively, where mod represents a modulo operation. This is in Fig. 7 for the example ϕ = 65° and M = 8. The angular positions closest to the 65° angle are the positions m = 1 and m + 1 = 2. The position m + 90° corresponds to the angular position (8 + 1 + 2) mod 8 = 3, the position m - 90° corresponds to the angular position (8 + 1 - 2) mod 8 = 7. The position m + 1 + 90° corresponds to the angular position (8 + 2 + 2) mod 8 = 4, the position m + 1 - 90° corresponds to the angular position (8 + 2 - 2) mod 8 = 0.
[0122] If the head angle ϕ corresponds to one of the predefined N angular positions, the audio signals S B1,L and S B1,R of the stereo group (stereo signal B1) according to the mixing factors 802-2 assigned to the corresponding angular position to the channels L(ϕ) B1 and R(ϕ) B1If, however, the head angle ϕ does not correspond to any of the predefined N angle positions, the mixing factors for the stereo group (stereo signal B1) are first blended according to the head angle ϕ and the blended mixing factors are used to distribute the audio signals S B1,L and S B1,R the stereo group (stereo signal B1) to the channels L(ϕ) B1 and R(ϕ) B1 to be distributed (cf. amplifiers 840-1, 840-2, 840-3, 840-4 and adder elements 816, 818). The channels L(ϕ) B1 and R(ϕ) B1 form a resulting, head angle-dependent stereo signal, which corresponds to the stereo signal with the channels L(ϕ) array and R(ϕ) array is added (cf. adding elements 804, 806). The channels L(ϕ) B1 and R(ϕ) B1 can form as follows: L(ϕ)B1=SB1,L⋅i1,L(ϕ)+SB1,R⋅i2,L(ϕ) R(ϕ)B1=SB1,L⋅i1,R(ϕ)+SB1,R⋅i2,R(ϕ) with the blended mixing factors i j,1,L , ij,1,R , i j,2,L , i j,2,R , i j,L , i j,R , from the metadata 802-2 for the stereo group (stereo signal B1): i1,L(ϕ)=in,1,L⋅(1−b)+in+1,1,L⋅b i1,R(ϕ)=in,2,L⋅(1−b)+in+1,2,L⋅b i2,R(ϕ)=in,1,R⋅(1−b)+in+1,1,R⋅b i2,R(ϕ)=in,2,R⋅(1−b)+in+1,2,R⋅b where the indices n and n + 1 denote the two angular positions closest to the head angle ϕ and with b=ϕ−n⋅360° / N360° / N, where n ∈ 0, 1,.., N - 1 and the value of n is chosen such that: ϕ ≥ n · 360° / N and ϕ ≤ (n + 1) - 360° / N. Where N denotes the number of angular perspectives (e.g. N=8 or N=12). Analogous to the explanations regarding the determination of the angular positions m ± 90° and m + 1 ± 90° for the audio signals of Layer A with respect to Fig. 7, here n ± 90° and n + 1 ± 90° denote the two angular positions of the N angular positions, which are shifted by 90° to the left and right, respectively, with respect to the angular positions n and n + 1. For N = 8, the angle between the angular positions would be 45°, so that n ± 90° and n + 1 ± 90° correspond to the additional audio signals of layer B with the angular positions (N + n ± 2) mod N and (N + n + 1 ± 2) mod N, respectively, where mod represents a modulo operation. D
[0123] The range of the mixing factors i n,1,L , i n,1,R , i n,2,L , i n,2,R, can, for example, be between -2 (inclusive) and +2 (inclusive). A factor of 2 can correspond to a gain of 6 dB, a factor of 1 corresponds to 0 dB gain, a factor of 0 to signal cancellation (multiplication by 0), a factor of -1 corresponds to 0 dB gain with an additional 180° phase shift of the signal, and a factor of -2 corresponds to a gain of 6 dB with a 180° phase shift of the respective channel. The value range of the mixing factors can also be specified in 0.01 steps, for example.
[0124] For mono signals S Bi In layer B, only two blending factors i are used in the metadata set. j,L , i j,R for each of the N angular positions j ∈ [0,1,2, ..., N - 1] and the head angle-dependent stereo signal is formed as follows: L(ϕ)Bi=SBi⋅(in,L⋅(1−b)+in+1,L⋅b) R(ϕ)Bi=SBi⋅(in,R⋅(1−b)+in+1,R⋅b) Here too, the mixing factors can be blended.
[0125] If layer B contains several such stereo groups (and / or mono signals), these are processed in player 122 like the stereo signal B1 (or as shown for the mono signal) and the resulting head angle-dependent stereo signals are added (cf. adding elements 820, 822).
[0126] If Layer B contains one or more HL stereo signals, such as the HL stereo signal (HL Stereo HL1) with the audio signals S HL1,L and S HL1,R , these are combined with the other signals (here: L(ϕ) array and R(ϕ) array and L(ϕ) B1 and R(ϕ) B1 ) are added (see adding elements 804, 806 and adding elements 824, 826).
[0127] The resulting stereo signal L(ϕ) Kopf and R(ϕ) Kopf forms the immersive output stereo signal that is fed to the headphones 104. Optionally, the stereo signal L(ϕ) Kopf and R(ϕ) Kopfa limiter 828. The limiter 828 can, for example, prevent overloads of the stereo signal L(ϕ) Kopf and R(ϕ) Kopf The limiter 828 can be configured via system metadata 830. This system metadata 830 can define the following parameters for the limiter 828: Limiter Treshold T Release R Gain G
[0128] Where T is the threshold of the control range in dB below full scale (0 dBFS), from which the signal is limited and is specified in 0.1 dB steps; R is the release time of the limiter in milliseconds (ms), which determines how long the control process returns to normal after adjustment, and Gain is the level increase or reduction before the limiter is adjusted (e.g. in 0.1 dB steps).
[0129] In an alternative embodiment, the stereo groups (or mono signals) in layer B in the producer 118 of the recording-side part 110 of the system 100 can already be panned for all N angular positions and all N fully panned stereo signals S j,L and S j,R (with j ∈ [0,1,2, ..., N - 1]) for all N angular positions j from the producer 118 in the stream 130. In this case, the metadata 802-2 could be omitted. Instead, two stereo signals S n and S n+1 for the angular positions n and n + 1, which are closest to the head angle ϕ, are directly blended with each other linearly. L(ϕ)Bi=Sn,L⋅(1−b)+Sn+1,L⋅b R(ϕ)Bi=Sn,R⋅(1−b)+Sn+1,R⋅b where the indices n and n + 1 denote the two angular positions closest to the head angle ϕ and with b=ϕ+n⋅360° / N360° / N, where n ∈ 0, 1,.., N - 1 and the value of n is chosen such that: ϕ ≥ n · 360° / N and ϕ ≤ (n + 1) - 360° / N. N denotes the number of angle perspectives (N = 8 or N = 12). This places lower demands on the performance of the hardware implementing Player 122. Only the N stereo channels in Layer B are blended, but no equalizers are used. This reduces the CPU load, but at least N stereo channels are required in Layer B.
[0130] Another exemplary implementation of the player 122 is shown in Fig. 9 ( Fig. 9A-E). This implementation offers, for example, the possibility of selecting or setting a variable distance to the sound source. For this purpose, the data format provides the player 122 with additional metadata, as explained in more detail below. As already mentioned in connection with Fig. 8, it can also be assumed here that the player 122 receives audio and metadata in the form of a stream 130 or alternatively obtains it from a file that is present in a corresponding data format with layers A, B and C. In the implementation example shown, it is again assumed by way of example that layer A comprises a number of M audio channels with the signals S(j) with j ∈ 0, 1,..., M - 1 (see (1) a. of the data format), which are assigned to M predetermined angular positions (1 to M). The audio channels of layer A can have been recorded with a microphone array 112. Layer B comprises a stereo group (stereo signal B1 - see (2) a. of the data format) with the audio signals S SB1,L and S SB1,R , as well as an HL stereo signal (HL Stereo HL1 - see (2) b. of the data format) with the audio signals S HL1,L and S HL1,R. In layer C for the metadata, (optional) system-related metadata (see (3) e. of the data format) and metadata 802-2 for the stereo group of layer B (metadata stereo signal B1 - see (3) a. of the data format) are included.
[0131] As in Fig. 8 will also be implemented in Fig. 9 The M audio signals are rendered in the audio blender 810 according to the head angle ϕ in the tracking signal, so that the currently required neighboring stereo pairs are blended, provided the head angle ϕ does not correspond to any of the specified angular positions. Optionally, the reference direction (orientation) of the microphone arrays 112 and the N angular positions of the audio channels in Layer A can be subsequently aligned to one another in the audio blender using a rotation offset 960 specified in the metadata. Furthermore, also optionally, a divergence control 960 can be specified in the metadata, which serves to (subsequently) reduce or increase the binaural angle of the microphones of the microphone array 112 (normally 180°). This function can also be implemented in the audio blender 810.
[0132] The audio blender 810 feeds the output signal to an RF spectrum interpolator 910. In some cases, the addition of the audio signals from Layer B (e.g., at very low levels) may not be sufficient to mask any comb filter effects that may occur in the audio signals from Layer A. A high-frequency spectrum interpolator 910 can prevent phase cancellations.
[0133] An exemplary RF spectrum interpolation and crossfading 1002 is shown in Fig. 10. For example, using a high-pass filter 1006 and a low-pass filter 1004, the two signals S to be blended in the two channels (L and R) n,L and S n+1,L or S n,R and S n+1,R at a given cut-off frequency at which the cancellation takes place (for example, 2.5 kHz in the case previously discussed with respect to Fig. 3 discussed example) into a low frequency component and a high frequency component. The non-critical low frequency components are blended 1008 (as in relation to Fig. 8). The frequency components above the cutoff frequency are blended 1012 using an FFT 1010 (or DFT, DCT, etc.) without their phases. In the next step, the current phase of the signal is impressed on the amplitude sum with the larger mixing factor i. This is done by appropriately parameterizing the inverse FFT 1016 (or inverse DFT, DCT, etc.). Afterwards, the low and high frequency ranges are mixed / added again 1018 to produce a stereo signal L(ϕ). array and R(ϕ) array to generate. The Fig. 10 shows the signal processing as an example for obtaining the signal L(ϕ) array of the left channel. A corresponding signal processing also takes place to obtain the signal R(ϕ) array of the right channel.
[0134] The RF spectrum interpolation 910 can prevent dips in the high-frequency range when the user turns their head. The RF spectrum interpolation 910 is particularly suitable, for example, for transmissions / recordings in which the microphone array is used exclusively. The RF spectrum interpolation 910 can, for example, be selectively activated or deactivated in the player 122. To do this, the producer 118 can, for example, indicate the activation and deactivation of the RF spectrum interpolation 910 in the metadata for Layer A. Alternatively or additionally, it is possible for the listener to activate or deactivate the RF spectrum interpolation 910 using the GUI of the player 122. In this case, specifications from the metadata can also be "overwritten" by the listener.
[0135] In the exemplary implementation, the binaural stereo signal of the channels L(ϕ) array and R(ϕ) arraya distance fader 912, 914. As explained, the metadata (see (3) f.) for Layer A can also contain one or more gain / attenuation factors 964 for the amplification / attenuation of the channels L(ϕ) array and R(ϕ) array A gain / attenuation factor (p medium , p far ) p close ) can be specified for each of several predefined (virtual) distances (e.g., Close, Medium, Far). In addition to the gain / attenuation factor(s) 964 (for a respective distance), one or more blending curves (e.g., their curve shape) can optionally be defined or specified, which can be used to blend 980 two gain / attenuation factors 964 when a (virtual) distance between predefined (virtual) distances is to be set. Medium Close Far Kurvenform p medium p close p far 1-5
[0136] The blend curves are used to blend the gain / attenuation factors (Gain) 964: p(d) = p i - f(1 - d) + p j · f(d) where d is the set distance between the two directly adjacent distances (Close and Medium or Medium and Far), f(x) describes the curve shape, and p i and p j are the gain factors of the two directly adjacent distances. f(1 - d) = 0 if the set distance d is equal to the distance of the gain factor p j corresponds, or (d) = 0, if the set distance d is the distance of the gain factor p i The resulting gain factor p(d) is distributed evenly across the left channel L(ϕ) array and right channel R(ϕ) array applied.
[0137] The crossfade curves (or functions) can, for example, be a selection of one or more functions from the group: linear functions, exponential functions, cosine functions, sine functions, etc. The crossfade curves allow for various transitions between the distance levels and can also be selected differently for the left and right channels. In this case, a separate value of p(d) would be determined for each channel and applied to that channel.
[0138] A corresponding regulation 938, 942 of the distance can also be used for the stereo signals generated from the audio channels of Layer B (e.g. L B1 , R B1 ) or HL audio channels (e.g. L HL1 , R HL1) can be used. For this purpose, the metadata (see (3) c.) for Layer B can contain one or more gain / attenuation factors 918, which provide a corresponding blending 932 of the gain / attenuation factors 918 for a desired distance d in an analogous manner as in connection with the distance blending of the binaural stereo signals L(ϕ) array and R(ϕ) array described. Here, too, it is conceivable that the amplifiers 938, 942 apply the blended gain factor p(d) evenly to the left channel and right channel of the respective input stereo signal.
[0139] With distance blending 932, 938, for example, the level of a solo instrument (corresponding to an additional stereo signal of Layer B (e.g., stereo signal B1)) can increase less sharply between the distances from Medium to Close than the orchestral instruments (received as other additional stereo signals of Layer B) to avoid excessive presence. Likewise, the solo instrument can decrease in level more slowly at greater distances so that it can still be heard clearly at the distant position Far.
[0140] On the other hand, a signal's level can increase as one moves away from the virtual sound body. For example, the distance blend 942 can be used to slightly amplify an HL stereo signal from a room microphone system as the distance increases, simulating a more convincing increase in distance.
[0141] Optionally, the player 122 can also apply a distance-dependent equalizer setting 920, 966 to each additional stereo signal / HL signal of layers A and B (see equalizers 914, 940, and 944). For each of several predefined (virtual) distances (e.g., close, medium, far), a corresponding equalizer setting 920, 966 can be specified in the metadata (see (3) d.). The equalizer settings 920, 966 can, in particular, affect two or more bands of an equalizer 914, 940, 944. The player 122 can use the equalizer settings 920, 966, for example, in addition to the level changes 912, 938, 942. The equalizer settings 920, 966 can be used, for example, to simulate the attenuation of the air at high frequencies at long distances. For each distance (e.g.Close, Medium and Far) in the equalizer settings 920, 966 two or more sets of EQ parameters (for 2 or more EQ channels) can be specified, which define the frequency, an amplification factor (Gain) and a quality factor (Q), as well as a filter type (e.g. Bell, Tail or Cut Off):. Entfernung Medium Close Far Band EQ 1m EQ 2m EQ 1e EQ 2c EQ 1f EQ 2f Frequenz F 1m F 2m F 1c F 2c F 1f F 2f Gain G 1m G 2m G 1c G 2c G 1f G 2f Güte Q 1m Q 2m Q 1c Q 2c Q 1f Q 2f Kurvenform B, T, C 1m B, T, C 2m B, T, C 1c B, T, C 2c B, T, C 1f B, T, C 2f
[0142] Equalizers 914, 949, and 944 apply the respective equalizer settings to both channels of the corresponding input stereo signals. Here, too, the individual parameters of equalizer settings 920 and 966 for a distance d can be linearly blended / averaged into each other in the manner already described to obtain the corresponding parameters applied by equalizers 914, 949, and 944 for the distance d.
[0143] As already mentioned in connection with Fig. 8, are also used in the player 122 of the Fig. 9 the audio signals of each stereo group ((e.g. S B1,L and S B1,Rfor stereo signal B1) according to the mixing factors 802-2 assigned to the corresponding angular position to a left channel and a right channel (e.g. channels L(ϕ) B1 and R(ϕ) B1 ). If, however, the head angle ϕ does not correspond to any of the predefined N angle positions, the mixing factors for the stereo group (e.g. for stereo signal B1) are first blended according to the head angle ϕ and the blended mixing factors are used to mix the audio signals S B1,L and S B1,R the stereo group (stereo signal B1) to the channels L(ϕ) B1 and R(ϕ) B1 to distribute (cf. amplifiers 840-1, 840-2, 840-3, 840-4 and adder elements 816, 818). The channels L(ϕ) B1 and R(ϕ) B1 form a resulting, head-angle-dependent stereo signal. The channels L(ϕ) B1 and R(ϕ) B1can optionally be fed to equalizers 936-1 and 936-2, respectively. These are also referred to here as "angular position equalizers." The metadata includes, for example, dual-band equalizer settings 916-1 and 916-2 for channels L(ϕ) B1 and R(ϕ) B1 (or for each stereo signal of Layer B). The equalizer settings 916-1, 916-2 are specified for both channels for each of the N angular positions. The equalizer settings can be parametric or semi-perimetric settings. In particular, the equalizer settings can affect three or more bands of an equalizer 936-1, 936-2. These equalizer settings can include a filter for the left channel L(ϕ) for each of the N angular positions. B1 and the right channel R(ϕ) B1Equalizers 936-1 and 936-2 can be implemented in three bands. Equalizers 936-1 and 936-2 can, for example, simulate shadows of the ear facing away from the sound source. A filter type can be selected for each band (e.g., Bell, Tail, or Cut Off). The distance from the sound source does not affect the equalizer settings 916-1 and 916-2. For three-band equalizers, metadata 916-1 and 916-2 can contain the following parameters for each of the N angle perspectives: Linker Kanal Rechter Kanal Band EQ 1l EQ 2l EQ 3l EQ 1r EQ 2r EQ 3r Frequenz F 1l F 2l F 3l F 1r F 2r F 3r Gain G 1l G 2l G 3l G 1r G 2r G 3r Güte Q 1l Q 2l Q 3l Q 1r Q 2r Q 3r Kurvenform B, T , C 1l B, T , C 2l B, T , C 3l B, T , C 1r B, T, C 2r B, T, C 3r
[0144] The number of parameters in metadata 916-1, 916-2 also depends accordingly on the number of bands of equalizers 936-1, 936-2. The output signals of equalizers 936-1, 936-2 can be fed to amplifier 938 for subsequent optional distance adjustment.
[0145] If layer B contains multiple stereo groups (and / or mono signals), these are processed in player 122 as described for stereo signal B1 (or as shown for the mono signal), and the resulting head-angle-dependent stereo signals are added (see adding elements 820, 822). The two channels of this sum signal can also be fed to an optional stereo base control 946 (MS wide / narrow). This can be used, for example, to adapt the overall base width of the sum signal of the adding elements 820, 822 to the listener's head diameter. Metadata 954 can contain the settings for the stereo width control 946.
[0146] If Layer B contains one or more HL stereo signals, such as the HL stereo signal (HL Stereo HL1) with the audio signals S HL1,L and S HL1,R , these are combined with the other signals (here: L(ϕ) array and R(ϕ) array and L(ϕ) B1 and R(ϕ) B1) are added (cf. adding elements 804, 806 and adding elements 824, 826). The resulting stereo signal can be fed to the optional output equalizer 952. The output equalizer 952 can be, for example, a multi-band, in particular 3- or 4-band, equalizer. The output equalizer 952 can be used to adapt the stereo signal to the headphones 104 used. The output equalizer 952 can, for example, linearize the headphones 104 used. If necessary, presets of frequently used headphone models can be made available in the player 122, which can be selected via the GUI of the player 122. Additionally or alternatively, corresponding equalizer settings 950 can also be provided in the metadata (cf. (3) g.) that the player 122 uses. These equalizer settings 950 can be set for each band EQ6, EQ7, EQ8, etc.The equalizer settings 950 can define the frequency, an amplification factor (Gain) and a quality factor (Q), as well as a filter type (e.g. Bell, Tail or Cut Off) for each band (here for a 4-band equalizer):. Equalizer-Parameter Band EQ 6 EQ 7 EQ 8 EQ 9 Frequenz F6 F7 F8 F9 Gain G6 G7 G e G9 Güte Q6 Q7 Q8 Q9 Kurvenform B, T, C6 B, T, C7 B, T, C8 B, T, C9
[0147] The stereo output signal of the output equalizer 952 can also be fed to the optional limiter 828, which can be used to prevent clipping. The limiter 823 can be adjusted via system metadata (see (3) g.). As described in connection with Fig. As described in section 8, potential peak levels above -1 dBFS are intercepted and reduced. This is achieved with very low latency.
[0148] The output signal of the limiter 828 with the channels L(ϕ) Kopf and R(ϕ) Kopfforms the immersive output stereo signal that is fed to headphones 104. The result of the mixing of audio signals from Layer A and additional signals in Layer B is a spatial, three-dimensional auditory impression with strong AKL and presence, which is further enhanced by head movements. The ITDs of the array, i.e. the phase difference between the left and right microphones up to approximately 1.5 kHz, are particularly crucial for AKL and the perception of spatial depth. The presence of a sound body is achieved by the additional stereo audio signals in Layer B. In addition to location, the audio signals from Layer B play a decisive role in the "sound," especially when main microphones are used. The described system, in particular the data format and the player 130, make it possible to create a sonic quality and presence of the headphone signal that cannot be achieved with HRFT-based processes such as Dolby Atmos.Tangible proximity and intimacy, opulent sonic quality, and enveloping spatiality can be achieved simultaneously. As described above, the listener can choose their distance from the sound body and decide between richness of detail and spatiality.
[0149] In addition to concert recordings, live broadcasts with and without video or audio-only recordings, and music elements in VR games, many other applications are conceivable in the context of museums, installations, and gaming. For example, as a walk-in concert event in a VR game or in an empty concert hall. The headphones 104 can also be part of VR glasses or the like.
Claims
[1] Audio container for generating an output stereo signal (808) with the channels L and R based on M audio channels (A), with M ≥ 4, which are assigned to M predetermined angular perspectives lying in a plane, and at least one further mono or stereo audio channel (B1), wherein the audio container defines the following layers: a first layer for the M audio channels (A) assigned to the M angle perspectives; a second layer for the at least one mono or stereo audio channel (B1); a third layer with first metadata containing a metadata set for each of the at least one mono or stereo audio channel (B1) of the second layer, wherein each metadata set defines the mixing ratios of the respective mono or stereo audio channel (B1) of the second layer for each of predetermined N angular perspectives, with N ≥ 6. [2] The audio container of claim 1, wherein the first metadata corresponds to a first predetermined virtual distance from a sound source, and the third layer further comprises second metadata and / or third metadata corresponding to a second predetermined virtual distance and a third predetermined virtual distance from the sound source; wherein the second predetermined virtual distance is smaller than the first predetermined virtual distance from the sound source and the third predetermined virtual distance is greater than the first predetermined virtual distance from the sound source; and wherein the second metadata and / or the third metadata contain / contain a metadata set for each of the at least one mono or stereo audio channel (B1) of the second layer, wherein each metadata set defines the mixing ratio of the respective further mono or stereo audio channel (B1) of the second layer for the predetermined N angular perspectives. [3] Audio container according to claim 1 or 2, wherein each meta-data record for a stereo audio channel of the second layer identifies four parameter values for each of the N angular perspectives, the four parameter values for a respective angular perspective specifying: - a first mixing factor (i 1L ) for the first channel of the stereo audio channel of the second layer and a second mixing factor (i 2L ) for the second channel of the stereo audio channel of the second layer for the channel L of the output stereo signal (808), and - a third mixing factor (i 1R ) for the first channel of the stereo audio channel of the second layer and a fourth mixing factor (i 2R ) for the second channel of the stereo audio channel of the second layer for the channel R of the output stereo signal (808). [4] Audio container according to one of claims 1 to 3, wherein each meta-data record for a mono audio channel of the second layer identifies two parameter values for each of the N angular perspectives, the two parameter values for a respective angular perspective specifying: - a first mixing factor (i L ) of the mono audio channel of the second layer for the channel L of the output stereo signal (808), and - a second mixing factor (i R ) of the mono audio channel of the second layer for the channel R of the output stereo signal (808). [5] Audio container according to claim 3 or 4, wherein each meta-data record further comprises one or more of the following information: for each of the N angular perspectives, parametric equalizer settings to simulate a head-related or outer ear transfer function, HRTF, for both channels of a stereo signal generated from a mono or stereo audio channel (B1) of the second layer; and / or for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying the stereo signal generated from the mono or stereo audio channel (B1) of the second layer, and one or more crossfading curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source; and / or for the at least three specified virtual distances from a sound source, a parametric equalizer setting for simulating distance-dependent sound absorption by air for the stereo signal generated from the mono or stereo audio channel (B1) of the second layer. [6] Audio container according to one of claims 1 to 5, wherein the third layer further comprises metadata for the M audio channels (A) of the first layer, wherein the metadata for the M audio channels (A) further comprises one or more of the following information: a rotation offset for aligning the M angular perspectives with respect to a sound source represented by the M audio channels (A) of the first layer; a divergence control parameter for decreasing or increasing the binaural angle of the M audio channels (A) of the first layer; for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying a stereo signal generated from the M audio channels (A), and one or more crossfading curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source; and / or for the at least three specified virtual distances from a sound source, a parametric equalizer setting for simulating distance-dependent sound absorption by air for the stereo signal generated from the M audio channels (A). [7] Audio container according to one of claims 1 to 6, wherein the audio container further comprises information indicating for each audio channel of the second layer whether it is a mono channel or a stereo channel. [8] Audio container according to one of claims 1 to 7, wherein the audio container further comprises in a fourth layer or as part of the second layer at least one head-locked stereo audio channel (HL1), both channels of which are assigned to the L and R channels of the output stereo signal (808). [9] The audio container of claim 8, wherein the third layer further comprises metadata for the at least one head-locked stereo audio channel (HL1), wherein the metadata for the at least one head-locked stereo audio channel (HL1) further comprises one or more of the following information: for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying or attenuating the head-locked stereo audio channel (HL1), and one or more crossfading curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source; and / or for at least three specified virtual distances from a sound source, parametric equalizer setting for simulating distance-dependent sound absorption by air for the head-locked stereo audio channel (HL1). [10] Audio container according to one of claims 1 to 9, wherein the M angular perspectives correspond to the respective angles of M microphones in a spherical microphone array. [11] Audio container according to one of claims 1 to 10, wherein the M angular perspectives correspond to the angles i · 360° / M with i ∈ [0,1,2, ..., M - 1] and / or wherein the M angular perspectives correspond to the angles i · 360° / M with i E [0,1,2, ...,M - 1]. [12] Audio container according to one of claims 1 to 11, wherein the M audio channels are mono channels and two of the M audio channels (A), whose angular perspectives are offset by 180° from each other, form one of M / 2 binaural stereo channels. [13] Audio container according to one of claims 1 to 12, wherein each further stereo audio channel of the second layer is a stereo channel. [14] A method for generating an output stereo signal (808) having L and R channels for playback with a headphone (104), comprising: Receiving M audio channels (A), with M ≥ 4, which are assigned to M given angular perspectives lying in a plane Receiving at least one further mono or stereo audio channel (B1); Receiving metadata containing a metadata set for each of the at least one further mono or stereo audio channel (B1), wherein each metadata set defines the mixing ratios of the respective mono or stereo audio channel (B1) for N angular perspectives, with N ≥ 6; Determining a head angle ϕ of a listener wearing the headphones (104) in the plane relative to one of the N or M angular perspectives defining a reference direction, wherein the absolute head angle ϕ is determined based on sensor signals from at least one sensor; Generating a first stereo audio signal having a channel L and a channel R, wherein the channel L of the stereo audio signal is generated by a linear crossfading of two audio channels of the M audio channels (A) which are assigned angular perspectives that are closest to the head angle ϕ - 90° and the channel R of the stereo audio signal is generated by a linear crossfading of two audio channels of the M audio channels (A) which are assigned angular perspectives that are closest to the head angle ϕ + 90°; for each additional mono or stereo audio channel (B1), linear averaging of the mixing ratios (i 1Ln , i 1Lnn , i 1Rn , i 1Rnn , i 2Ln , i 2Lnn , i 2Rn , i 2Rnn) specified in those two metadata sets for the respective mono or stereo audio channel (B1) that are assigned to the two angular perspectives of the N angular perspectives (n, nn) that are closest to the absolute head angle ϕ in order to generate blended mixing ratios (i 1LΦ , i 1RΦ , i 2LΦ , i 2RΦ ); for each further mono or stereo audio channel (B1), generating a second stereo audio signal with a channel L and a channel R based on the respective further mono or stereo audio channel (B1) using the averaged mixing ratios (i 1LΦ , i 1RΦ , i 2LΦ , i 2RΦ ); Generating the output stereo signal (808) by summing the first stereo audio signal with the at least one generated second stereo audio signal. [15] The method of claim 14, wherein the at least one further mono or stereo audio channel (B1) is or contains a stereo audio channel, and the meta-data set for the stereo audio channel identifies four parameter values for each of the N angular perspectives, where the four parameter values indicate a respective angle perspective: - a first mixing factor (i 1L ) for the first channel of the stereo audio channel and a second mixing factor (i 2L ) for the second channel of the stereo audio channel for the channel L of the second stereo audio signal, and - a third mixing factor (i 1R ) for the first channel of the stereo audio channel and a fourth mixing factor (i 2R ) for the second channel of the stereo audio channel for the channel R of the second stereo audio signal; where the linear averaging of the mixing ratios for the stereo audio channel specified in the two meta-data sets associated with the two angular perspectives of the N angular perspectives (n, nn) closest to the absolute head angle ϕ comprises: - respective linear crossfading, depending on the head angle ϕ, of the first two mixing factors (i 1Ln , i 1Lnn ), the two second mixing factors (i 2Ln , i 2Lnn ), the two third mixing factors (i 1Rn , i 1Rnn ) and the two fourth mixing factors (i 2R , i 2Rnn ) specified in the two meta-data sets to create a blended first blending factor (i 1LΦ ), a blended second mixing factor (i 2RΦ ), a blended third mixing factor (i 1RΦ ) and a blended fourth mixing factor (i 2RΦ ) and wherein generating the second stereo audio signal based on the stereo audio channel comprises: - Generating the channel L of the second stereo signal by multiplying the first channel of the stereo audio channel by the blended first mixing factor (i 1LΦ ) and the second channel of the stereo audio channel with the blended second mixing factor (i 2LΦ ) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel; and - Generating the channel R of the second stereo signal by multiplying the first channel of the stereo audio channel by the blended third mixing factor (i 1RΦ ) and the second channel of the stereo audio channel with the blended fourth mixing factor (i 2RΦ ) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel. [16] The method of claim 15, further comprising: for each of the N angular perspectives, receiving a parametric equalizer setting to approximate a head-related or outer ear transfer function, HRTF, for the L and R channels of the second stereo signal; Filtering the L channel of the second stereo signal based on the parametric equalizer settings, which are faded in and out linearly according to the two angle perspectives of the N angle perspectives; and Filtering the R channel of the second stereo signal based on the parametric equalizer settings that correspond to the two angle perspectives of the N angle perspectives, which are faded in and out linearly. [17] The method of claim 15 or 16, further comprising: for each of at least three given virtual distances from the sound source, receiving a gain factor (g near , g medium , g far) for amplifying the second stereo signal, and one or more crossfade curves for calculating a gain factor (g dist ) for a virtual distance between two of the given virtual distances from a sound source; and Amplifying the second stereo signal depending on a selected virtual distance from the sound source based on one of the gain factors (g near , g medium , g far ) or a gain factor calculated using one of the crossfade curves (g dist ). [18] A method according to any one of claims 15 to 17, further comprising: for at least three specified virtual distances from the sound source, receiving a parametric equalizer setting (EQ near , EQ medium , EQ far ) for the second stereo signal; and Changing the sound image of the second stereo signal according to a selected virtual distance based on the received parametric equalizer settings (EQ near , EQ medium , EQ far ) for the second stereo signal. [19] A method according to any one of claims 14 to 18, further comprising: Receiving one head-locked stereo audio channel (HL1); for each of at least three given virtual distances from the sound source, receiving a gain factor (g HL_near , g HL_medium , g HL_far ) for amplifying the head-locked stereo audio channel (HL1), and one or more crossfading curves for calculating a gain factor (g HL_dist ) for a virtual distance between two of the given virtual distances from a sound source; and Amplifying the head-locked stereo audio channel (HL1) depending on a selected virtual distance from the sound source based on one of the gain factors (g HL_near , g HL_medium , g HL_far ) or a gain factor calculated using one of the crossfade curves (g HL_dist ). [20] The method of claim 19, further comprising: for at least three specified virtual distances from the sound source, receiving parametric equalizer settings (EQ HL_near , EQ HL_medium , EQ HL_far ) for the head-locked stereo audio channel (HL1); and Changing the sound image of the head-locked stereo audio channel (HL1) according to a selected virtual distance based on the received parametric equalizer settings (EQ HL_near , EQ HL_medium , EQ HL_far ) for the head-locked stereo audio channel (HL1). [21] The method of any one of claims 14 to 20, further comprising interpolating the high frequency spectrum of the first stereo audio signal. [22] A method according to any one of claims 14 to 21, further comprising: for each of at least three given virtual distances from the sound source, receiving a gain factor (g array-near , S array_medium , g array_far ) for amplifying the first stereo audio signal, and one or more crossfading curves for calculating a gain factor (g array_dist ) for a virtual distance between two of the given virtual distances from a sound source; and Amplifying the first stereo audio signal depending on a selected virtual distance from the sound source based on one of the gain factors (g array_near , g array_medium , g array_far ) or a gain factor calculated using one of the crossfade curves (g array_dist ). [23] The method of claim 22, further comprising: for at least three specified virtual distances from the sound source, receiving a parametric equalizer setting (EQ array_near , EQ array_medium , EQ array_far ) for the first stereo audio signal; and Changing the sound image of the first stereo audio signal according to a selected virtual distance based on the received parametric equalizer settings (EQ array_near , EQ array_medium , EQ array_far ) for the first stereo audio signal. [24] A computer-readable medium storing instructions which, when executed by a computing unit of a device, cause the device to carry out the method of any one of claims 14 to 23. [25] Device with a computing unit adapted to carry out the method according to one of claims 14 to 23. [26] Apparatus according to claim 25, wherein the apparatus is adapted to supply the generated output stereo signal having the L and R channels to a headphone (104) for reproduction. [27] Apparatus according to claim 25 or 26, adapted to receive and process one or more audio containers according to any one of claims 1 to 13 as a continuous data stream. [28] Device according to one of claims 25 to 27, wherein the computing unit is adapted to perform the method according to one of claims 14 to 23 based on the data in a plurality of audio containers according to one of claims 1 to 13. [29] Device according to one of claims 25 to 28, wherein the device includes one or more sensors which provide the sensor signals for determining the head angle ϕ of the listener wearing the headset (104) in the plane. [30] The device of claim 29, wherein the one or more sensors comprise one or more of the following sensors, accelerometer, gyroscope or magnetic field sensor, to detect a change in the orientation of the device; and / or wherein the one or more sensors further comprise a camera that detects a head movement of the head of the listener wearing the headset (104). [31] Device according to claim 30, wherein the computing unit is adapted to detect a change in the head angle of the listener from the sensor signals. [32] A system comprising a device according to any one of claims 25 to 31 and a headphone (104) for reproducing an output stereo signal (808) generated by the device. [33] The system of claim 32, wherein the device is implemented in a smartphone or tablet computer, or the headset (104).
Citation Information
Patent Citations
Hearing aid
US20110091056A1
Ambisonic depth extraction
US20190313200A1
Dynamic binaural sound capture and reproduction
US7333622B2