Device for generating an immersive stereo signal for playback via headphones and data format for the transmission of audio data
The use of a microphone array and metadata-driven mixing for headphones addresses issues in existing immersive audio methods, enhancing sound localization and immersion by accurately positioning sound sources based on head angles and distances.
Patent Information
- Application Number
- PCT/EP2024/084873
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-02
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-10
AI Technical Summary
Existing methods for generating immersive stereo audio for headphones, such as those using head-related transfer functions (HRTFs) and motion-tracked binaural recordings, fail to provide an optimal binaural sound experience due to issues like reduced low-frequency levels, altered sound quality, and inaccurate localization of sound sources, particularly in frontal planes.
A method involving a microphone array with M microphones arranged in a plane, capturing audio signals based on head angles, and generating stereo signals using metadata-driven mixing ratios and additional audio channels to enhance distance and direction localization, combined with metadata transmission for precise sound positioning.
Improves the binaural listening experience by providing accurate distance and direction localization of sound sources, enhancing the immersive audio experience with precise sound positioning and reduced comb filter effects.
Smart Images

Figure EP2024084873_10072025_PF_FP_ABST
Abstract
Description
[0001]Device for generating an immersive stereo signal for reproduction via headphones and data format for transmitting audio data. Technical field of the invention. The invention relates to the generation of stereo signals, in particular, of immersive and / or binaural stereo signals for reproduction via headphones. In some embodiments of the invention, the stereo signal is generated as a function of a detected head angle of the listener wearing the headphones. Furthermore, the invention relates to an audio container format for transmitting audio information for generating the stereo signal, as well as computer-readable media, devices, and systems for generating and / or reproducing the stereo signal. Technical background. The stereo audio format with two channels (often referred to as left and right channels) has long been standard in the audio industry. Multi-channel audio formats such as 5.1 and 7.1.1 have now become established in certain application areas, such as cinemas or home cinemas. Immersive multi-channel audio formats such as Dolby Atmos, Auro 3D, MPEG-H, DTS:X, NHK 22.2, etc. expand the horizontal dimension of the audio format to include the reproduction of sound sources from above. They require greater technical complexity and additional speakers or soundbars. The term "immersive audio" stands for a variety of formats in which sound emerges from all directions in the room, thus immersing the listener deeper into the "action". The creation of an immersive listening experience using headphones dates back to the 1980s with the introduction of artificial head stereophony, which, however, did not gain market acceptance. Due to the ubiquity of headphones and in-ears, there are renewed efforts to make immersive music experience possible using headphones.Common solutions typically use head-related transfer functions (HRTFs). HRTFs describe the complex filtering effect of the head, outer ear (pinna), and torso on sound arriving at the inner ear from different directions. These methods convolve the loudspeaker signals with the corresponding directional HRTFs, allowing the listener to localize them more or less accurately from that direction. In addition to generic HRTFs, individual HRTFs can also be generated, which can usually better match the listener's own listening experience. Some methods allow headphones to be tracked for rotation and tilt around the listener's head axis. With Apple Spatial Audio, the Dolby Atmos stream is converted to 7.1.2 format, and the HRTFs are then dynamically adjusted to the current speaker positions.All HRTF-based methods for creating an immersive listening experience using headphones involve a transformation from a playback for loudspeakers to a playback for headphones: Instead of a performance space, such as a concert hall, one hears the renderer's virtual listening room through the headphones, and instead of the instruments, the loudspeaker signals. Instead of sound source positions, the loudspeaker positions and their phantom sound sources are transmitted. The convolutions with the generic or individual HRTFs alter the original sound, which leads to a reduction in level, especially in the low-frequency range; the sound is often perceived as thin. The distance of the sound sources from the listener is determined during production but cannot be less than the specified loudspeaker plane, and satisfactory localization of sources in the frontal plane is impossible for most people.Another approach to generating a motion-dynamic and binaural audio signal that records binaurally at the recording location, i.e., actually head-related, is described in US Patent No. 7,333,622 B2, which is incorporated herein by reference. US Patent No. 7,333,622 B2 describes a method for capturing and playing back live or recorded 3D audio exclusively for headphone playback. This method is called "motion-tracked binaural" (MTB). It uses a microphone array with multiple microphones, a head-tracking headset, and special signal processing to combine the signals recorded by the microphones into a binaural audio signal that follows the rotation of the listener's head, creating a counter-rotating sound field. The microphone array is shaped like a sphere roughly the size of a head, in which 8, 16, or 32 individual microphones (microphone capsules) are flush-mounted and evenly arranged in a plane.Alexander Lindau and Sebastian Roos, "Perceptual evaluation of discretization and interpolation for motion-tracked binaural (MTB) recordings," report of the 26th Tonmeistertagung, pages 660-701, were able to show that when using high-frequency spectrum interpolation (HF-SP), doubling or quadrupling the number of microphones from eight does not result in a measurable improvement in localization and timbre. Summary of the invention: Even if a fairly high degree of realism is achieved, the solution provided in US 7,333,622 B2 still does not provide the listener with an optimal binaural sound result. In the solution provided in US 7,333,622 B2 (as well as in some embodiments of the invention), two opposing microphones form a binaural audio pair.The spherical separator of the microphone array, on which the eight microphones are arranged in a horizontal plane, generates characteristic ILD (Interaural Level Differences) and ITD (Interaural Time Differences), similar to natural hearing. This microphone array is positioned at a suitable distance in front of a sound body. Its function is that of out-of-head localization (AKL) and the location of sound sources during headphone playback. In sound engineering (and in the embodiments), sound sources are referred to as musical sound bodies, e.g., a piano, string quartet, orchestra, or band. As with an artificial head, sound sources are located around the head. Two opposing microphones form an audio pair, which, as with an artificial head, represents the ears. With eight microphones, the audio angle pairs shown in Fig. 11b result: 0°, +45°, +90°, +135°, 180°, -135°, -90°, and -45°. with 0° as the reference axis, e.g.the central, horizontal line of sight to the center of the stage or a head position frontal to the sound body. The corresponding microphones of a stereo pair are offset by -90° (left ear) and +90° (right ear) to the head angle, as illustrated in Fig. 11b for a head angle of 0° and -45°. The term “head angle” is understood here to be the angle of head rotation in a plane relative to a reference axis. This reference axis of head rotation can coincide with the reference axis relative to which M angular positions are defined, or it can have an offset (e.g. configurable) to it. If an audio pair of the microphone array is transmitted to headphones – i.e. one microphone to each channel of the headphones – a sound source creates a realistic out-of-head localization (AKL) for the listener, the horizontal angle of which corresponds to natural perception.Out-of-head localization means that the sound source is perceived outside of one's head. In contrast, in-head localization (IHL) occurs, where sources are localized inside one's head. This is often experienced with normal stereo recordings or with monophonic signals heard through headphones. Due to the lack of ear cups on the microphone array in the solution of US patent US 7,333,622 B2, the diffuse sound is reproduced at the expense of the direct sound component, which reduces the frontal presence. Sound sources appear further away than is the case with natural hearing. To compensate for this excessive distance impression, one could move the microphone array closer to the sound source to receive more direct sound. However, this makes the sound source appear wider than the perceived proximity would suggest.One object of the invention is to provide solutions that improve the binaural listening experience in order to create the most natural listening impression possible with precise distance localization. A further object of the invention is to provide solutions that improve the binaural listening experience in order to create the most natural listening impression possible with precise direction and distance localization. One aspect of the invention is therefore to improve the distance localization of the sound source (e.g., an orchestra, a band, etc. on a stage). In addition to M audio channels (with ^^ ≥ 4 ) assigned to angular perspectives, one or more further mono or stereo audio channels (e.g., from spot microphones and / or main microphones used in the conventional recording of the sound source) are received.A stereo audio signal is generated from the M audio channels (also Layer A) depending on a detected / determined head angle of the listener, to which a further stereo audio signal is mixed, which is also generated depending on the detected / determined head angle and optionally depending on a selectable (virtual) distance from the sound source from the further one or more mono or stereo audio channels. The respective mixing ratio with which a further stereo audio signal generated from a further mono or stereo audio channel is mixed to the stereo signal generated from the M audio channels can depend on the detected / determined head angle of the listener and optionally on a selectable (virtual) distance from the sound source. A further aspect of the invention relates to a transmission format or a data structure, e.g., an audio container, by means of which the information required for binaural headphone playback (e.g.,Metadata) and audio signals of the individual audio channels can be transmitted. In order to implement the mixing of the additional mono or stereo audio channel(s), metadata is transmitted using a data structure / audio container, which contains a metadata set for each of the at least one additional mono or stereo audio channel. Each metadata set identifies the mixing ratios of the respective mono or stereo audio channel of the second layer (also Layer B) for predefined N angular perspectives. The N angular perspectives of the additional mono or stereo audio channel(s) (Layer B) can result from practical considerations. For example, when panoramas of the mono and / or stereo channels (Layer B), the number N of angular perspectives should not be less than 8 (i.e. ^^ ≥ 8) in order to achieve a continuous, flowing panoramic movement against the head rotation. Also, in the case of a frontal sound event (i.e.the sound source is mostly located at 0° as the reference direction) +90° / -90° must be included as the angular position, since this is where the panorama is most critical. The number N can therefore be chosen, for example, as ^^ = (^^ + 1) ∙ 4,^^^^^^ ^^ ∈[1,2,3, … ]. The N angular perspectives of layer B can match the M angular perspectives of layer A (ie ^^ = ^^); however, this is not mandatory, i.e. ^^ can also be used. ^^ apply. In some embodiments, ^^ = ^^ =8 can apply. If the mixing of the further mono or stereo audio channel(s) is to be distance-dependent, several such metadata sets can be provided for each further mono or stereo audio channel, which define the mixing ratio(s) of the respective mono or stereo audio channel of the second layer for each of the predetermined N angular perspectives for different (in particular three) predetermined virtual distances from the sound source. A further aspect of the invention relates to the implementation of the above aspects and their embodiments described herein in software and / or hardware, and to a playback system that generates a stereo audio signal that is played back by means of headphones and enables a binaural listening experience.Some embodiments provide a method for generating an output stereo signal (in particular a binaural stereo signal) with the L and R channels for playback with headphones. The method can be carried out by a device for generating an output stereo signal with the L and R channels, which device comprises a computing unit (e.g., a processor). The method includes receiving M audio channels, with ^^ ≥ 4, which are assigned to M predetermined angular perspectives lying in a plane, and receiving at least one further mono or stereo audio channel. Furthermore, metadata is received which contains a metadata set for each of the at least one further mono or stereo audio channel.The metadata serves to acoustically position the auditory event of the respective additional mono or stereo audio channel relative to an auditory event generated from the M audio channels for the individual M angular positions (similar to "panoramicization"). For this purpose, for example, each metadata set can define the mixing ratio of the respective mono or stereo audio channel for each of the N angular perspectives. The method further comprises determining a head angle ^^ of a listener wearing the headphones. The head angle ^^ can be determined in a plane relative to one of the N angular perspectives and / or the M angular perspectives that define a reference direction. The determination of the head angle ^^ can be based on sensor signals from at least one sensor. In one embodiment, the one or more sensors form a head-tracking sensor system, the sensor signal(s) of which enables the determination of the head angle ^^.The method generates a first binaural stereo audio signal (hereinafter also binaural audio signal) with one channel L and one channel R. The channel L of the first stereo audio signal is generated, for example, by a linear crossfading of two audio channels of the M audio channels (in layer A), which are assigned angular perspectives that are closest to the head angle ^^ − 90°. Accordingly, the channel R of the first stereo audio signal is generated, for example, by a linear crossfading of two audio channels of the M audio channels, which are assigned the M angular perspectives that are closest to the head angle ^^ + 90°. It is assumed here that the ears of the head are at an angle of −90° or +90° relative to the head angle ^^. If the head angle ^^ ± 90° corresponds to one of the M angle positions, the audio channels of the two angle positions (^^ − 90° and ^^ + 90°) can be used for the L and R channels of the first stereo audio signal (ieCrossfading of audio channels is not necessary in this case). In order to mix the one or more mono or stereo audio channels into the first stereo audio signal, the mixing ratios are first determined for each additional mono or stereo audio channel (e.g. for a stereo audio channel: mixing ratios i. 1Ln , i 1Lnn , i 1Rn , i 1Rnn , i 2Ln , i 2Lnn , i 2Rn , i 2Rnn for angle perspectives n and nn from the N angle perspectives) specified in those two metadata sets for the respective mono or stereo audio channel, which are assigned to the two angle perspectives n and nn of the N angle perspectives that are closest to the absolute head angle ^^, (linearly) crossfaded to create crossfaded mixing ratios (e.g. for a stereo audio channel: crossfaded mixing ratios: i 1Lф =f(i 1Ln , i 1Lnn ), i 1Rф =f(i 1Rn , i 1Rnn),i2Lф=f(i2Ln, i2Lnn), i2Rф=f(i2Rn, i2Rnn), where f(x,y) represents a (linear) function of the mixing ratios x and y). The (linear) crossfading can, for example, combine the respective two mixing ratios according to the respective (angular) distance of the head angle ^^ from the two nearest angular positions n and nn (nearest neighbor n and next-nearest neighbor nn). For each additional mono or stereo audio channel, a second stereo audio signal with an L channel and an R channel is also generated. The second stereo audio signal can be generated based on the respective additional mono or stereo audio channel using the crossfaded mixing ratios. Subsequently, the output stereo signal can be generated by mixing (summing) the first stereo audio signal with the at least one generated second stereo audio signal.In a further embodiment, the at least one further mono or stereo audio channel includes a stereo audio channel or is a stereo audio channel. The metadata set for the stereo audio channel identifies four parameter values for each of the N angular perspectives. The four parameter values specify, for a respective angular perspective of the N angular perspectives: ^a first mixing factor (i1L) for the first channel of the stereo audio channel and a second mixing factor (i2L) for the second channel of the stereo audio channel for channel L of the second stereo audio signal, and^ a third mixing factor (i1R) for the first channel of the stereo audio channel and a fourth mixing factor (i. 2R) for the second channel of the stereo audio channel for the channel R of the second stereo audio signal; The above-described (linear) crossfading of the mixing ratios depending on the head angle ^^ for such a stereo audio channel includes a respective linear crossfading of the two first mixing factors (i 1Lф =f(i 1Ln , i 1Lnn )), the two second mixing factors (i 2Lф =f(i 2Ln , i 2Lnn )), the two third mixing factors (i 1Rф =f(i 1Rn , i 1Rnn )) and the two fourth blending factors (i2Rф=f(i2Rn, i2Rnn)) specified in the two metadata sets to form a blended first blending factor (i1Lф), a blended second blending factor (i 2Lф), a cross-faded third mixing factor (i1Rф) and a cross-faded fourth mixing factor (i2Rф). Accordingly, this can comprise generating the second stereo audio signal based on the one stereo audio channel: generating the channel L of the second stereo signal by multiplying the first channel of the stereo audio channel by the cross-faded first mixing factor (i1Lф) and the second channel of the stereo audio channel by the cross-faded second mixing factor (i 2Lф ) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel; and generating the channel R of the second stereo signal by multiplying the first channel of the stereo audio channel by the blended third mixing factor (i 1Rф ) and the second channel of the stereo audio channel with the blended fourth mixing factor (i 2Rф) and subsequent addition of the first channel of the stereo audio channel and the second channel of the stereo audio channel . In a further embodiment, the method further comprises, for each of the N angular perspectives, receiving a parametric equalizer setting for the L and R channels of the second stereo signal. For example, the parametric equalizer setting for a respective one of the two L and R channels can approximate or simulate a head-related or outer ear transfer function (HRTF). The equalizer setting can be a multi-band equalizer setting, in particular an equalizer setting for at least three frequency bands. Based on the equalizer settings, the L channel of the second stereo signal is filtered. During filtering, the two equalizer settings assigned to the two angular perspectives n and nn that are closest to the absolute head angle ^^ can be used.These two equalizer settings are faded in or out linearly according to the distance of the head angle ^^ from the respective one of the two angular perspectives n and nn. Channel R of the second stereo signal is also filtered accordingly: Here, too, the two equalizer settings assigned to the two angular perspectives n and nn that are closest to the absolute head angle ^^ can be used for filtering. These two equalizer settings are faded in or out linearly according to the distance of the head angle ^^ from the respective one of the two angular perspectives n and nn. In a further embodiment, a gain factor (gnear, gmedium, gfar) is received for each of at least three predetermined virtual distances from the sound source. The gain factor (g. near , g medium , g far) is provided for amplifying the second stereo signal. Likewise, one or more crossfade curves can be received for calculating a gain factor (gdist) for a virtual distance between two of the predefined virtual distances from a sound source. If no crossfade curve(s) are included in the received data, a predefined crossfade curve can be used (e.g. a linear crossfade). The second stereo signal can be amplified accordingly depending on a selected virtual distance from the sound source based on one of the gain factors (gnear, gmedium, gfar) (if the selected virtual distance corresponds to one of the predefined distances) or a gain factor (gdist) calculated using one of the crossfade curves (if the selected virtual distance does not correspond to any of the predefined distances).In a further embodiment, a respective parametric equalizer setting (EQ) is used for at least three predetermined virtual distances from the sound source. near , EQmedium, EQfar) for the second stereo signal. This can be used to adjust the sound image of the second stereo signal according to a selected virtual distance based on the received parametric equalizer settings (EQ near , EQ medium , EQfar) for the second stereo signal. For example, using the equalizer settings, the sound of the second stereo signal can be adjusted to simulate the high-frequency attenuation of the air at great distances. In a further embodiment, a head-locked stereo audio channel is also received. In addition, for each of at least three predetermined virtual distances from the sound source, an amplification factor (g HL_near , g HL_medium , g HL_far) are received. The gain factor (gHL_near, gHL_medium, gHL_far) can be used to amplify the head-locked stereo audio channel. Furthermore, one or more crossfade curves can optionally be used to calculate a gain factor (g HL_dist ) for a virtual distance between two of the predefined virtual distances from a sound source. The head-locked stereo audio channel can be adjusted depending on a selected virtual distance from the sound source based on one of the gain factors (gHL_near, gHL_medium, gHL_far) (if the selected virtual distance corresponds to one of the predefined distances) or a gain factor calculated using one of the crossfade curves (g HL_dist) (if the selected virtual distance does not correspond to any of the predefined distances). In a further embodiment, a respective parametric equalizer setting (EQHL_near, EQ HL_medium, EQ HL_far) for the head-locked stereo audio channel can be received for the at least three predefined virtual distances from the sound source. These can be used to effect a change in the sound image of the head-locked stereo audio channel according to a selected virtual distance. For example, with the aid of the equalizer settings, the sound image of the head-locked stereo audio channel can be adjusted such that a stronger spatial effect is achieved by boosting the Blauert band by 1 kHz. Optionally, according to a further embodiment of the invention, an interpolation of the high-frequency spectrum of the first stereo audio signal can be implemented.This can be used to compensate for comb filter effects that occur due to the phase difference in the M audio channels as a result of the spatial offset of the microphones when recording the sound source. The interpolation of the high-frequency spectrum of the first stereo audio signal can comprise high-pass filtering of the signals, decomposition of the high signal components into amplitudes and phases, and amplitude addition and introduction of the phase of a microphone. This signal is then added again to the low-frequency signal component. In a further embodiment, an amplification factor (garray_near, garray_medium, garray_far) is received for each of at least three predetermined virtual distances from the sound source. The amplification factor (garray_near, garray_medium, garray_far) can be used to amplify the first stereo audio signal.Furthermore, one or more crossfading curves for calculating a gain factor (garray_dist) for a virtual distance between two of the predefined virtual distances from a sound source can optionally be received. Furthermore, the first stereo audio signal can be adjusted as a function of a selected virtual distance from the sound source based on one of the gain factors (garray_near, garray_medium, g. array_far) (if the selected virtual distance corresponds to one of the predetermined distances) or a gain factor calculated using one of the crossfade curves (garray_dist) (if the selected virtual distance does not correspond to any of the predetermined distances). In a further embodiment, a respective parametric equalizer setting (EQarray_near, EQ array_medium, EQ array_far) for the first stereo audio signal is received for the at least three predetermined virtual distances from the sound source. The sound image of the first stereo audio signal can be changed according to a selected virtual distance based on the received parametric equalizer settings (EQarray_near, EQ array_medium, EQ array_far) for the first stereo audio signal.Further embodiments relate to one or more computer-readable medium(s) storing instructions which, when executed by a computing unit of a device, cause the device to carry out the method described herein for generating an output stereo signal. Another embodiment relates to a device adapted to carry out the method described herein for generating an output stereo signal. For this purpose, the device can be equipped with a computing unit that executes the method for generating an output stereo signal according to any of the embodiments of the method described herein. The device can be, for example, a user terminal, in particular a smartphone or tablet computer. The device can be adapted accordingly to supply the generated output stereo signal with the L and R channels to a headphone for playback.Alternatively, the device can also be integrated into headphones. A further embodiment relates to a device adapted to receive and process one or more audio containers according to any one of the embodiments described below to generate the output stereo signal. The audio containers can be received as part of a continuous data stream (streaming). In a further embodiment, the computing unit of the device is adapted to carry out the method for generating an output stereo signal according to any one of the embodiments described herein based on the data in a plurality of audio containers according to any one of the embodiments described herein below.In one embodiment, the device can contain one or more sensors that at least partially provide the sensor signals for determining the head angle ^^ of the listener wearing the headphones in the plane. The head angle ^^ is specified relative to one of the M angular perspectives of the stereo signal of layer A or relative to one of the N angular perspectives of the further mono or stereo signal(s) (layer B). Optionally, it is therefore conceivable that at least one further sensor signal from a sensor that is not part of the device itself is taken into account for determining the head angle ^^ of the listener. The one or more sensors can, for example, comprise one or more of the following sensors: acceleration sensor, gyroscope, or magnetic field sensor, in order to detect a change in the orientation of the device.Additionally or alternatively, the one or more sensors may comprise a camera to detect a head movement of the listener wearing the headphones. The computing unit of the device may, for example, be further adapted to detect a change in the listener's head angle from the sensor signals. Alternatively, the listener's head angle may also be provided to the device directly by a head-tracking sensor system. Furthermore, embodiments relate to a system that may comprise a device according to any of the embodiments described herein and headphones for reproducing an output stereo signal generated by the device. A further aspect of the invention relates to a data structure and, in particular, to an audio container.Embodiments therefore also relate to an audio container for generating an output stereo signal with the L and R channels based on M audio channels, with ^^ ≥ 4, which are assigned to M predetermined angular perspectives lying in a plane, and at least one further mono or stereo audio channel. The audio container can define the following layers: ^a first layer (layer A) for the M audio channels assigned to the M angular perspectives;^ a second layer (layer B) for the at least one mono and / or stereo audio channel and optionally further mono and / or stereo audio channels assigned to N angular perspectives (e.g. ^^ ≥ 8);^ a third layer with metadata;^ a fourth layer (optional) for one or more HL stereo signals (note: the HL stereo signals can also be considered part of layer B).The metadata serves to acoustically position the auditory event of the respective further mono or stereo audio channel relative to an auditory event (similar to a "panorama"). The metadata can contain first metadata that contains a metadata set for each of the at least one mono or stereo audio channel of the second layer. Each metadata set can define the mixing ratio or mixing ratios of the respective mono or stereo audio channel of the second layer for each of the N predefined angular perspectives. In a further embodiment of the audio container, the first metadata corresponds to a first predefined virtual distance from a sound source. The third layer can further comprise second metadata and / or third metadata that correspond to a second predefined virtual distance and a third predefined virtual distance from the sound source.The second specified virtual distance can be smaller than the first specified virtual distance from the sound source and the third specified virtual distance can be greater than the first specified virtual distance from the sound source. The second metadata and / or the third metadata for each of the at least one mono or stereo audio channel of the second layer can further contain a metadata set. Each metadata set can define the mixing ratio of the respective further mono or stereo audio channel of the second layer for the specified N angular perspectives. In a further embodiment, each metadata set for a stereo audio channel of the second layer (B) identifies four parameter values for each of the N angular perspectives.The four parameter values can specify for a respective angular perspective: ^a first mixing factor (i1,L) for the first channel of the stereo audio channel of the second layer and a second mixing factor (i2,L) for the second channel of the stereo audio channel of the second layer for the channel L of the output stereo signal, and^ a third mixing factor (i1,R) for the first channel of the stereo audio channel of the second layer and a fourth mixing factor (i. 2,R) for the second channel of the stereo audio channel of the second layer for channel R of the output stereo signal. In another embodiment, each metadata record for a mono audio channel of the second layer identifies two parameter values for each of the N angular perspectives. The two parameter values for a respective angular perspective specify: ^a first mixing factor (iL) of the mono audio channel of the second layer for channel L of the output stereo signal, and ^a second mixing factor (iR) of the mono audio channel of the second layer for channel R of the output stereo signal. In a further embodiment, each metadata record may further comprise one or more of the following information for each of the N angular perspectives: parametric equalizer settings.The equalizer settings can, for example, relate to a simulation or approximation of a head-related or outer-ear transfer function (HRTF) for the two channels of a stereo signal generated from a mono or stereo audio channel of the second layer. Alternatively or additionally, each metadata set can contain a gain factor for amplifying the stereo signal generated from the mono or stereo audio channel of the second layer for each of at least three predefined virtual distances from a sound source. Additionally, one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predefined virtual distances from a sound source can be included in each metadata set.Furthermore, each metadata set can additionally or alternatively contain, for the at least three predetermined virtual distances from a sound source, a parametric equalizer setting for simulating distance-dependent sound absorption by air for the stereo signal generated from the mono or stereo audio channel of the second layer. Applying the equalizer settings depending on a selected virtual distance from the sound source can, for example, result in an attenuation of the treble of the stereo signal generated from the mono or stereo audio channel of the second layer in order to simulate air absorption of the sound waves. In a further embodiment, the third layer can further comprise metadata for the M audio channels of the first layer.The metadata for the M audio channels can further include one or more of the following information: ^A rotation offset for aligning the M angular perspectives with respect to a sound source, which are represented by the M audio channels of the first layer. If the sound source was recorded, for example, with a microphone array with M microphones arranged in a plane, the alignment angle of the microphone array with respect to the sound source could be subsequently set or adjusted using the rotation offset. ^A divergence-controlling parameter. This parameter can be used, for example, to decrease or increase a binaural angle of the M audio channels of the first layer.^For each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying a stereo signal generated from the M audio channels, and optionally further one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source. ^For the at least three predetermined virtual distances from a sound source, a parametric equalizer setting, e.g. for simulating distance-dependent sound absorption by air for the stereo signal generated from the M audio channels. In further embodiments, the audio container further comprises information indicating for each audio channel of the second layer whether it is a mono channel or a stereo channel.In further embodiments, the audio container further comprises, in a fourth layer or as part of the second layer (B), at least one head-locked stereo audio channel whose channels are assigned to the L and R channels of the output stereo signal. Optionally, the audio container can further comprise, for example as part of the third layer, metadata for the at least one head-locked stereo audio channel, which comprises one or more of the following information: ^ For each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying or attenuating the head-locked stereo audio channel, and one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source. ^ For the at least three predetermined virtual distances from a sound source, a parametric equalizer setting.The application of the equalizer settings as a function of a selected virtual distance from the sound source can, for example, result in an attenuation of the treble of the head-locked stereo audio channel in order to simulate the air absorption of the sound waves. In the various embodiments described herein, the M angular perspectives can correspond to the respective angles of M microphones in a omnidirectional microphone or microphone array that is or was used to record the sound source. The M angular perspectives can lie in one plane (e.g., horizontal plane). The M angular perspectives can, for example, correspond to the angles ^^ ∙360°⁄ ^^ with ^^ ∈ [0,1,2, … , ^^− 1]. In the embodiments described herein, it is assumed that: ^^ ≥ 4 .In principle, it is advisable to choose an even number M of angular perspectives, since this way two angular perspectives are offset by 180° from each other and the two corresponding audio signals can form a stereo pair / stereo channel. Thus, in an even number of M audio channels (which can be referred to as mono channels), two of the M audio channels whose angular perspectives are offset by 180° from each other can form one of ^^ / 2 stereo channels (especially bi-aural stereo channels). In principle, however, an odd number of M audio channels can also be used, where no two microphones are offset by 180° from each other. This means that a microphone on one side (e.g., left) is assigned the crossfade of two microphones on the other side (here: right).The embodiments are not limited to the use of a microphone array with in-plane microphones for recording the M audio channels of a sound source. In principle, other microphone arrangements can also be used for recording, optionally together with appropriate recording-side audio signal processing, to generate the M audio channels of a sound source (first layer in the audio container format). For example, it is also possible to use a microphone array on the recording side in which microphones are arranged in two mutually normal planes (horizontal plane and vertical plane), forming what are essentially two orthogonally aligned microphone arrays.It is also possible to record the sound source using a microphone arrangement of fewer or (significantly) more than M microphones, and to generate the M audio channels of a sound source (first layer (A) in the audio container format) from the respective microphone signals through signal processing. Each additional mono or stereo signal that is mixed into the stereo audio signal formed from the M audio channels can, for example, be recorded with a single main or support microphone. However, it is equally conceivable to use multiple microphones to generate a respective mono or stereo signal, e.g., by means of pre-mixing or signal processing of the individual microphone signals on the recording side. The N angular perspectives can lie in one plane (e.g., horizontal plane). The N angular perspectives can, for example, correspond to the angles ^^ ∙ 360°⁄ ^^ with ^^ ∈ [0,1,2, … , ^^ − 1]. In the embodiments described herein, it is assumed that: ^^ ≥6 .In principle, it is advisable to choose the number N as a multiple of 4 (e.g.: 8 or 12) as angular perspectives, since this includes the critical angular positions 0°, + / - 90° and 180°. Description of the figures Fig. 1 shows a system 100 according to an embodiment of the invention, which illustrates the recording / live recording, the processing of the recording and the generation of metadata, the streaming up to the playback via headphones; Fig. 2 shows an example numbering of angular positions in a microphone array 112 and the assignment of microphone signals to a headphone signal; Fig. 3 illustrates by way of example comb filter effects that can arise when recording a sound source with a microphone array 112 when cross-fading two adjacent microphone capsules of the microphone array 112 due to their distance;4 illustrates, by way of example, a localization of the sound sources on a line between the ears (in-head localization – IKL) in a pure stereo recording via headphones; Fig. 5 illustrates, by way of example, a localization of the sound sources outside the head (out-of-head localization – AKL), which occurs when microphone array signals and additional support microphone signals are mixed; Fig. 6 shows, by way of example, a mixture of the mixing factors according to the listener's head angle for panning the stereo group from Layer B, according to an embodiment consisting of four support microphones; Fig. 7 shows, by way of example, the determination of the two most obvious angular positions for a head angle; Fig. 8 shows a schematic representation of the functions of the renderer / player 122 according to an exemplary embodiment; Fig. 9 shows a further schematic representation of the functions of the renderer / player 122 according to a further exemplary embodiment; Fig.10 shows an exemplary RF spectrum interpolation and crossfading 1002 that can be used in the embodiments; Figs. 11a and 11b show an exemplary numbering of angular positions in a microphone array 112 and the assignment of pairs of microphone signals to a headphone signal for the head angles ^^ = 0° (reference direction) and ^^ = 45°; Fig. 12a shows the localization of a sound source at a head position of -45° and the corresponding angular position of a microphone array 112 with the transmitted microphone signals; and Fig. 12b shows the same sound source and its localization at a head position between 0° and -45° with crossfaded microphone signals corresponding to the head position. Detailed Description: The invention relates to the generation of immersive and / or binaural stereo signals for playback via headphones.In the following, embodiments of the invention are described with reference to a system for recording, transmitting, and reproducing audio. The different aspects of recording, transmitting (e.g., data format), and reproducing (e.g., generating stereo signals for headphones) of audio each form individual aspects of the invention, but also in any combination. In some embodiments of the invention, the stereo signal for the headphones is generated depending on a detected head angle of the listener wearing the headphones. Furthermore, the invention relates to an audio container format for transmitting audio information for generating the stereo signal in a corresponding device (also called a "renderer" or "player"), as well as computer-readable media, devices, and systems for generating and / or reproducing the stereo signal. In audio recordings of sound sources, e.g.,At a concert, the sound source is typically recorded by a main microphone system. For example, several microphones are used as the main microphone system, which are appropriately positioned in a recording room. Additionally, spot microphones are usually positioned near the instruments and mixed into the main microphone system. By increasing the mixing level of a spot microphone, the acoustic presence and proximity of the respective instrument can be increased. The inventor has discovered that this method, which is also applicable to dummy head stereophony, can also be combined with a microphone array as the main microphone system.A microphone array that can be used in some embodiments of the invention to record the sound source can, for example, be a spherical, head-sized separating body, on which ideally M=8 microphones (microphone capsules with pressure transducers) are mounted in a ring on the same plane and at the same distance from each other. The microphones are thus located in a (horizontal) plane with an angle of 45° (= 360° / 8). In particular, the microphone array described in US 7,333,622 B2 can be used, but the invention is not limited thereto. The microphone array can be placed at a suitable distance in front of the sound body. The exact positioning of the microphone array depends in particular on the local acoustic conditions of the recording room (e.g., concert hall, open-air stadium, etc.) and is usually determined by the recording director / sound engineer and / or sound engineer. In a concert hall, the microphone array could, for example,placed in the center of the second or third row in the auditorium of a concert hall. The invention is not limited to a specific arrangement / positioning of the microphone array in the recording room or to a specific design of the microphone array. However, the recording technology and the main microphone system should be capable of capturing the spatial sound in such a way that individual stereo groups from the mixing console can be dynamically mixed into the array signals. All M (e.g. 8) microphone array signals are transmitted to a playback unit (e.g. renderer / player), where they are blended according to the head position. If the head is turned, a tracking sensor ensures that the playback unit blends neighboring microphones accordingly; the array microphones are essentially tracked to the listener's ears. This makes the sound source appear to be fixed in the sound space when listening through headphones.By slight movements of the head, a front-to-back orientation is immediately established. Furthermore, in addition to the microphone array, additional main and support microphones are positioned in the recording room. Here, too, the number of main and / or support microphones used and / or the positioning of the microphones is not limited to a specific design; both the number (e.g., string quartet vs. symphony orchestra) and the positioning of the main and / or support microphones, as well as the optional combination of several microphone signals into stereo signals, can be determined by the recording manager / sound engineer and / or sound engineer. In practice, stereo signals or stereo groups with similar signals, e.g., main microphone signals, string support microphones, drum microphones, etc., formed at the mixing console and / or a digital audio workstation (DAW), are suitable. These are normally panned accordingly for a stereo presentation via loudspeakers, i.e.The phantom sound source positions of the individual signals between the loudspeakers correspond to the position on the stage or that of the main microphone arrangement. The DAW can be software for recording, editing, and producing audio that runs on a processing unit (e.g., computer, laptop, tablet, etc.). These main and / or support signals are also transmitted to the playback unit (e.g., renderer / player) to be mixed with the audio signals from the microphone arrays. By increasing the mixing level of the support signals, the acoustic presence and proximity or directional localization of the respective sound source (e.g., the instrument or a group of instruments) can be increased. It is also possible to use mono signals as support signals (or mono and stereo signals) as support signals. If necessary, the main and support signals can be delayed by the sound propagation times to the microphone array. In order to provide the playback unit (e.g.,In order to enable the renderer / player to mix the support signals depending on the respective head rotation of the listener, associated metadata is transmitted which determines the mixing levels and thus also the panoramas for all angular positions. Fig. 1 shows a system 100 according to an embodiment of the invention. The system 100 consists of a recording-side part 110 which detects a sound source with a microphone array 112 with M microphones and a main microphone system and further support microphones and provides the audio data and the associated metadata in a data format for transmission (e.g. by means of streaming) and / or storage. The system 100 also comprises a playback-side part 120 which generates an immersive stereo signal for headphone playback, wherein the stereo signal adaptively follows the head movement of the listener.The recording-side part 110 of the system 100 can comprise a microphone array 112 with M=8 or more or fewer microphone capsules, by means of which the sound source is detected and corresponding M microphone array signals are generated. The microphone capsules can be arranged in a plane (e.g., on a circular line) and assigned corresponding M angular perspectives. For example, the microphone capsules can be arranged in a circle, wherein the angle between two adjacent microphone capsules is 360° / M. Advantageously, M can be selected such that two microphone capsules opposite one another on the circular line form a 180° angle (i.e., ^^ ∙ 360°⁄ ^^ =180°,^^^^^^ ^^ < ^^, ^^ ∈ ℕ ^^^^^^ ^^ ∈ ℕ, where ℕ denotes the set of natural numbers excluding 0). The audio signals of such a pair of opposing microphone capsules can form a stereo pair.The microphone array 112 is placed at a suitable distance in front of the sound box, relative to conventional microphones for a stereo or Dolby Atmos recording. This could be, for example, the second or third row in the audience. The microphone array 112 can, for example, comprise a spherical, approximately head-sized separator, on which the M microphones are mounted in a ring at the same level and at the same distance from each other. It is possible to use a microphone array 112 as in US patent specification US 7,333,622 B2. M / 2 audio or stereo pairs can be formed from the signals of the M microphones of the microphone array 112. For each angle ^^^^ = ^^ ∙ 360° / M with ^^ ∈[0,1,2, … ,^^ − 1], the microphones at the orthogonal angles (^^^^ − 90° and ^^^^ + 90°) can form an audio or stereo pair that can be fed to the headphones as part of the binaural stereo signal. For example, if one numbers them as in Fig. 2, or Fig. 11a and Fig.In Fig. 11b, the microphones are arranged clockwise, with No. 0 for the microphone pointing in the reference direction ^^ = 0° (e.g., pointing directly toward the sound body). For the reference direction ^^ = 0°, microphones No. 6 (left) and No. 2 (right) form such a stereo pair that can be transmitted to the headphones (see Fig. 2). For backward head angles, the microphones are swapped accordingly (e.g., 180°: No. 2: left and No. 6: right). When crossfading two adjacent stereo pairs, e.g., from 0° to -45° for a (tracked) head angle ^^ between these two angles, an acoustic counter-rotation of the transmitted acoustic environment is perceived. This is illustrated exemplarily in Fig. 11b. The head angle ^^ = 0° defines the reference direction of the head angle, for example the head position frontal to the sound body. Furthermore, it can be assumed here, for example, that the angular perspectives ^^^^ = ^^ ∙360° / M have the same reference direction as the head angle ^^.In the embodiments described herein, it is also assumed that the M angular perspectives ^^^^ correspond to the angular positions of the M microphones of the microphone array 112 (e.g., with regard to number and / or orientation). However, this is not mandatory. The orientation of the microphone array 112 to the M angular positions ^^^^ could, for example, also be subsequently corrected or adjusted using a rotation offset in the metadata. Furthermore, it is also conceivable that the number of microphones of the microphone array 112 is smaller or larger than the number M of angular perspectives ^^^^. In this case, the smaller or larger number of microphone signals could be converted into the corresponding M audio signals for the predetermined M angular positions using suitable signal processing (e.g., in a DAW or producer 118).It is also conceivable that the microphones in microphone array 112 are not all provided with the same angle and / or that fewer than 8 microphones are provided in microphone array 112 (e.g., 4, 5, or 6 microphones). For example, the range (angle) of the head angle to which the headphone signal is adapted could not be 360°. For example, in microphone array 112 in Fig. 2, microphones 0 and 4 (0° and 180°) could not be provided, so that a head angle range of < 180° can be covered. For ease of understanding and purely as an example, it is assumed below that M microphone array signals of microphone array 112 are detected and corresponding M audio signals are transmitted or recorded. The M microphone array signals as well as the audio signals of the main and / or support microphones 114 are fed to a mixer 116.The mixing console 116 can be implemented in hardware or by means of a DAW on a computing unit, or as a combined software and hardware solution. Using the mixing console 116, the signals from multiple support microphones 114 can optionally be combined into a stereo or mono signal before the resulting support signal is fed to the producer 118. The producer 118 can, for example, be implemented using a software solution installed on a computing unit (e.g., a computer). Optionally, the functions of the mixing console 116 and the producer 118 can also be combined in a software solution. The M audio signals of the microphone array 112 as well as the audio signals of the support microphones 114 (which may have been combined in groups into individual stereo support signals using the mixing console 116) are fed to the producer 118. The producer 118 can, for example, be a stand-alone application or a plug-in for DAWs.The producer 118 generates metadata and combines it with the audio data in a data format. In order to implement the mixing of one or more audio signals based on the signals from the support microphones 114, the metadata is transmitted using a data structure, which contains a metadata set for each of these support signals or the associated channel. Each metadata set identifies the mixing ratios of the respective support signal for each of N predefined angular perspectives. If the mixing of the support signals is to be distance-dependent, several such metadata sets can be provided for each support signal, which define the mixing ratio(s) of the respective support signal for each of the N predefined angular perspectives for different (in particular three) predefined virtual distances from the sound source. All metadata (e.g.Mixing ratios and EQ settings for each additional mono or stereo channel per N angle perspective) can be entered manually by the sound engineer / sound mixer via a GUI on the producer. If a microphone array is permanently installed at one location (e.g., a concert hall), only a few parameters will need to be changed for different concerts. The N angle perspectives of the additional mono or stereo audio channel(s) (Layer B) can result from practical considerations. The number N can therefore be chosen, for example, as ^^ = (^^ + 1) ∙ 4,^^^^^^ ^^ ∈ [1,2,3, … ]. The N angle perspectives of the additional mono or stereo audio channel(s) can correspond to the M angle perspectives of the microphone array 112 (i.e., ^^ = ^^); however, this is not mandatory; i.e., ^^ ^^ can also apply. In the following, it is assumed purely for example that M=8 and N=8. The N angle perspectives can be in one plane (e.g.horizontal plane). The N angular perspectives can, for example, correspond to the angles ^^ ∙ 360°⁄ ^^ with ^^ ∈[0,1,2, … ,^^ − 1]. In the embodiments described herein, it is assumed that: ^^ ≥ 6. In principle, it is advisable to select the number N as a multiple of 4 (e.g., 8 or 12) as angular perspectives, since this covers the critical angular positions 0°, + / - 90°, and 180° for a frontal sound event (0°). The audio signals and metadata generated by the producer 118 are sent, for example, as a stream 130 to a playback-side part 120 of the system 100 and / or recorded on a storage medium. The audio signals include the M array signals of the microphone array 112 and the additional stereo and / or mono signals originating from a mixing console 112 (or the DAW). The producer 118 can be optimized specifically for working on a live production.It may be advisable to process as few stereo groups as possible in the producer 118 in order to keep the data rate and the additional computing effort as low as possible. The producer 118 can provide a graphical user interface (GUI) by means of a display device (not shown), by means of which the operator can generate the metadata that describes, for example, the behavior of the mono and / or stereo signals based on the main and / or support microphones 114 at different head positions of the listener, as will be explained in more detail below. Optionally, it is also possible to integrate the functions of the renderer / player 122 into the producer 118. Using the functions of the renderer / player 122 (or an emulation of its functionality), the settings made by the operator (e.g., the set values in the metadata, which, for example,describe the behavior of the mono and / or stereo signals based on the support microphones 114 at different head positions of the listener). The producer 118 can enable the operator to freely adjust the head position (e.g., head angle) using the GUI. The immersive audio signal for the headphones 104 (also output stereo signal) corresponding to the set head position, which is generated by the functions of the renderer / player 122 in the producer 118, can thus be "pre-listened to" by the operator with headphones (e.g., during a recording) or "listened to" live (e.g., during a livestream of the recording). The producer 118 processes the M audio signals supplied by the microphone array 112 and the additional audio signals generated in the mixer 116 based on the signals from the main and / or support microphones 114, and generates digital audio data in a predefined data format.One aspect of the invention relates to such a data or transmission format or a data structure, e.g. audio container, by means of which the information necessary for binaural headphone playback can be stored and / or transmitted. The data format can be a streaming-capable container format for multimedia, video and / or audio data, by means of which the M audio signals of the microphone array 112, the further audio signals generated in the mixer 116 based on the signals of the main and / or support microphones 114, and the associated metadata can be transmitted. The data format can, for example, be one based on ISO / IEC 14496-12:2022 (ISO base media file format), MPEG:H, EBU Audio Definition Model (ADM), etc., or be implemented as an extension thereof. However, it is also possible to use other data formats orContainer formats suitable for transmitting and / or storing audio signals and metadata can be used and / or expanded. The metadata can be specified in an XML-based or XML-like syntax in the data format. The data format (or data structure) can define the following layers: ^A first layer (Layer A) for M audio channels. These M audio channels transport corresponding M audio signals, which were recorded, for example, with a microphone array 112. The M audio channels or signals are assigned to M angular perspectives. ^A second layer (Layer B) for the at least one main and / or support signal channel, hereinafter referred to simply as support signal channel. Each support signal channel contained in Layer B can be either a mono audio channel or a stereo audio channel. Each support signal channel contains a mono or stereo signal, which was recorded, for example, with one or more support microphones 114.^A third layer (Layer C) with metadata containing a metadata set for at least one supporting signal channel in the second layer. Each metadata set identifies the mixing ratios of the respective supporting signal for each of the predefined N angular perspectives of the signals from Layer B. If the mixing of the supporting signals is to be distance-dependent, several such metadata sets can be provided for each supporting signal, which define the mixing ratio(s) of the respective supporting signal for each of the predefined N angular perspectives for different (in particular three) predefined virtual distances from the sound source. The metadata for Layer B can be used to acoustically position the auditory event of the respective supporting channel relative to an auditory event generated from the M audio channels for the individual M angular positions (Layer A) (similar to a "panoramaization").In a further fourth layer (Layer D) or as part of the second layer (B), the data format can further contain headlocked (HL) stereo signals. These HL stereo signals can be transmitted to the headphones without panning, i.e., with a fixed LR assignment. The HL stereo signals can, for example, contain direction-independent signals (e.g., reverberation, moderation, etc.). The metadata in Layer C can contain further metadata for Layer A and / or Layer B and / or Layer D. For example, the metadata for different (virtual) distances from the sound source can contain a metadata set for each (virtual) distance. Alternatively or additionally, the metadata can also contain distance levels and / or equalizer settings that are used by the renderer / player 122 to generate an immersive stereo signal for the headphones.The renderer / player 122 can be implemented, for example, as software or firmware that runs in a device with a computing unit (e.g., DSP, processor, SoC, SiP, ASIC, etc.). The renderer / player 122 could, for example, be implemented in a smartphone, tablet, laptop, or other device with the ability to integrate tracking data (e.g., camera and gyroscope), which generates the generated headphone audio signal and feeds it to the headphones. The renderer / player 122 can be implemented as a web app or an app on a smart device, or can be used in a gaming environment. The audio and metadata can be transferred to the renderer / player 122 as a stream 130 or a file in a data format described herein. The renderer / player 122 blends the audio signals of layer A according to the head angle (of the listener) between the M angular positions, preferably in “real time” (e.g. in less than 20 ms, preferably less than 10 ms).At the same time, the additional microphone signals from Layer B are mixed in depending on the head posture (e.g., head angle) and optionally also a (virtual) distance according to the metadata from Layer C. A tracking system 124 is used to detect the head posture. The tracking system 124 is configured for dynamic detection of head movements or "head detection" (in particular in a (horizontal) plane or for detecting the head angle or a change in the head angle) and can also be referred to as a head tracking system. The sensors of the tracking system 124 can also be integrated into the device implementing the renderer / player 122. Alternatively, this sensor technology of the tracking system 124 can also be implemented in the headphones 104. For example, (in-ear) headphones such as Apple's AirPods Max & AirPod Pro, JBL's Quantum One, HyperX's Cloud Orbit S, etc.Head tracking functions that an exemplary tracking system 124 implements and that can be used in some embodiments. A headset with a built-in tracking system can transmit full 360° head movements to the renderer / player 122. Combined with a (smartphone) camera, for example, the (virtual) distance to a virtual stage can also be simulated or adjusted. Alternatively, the tracking system 124 can be implemented as a single device that can be attached to the headset, for example (e.g., Waves Nx Head Tracker for headphones from Waves Inc.). In a further alternative, the sensors of the tracking system 124 can be implemented both in the device that supports the renderer / player 122 and in the headset 124. The tracking system 124 is configured to determine the (current) head angle ^^ or a change in the head angle ^^ and to feed it to the renderer / player 122.The angular resolution of the tracking system 124 should be less than or equal to 5° in order to track the listener's head movement as accurately as possible and to enable the renderer / player 122 to generate an immersive stereo signal for the headphones 104 that follows the head position as accurately as possible. In some embodiments, the tracking system 124 enables an angular resolution of 5° or less, preferably 2° or even less. Various systems can be used for head tracking, e.g., camera(s) and image processing, gyroscope(s), magnetometer, acceleration sensor(s), ultrasonic sensor(s) and / or infrared sensor(s), and three-point measurement systems using radio waves. Gyroscopes and acceleration sensors can be used to detect the rotational movements and accelerations of a device / headphone. By integrating the signals from these sensors, the relative position or orientation of the device / headphone in space can be determined.When a gyroscope is combined with a camera, the tracking data is added together, making 360° tracking with a moving head possible. Magnetometers can be used to measure the magnetic field around the device / headphone. By analyzing changes in the magnetic field, the movements and orientation of the device / headphones can be tracked. Ultrasonic or infrared sensors can use ultrasonic or infrared signals to determine the position of the head or headphones in space. This usually involves placing several sensors in the environment to receive the signals and calculate the position of the headphones. Camera(s) and image processing can also be used to track the listener's head rotation based on the images captured by the camera. When using a smartphone camera, the smartphone can be placed or held in front of the listener, like a selfie.Using the camera, software (e.g., a facial recognition program) can detect head posture within a range of less than + / - 90° from the reference direction (0°). At the same time, the camera can measure the distance of the listener from the smartphone and convert this into a (virtual) distance from the virtual stage. To improve tracking accuracy, signals from multiple and / or different sensors / sensor types can be used. Alternatively or additionally, it is also possible for all movement parameters relating to head rotation / orientation (particularly the head angle) and / or the virtual distance from the listener to be manually adjusted via the graphical user interface (GUI) of the renderer / player 122. For example, a frontal view of a (sketchy) head can be shown on the display. The listener / user could then turn their head left or right with one or more fingers to adjust the head angle.Moving the head up or down with one or more fingers could be used to adjust the listener's distance from a virtual stage. Furthermore, it is possible for the renderer / player 122, together with the tracking system 124, to enable gesture control of the renderer / player 122. The gestures can be used, for example, to: (1) fix the current head position (distance and angle of rotation) ("freeze"), e.g., by nodding twice, (2) release the "freeze" by nodding three times, (3) set the head angle and / or a virtual distance to a default value (0° and distance to a predefined "sweet spot" - e.g., the distance to the medium), start or pause / stop audio playback, etc. The individual corresponding functions can be implemented alternatively or additionally using the control elements on the GUI.As already mentioned, the renderer / player 122 can crossfade the audio signals of layer A according to the head angle (of the listener) between the M angular positions (e.g., in real time). At the same time, the additional microphone signals of layer B are mixed in depending on the head posture (e.g., head angle) and optionally also a (virtual) distance according to the metadata from layer C. By mixing in the additional microphone signals of layer B, it is possible, above certain levels, to mask comb filter effects that arise when crossfading neighboring microphones of the microphone array 112 without the need for complex signal processing. As illustrated by way of example in Fig. 3, the spatial offset between neighboring microphones in the microphone array 112 causes the incoming sound to be received with different phase positions.For example, at a distance of 6.8 cm between two microphone capsules on the sphere's surface, the passing sound is maximally attenuated at approximately 2.5 kHz when the capsules are mixed, which can be masked by mixing in the additional microphone signals from Layer B. In comparison, in US 7,333,622 B2, such comb filter effects are compensated for using more complex signal processing that includes high-pass filtering of the critical signals, decomposition of the high signal components into amplitudes and phases, amplitude addition, and introduction of the phase of a microphone. In principle, it is also possible in some embodiments to use the signal processing described in US 7,333,622 B2 to compensate for comb filter effects. The decisive effect of mixing in the additional microphone signals from Layer B serves to increase presence.When a normal stereo recording is played back solely through headphones, the sound sources are localized on a line between the ears (in-head localization – ICL), as illustrated in Fig. 4. Only when the omnidirectional array signals of Layer A and the possibly delayed, panned, and possibly regulated microphone signals of Layer B are mixed do the sound events migrate out of the head (out-of-head localization – AKL) and appear to the listener with a more vivid, closer position and presence, as illustrated in Fig. 5. In the following, purely as an example and for easier understanding, a string quartet on a stage is considered, in which the first violin is positioned on the left and the cello on the far right of the stage, with the two other musicians in between. Fig. 6 shows this sketchily. The reference direction here is ^^ = 0°.When recording the string quartet, for example, a microphone array 112 with 8 microphones can be positioned in the center of the auditorium at the height of the second row of seats, with microphone number 0 directed toward the stage (in the reference direction ^^ = 0°). Each of the four instruments also has its own spot microphone. In this example, the recording-side part 110 of the system 100 can transmit 8 microphone signals from the microphone array 112 in four stereo pairs (Layer A) in the stream 130; for example, for the head angle 0° (microphones 6 & 2 for left & right), -45° (microphones 5 & 1 for left & right), +45° (microphones 7 & 3 for left & right), and -90° (microphones 4 & 0 for left & right). By swapping the stereo channels in the renderer / player 122 you get the four other head angles +90° (microphones 0 & 4 for left & right), 180° (microphones 2 & 6 for left & right) and +135° (microphones 1 & 5 for left & right) and -135° (microphones 3 & 7 for left & right).The audio signals from the support microphones are combined as a stereo group with panned audio signals from the instruments and transmitted as a single stereo audio signal in Layer B: For a head angle of ^^ = 0°, the level of the first violin is highest in the left channel; the cello signal is highest on the right. This corresponds to the stereo group signal in a normal loudspeaker recording. In the producer 118, the stereo group is panned, adjusted, and if necessary delayed (optional) at given angular positions in N head positions, so that the musicians appear in the "correct" positions for each head angle in the immersive stereo signal for the headphones 104, which can be generated by the renderer / player 122. These settings (parameters) made in the producer 118 for each of the N angular positions are transmitted as a metadata record in the metadata (Layer C) of the stream 130 for the stereo group (in Layer B).When the head is rotated, the renderer / player 122 does not blend the stereo signal of Layer B. Instead, the parameter values in the metadata for panoramaizing the corresponding N angular positions are averaged according to the listener's head angle and used to generate a stereo signal from the stereo signal of Layer B. With appropriate parameter settings, the panoramas "move" in the opposite direction to the head movement, and the sound source positions essentially "remain" in place. The stereo groups can be prepared for the producer 116 at the mixing console 116 (or in the DAW). It may be useful for the audio signals from the microphone array 112 to arrive at the listener's ears simultaneously or slightly earlier (law of the first wave front).In the example of the string quartet, the signals from the support microphones can therefore be delayed by the mixing console 116 (or in the DAW) by at least the propagation time of the distance to the microphone array 112 before being fed to the producer 118. It may also be sufficient to delay a stereo group of support signals as a whole. A corresponding temporal adjustment of the propagation times can, however, also be implemented, for example, in the producer 118 and can be set via a user interface. In practice, angular positions have proven useful for the additional mono and stereo signals from layer BN=8 in order to obtain stable localizations of the sound sources during head rotation. A smaller number of angular positions can lead to lateral and central deflections of the sound source positions during head rotation. A larger number of angular positions can also be used. Increasing the number of angular positions (e.g.However, increasing the angular position (to 12 angular positions in 30° increments) does not significantly stabilize the localization, but does result in increased workload in generating the metadata for Layer B. Eight angular positions also have the advantage that acoustic events can be easily understood in 45° increments in order to make the panorama settings. In the following, N=8 angular positions are therefore assumed for the panoramaization of the Layer B signals. The panoramas are displayed in Producer 118 for each stereo channel using four mixing factors, ^^1,^^ , ^^1,^^ , ^^2,^^ and ^^2,^^ (also mixing ratio) per angular position. With ^^ as the value of the mixing factor, 1 is the left channel, 2 is the right channel of the stereo signal, and L and R are the targets for the headphone channels L and R of the headphone signal generated in Renderer / Player 122, on which the respective channels. and ^^2 of the stereo channel are rendered with the mixing factors ^^. Thus, each of the N=8 angular positions is assigned a metadata set with the four mixing factors ^^1,^^ , ^^1,^^ , ^^2,^^ and ^^2,^^ and transmitted in the metadata of Layer C in Stream 130. For each additional stereo support signal, the following mixing factors are transmitted as a metadata set in the metadata of Layer C: where ^^ denotes the angular position ^^ (^^ ∈ 0, 1, .. , ^^ − 1). Index 1 denotes channel ^^1 of the support signal, index 2 denotes channel ^^2 of the support signal, and indices ^^ and ^^ denote the headphone channels L and R. In the renderer / player 122, the mixing factors of the angular positions are read from the metadata of Layer C and linearly averaged according to the head angle, depending on the head position. The mixing factors thus obtained are then rendered with the signals ^^1 and ^^2 of the stereo channel from Layer B. With the appropriate selection of the mixing ratios, the impression of opposing sound sources with a rotating head is created, albeit still as ICL.In parallel, the M audio signals of the microphone arrays 112 from Layer A of the stream 130 are blended in the renderer / player 122 to form a binaural stereo signal and summed with the output of the rendered stereo signals from Layer B, thus generating the headphone signal for the headphones 104 as the output stereo signal. Remaining with the example of the string quartet, Fig. 6 shows an example of a mix of the mixing factors according to the listener's head angle for panning the stereo group generated from the four spot microphones 602, 604, 606, 608. As shown in Fig. 6 above, the level of the signal from the first spot microphone 602 of the first violin is approximately 18 dB stronger on the left channel than on the right channel of the stereo group. The second violin (support microphone 604) is panned with a difference of 6 dB in favor of the left channel and appears on about 50% of the left side (in loudspeaker playback).The level of the cello (spot microphone 608) is approximately 18 dB higher in the right channel than in the left channel of the stereo group. The viola (spot microphone 606) is panned with a 6 dB difference in favor of the right channel and appears approximately 50% of the right side (the loudspeaker base). This spot microphone stereo mix serves as the starting point and is assigned to a head angle of 0°. In this mix, the instruments appear acoustically to the listener as shown in graph 610, but still as IKL. The mixing factors for the angle position = 0° are therefore, for example, as follows: ^^2,^^(^^ = 0°) = 0, ^^2,^^(^^ = 0°) = 1For a head angle of ^^ = 0°, the spot microphone stereo mix ("original mix") is mixed with these mixing factors to the corresponding stereo signals of the microphone array 112 for the head angle 0° (microphone 6 (left) and 2 (right)). Now the sound event migrates out of the head and appears three-dimensional in front of the listener. For the angular position = -45°, the level of the right stereo channel of the spot microphone stereo mix is raised and that of the left channel is lowered and distributed across both channels L and R, so that the instruments appear acoustically "shifted" to the right onto channel R to the listener, as shown in graphic 612. The mixing factors for the angular position = -45° are, for example: ^^1,^^(^^ = −45°) = 0.5, ^^1,^^(^^ = −45°) = 0,5, 1,2For a head angle of ^^ = −45°, the spot microphone stereo mix is mixed with these mixing factors to the corresponding stereo signals of the microphone array 112 for the head angle -45° (microphone 5 (left) and 1 (right)). For a head angle that does not correspond to any of the predefined angular positions, for example ^^ = −22.5°, the mixing factors of the two predefined angular positions closest to the head angle - i.e. here the angular positions 0° and -45° - are linearly blended (see Fig. 6, right side), resulting in: ^^1,^^(^^ = −22.5°) = 0.75, ^^1,^^(^^ = −22.5°) = 0.25, ^^2,^^(^^ = −22.5°) = 0, ^^2,^^(^^ = −22.5°) = 1.1. The instruments are acoustically "shifted" slightly to the right to channel R for the listener, as shown in graphic 614, but still with signal components on channel L. The blended metadata is rendered with the stereo signal.In Layer A, depending on the head position, it is not the metadata but the M audio signals themselves that are blended. The (here, for example, M=8) audio signals from microphone array 112 in Layer A are blended linearly (microphones 6 and 7 for the left channel, and microphones 2 and 3 for the right channel). For a head angle of ^^ = −22.5°, the spot microphone stereo mix is mixed with the averaged mixing factors to the blended stereo signals from microphone array 112 for a head angle of -45°. The linear mixing of the mixing ratios of the spot microphone stereo mix (Layer B) for a head angle ^^ can generally be represented as follows: where the indices n and n+1 denote the two angular positions closest to the head angle ^^ and with ^^ = , where ^^ ∈ 0, 1, .. , ^^ − 1 and the value of ^^ is chosen such that: ^^ ≥ ^^ ∙ 360° / ^^ and ^^ ≤ (^^ + 1) ∙ 360° / ^^. Here, ^^ denotes the number of angular perspectives (e.g., N=8 or N=12). The two angular positions n and n+1 can also be referred to as the angular positions n and nn (nearest neighbor n and next-nearest neighbor nn). Which of the angular positions n or n+1 forms the nearest neighbor n or the next nearest neighbor depends on the head angle ^^. Accordingly, the stereo support signal of layer B for a head angle ^^ is linearly averaged as follows: where ^^(^^) and ^^(^^) represent the left and right channels of the blended stereo signal, respectively, and ^^1 and ^^2 represent the left and right channels of the stereo support signal of Layer B. In addition, further metadata can be transmitted in Layer C to simulate variable distances to the sound body. This allows the listener to choose a virtual distance to the sound body, e.g., from the conductor's position via the "best seat" to the back of the hall. With the distance, the panoramas of the stereo signals in the frontal area also change for the listener, since, for example, the aperture angle of an ensemble decreases with distance. By appropriate panoramas for different distances, these relationships can be credibly recreated. Accordingly, for example, a respective metadata set can be ^^^^,2,^^ , ^^^^,2,^^ for (at least) two, advantageously for three (or more) different virtual distances (positions) to the sound source in Layer C. These additional metadata sets can be generated accordingly in the producer 118. For example, three different metadata sets, each containing the mixing ratios ^^^^,1,^^ , ^^^^,2,^^, ^^^^,2,^^ for the N angular positions, for three virtual distances (e.g., Close, Medium, Far) for the respective stereo signals of Layer B are provided in the metadata. In the renderer / player 122, virtual distances that lie between two of the defined distances (e.g., distance between Close and Medium or between Medium and Far) can be calculated by linear averaging (analogous to the averaging of the mixing ratios for an angular position ^^) of the (possibly averaged) mixing ratios for the respective angular position ^^: where d denotes the set virtual distance between the two specified distances D1 (e.g., Close or Medium) and D2 (e.g., Medium or Far). The values ^^(^^) denote the already averaged mixing ratios for the head angle ^^, as described above. Accordingly, the stereo support signal of Layer B for a head angle ^^ and a distance d is linearly averaged as follows: ^^(^^, ^^) = ^^1 ∙ ^^1,^^ (^^,^^) + ^^2 ∙ ^^2,^^ (^^,^^) Alternatively, the mixing ratios can first be averaged depending on the selected / desired virtual distance (analogous to the averaging of the mixing ratios for an angular position ^^) and then a further averaging of the distance-dependent mixing ratios thus obtained can be carried out for any angular position ^^ (if the angular position ^^ does not correspond to any of the predetermined angular positions). In the embodiments described above, it was assumed that the audio signals from the support microphones are combined into one or more stereo groups (i.e., one or more stereo signals with left and right channels) and that the individual stereo groups are transmitted in Layer B of the data format (e.g., audio container) in stream 130.It is also possible for the signals from the support microphones to be transmitted individually or in groups as mono signals in Layer B. In this case, the mixing factors of a metadata set describe how the mono signal in Layer B is distributed between the two channels to the left and right of the stereo signal generated from the mono signal in the renderer / player 112. If the data format supports both stereo signals and mono signals in Layer B, the metadata can further indicate whether a respective channel in Layer B is a mono channel or a stereo channel. In further embodiments, the metadata in Layer C can also contain equalizer settings that, for example, simulate treble attenuation due to air absorption for a predetermined (virtual) distance.The equalizer settings can – similar to the metadata sets for the mixing factors – also be specified for at least two, for example, three or more different virtual distances, in order to simulate, for example, the treble attenuation caused by air absorption for different virtual distances. The equalizer settings fade in and out linearly. The metadata for equalizer settings can be specified for each stereo channel of Layer B and / or for each of the N predefined angular positions. The equalizer settings for the left and right headphone channels can be set individually in Producer 118, for example, to simulate head shadowing at one ear when the head is rotated.Accordingly, a metadata set for the equalizer settings for each stereo channel of Layer B and / or for each of the N predefined angular positions contains equalizer settings for the left and right channels of the stereo signal generated from the stereo channel of Layer B in the renderer / player 122. Embodiments of the channel layout or the data structure of the stream 130 take into account, as described above, the various degrees of freedom of the listener: On the one hand, the additional support signals in Layer B and the associated metadata sets of the mixing ratios for the different predefined N angular positions of Layer C allow consideration of a head rotation of the listener (up to 360° rotation) when generating an immersive headphone signal in the renderer / player 122.Furthermore, specifying metadata sets of the mixing ratios for different distances from the sound body allows for consideration of a (virtual) distance of the listener from the sound body. The producer 118 in the recording-side part 110 of the system 100 is configured to provide or generate the audio signals of the microphone array 112 (Layer A) and the support microphones 114 (Layer B), as well as the associated metadata (Layer C). The metadata is generally not time-critical and can be changed during a transmission. For example, the channel layout or data structure of the stream 130 can be structured as follows: (1) Layer A: a. M audio channels. The M audio channels in M mono tracks. The M audio channels can be generated with a microphone array 112 (e.g. with M=8 with 8 microphones for the angular perspectives 0°, - 45°, +45°, +90°, -90°, -135°, +135° and 180°).(2) Layer B:a.Audio channels B1, ..., Bm: Any number of audio channels for additional audio signals, which can be mono and / or stereo signals. For example, mixed signals from 5.1 to Dolby Atmos beds can be transmitted in mono format. Possible application: Transmission of stereo support signals and / or stereo main microphone signals. b. Audio channels C1, ..., Ck (optional): Any number of additional channels for headlocked (HL) stereo signals. These signals are not further panned in the renderer / player 122, but are mixed into the headphone signal in the renderer / player 122. Application: Stereo reverb signals without directional information, presentation, etc. (3) Layer C: a. Metadata record for each audio channel in Layer B. The metadata record 802-2 contains N mixing factors for panning for the N angular positions for the respective audio channel B1, ..., Bm from Layer B.Optionally, several such metadata sets 802-1, 802-2, 802-3 can be included in the metadata, with each metadata set being assigned to a predefined (virtual) distance (e.g., Close, Medium, Far). It is also possible for several audio channels from Layer B to share one or more metadata sets. b. Equalizer Settings (Optional) Dual-band equalizer settings 916-1, 916-2 for the L channel and R channel of a stereo signal generated in the renderer / player 122 from the respective audio channel B1, ..., Bm from Layer B. The dual-band equalizer settings 916-1, 916-2 are specified for both channels for each of the N angular positions. The equalizer settings can be parametric or semi-perimetric settings. The equalizer settings may affect three or more bands of an equalizer 936-1, 936-2.c.Amplification factors (optional) Gain / attenuation factor 918 for the amplification / attenuation of a stereo signal generated in the renderer / player 122 from the respective audio channel B1, ..., Bm from Layer B or for a respective HL signal from Layer B. An amplification / attenuation factor can be specified for each of several predefined (virtual) distances (e.g., Close, Medium, Far). The amplification / attenuation factor 918 or distance-dependent amplification / attenuation factors 918 can be defined for each audio channel of Layer B. In addition to the gain / attenuation factor 918 (for a respective distance), one or more blending curves 932 can optionally be defined or specified, which can be used to average two gain / attenuation factors 918 when a (virtual) distance is to be set between specified (virtual) distances.d.Distance-Dependent Equalizer Settings (Optional) Equalizer settings 920 for a stereo signal generated in the renderer / player 122 from the respective audio channel B1, ..., Bm from Layer B or for a respective HL signal from Layer B or equalizer settings 966 for a binaural stereo signal generated from the M audio channels of Layer A, wherein corresponding equalizer settings 920, 966 are specified for each of several predetermined (virtual) distances (e.g., Close, Medium, Far). The equalizer settings 920, 966 can, in particular, relate to two or more bands of an equalizer 914, 940, 944.e. System-related metadata (optional) This metadata 954 contains, for example, settings for stereo width controls 946 after summation 820, 822 of the stereo signals generated from the audio signals B1, ..., Bm of Layer B (e.g., head size adjustment). Alternatively or additionally, settings 950, 830 for output equalizer 952 and / or limiter 828f.Metadata for Layer A (optional) The metadata for Layer A can contain one or more of the following information: ^A rotation offset 960 for the subsequent "alignment" of the microphone array 112. ^Divergence control 960 for (subsequent) reducing or increasing the binaural angle of the microphones of the microphone array 112 (normally 180°). ^Gain / attenuation factor 964 for the amplification / attenuation of a stereo signal that is generated in the renderer / player 122 from the M audio signals of Layer A. A gain / attenuation factor can be specified for each of several predefined (virtual) distances (e.g., Close, Medium, Far).In addition to the gain / attenuation factor 964 (for a respective distance), one or more blending curves can optionally be defined or specified, which can be used to blend two gain / attenuation factors 964 when a (virtual) distance between predefined (virtual) distances is to be set. Figure 8 shows a schematic and exemplary representation of the functions of the renderer / player 122 (hereinafter referred to as player 122). The player 122 can, for example, receive audio and metadata in the form of a stream 130 or, alternatively, from a file that is present in a corresponding data format with layers A, B, and C. In the implementation example shown, layer A comprises a number of M audio channels with the signals ^^(^^) with ^^ ∈ 0,1, ... ,^^ − 1 (see (1) a. of the data format), which are assigned to M predetermined angular positions (1 to M).The audio channels of Layer A can have been recorded with a microphone array 112. Layer B comprises a stereo group (stereo signal B1 - see (2) a. of the data format) with the audio signals ^^^^^^1,^^ and ^^^^^^1,^^ , as well as an HL stereo signal (HL Stereo HL1 - see (2) b. of the data format) with the audio signals ^^^^1,^^ and ^^^^^^1,^^. Layer C for the metadata contains (optional) system-related metadata (see (3) e. of the data format) and metadata 802-2 for the stereo group of Layer B (metadata stereo signal B1 - see (3) a. of the data format). The audio blender 810 of the player 122 receives a tracking signal (e.g., continuously, at predetermined time intervals) from the tracking system 124, which indicates at least the current head angle ^^ of the listener wearing the headphones 104 with which the immersive output stereo signal 808 is reproduced.The M audio signals are rendered in the audio blender 810 according to the head angle ^^ in the tracking signal, so that the currently required adjacent stereo pairs are blended, provided the head angle ^^ does not correspond to any of the predefined angle positions. For a head angle of, for example, -22.5°, the audio signals at angle positions 6 and 7 for the left headphone channel and 2 and 3 for the right headphone channel are linearly blended 812:. ^^(^^)^^^^^^^^^^ = ^^(^^ + 90°) ∙ (1 − ^^) + ^^(^^ + 90°) ∙ ^^where the indices m and m+1 denote the two angular positions that are closest to the head angle ^^ and with ^^ = , where ^^ ∈ 0, 1, .. , ^^− 1 and the value of ^^ is chosen such that: ^^ ≥ ^^ ∙ 360° / ^^ and ^^ ≤ (^^ + 1) ∙ 360° / ^^. Here, ^^ denotes the number of angular perspectives (e.g. M=6 or M=8).^^ ± 90° and ^^ + 1 ± 90° denote the two angular positions which are shifted by 90° to the left or right with respect to the angular positions ^^ and ^^ + 1. For M=8, the angle between the angular positions would be 45°, so that the ^^ ± 90° or ^^ + 1 ± 90° corresponds to the audio signals with the angular positions (^^ + ^^ ± 2) ^^^^^^ ^^ or (^^ + ^^ + 1 ± 2) ^^^^^^ ^^, where mod represents a modulo operation. This is illustrated in Fig. 7 for the example ^^ = 65° and M=8. The angular positions closest to the 65° angle are the positions ^^ = 1 and ^^ + 1 = 2. The position m+90° corresponds to the angular position (8 + 1 + 2) ^^^^^^ 8 = 3, the position m−90° corresponds to the angular position (8 + 1 − 2) ^^^^^^ 8 = 7.The position m+1 + 90° corresponds to the angular position (8 + 2 + 2) ^^^^^^ 8 = 4, the position ^^ + 1 − 90° corresponds to the angular position (8 + 2 − 2) ^^^^^^ 8 = 0. If the head angle ^^ corresponds to one of the predefined N angular positions, the audio signals ^^^^1,^^ and ^^^^1,^^ of the stereo group (stereo signal B1) can be distributed to the channels ^^(^^)^^1 and ^^(^^)^^1 according to the mixing factors 802-2 assigned to the corresponding angular position. If, however, the head angle ^^ does not correspond to any of the predefined N angular positions, the mixing factors for the stereo group (stereo signal B1) are first blended according to the head angle ^^ and the blended mixing factors are used to distribute the audio signals ^^^^1,^^ and ^^^^1,^^ of the stereo group (stereo signal B1) accordingly to the channels ^^(^^)^^1 and ^^(^^)^^1 (cf. amplifiers 840-1, 840-2, 840-3, 840-4 and adding elements 816, 818).Channels ^^(^^)^^1 and ^^(^^)^^1 form a resulting, head-angle-dependent stereo signal, which is mixed with the stereo signal with channels ^^(^^)^^^^^^^^^^ and ^^(^^)^^^^^^^^^^ (see adding elements 804, 806). Channels ^^(^^)^^1 and ^^(^^)^^1 can be formed as follows: with the blended mixing factors ^^^^,1,^^, ^^^^,1,^^ , ^^^^,2,^^ ^^^^,^^, ^^^^,^^ from the metadata 802-2 for the stereo group (stereo signal B1): where the indices n and n+1 denote the two angular positions closest to the head angle ^^ and with ^^ = , where ^^ ∈ 0, 1, .. , ^^ − 1 and the value of ^^ is chosen such that: ^^ ≥ ^^ ∙ 360° / ^^ and ^^ ≤ (^^ + 1) ∙ 360° / ^^. ^^ denotes the number of angular perspectives (e.g., N=8 or N=12). Analogous to the explanations regarding the determination of the angular positions ^^ ± 90° and ^^ + 1 ± 90° for the audio signals of Layer A with reference to Fig. 7, here ^^ ± 90° and ^^ + 1 ±90° denote the two angular positions of the N angular positions that are shifted by 90° to the left or right with respect to the angular positions n or n+1. For N=8, the angle between the angular positions would be 45°, so that the ^^ ± 90° or ^^ + 1 ±90° corresponds to the additional audio signals of layer B with the angular positions (^^ + ^^ ±2) ^^^^^^ ^^ or (^^ + ^^ + 1 ± 2) ^^^^^^ ^^, where mod represents a modulo operation. The value range of the mixing factors ^^^^,1,^^ , ^^^^,2,^^, ^^^^,2,^^ can, for example, be between -2 (inclusive) and +2 (inclusive).A factor of 2 can correspond to a gain of 6 dB, a factor of 1 corresponds to 0 dB gain, a factor of 0 to signal cancellation (multiplication by 0), a factor of -1 corresponds to 0 dB gain with an additional 180° phase shift of the signal, and a factor of -2 corresponds to a gain of 6 dB with a 180° phase shift of the respective channel. The value range of the mixing factors can, for example, also be specified in 0.01 steps. For mono signals ^^^^^^ in Layer B, only two mixing factors ^^^^,^^, ^^^^,^^ are signaled in the metadata record for each of the N angular positions ^^ ∈ [0,1,2, … ,^^ − 1], and the head angle-dependent stereo signal is formed as follows:. Here, too, the mixing factors can be blended. If layer B contains several such stereo groups (and / or mono signals), these are processed in player 122 like the stereo signal B1 (or as shown for the mono signal), and the resulting head-angle-dependent stereo signals are added (see adding elements 820, 822). If layer B contains one or more HL stereo signals, such as the HL stereo signal (HL Stereo HL1) with the audio signals ^^^^^^1,^^ and ^^^^^^1,^^, these are added to the other signals (here: ^^(^^)^^^^^^^^^^ and ^^(^^)^^^^^^^^^^ as well as ^^(^^)^^1 and ^^(^^)^^1) (see adding elements 804, 806 and adding elements 824, 826). The resulting stereo signal ^^(^^)^^^^^^^^ and ^^(^^)^^^^^^^^ forms the immersive output stereo signal that is fed to the headphones 104. Optionally, the stereo signal ^^(^^)^^^^^^^^ and ^^(^^)^^^^^^^^ can be fed to a limiter 828.The Limiter 828 can, for example, intercept clipping of the stereo signal ^^(^^)^^^^^^^^ and ^^(^^)^^^^^^^^. The Limiter 828 can be adjusted via System Metadata 830. This System Metadata 830 can define the following parameters for the Limiter 828: Limiter T. reshold T Release R Gain GHere, T is the threshold of the control range in dB below full scale (0 dBFS), above which the signal is limited, and is specified in 0.1 dB steps; R is the limiter's release time in milliseconds (ms), which determines how long the control process rolls back after adjustment, and gain is the level swing or reduction before the limiter is adjusted (e.g., in 0.1 dB steps). In an alternative embodiment, the stereo groups (or mono signals) in layer B in producer 118 of the recording-side part 110 of system 100 can already be panned for all N angular positions, and all N fully panned stereo signals ^^^^,^^ and ^^^^,^^ (with ^^ ∈ [0,1,2, … , ^^ − 1]) for all N angular positions ^^ can be transmitted from producer 118 in stream 130. In this case, metadata 802-2 could be omitted.Instead, in the player 122, two stereo signals ^^^^and ^^^^+1 for the angular positions n and n+1 that are closest to the head angle ^^ can be directly blended with each other in a linear manner. where the indices n and n+1 denote the two angular positions closest to the head angle ^^ and with ^^ = , where ^^ ∈ 0, 1, .. , ^^ − 1 and the value of ^^ is chosen such that: ^^ ≥ ^^ ∙ 360° / ^^ and ^^ ≤ (^^ + 1) ∙ 360° / ^^. Here, ^^ denotes the number of angular perspectives (N=8 or N=12). This places lower demands on the performance of the hardware implementing the player 122. Only the N stereo channels in layer B are blended, but no equalizers are used. This can reduce the CPU load, but at least N stereo channels are required in layer B. Another example implementation of the player 122 is shown in Fig. 9. This implementation offers, for example, the possibility of selecting or setting a variable distance to the sound source. For this purpose, the data format provides the player 122 with additional metadata, as explained in more detail below. As already mentioned in connection with Fig.8, it can also be assumed here that the player 122 receives audio and metadata in the form of a stream 130 or alternatively obtains it from a file that is present in a corresponding data format with layers A, B and C. In the implementation example shown, it is again assumed by way of example that layer A comprises a number of M audio channels with the signals ^^(^^) with ^^ ∈ 0,1, ... , ^^ − 1 (see (1) a. of the data format), which are assigned to M predetermined angular positions (1 to M). The audio channels of layer A can have been recorded with a microphone array 112. Layer B comprises a stereo group (stereo signal B1 – see (2) a. of the data format) with the audio signals ^^^^^^1,^^ and ^^^^^^1,^^ , as well as an HL stereo signal (HL Stereo HL1 – see (2) b. of the data format) with the audio signals ^^^^^^1,^^ and ^^^^^^1,^^. Layer C for the metadata contains (optional) system-related metadata (see (3) e.of the data format) and metadata 802-2 for the stereo group of layer B (metadata stereo signal B1 - see (3) a. of the data format). As in Fig. 8, in the implementation in Fig. 9, the M audio signals are rendered in the audio blender 810 according to the head angle ^^ in the tracking signal, so that the currently required neighboring stereo pairs are blended, provided the head angle ^^ does not correspond to any of the predefined angular positions. Optionally, the reference direction (orientation) of the microphone arrays 112 and the N angular positions of the audio channels in layer A can be subsequently aligned to one another in the audio blender using a rotation offset 960 specified in the metadata. Furthermore, also optionally, a divergence control 960 can be specified in the metadata, which serves to (subsequently) reduce or increase the binaural angle of the microphones of the microphone array 112 (normally 180°).This function can also be implemented in the audio blender 810. The audio blender 810 feeds the output signal to an RF spectrum interpolation 910. In some cases, the addition of the audio signals from Layer B (e.g., at very low levels) may not be sufficient to mask any comb filter effects that may occur in the M audio signals of Layer A. A high-frequency spectrum interpolation 910 can prevent phase cancellations. An exemplary RF spectrum interpolation and crossfading 1002 is shown in Fig. 10. For example, using a high-pass filter 1006 and a low-pass filter 1004, the two signals to be crossfaded in the two channels (L and R) ^^^^,^^ and ^^^^+1,^^ or ^^^^,^^ and ^^^^+1,^^ can be split into a low frequency component and a high frequency component at a given cutoff frequency at which the cancellation takes place (for example, 2.5 kHz in the example previously discussed with reference to Fig.3).The non-critical low frequency components are crossfaded 1008 (as described with reference to Fig. 8). The frequency components above the cutoff frequency are crossfaded 1012 using an FFT 1010 (or DFT, DCT, etc.) to obtain the amplitudes without their phases. In the next step, the current phase of the signal is impressed on the amplitude sum with the larger mixing factor i. This is done by appropriately parameterizing the inverse FFT 1016 (or inverse DFT, DCT, etc.). The low and high frequency ranges are then mixed / added again 1018 to generate a stereo signal ^^(^^)^^^^^^^^^^ and ^^(^^)^^^^^^^^^^. Fig. 10 shows the signal processing as an example for obtaining the signal ^^(^^)^^^^^^^^^^ of the left channel. Corresponding signal processing also takes place to obtain the signal ^^(^^)^^^^^^^^^^ of the right channel. The HF spectrum interpolation 910 can prevent dips in the high frequency range when the head is turned.The RF spectrum interpolation 910 is particularly suitable, for example, for transmissions / recordings in which exclusively the microphone array is used. The RF spectrum interpolation 910 can, for example, be selectively activated or deactivated in the player 122. To do this, the producer 118 can, for example, indicate the activation and deactivation of the RF spectrum interpolation 910 in the metadata for Layer A. Alternatively or additionally, it is possible for the listener to activate or deactivate the RF spectrum interpolation 910 using the GUI of the player 122. In this case, specifications from the metadata can also be "overwritten" by the listener. In the exemplary implementation, the binaural stereo signal of the channels^^(^^)^^^^^^^^^^ and ^^. fed to a distance blender 912, 914. As explained, the metadata (see (3) f.) for Layer A can also specify one or more gain / attenuation factors 964 for the amplification / attenuation of the channels ^^(^^)^^^^^^^^^^ and^^(^^)^^^^^^^^^^. A gain / attenuation factor (^^^^^^^^^^^^^^ , ^^^^^^^^ , ^^^^^^^^^^^^) can be specified for each of several predefined (virtual) distances (e.g., Close, Medium, Far). In addition to the gain / attenuation factor(s) 964 (for a respective distance), one or more blending curves (e.g., their curve shape) can optionally be defined or specified, which can be used to blend 980 two gain / attenuation factors 964 when a (virtual) distance between specified (virtual) distances is to be set. Medium Close Far Curve shape medium p close p far1-5 The crossfade curves are used for a blend of the gain / attenuation factors (Gain) 964: ^^(^^) = ^^^^ ∙ ^^(1 − ^^) + ^^^^ ∙ ^^(^^) where ^^ denotes the set distance between the two directly adjacent distances (Close and Medium or Medium and Far), ^^(^^) describes the curve shape, and ^^^^ and ^^^^ are the gain factors of the two directly adjacent distances. ^^(1 − ^^) = 0 if the set distance ^^ corresponds to the distance of the gain factor ^^^^, or (^^) = 0 if the set distance ^^ corresponds to the distance of the gain factor ^^^^. The resulting gain factor ^^(^^) is applied evenly to the left channel. ^^ and right channel ^^(^^)^^^^^^^^^^. The crossfade curves (or functions) can, for example, be a selection of one or more functions from the group: linear functions, exponential functions, cosine functions, sine functions, etc. The crossfade curves enable various transitions between the distance levels and can also be selected differently for the left channel and the right channel. In this case, a separate value ^^(^^) would be determined for each channel and applied to the channel. A corresponding control 938, 942 of the distance can also be used for the stereo signals generated from the audio channels of Layer B (e.g. ^^^^1,^^^^1) or HL audio channels (e.g. ^^^^^^1,^^^^^^1). The metadata (see (3) c.) for layer legs or multiple gain / attenuation factors 918, which implement a corresponding blending 932 of the gain / attenuation factors 918 for a desired distance ^^ in a manner analogous to that described in connection with the distance blending of the binaural stereo signals ^^(^^)^^^^^^^^^^ and ^^(^^)^^^^^^^^^^. Here, too, it is conceivable that the amplifiers 938, 942 apply the blended gain factor ^^(^^) evenly to the left channel and right channel of the respective supplied stereo signal. In the case of distance blending 932, 938, for example, the level of a solo instrument (corresponding to an additional stereo signal of Layer B (e.g., stereo signal B1)) may, for example, increase less between the distances from Medium to Close than the orchestral instruments (received as other additional stereo signals of Layer B) in order to avoid over-presence.Likewise, the solo instrument's level can decrease more slowly at greater distances so that it can still be heard clearly at a distance. On the other hand, a signal's level can increase as one moves away from the virtual sound body. Distance blending 942 can, for example, be used to slightly amplify an HL stereo signal from a room microphone system as the distance increases, in order to simulate a more convincing increase in distance. Optionally, distance-dependent equalizer settings 920, 966 can also be applied in player 122 to each additional stereo signal / HL signal of layers A and B (see equalizers 914, 940, and 944). For each of several predefined (virtual) distances (e.g., Close, Medium, Far), a corresponding equalizer setting 920, 966 can be specified in the metadata (see (3) d.).The equalizer settings 920, 966 can particularly relate to two or more bands of an equalizer 914, 940, 944. The player 122 can use the equalizer settings 920, 966, for example, in addition to the level changes 912, 938, 942. The equalizer settings 920, 966 can be used, for example, to simulate the attenuation of air frequencies at great distances. For each distance (e.g., Close, Medium, and Far), two or more sets of EQ parameters (for two or more EQ channels) can be specified in the equalizer settings 920, 966, which include the frequency, a gain factor (Gain), a quality factor (Q), as well as a filter type (e.g.,Bell, Tail or Cut Off): Distance Medium Close FarBand ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^Frequency ^^1^^ ^^2^^ ^^1^^ ^^2^^ ^^1^^ ^^2^^Gain ^^1^^ ^^2^^ ^^1^^ ^^2^^ ^^1^^ ^^2^^Quality ^^1^^ ^^2^^ ^^1^^ ^^2^^ ^^1^^ ^^2^^Curve shape ^^,^^, ^^1^^ ^^,^^, ^^2^^ ^^, ^^, ^^1^^ ^^,^^, ^^2^^ ^^,^^, ^^1^^ ^^, ^^,^^2^^The 914, 949 and 944 equalizers apply the respective equalizer settings to both channels of the corresponding input stereo signals. Here, too, the individual parameters of the equalizer settings 920, 966 for a distance ^^ between can be linearly blended / averaged in the manner already described in order to obtain the corresponding parameters for the distance ^^ to be applied by the equalizer 914, 949 and 944. As already described in connection with Fig. 8, the audio signals of each stereo group (e.g.^^^^1,^^ and ^^^^1,^^ for stereo signal B1) are distributed to a left channel and a right channel (e.g., channels ^^(^^)^^1 and ^^(^^)^^1) according to the mixing factors 802-2 assigned to the corresponding angular position. If, however, the head angle ^^ does not correspond to any of the predefined N angular positions, the mixing factors for the stereo group (e.g., for stereo signal B1) are first blended according to the head angle ^^, and the blended mixing factors are used to distribute the audio signals ^^^^1,^^ and ^^^^1,^^ of the stereo group (stereo signal B1) accordingly to the channels ^^(^^)^^1 and ^^(^^)^^1 (cf. amplifiers 840-1, 840-2, 840-3, 840-4 and adding elements 816, 818). Channels ^^(^^)^^1 and ^^(^^)^^1 form a resulting, head-angle-dependent stereo signal. Channels ^^(^^)^^1 and ^^(^^)^^1 can optionally be fed to equalizers 936-1 and 936-2, respectively. These are also referred to here as "angular position equalizers."The metadata includes, for example, dual-band equalizer settings 916-1, 916-2 for channels ^^(^^)^^1 and ^^(^^)^^1 (or for each stereo signal of Layer B). The equalizer settings 916-1, 916-2 are specified for both channels for each of the N angular positions. The equalizer settings can be parametric or semi-perimetric settings. In particular, the equalizer settings can affect three or more bands of an equalizer 936-1, 936-2. These equalizer settings can specify a filter for the left channel ^^(^^)^^1 and the right channel ^^(^^)^^1 for each of the N angular positions. The equalizers 936-1, 936-2 can be implemented in three bands. The equalizers 936-1, 936-2 can, for example, simulate shadows of the ear facing away from the sound source. A filter type (e.g., bell, tail, or cutoff) can be selected for each band. Distance from the sound source does not affect the equalizer settings 916-1 and 916-2.The metadata 916-1, 916-2 can contain the following parameters for three-band equalizers for each of the N angle perspectives: Left channel Right channelBand ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^ ^^^^^^^^Frequency ^^1^^ ^^2^^ ^^3^^ ^^1^^ ^^2^^ ^^3^^Gain ^^1^^ ^^2^^ ^^3^^ ^^1^^ ^^2^^ ^^3^^Quality ^^1^^ ^^2^^ ^^3^^ ^^1^^ ^^2^^ ^^3^^Curve shape ^^, ^^,^^1^^ ^^, ^^,^^2^^ ^^, ^^,^^3^^ ^^, ^^,^^1^^ ^^, ^^,^^2^^ ^^, ^^,^^3^^ The number of parameters in the metadata 916-1, 916-2 depends accordingly on the number of bands of the Equalizers 936-1, 936-2. The output signals of equalizers 936-1, 936-2 can be fed to amplifier 938 for subsequent optional distance adjustment. If layer B contains multiple stereo groups (and / or mono signals), these are processed in player 122 as described for stereo signal B1 (or as shown for the mono signal), and the resulting head-angle-dependent stereo signals are added (seeAdding elements 820, 822). The two channels of this sum signal can also be fed to an optional stereo base control 946 (MS wide / narrow). This can be used, for example, to adapt the overall base width of the sum signal of the adding elements 820, 822 to the head diameter of the listener. Metadata 954 can contain the settings for the stereo width control 946. If layer B contains one or more HL stereo signals, such as the HL stereo signal (HL Stereo HL1) with the audio signals ^^^^^^1,^^ and ^^^^^^1,^^, these are added to the other signals (here: ^^(^^)^^^^^^^^^^ and ^^(^^)^^^^^^^^^^ as well as ^^(^^)^^1 and ^^(^^)^^1) (cf. adding elements 804, 806 and adding elements 824, 826). The resulting stereo signal can be fed to the optional output equalizer 952. The output equalizer 952 can, for example, be a multi-band, particularly a 3- or 4-band, equalizer.The output equalizer 952 can be used to adapt the stereo signal to the headphones 104 used. The output equalizer 952 can, for example, perform a linearization of the headphones 104 used. If necessary, presets of frequently used headphone models can be made available in the player 122, which can be selected via the GUI of the player 122. Additionally or alternatively, corresponding equalizer settings 950 can be provided in the metadata (cf. (3) g.) that the player 122 uses. These equalizer settings 950 can be EQ6, EQ7, EQ8, etc. for each band. The equalizer settings 950 can specify the frequency, an amplification factor (gain) and a quality factor (Q), as well as a filter type (e.g.Bell, Tail or Cut Off) can be defined (here for a 4-band equalizer): Equalizer Parameter Band ^^^^^^ ^^^^^^ ^^^^^^ ^^^^^^Frequency ^^6 ^^7 ^^8 ^^9Gain ^^6 ^^7 ^^8 ^^9Quality ^^6 ^^7 ^^8 ^^9Curve Shape ^^,^^, ^^6 ^^,^^, ^^7 ^^,^^, ^^8 ^^, ^^,^^9The stereo output signal of the output equalizer 952 can also be fed to the optional limiter 828, which can be used to absorb overloads. The limiter 823 can be adjusted via system metadata (cf. (3) g.). As described in connection with Fig. 8, possible level peaks above -1 dBFS are intercepted and regulated down. This is ensured with very low latency. The output signal of the limiter 828 with the channels ^^(^^)^^^^^^^^ and ^^(^^)^^^^^^^^ forms the immersive output stereo signal that is fed to the headphones 104.The result of mixing audio signals from Layer A with additional signals in Layer B is a spatial, three-dimensional auditory impression with strong AKL and presence, which is further enhanced by head movements. The array's ITDs, i.e. the phase difference between the left and right microphones up to approximately 1.5 kHz, are particularly crucial for AKL and the perception of spatial depth. The presence of a sound body is achieved by the additional stereo audio signals in Layer B. In addition to localization, the audio signals from Layer B play a decisive role in the "sound," especially when main microphones are used. The described system, in particular the data format and the Player 130, make it possible to create a sonic quality and presence of the headphone signal that cannot be achieved with HRFT-based processes such as Dolby Atmos. Tangible proximity and intimacy, opulent sound quality and enveloping spatiality can be achieved simultaneously.As described, the listener can choose their distance from the sound body and decide between richness of detail and spatiality. In addition to concert recordings, live broadcasts with and without video or audio-only recordings, and music elements in VR games, many other applications are conceivable in the context of museums, installations, and gaming. For example, as a walk-in concert event in a VR game or in an empty concert hall. The headphones 104 can also be part of a VR headset or the like.
Claims
Patent claims 1. Audiocontainer für die Erzeugung eines Ausgangs-Stereosignals mit die Kanälen L und R basierend auf M Audiokanälen, mit ^^ ≥ 4, die M in einer Ebene liegenden predetermined angular perspectives are assigned, and at least one further mono or stereo audio channel, wherein the audio container defines the following layers: a first layer for the M audio channels assigned to the M angular perspectives; a second layer for the at least one mono or stereo audio channel; a third layer with first meta data, which contains a meta data record for each of the at least one mono or stereo audio channel of the second layer, wherein each meta data record contains the mixing ratios of the respective mono or stereo audio channel of the second layer for each of predetermined N angular perspectives, m it ^^ ≥ 6, definiert.
2. Audiocontainer nach Anspruch 1, wobei die ersten Meta-Daten einem erstenpredetermined virtual distance from a sound source, and the third layer further comprises second metadata and / or third metadata corresponding to a second predetermined virtual distance and a third predetermined virtual distance from the sound source; wherein the second predetermined virtual distance is smaller than the first predetermined virtual distance from the sound source and the third predetermined virtual distance is greater than the first predetermined virtual distance from the sound source; and wherein the second metadata and / or the third metadata contain / contain a meta data set for each of the at least one mono or stereo audio channel of the second layer, each meta data set defining the mixing ratio of the respective further mono or stereo audio channel of the second layer for the predetermined N angular perspectives.
3. Audiocontainer nach Anspruch 1 oder 2, wobei jeder Meta-Datensatz für einenStereo audio channel of the second layer identifies four parameter values for each of the N angular perspectives, where the four parameter values for a respective angular perspective indicate: ^ einen ersten Zumischfaktor (i1L) für den ersten Kanal des Stereo- audio channel of the second layer and a second mixing factor (i2L) for the second channel of the stereo audio channel of the second layer for the channel L of the output stereo signal, and ^ einen dritten Zumischfaktor (i1R) für den ersten Kanal des Stereo- audio channel of the second layer and a fourth mixing factor (i2R) for the second channel of the stereo audio channel of the second layer for the channel R of the output stereo signal.
4. Audiocontainer nach einem der Ansprüche 1 bis 3, wobei jeder Meta-Datensatz für a mono audio channel of the second layer identifies two parameter values for each of the N angle perspectives, where the two parameter values specify a respective angle perspective: ^ einen ersten Zumischfaktor (iL) des Mono-Audiokanals des zweiten Layer for the L channel of the output stereo signal, and ^ einen zweiten Zumischfaktor (iR) des Mono-Audiokanals des zweiten Layer for channel R of the output stereo signal.
5. Audiocontainer nach Anspruch 3 oder 4, wobei jeder Meta-Datensatz ferner eineor more of the following information: for each of the N angular perspectives, parametric equalizer settings to simulate a head-related or outer ear transfer function (HRTF) for the two channels of a stereo signal generated from a mono or stereo audio channel of the second layer; and / or for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying the stereo signal generated from the mono or stereo audio channel of the second layer, and one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source; and / or for the at least three predetermined virtual distances from a sound source, a parametric equalizer setting for simulating distance-dependent sound absorption by air for the stereo signal generated from the mono or stereo audio channel of the second layer.
6. Audiocontainer nach einem der Ansprüche 1 bis 5, wobei der dritte Layer fernerMeta data for the M audio channels of the first layer, wherein the meta data for the M audio channels further comprise one or more of the following information: a rotation offset for aligning the M angular perspectives with respect to a sound source represented by the M audio channels of the first layer; a divergence-controlling parameter for reducing or increasing the binaural angle of the M audio channels of the first layer; for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying a stereo signal generated from the M audio channels, and one or more cross-fading curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source;and / or for the at least three specified virtual distances from a sound source, a parametric equalizer setting for simulating distance-dependent sound absorption by air for the stereo signal generated from the M audio channels; 7. Audiocontainer nach einem der Ansprüche 1 bis 6, wobei der Audiocontainer ferner Includes information that indicates for each audio channel of the second layer whether it is a mono channel or a stereo channel.
8. Audiocontainer nach einem der Ansprüche 1 bis 7, wobei der Audiocontainer ferner in a fourth layer or as part of the second layer comprises at least one head-locked stereo audio channel, both channels of which are assigned to the L and R channels of the output stereo signal.
9. Audiocontainer nach Anspruch 8, wobei der dritte Layer ferner Meta-Daten für denat least one head-locked stereo audio channel, wherein the metadata for the at least one head-locked stereo audio channel further comprises one or more of the following information: for each of at least three predetermined virtual distances from a sound source, a gain factor for amplifying or attenuating the head-locked stereo audio channel, and one or more crossfade curves for calculating a gain factor for a virtual distance between two of the predetermined virtual distances from a sound source; and / or for the at least three predetermined virtual distances from a sound source, parametric equalizer settings for simulating distance-dependent sound absorption by air for the head-locked stereo audio channel.
10. Audiocontainer nach einem der Ansprüche 1 bis 9, wobei die M Angle perspectives correspond to the respective angles of M microphones in an omnidirectional microphone array.
11. Audiocontainer nach einem der Ansprüche 1 bis 10, wobei die M Winkelperspektiven den Winkeln ^^ ∙ 360°⁄ ^^ mit ^^ ∈ [0,1,2, … ,^^ − 1] entsprechen und / oder wobei die M Winkelperspektiven den Winkeln ^^ ∙ 360°⁄ ^^ mit ^^ ∈ [0,1,2, … , ^^ − 1] entsprechen.
12. Audiocontainer nach einem der Ansprüche 1 bis 11, wobei die M AudiokanäleMono channels and two of the M audio channels, whose angle perspectives are m 180° zueinander versetzt sind, einen von ^^ / 2 binauralen Stereokanälen bilden.
13. Audiocontainer nach einem der Ansprüche 1 bis 12, wobei jeder weitere Stereo- Audio channel of the second layer is a stereo channel.
14. Verfahren zur Erzeugung eines Ausgangs-Stereosignals mit die Kanälen L und R for playback with headphones, comprising: E mpfangen von M Audiokanälen, mit ^^ ≥ 4, die M in einer Ebene liegenden predetermined angle perspectives are assigned Receiving at least one further mono or stereo audio channel; Receiving metadata containing a metadata record for each of the at least one further mono or stereo audio channel, each meta- D atensatz für N Winkelperspektiven, mit ^^ ≥ 6, das Zumischverhältnisse des respective mono or stereo audio channel; B estimmen eines Kopfwinkels ^^ eines den Kopfhörer tragenden Hörers in der Plane relative to one of the N or M angular perspectives, which has a reference direction d efiniert, wobei der absolute Kopfwinkel ^^ basierend auf Sensorsignalen at least one sensor; generating a first stereo audio signal with a channel L and a channel R, wherein the channel L of the stereo audio signal is generated by a linear crossfading of two audio channels of the M audio channels, to which angular perspectives z ugeordnet sind, die dem Kopfwinkel ^^ − 90° am nächsten liegen und der KanalR of the stereo audio signal is generated by a linear crossfading of two of the M audio channels, which are assigned angle perspectives that d em Kopfwinkel ^^ + 90° am nächsten liegen; for each additional mono or stereo audio channel, linear averaging of the mixing ratios (i1Ln, i1Lnn, i1Rn, i1Rnn, i2Ln, i2Lnn, i2Rn, i2Rnn) contained in those two metadata sets for the respective mono or stereo audio channel are given, which correspond to the two angular perspectives of the N angular perspectives (n, n n) zugeordnet sind, die dem absoluten Kopfwinkel ^^ am nächsten liegen, um to generate cross-faded mixing ratios (i1Lф, i1Rф, i2Lф, i2Rф); for each additional mono or stereo audio channel, generating a second stereo audio signal with an L channel and an R channel based on the respective additional mono or stereo audio channel using the averaged mixing ratios (i 1Lф , i 1Rф , i 2Lф , i 2Rф); generating the output stereo signal by summing the first stereo audio signal with the at least one generated second stereo audio signal.
15. Verfahren nach Anspruch 14, wobei der mindestens eine weitere Mono- oder Stereo audio channel is or contains a stereo audio channel, and the meta-data record for the stereo audio channel identifies four parameter values for each of the N angular perspectives, where the four parameter values for a respective angular perspective indicate: ^ einen ersten Zumischfaktor (i1L) für den ersten Kanal des Stereo- audio channel and a second mixing factor (i 2L ) for the second channel of the stereo audio channel for the channel L of the second stereo audio signal, and ^ einen dritten Zumischfaktor (i1R) für den ersten Kanal des Stereo- audio channel and a fourth mixing factor (i 2R) for the second channel of the stereo audio channel for the channel R of the second stereo audio signal; wherein the linear averaging of the mixing ratios for the stereo audio channel specified in the two meta-data sets associated with the two angular perspectives of the N angular perspectives (n, nn) closest to the absolute head angle ^^ comprises: ^ jeweiliges lineares Überblenden, in Abhängigkeit vom Kopfwinkel ^^, der beiden ersten Zumischfaktoren (i1Ln, i1Lnn), der beiden zweiten Zumischfaktoren (i2Ln, i2Lnn), der beiden dritten Zumischfaktoren (i1Rn, i1Rnn) und der beiden vierten Zumischfaktoren (i2R, i2Rnn), die in den zwei Meta- data sets are specified to provide a blended first mixing factor (i1Lф), a blended second mixing factor (i2Rф), a blended third mixing factor (i 1Rф ) and a blended fourth mixing factor (i 2Rф ); and wherein generating the second stereo audio signal based on the stereo audio channel comprises: ^ Erzeugen des Kanals L des zweiten Stereosignals durch Multiplikation des first channel of the stereo audio channel with the crossfaded first mixing factor (i1Lф) and the second channel of the stereo audio channel with the crossfaded second mixing factor (i 2Lф) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel; and ^ Erzeugen des Kanals R des zweiten Stereosignals durch Multiplikation des first channel of the stereo audio channel with the crossfaded third mixing factor (i1Rф) and the second channel of the stereo audio channel with the crossfaded fourth mixing factor (i 2Rф ) and then adding the first channel of the stereo audio channel and the second channel of the stereo audio channel.
16. Verfahren nach Anspruch 15, ferner umfassend: for each of the N angular perspectives, receiving a parametric equalizer setting to approximate a head-related or outer ear transfer function (HRTF) for the L and R channels of the second stereo signal; filtering the L channel of the second stereo signal based on the parametric equalizer settings, which are faded in and out linearly according to the two angular perspectives of the N angular perspectives; and Filtering the R channel of the second stereo signal based on the parametric equalizer settings that correspond to the two angle perspectives of the N angle perspectives, which are faded in and out linearly.
17. Verfahren nach Anspruch 15 oder 16, ferner umfassend: for each of at least three predetermined virtual distances from the sound source, receiving a gain factor (gnear, gmedium, gfar) for amplifying the second stereo signal, and one or more crossfading curves for calculating a gain factor (g dist ) for a virtual distance between two of the predetermined virtual distances from a sound source; and amplifying the second stereo signal as a function of a selected virtual distance from the sound source based on one of the gain factors (gnear, gmedium, gfar) or a gain factor calculated using one of the crossfading curves (g dist ).
18. Verfahren nach einem der Ansprüche 15 bis 17, ferner umfassend:for at least three predefined virtual distances from the sound source, receiving a parametric equalizer setting (EQnear, EQmedium, EQfar) for the second stereo signal; and changing the sound image of the second stereo signal according to a selected virtual distance based on the received parametric equalizer settings (EQ near , EQ medium , EQ far ) for the second stereo signal.
19. Verfahren nach einem der Ansprüche 14 bis 18, ferner umfassend: Receiving a head-locked stereo audio channel; for each of at least three predetermined virtual distances from the sound source, receiving a gain factor (g HL_near , g HL_medium , g HL_far ) for amplification of the head-locked stereo audio channel, and one or several crossfade curves for calculating a gain factor (g HL_dist) for a virtual distance between two of the predetermined virtual distances from a sound source; and amplifying the head-locked stereo audio channel in dependence on a selected virtual distance from the sound source based on one of the gain factors (g HL_near , g HL_medium , g HL_far ) or a gain factor calculated using one of the blending curves (gHL_dist).
20. Verfahren nach Anspruch 19, ferner umfassend: for the at least three specified virtual distances from the sound source, receiving parametric equalizer settings (EQHL_near, EQ HL_medium, EQ HL_far) for the head-locked stereo audio channel; and changing the sound image of the head-locked stereo audio channel according to a selected virtual distance based on the received parametric equalizer settings (EQHL_near, EQHL_medium, EQHL_far) for the head-locked stereo audio channel.
21. Verfahren nach einem der Ansprüche 14 bis 20, ferner umfassend eine Interpolation of the high-frequency spectrum of the first stereo audio signal.
22. Verfahren nach einem der Ansprüche 14 bis 21, ferner umfassend:for each of at least three predetermined virtual distances from the sound source, receiving a gain factor (garray_near, garray_medium, garray_far) for amplifying the first stereo audio signal, and one or more crossfading curves for calculating a gain factor (g array_dist ) for a virtual distance between two of the predetermined virtual distances from a sound source; and amplifying the first stereo audio signal as a function of a selected virtual distance from the sound source based on one of the gain factors (garray_near, garray_medium, garray_far) or a gain factor calculated using one of the crossfading curves (garray_dist).
23. Verfahren nach Anspruch 22, ferner umfassend:for the at least three predetermined virtual distances from the sound source, receiving a parametric equalizer setting (EQarray_near, EQ array_medium, EQ array_far) for the first stereo audio signal; and changing the sound image of the first stereo audio signal according to a selected virtual distance based on the received parametric equalizer settings (EQarray_near, EQ array_medium, EQ array_far) for the first stereo audio signal.
24. Computer-lesbares Medium, das Befehle speichert, die, wenn sie von einer computing unit of a device to cause the device to carry out the method according to one of claims 14 to 23.
25. Vorrichtung mit einer Recheneinheit, die angepasst ist das Verfahren nach einem to carry out claims 14 to 23.
26. Vorrichtung nach Anspruch 25, wobei die Vorrichtung angepasst ist, das erzeugte Output stereo signal with the L and R channels to be fed to headphones for playback.
27. Vorrichtung nach Anspruch 25 oder 26, die angepasst ist, um einen oder mehrere To receive and process audio containers according to one of claims 1 to 13 as a continuous data stream.
28. Vorrichtung nach einem der Ansprüche 25 bis 27, wobei die Recheneinheit, diethe method according to one of claims 14 to 23 is adapted to be carried out based on the data in a plurality of audio containers according to one of claims 1 to 13.
29. Vorrichtung nach einem der Ansprüche 25 bis 28, wobei die Vorrichtung einen oder contains several sensors that provide the sensor signals for determining the K opfwinkels ^^ des den Kopfhörer tragenden Hörers in der Ebenebereitstellen.
30. Vorrichtung nach Anspruch 29, wobei der eine oder die mehreren Sensoren einen or more of the following sensors, accelerometer, gyroscope, or magnetic field sensor, to detect a change in the orientation of the device; and / or wherein the one or more sensors further comprise a camera that detects a head movement of the head of the listener wearing the headphones.
31. Vorrichtung nach Anspruch 30, wobei die Recheneinheit angepasst ist, aus den Sensor signals to detect a change in the listener's head angle.
32. System mit einer Vorrichtung nach einem der Ansprüche 25 bis 31 und einem Headphones for reproducing an output stereo signal generated by the device.
33. System nach Anspruch 32, wobei die Vorrichtung in einem Smartphone oder tablet computer, or headphones.
Citation Information
Patent Citations
Dynamic binaural sound capture and reproduction
US7333622B2
Dynamic binaural sound capture and reproduction in focued or frontal applications
US20080056517A1
Object Clustering for Rendering Object-Based Audio Content Based on Perceptual Criteria
US20150332680A1
Audio Signal Rendering
US20200280816A1
Audio processing method, apparatus, system, and storage medium
US20230156403A1