Apparatus, method and computer program for retrieving an acoustic room information and audio signal processor using the retrieved acoustic room information
Patent Information
- Application Number
- PCT/EP2025/060950
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2025-04-22
- Publication Date
- 2025-11-27
AI Technical Summary
Existing binaural synthesis algorithms for audio rendering are computationally expensive, require complex room geometric data acquisition, and suffer from high motion-to-sound latency, making them unsuitable for real-time operation on mobile and wearable devices.
A method for retrieving acoustic room information using a low-resolution test fingerprint, decomposed into sub-bands, with energy-related characteristics calculated and parameterized to select a high-quality RIR from a database optimized for binaural audio synthesis, utilizing psychoacoustic weighting and vector databases for efficient retrieval.
Enables high-quality binaural audio rendering on resource-constrained devices by selecting perceptually similar RIRs from a database, reducing computational complexity and latency, and improving sound localization.
Smart Images

Figure EP2025060950_27112025_PF_FP_ABST
Abstract
Description
[0001] Apparatus, Method and Computer Program for Retrieving an Acoustic Room Information and Audio Signal Processor Using the Retrieved Acoustic Room Information
[0002] Specification
[0003] The present invention relates to an apparatus, method and computer program for retrieving an acoustic room information and additionally to an audio signal processor that uses the retrieved acoustic room information. The usage of the room information can be for an audio reproduction such as binaural production via headphones or speakers. The present invention relates to the processing of digital audio signals together with acoustic data describing the acoustic environment.
[0004] State of the art binaural audio rendering systems allow users to simulate and listen to virtual sound sources, which are precisely localizable in space. The simulated sounds seem to originate from outside of the head, which is called “externalization”. With an appropriate system, binaurally rendered sound sources can be perceived at a stable position in space and seem to have similar acoustic properties to real sound sources. This can make them virtually indistinguishable from real sound sources.
[0005] A number of Binaural Synthesis methods and algorithms exist, which can be used to achieve externalization. They have in common, that they aim to approximate filter effects that the sound is subjected to on its simulated path to the listeners ear. The combined filters of the system, consisting of a sound source, the acoustic influence of a virtual or real environment and its geometry, the listeners head and body and potentially other influences on the sound, caused by the environment, are called a Binaural Room Impulse Response (BRIR).
[0006] The two main components of a BRIR are the Head Related Transfer Function (HRTF) and the Room Impulse Response (RIR). The HRTF encodes the measured or approximated filter effects of the head, torso and outer ear of a human. As such, it is dependent on the listeners head and geometry and the relative position and rotation of head and sound source.
[0007] The RIR encodes filter effects of the Room, i.e. reflection, diffraction and shadowing of the sound, introduced by room geometry. It is dependent on the room geometry and the positions and rotations of the listener and sound source inside of the room. (Room hereby refers to any kind of environment, not limited to buildings.)
[0008] Simulating these effects is often done by means of complex simulations or more lightweight approximations, which require a complex room geometric model to simulate convincing room impulse responses. Depending on the used binaural synthesis algorithm, current state of the art algorithms often have to make a trade-off between computational complexity, limiting the lower size bound of target systems, or effectiveness of the simulation, often resulting in badly localizable or completely in-head localized sound sources.
[0009] These devices require room geometric data of the current room, including reflective surfaces, their absorption and scattering coefficients. This data is hard to acquire, especially in Augmented Reality (AR) settings, where the use of the device is not limited to a single room. Acquiring it is usually not feasible, even for trained users and measuring it automatically is a daunting task.
[0010] Depending on the employed Binaural Synthesis algorithms and techniques, these processes can be very computing intensive and time consuming. However, processing power is often limited on the target devices. For instance, the binaural rendering might be deployed on “T rue Wireless Earbuds” or similar smart headphones or wearables, which only provide very limited processing power to provide an adequate battery life.
[0011] These devices are often wirelessly coupled with other devices, like a smartphone, via Bluetooth or a similar wireless protocol. However, these connections introduce an additional delay by requiring coding, conversion and transmission over the air. This delay usually far exceeds the required maximum motion-to-sound latency, which is required to achieve externalization. Motion-to-sound latency here describes the time frame, which the binaural audio system requires to auralize acoustic changes, caused by a user’s head movement. The exact audibility threshold for motion-to-sound latency varies and is dependent on the listener, the signal employed and the acoustics of the environment. A latency of at most 50 ms has been determined as a valid threshold, which is inaudible to most users under most circumstances.
[0012] In order to produce convincing virtual sound sources, binaural signals and binaural filters are usually updated at this high rate. Depending on the employed binaural synthesis methods, this results in a computational complexity, which is often too high for mobile and wearable devices. Instead, such devices are often cable connected to another computing device, which handles these calculations.
[0013] The publication “Proof of Concept of a Binaural Renderer with Increased Plausibility” by II. Sloma, et al., DAGA 2023 Hamburg, pages 208-211 describes a proof of concept demo showcasing the comparison of a real loudspeaker setup and headphones based rendering in a given room. The room acoustic processing has been included and is processed in runtime. Particularly, Binaural Room Impulse Responses (BRIRs) are calculated in realtime, based on a single omnidirectional Room Impulse Response (RIR). A very basic room geometric model as well as the positions of the sound sources and microphone need to be captured. From that, the Directions of Arrival (DOAs) of the direct sound and early reflections are estimated by a simplified image source model. The RIR is processed in segments and appropriately convolved with generic HRTF filters. Late reverberation is simulated by noise shaping. This algorithm allows a 6DoF rotation and translation. The Spatial Decomposition Method is discussed. This method uses one measurement microphone and six electret condenser microphones. It is assumed, that the sound field consists of a sequence of individual acoustic events. They can be described with the captured RIRs and the captured DOAs. In the post-processing, the HRIRs are calculated for the measurement position with a 3DoF rotation and a generic HRTF filter.
[0014] The publication “Creation of Auditory Augmented Reality Using a Position-Dynamic Binaural Synthesis System - Technical Components, Psychoacoustic Needs, and Perceptual Evaluation”, S. Werner, et al., Applied Sciences, 2021 , 11 , 1150, discloses a positiondynamic binaural synthesis system that is used to synthesize the ear signals for a moving listener. The goal is the fusion of the auditory perception of the virtual audio objects with the real listening environment. For each possible position of the listener in the room, a set of binaural room impulse responses (BRIRs) congruent with the expected auditory environment is required to avoid room divergence effects. The required spatial resolution of the BRIR positions can be estimated by spatial auditory perception thresholds. Particularly, a specific position-dynamic binaural synthesis system relies on a pre-processing of the room geometry, a spatial resolution of reproduction, a listening position representation, a real-time processing block comprising get tracking data and processing and a convolution engine, and a filter creation block comprising the listening positions and the BRIR synthesis. The result of the BRIR synthesis are binaural filters that are used by the convolution engine in the real-time processing block for the purpose of position-dynamic binaural playback. A constant reverberation, an acoustically shaping, a synthesis approach adapting an initial time delay gap (ITDG), a sound source directivity and a real-time processing are discussed.
[0015] The publication “Binauralization of Omnidirectional Room Impulse Responses - Algorithm and Technical Evaluation”, C. Pdrschmann, et al., Proceedings of the 20thInternational Conference on Digital Audio Effects (DAFx-17), Edinburgh, UK, September 5-9, 2017, pp. 345-352 discloses a binauralization of omnidirectional room impulse responses algorithm which synthesizes BRIR data sets for dynamic auralization based on a single measured omnidirectional room impulse response (RIR). Direct sound, early reflections, and diffuse reverberation are extracted from the omnidirectional RIR and are separately spatialized. Spatial information is added according to assumptions about the room geometry and on typical properties of diffuse reverberation. The early part of the RIR is described by a parametric model. Modifications of the of the listener position can be considered. The late reverberation part is synthesized using binaural noise, which is adapted to the energy decay curve of the measured RIR. The direct sound frame starts with the onset of the sound and ends after 10 ms. The following time section is assigned to the early reflections and the transition towards the diffuse reverberation. Sections with strong early reflections are determined. Following this procedure, small window sections of the omnidirectional RIR are extracted describing the early reflections. The incidence directions of the synthesized reflections base on a spatial reflection pattern adapted from a shoebox room with non- symmetric positioned source and receiver. A fixed lookup-table containing the incidence directions is used. By this, a parametric model of the direct sound and the early reflections is created. Amplitude, incidence direction, delay and the envelope of each of the reflections are stored. By convolving each window section of the RIR with the HRIR (cp) of each of the directions, a binaural representation of the early geometric reflective part is obtained. To synthesize interim directions between the given HRIRs, interpolation in the spherical domain is performed. The early part of the single measured omnidirectional RIR contains the direct sound and strong early reflections. For this part, the directions of incidence are modelled reaching the listener from arbitrarily chosen directions. The late part of the RIR is considered being diffuse and is synthesized by convolving binaural noise with small sections of the omnidirectional RIR. By this, the properties of diffuse reverberation are approximated. The synthesized BRIRs can be adapted to shifts of the listener and freely chosen positions in the virtual room can be auralized.
[0016] It has been found that existing BRIR synthesis algorithms suffer from several disadvantages that render the processing computationally expensive, that result in a non-natural sound perception by a listener, that are problematic in adapting the system to a specific source characteristic, position or orientation or a specific listener position or orientation in an efficient way, or that can even forbid the system to be operating in real time. An additional disadvantage can be that artefacts are created that result in a reduced externalization of the sound impression which contributes to a non-natural and unpleasant feeling for the listener.
[0017] The international application PCT / EP2023 / 079658, published as WO 2024 / 089036 A1 on May 2, 2024 after the priority date of this application and before the application date of this application discloses an audio signal processor for generating a two-channel audio signal that comprises an interface for providing single-channel acoustic data describing an acoustic environment, a two-channel synthesizer for synthesizing two-channel acoustic data from the single-channel acoustic data using a listener position or rotation, and a sound generator for generating the two-channel audio signal from audio signal and the two- channel or acoustic data. The input interface is configured to acquire a raw representation related to the single-channel acoustic data and to derive the single-channel acoustic data using the raw representation and additional data stored in the audio signal processor or accessible by the audio signal processor.
[0018] The application describes to acquire, as the raw representation, an initial measurement of raw single-channel acoustic data, to derive a test fingerprint, to access a pre-stored database with an associated set of reference fingerprints, wherein each reference fingerprint is associated to a high-resolution single-channel acoustic data, wherein the high- resolution single-channel acoustic data has a higher resolution than the initial measurement, and to retrieve, from the pre-stored database the high-resolution singlechannel acoustic data having a reference fingerprint best matching with the test fingerprint.
[0019] The input interface is configured to determine, as the test fingerprint, at least one of the following parameters: RT60 (Reverberation Time 60), EDC (Energy Decay Curve), DRR (Direct Reverberation Ratio), and wherein the reference fingerprint comprises at least one of the parameters RT60, EDC, DRR.
[0020] In the meantime, it has been found that the way of calculating the test fingerprint on the one hand and the reference fingerprint on the other hand is of high importance, particularly when the raw input data is a low-resolution measurement of room impulse response. It is an object of the present invention to provide an improved concept for retrieving an acoustic room information from database or for generating an audio reproduction.
[0021] This object is achieved by an apparatus for retrieving an acoustic room information from a database in accordance with claim 1 , an audio signal processor in accordance with claim 21 , a method for retrieving an acoustic room information from a database in accordance with claim 33, a method of generating a two-channel audio signal in accordance with claim 34, or a computer program in accordance with claim 35.
[0022] The present invention is based on the finding that an optimum test fingerprint, suitable for a low-resolution test acoustic room information is obtained by decomposing the test acoustic room information into a plurality of acoustic room information sub-bands. A subband parameter for each acoustic room information sub-band of the plurality of acoustic room information sub-bands is calculated, wherein the sub-band parameter indicates an energy change over time in the acoustic room information sub-band. The test fingerprint, and, preferably also the reference fingerprint are based on the sub-band parameters for the acoustic room information sub-bands of all or at least a subset of the plurality of acoustic room information sub-bands obtained by the filter bank of the test finger processor.
[0023] Preferably, an energy-related characteristic over time is calculated and this energy-related characteristic over time is parameterized to obtain the sub-band parameter for the acoustic room information sub-band. The energy-related characteristic is approximated by calculating one or more curve parameters by fitting a pre-defined curve to the energy- related characteristic over time to obtain one or more calculated parameters, where these parameters are used to represent the sub-band parameters for the acoustic room information sub-band. In an embodiment, the pre-defined curve is a straight line and the one or more parameters is a slope of the straight line. In other embodiments, other predefined curves and corresponding other parameters describing these curves are usable as well. Such other pre-defined curves can be a parabolic curve, a logarithmic curve or an exponential curve. Parameters for the pre-defined curves to be fitted are shaping of stretching or compressing parameters of the parabolic curve, the logarithmic curve or the exponential curve. In a preferred embodiment, an energy decay curve is calculated from the acoustic room information sub-band and the energy decay curve is converted into a logarithmic scale, and the result is subjected to linear regression operation in order to obtain the slope parameter of the line obtained by the linear regression processing. In the further preferred embodiment, the sub-band parameters are weighted using a psychoacoustic characteristic so that the parameter for a first acoustic room information sub-band in which a listening perception is greater than in a second acoustic room information sub-band has a higher influence of the test fingerprint compared to an influence of a parameter for the second acoustic room information sub-band on the test fingerprint.
[0024] The psychoacoustic characteristic can, for example, be the Fletcher-Munson-curve or an equal loudness curve or an equal loudness contour in accordance with the International Standard ISO226 or the absolute listening threshold over frequency.
[0025] Preferred embodiments describe a system to provide a Room Impulse Response (RIR) from a database of synthetic RIRs to be used by an ’RIR-provider’ as part of a system for the auralization of binaural audio as described in PCT / EP2023 / 079658. This international application is incorporated in its entirety by reference. The system works by finding the best RIR for the desired acoustic situation from a database of RIRs which are suitable and specifically designed for binaural synthesis.
[0026] The main purpose of the system is to provide a high quality RIR specifically designed and optimized for the auralization of binaural audio from a low quality RIR that could be measured by a simple method not requiring professional equipment or expertise in the field.
[0027] This disclosure depicts a particular family of methods of finding a RIR, and possibly a Binaural Room Impulse Response (BRIR), that is perceptually similar to a given input RIR. The reason for the use of such a system is that there are methods that use RIRs to auralize convincing acoustic environments, however such RIRs need to be of very high quality, with a good signal to noise ratio and very few spectral distortions (caused by the recording equipment) over the whole audible frequency spectrum.
[0028] Recording RIRs for this purpose requires professional quality microphones and recording procedures that are very complex, time consuming and aren’t suitable for non-experts in the field.
[0029] Using simple microphones, such as the ones available in smartphones results in low quality RIRs with very low signal to noise ratio (SNR), non-linearity and a limited frequency range not suitable for high quality auralization.
[0030] To overcome the problems and limitations of acquiring RIRs, it’s possible to extract a ’’fingerprint” of RIRs, in other words, a set of values carrying the most important information of a particular RIR, and to use these values to find RIRs that are perceptually similar to each other. This set of low-dimensional parameters carries the most essential and important information of a RIR, is robust to noise and overcomes the frequency range limitation of low quality RIRs. Once the ’’fingerprint” parameters are calculated they can be used to select a perceptually similar high quality RIR designed specifically for auralization from a database, given a low-quality recorded RIR. For best results, it’s required that the compared recordings are recorded or simulated at the same distance from the sound source.
[0031] In the case of the Brandenburg Labs specific auralization and BRIR synthesis system, relevant parameters are mostly energy based, such as the bandwise energy over time, reverberation time (RT60), and others. The reason for that is that the BRIR synthesis modifies the frequency and phase content of an RIR in a way that fits the real room. The Direction of Arrival (DOA) of the reverberant parts, its phase content and therefore its interaural coherence are modeled by later binaural synthesis steps, just like a sources’ directivity as described in PCT / EP2023 / 079658.
[0032] It’s been seen through extensive experiments conducted by Brandenburg Labs that Energy Decay Curves (EDCs) carry the necessary information needed by the binaural audio rendering system described in PCT / EP2023 / 079658 to produce virtual sound sources that are perceived practically indistinguishable from real sound sources.
[0033] The proposed system processes EDCs in a way to extract a lower dimension set of values that are used to select the best match in a database of RIRs that are designed for binaural rendering purposes comprising different acoustic situations. This approach allows an efficient method to achieve the necessary sounding quality from a low quality measured RIR coming from different classes of sounds such as human clapping recorded by a smartphone or other automated and simple measurement methods.
[0034] A base RIR, that could be from a low quality measurement for example, is given to the system as an input. The system processes the input RIR and compares it to the synthetic RIRs from a database designed specifically for high quality binaural audio rendering. The best match is returned.
[0035] The dataset of synthetic RIRs is constructed using well-known simulation methods and contains impulse responses of a diverse range of rooms with different acoustic characteristics. These RIRs are designed using known parameters such as room geometry, absorption and scattering coefficients, distances from source to sink, that can be used for different binaural synthesis applications.
[0036] The set of values used by the selection mechanism can be pre-computed, bringing time and memory efficiency to the whole system.
[0037] In general terms the system can be implemented by calculating the band-wise EDCs of RIRs, logarithmic scaling the energy, finding suitable reference points in the logarithmically scaled energy curve to define a straight line ( e.g. -3dB from the maximum point as the first reference point and the -60 dB (RT60) point as the second reference point) and finally obtaining the slope of the linear regression line. This procedure is performed for multiple sub-bands and the slopes’ inclines are saved as the “fingerprint” set of values. A simple distance measure between the set of values, such as the L1 norm, is used to find the most similar RIRs.
[0038] In an embodiment, the present invention comprises an RIR Selection mechanism system being composed by 3 main parts:
[0039] • Computing the band-wise EDC for the RIRs.
[0040] • Calculating linear regression coefficients for each band of the logarithmically scaled EDC.
[0041] • Selecting the most suitable RIR from a database of RIRs based on the coefficients calculated in the step above.
[0042] This mechanism processes all the RIRs in the database as well as the base RIR used as input. The database can be constantly expanded and the processed values can be stored for speed and memory efficiency. The process can be further computationally improved by employing common space partitioning algorithms.
[0043] Preferred embodiments of the present invention are subsequently discussed with respect to the accompanying drawings, in which:
[0044] Fig. 1 illustrates a processing stage in accordance with an embodiment;
[0045] Fig. 2 illustrates an approximation of the 40-phon Fletcher-Munson curve for subbands k;
[0046] Fig. 3 illustrates a distance’s ranking stage; Fig. 4 illustrates an RIR waveform (clap recording in a room using a smartphone);
[0047] Fig. 5 illustrates energy decay curves;
[0048] Fig. 6 illustrates a slope calculation for sub-band k = 15;
[0049] Fig. 7 illustrates slopes for all sub-bands k;
[0050] Fig. 8 illustrates weighted slopes of sub-bands k;
[0051] Fig. 9 illustrates 21 x 5 vectors mirrored across the origin (pi and p2=) and query vector (q);
[0052] Fig. 10 illustrates a selection stage;
[0053] Fig. 11a illustrates a preferred embodiment of the apparatus for retrieving an acoustic room information from a database;
[0054] Fig. 11b illustrates a preferred implementation of the test fingerprint processor;
[0055] Fig. 12 illustrates a further preferred implementation of the test fingerprint processor;
[0056] Fig. 13 illustrates a preferred implementation of two-stage fingerprint matching process related to the Fig. 10 embodiment;
[0057] Fig. 14 illustrates a preferred implementation of the audio signal processor for generating a two-channel audio signal;
[0058] Fig. 15a illustrates a preferred implementation of the apparatus for retrieving an acoustic room information in the form of high resolution single-channel acoustic data;
[0059] Fig. 15b illustrates another preferred implementation of the apparatus for retrieving an acoustic room information in the form of high resolution single-channel acoustic data; Fig. 16 illustrates a preferred implementation of the two-channel synthesizer for Fig. 14;
[0060] Fig. 17 illustrates a preferred implementation for the calculation of the specular reflections;
[0061] Fig. 18 illustrates a preferred implementation for the calculation of the early reflection part for the first channel and the second channel using a specular part and diffuse part for a segment;
[0062] Fig. 19a illustrates a preferred implementation of the early reflection and / or the late reverberation part of the two-channel acoustic data from the high resolution single-channel acoustic information;
[0063] Fig. 19b illustrates a further embodiment of the procedure also illustrated in Fig. 19a; and
[0064] Fig. 20 illustrates a schematic representation of an implementation of the audio signal processor with two separate entities, i.e., for example a wearable device with a limited battery or computational resources on the one hand and another device with more computational resources or battery resources.
[0065] Fig. 11a illustrates a preferred implementation of the present invention in the form of an apparatus for retrieving a reference acoustic room information from a database. The apparatus comprises at test fingerprint processor 10 and a connected database processor 20.
[0066] The test fingerprint processor 10 is configured to receive test acoustic room information. In an embodiment, the test acoustic room information is a measured or synthesized room impulse response (RIR) as can, for example, be measured by a simple smartphone with the embedded smartphone microphone in response to a clap or any other loud and short acoustic event that can be easily generated by a user of the device, for example.
[0067] The test fingerprint processor 10 derives a test fingerprint from the test acoustic room information and forwards this test fingerprint to the database processor 20 that retrieves a reference acoustic room information having a reference fingerprint matching with the test fingerprint. The test fingerprint comprises sub-band parameters for energy changes over time and individual sub-bands calculated from the test acoustic room information.
[0068] The reference acoustic room information can be a multi-channel acoustic room information or, alternatively, a two-channel acoustic room information or, in a further preferred embodiment, a single-channel acoustic room information that is used by the two-channel synthesizer 200 of Fig. 14 to calculate a two-channel acoustic data such as a BRIR or a BRTF.
[0069] Although significant focus is given to the calculation of a single-channel acoustic room information in response to the processing of the test fingerprint processor 10 and the database processor 20 of Fig. 11a, it to be emphasize that, for other embodiments, the processing in Fig. 14 can be skipped and the already true two-channel acoustic data is retrieved from a database having stored such two-channel acoustic channel data.
[0070] Fig. 11 b illustrates a preferred implementation of the test fingerprint processor 10. The test fingerprint processor 10 comprises a filter bank 11 which is, for example, implemented as a fractional octave filter bank. However, other filter bank implementations such as an MDCT, MDST, an FFT or a QMF filter bank can be used as well as the case may be. However, the fractional octave filter bank is preferred due to the good matching with the psychoacoustic weighting discussed with respect to item 16 in Fig. 12 and discussed with respect to Fig. 2 and item 16 of Fig. 1.
[0071] The filter bank 11 provides a plurality of acoustic room information sub-bands 0, 1 , 2, ... , k due to a decomposition of the test acoustic room information input into the filter bank 11 .
[0072] The parameter calculator 12 calculates a sub-band parameter for each acoustic room information sub-band of the plurality of acoustic room information sub-bands, wherein the sub-band parameter indicates an energy change over time in the acoustic room information sub-band, wherein the test fingerprint is based on the sub-band parameters for the acoustic room information sub-bands.
[0073] Regarding the size of the filter bank, preferred embodiments refer to the implementation of a filter bank with 32 filter bank channels. However, it has been found that useful results are already obtained by using a filter bank having only four or more than four filter bank channels additionally, a preferred design of the filter bank for a low resource implementation has a number of filter bank channels between 10 and 20 such as 16 filter bank channels. A high resolution implementation can also comprise more than 32 filter bank channels such as 40- 50 filter bank channels.
[0074] The sub-band parameters indicating the energy change over time in a sub-band together represent the test fingerprint derived from the test acoustic room information.
[0075] In a preferred procedure, the parameter calculator 12 of Fig. 11 b receives the sub-band data and calculates an energy decay curve in a step 13. The result of step 13 is an energy decay curve for each sub-band as provided by the filter bank 11. In step 14, the energy decay curve is converted into a logarithmic scale and step 15 performs a curve fitting such as a linear regression to the result of the processing in step 14. As illustrated in Fig. 12, the conversion step 14 is optional and can be bypassed so that the energy decay curve is used for curve fitting without any logarithmic conversion. The curve fitting parameters will be different from linear parameters such as a slope of a line.
[0076] In step 16, a psychoacoustic weighting is performed and, as shown in Fig. 12, this psychoacoustic weighting is not only performed for a single sub-band, but for other subbands that are also processed via steps like steps 13-15 and these unweighted parameters from other sub-bands are also subjected to the psychoacoustic weighting 16 so that at the output of block 16, there are weighted sub-band parameters that together form the test fingerprint for the test acoustic room information.
[0077] In the course of computing the band-wise EDC for the RIRs, the system loads a RIR and calculates the EDC by using the well-known Schroeder’s decay function, which is the tail integral of the squared impulse response at time t.
[0078] The band-wise EDC is calculated by using a fractional octave filter bank to decompose the RIR into sub-bands and the EDC is calculated for each sub-band.
[0079] Sub-bands with low energy are discarded so the system is more robust and sensitive to sub-bands with an appropriate energy level. These computations are performed by the first two blocks in the diagram of Figure 1 illustrating a processing stage. An RIR 150 is input into an fractional octave filter bank 11 and the subbands are input into the EDC processing stage 13. The EDCs for the subbands are input into the log and slope block 14, 15. The coefficients are input into block 16 for the purpose of weighting to obtain the weighted coefficients per subband representing the test fingerprint (and the reference fingerprint).
[0080] A calculation of linear regression coefficients for each band of the EDC is carried out. For each sub-band the EDC is converted to decibels (dB), a space in which an exponential decay resembles a straight line. Two points are chosen to define a straight line, the start of the line at -3dB from the maximum and the end of the line at -60dB (RT60). Using the well- known equation of a straight line the incline or slope of a line is calculated. y x) — m ■ x + c , where m is the slope
[0081] To take into the account how humans perceive sound, these coefficients can be optionally weighted using the Fletcher-Munson equal loudness curve shown in Figure 2 illustrating an approximation of the 40-phon Fletcher-Munson Curve for Sub-bands k. This way the most significant coefficients according to human perception are given more importance in the selection process.
[0082] These calculated and weighted values are now the set of values to be used for the selection of the best suitable RIR within a database. These computations are performed by the last two blocks in the diagram of Figure 1.
[0083] In an embodiment, a selection of the most suitable RIR from a database is performed. The so called L1 (or Manhattan) pairwise distance between the weighted coefficients of the base RIR used as the input of the system and the weighted coefficients of all the RIRs in the database is calculated, and the RIR in the database with the smallest distance to the input base RIR is selected as the best match. The selection process is described by Figure 3 showing a distances ranking stage 46. The system checks (40) if an RIR has been previously processed (yes), if not (no) the RIR is processed in block 44 and the system can select the most similar RIR with the smallest distance. In block 150, the base RIR or test RIR is input and block 110 represents a RIR database. If the answer in block 40 is yes, the previously processed (and stored) coefficients are loaded in block 42.
[0084] In the following, a processing method example for an embodiment of the invention is described. This example shows the complete processing system for a RIR acquired by recording a clap inside a room using a smartphone as shown in Fig. 4 illustrating a RIR Waveform, i.e., a clap recording in a room using a Smartphone.
[0085] The RIR passes through a 3 bands per octave fractional octave filter band and the EDC is calculated for each sub-band as shown in Fig. 5 illustrating the energy decay curves (EDC).
[0086] The next step shows one of the sub-bands after being logarithmically scaled and the calculated slope of the linear regression line. See Fig. 6 showing a slope calculation for sub-band k=15.
[0087] The slopes are calculated for all sub-bands as shown in Figure 7.
[0088] These slopes can be optionally weighted using the 40-phon Fletcher-Munson curve and the RIR’s “fingerprint” is the set of values of weighted slopes of all sub-bands as shown in Figure 8.
[0089] Subsequently, further embodiments of the invention are described. The system provides a way to store pre-computed coefficients and to check for pre-processed RIRs resulting in an improvement in speed and memory efficiency. Apart from the L1 distance, there’s also the possibility to use other distances such as the L2 (Euclidean distance), the cosine distance and other well-known distances used for measuring proximity between vectors in a vector space. Different filter banks can be used to decompose the RIR into sub-bands.
[0090] Subsequently embodiments for storage and retrieval of similar acoustic fingerprints are illustrated. The audio processing operations performed by the inventive audio processor exemplarily described with respect to Figs. 14 to 20 require high-quality audio recordings to perform optimally. Unfortunately, not everyone has access to high-quality recording equipment. In order to combat this issue, the audio processor exemplarily described with respect to Figs. 14 to 20 employs an acoustic fingerprinting system to produce a fingerprint from the most pertinent acoustic parameters of a low-quality audio recording and uses that fingerprint to find a similar recording from a database of high-quality audio recordings. The purpose of this document is to outline the basic process of storing and retrieving acoustic fingerprints from a database of acoustic fingerprints. The functionality of the audio processor exemplarily described and how two-channel acoustic data are retrieved from the single channel acoustic data retrieved from the database are illustrated with respect to Figs. 14 to 20. Numerous database architectures exist with each architecture focusing on a specific type of data, e.g. relational or document-oriented, so when choosing a database architecture it is useful to choose the appropriate database for a particular data representation. For a use case, the most appropriate type of database is a vector database. In recent years, vector databases such as Qdrant and Pinecone have emerged as an efficient way of storing and comparing higher-dimensional data representations as a set of vectors, allowing for rapid calculation of similarity measures. Embodiments employ a vector database along with various advanced querying techniques to identify perceptually similar acoustic fingerprints.
[0091] The acoustic fingerprints consumed by the inventive audio processor exemplarily described with respect to Figs. 14 to 20 preferably represent room impulse responses (RIRs) and encapsulate the characteristics of specular and diffuse reflections at a given point in a 3- dimensional space. Specular and diffuse reflections are produced when a sound collides with a reflective surface and contribute to how one perceives an environment. Specular reflections occur when a sound collides with a reflective surface, producing a reflection with an angle equal to the angle of incidence on the opposite side of the surface normal. Specular reflections arriving at a listener’s ear have traveled the minimum distance between the listener and the sound source and provide useful perceptual cues about the geometry and spatial characteristics of the acoustic scene. Irregularities in the reflective surface will act as individual micro surfaces producing diffuse reflections, reflections with an angle that is not equal to the angle of incidence, meaning they can scatter in all directions. This scattering effect is dictated by the size of the irregularities, with irregularities that are large with respect to wavelength resulting in a larger level of diffusion.
[0092] An important distinction between specular and diffuse reflections is that specular reflections retain the phase and magnitude information from the sound that produced them, while diffuse reflections are not guaranteed to retain the phase and magnitude. The phase and magnitude differences of the diffuse reflections combine to produce a complex auditory masking effect, providing texture and color to an acoustic scene. Plausibly reproducing the interplay between specular and diffuse reflections produced by a particular environment is of importance for preferred embodiments. One technique that preferred embodiments use to achieve this is to search the fingerprint database for fingerprints with matching diffuse characteristics and synthesize the specular reflections. Due to the nature of diffuse reflections, they are computationally more expensive to synthesize than specular reflections and difficult to perfect, as a single specular reflection may produce numerous diffuse reflections. Regardless of their type, databases are able to efficiently retrieve data in part due to the ways in which they store data. Techniques such as sharding, partitioning, and indexing are used to organize data points in a way that reduces the size of the search space required to answer a query. Vector databases optimize these techniques with respect to distance calculations, making them ideal for similarity calculations.
[0093] Unlike relational databases that store data as a set of key / value pairs, vector databases store higher-dimensional data as a set of vector embeddings, a representation of the data that is optimized for machine learning operations. It is useful to define what a vector is within the given context. In the realms of physics and mathematics, the term vector generally refers to a quantity with magnitude and direction. In the context of machine learning, the term vector is used to represent multilinear data represented as a 1xn dimensional tensor. For the remainder of this document, when it is referred to vectors, it is exemplary referred to them within a machine learning context.
[0094] An additional benefit of vector databases is the ability to store vectors along with their metadata. Metadata can be stored as structured or unstructured data with some database management systems, e.g. PostgreSQL, coupling relational and vector databases to allow for more complex data querying and analysis.
[0095] By transforming an audio recording into an acoustic fingerprint, the original audio signal has been effectively condensed into its most useful characteristics. The mechanism by which the audio signal is transformed into a fingerprint may vary depending on which particular acoustic parameters are pertinent for a given task, including the energy decay curve (EDC), and further including but not limited to reverberation time (RT60), direct-to- reverberant ratio (DRR), etc. What is useful is that the output of the fingerprinting process can fit in a 1xn vector. Fingerprint vectors are stored in a vector database along with their metadata while their associated audio files are stored together on hard disk.
[0096] Retrieving a similar fingerprint from a database of acoustic fingerprints involves performing a nearest- neighbor search using some distance metric. Although the process appears to be straightforward, there is not always a deterministic solution. Choosing to use an exact method vs. an approximation for nearest-neighbor search and a choice of distance metric can change the results. In addition, different types of queries will produce different results. In this section, the process of retrieving a matching fingerprint from a database of acoustic fingerprints is briefly discussed. The nearest-neighbor problem is a well-known problem in computer science with applications in a myriad of fields. Donald Knuth posed a famous version of nearest-neighbor search as the post- office problem, with the task of identifying the closest post-office to a particular address. Apart from topology, nearest-neighbor search has applications in fields as varied as computer vision, spell checking, semantic search, and much more. A common classification technique in machine learning operations is k-means clustering, in which a data point is classified based on the classification of a majority of its k nearest neighbors.
[0097] The brute-force approach to solving the nearest-neighbor problem can be stated as the search for the point p* from the set of points J> = {p1,p2, --- , pn] that minimizes the distance to all dimensions of a query point q. The nearest-neighbor problem is defined by the following equation.
[0098] P piEp' - pi, Qi)
[0099] The brute-force approach to solving the nearest-neighbor problem involves iteratively calculating the distance between q, and p, for every p in J> and thus has a time complexity of O(dri) where n represents the number of points in J> and d represents the dimensionality of each point. Although a time complexity of O(n) is fine for many applications, in this particular case it is compounded by the dimensionality of d. As the complexity of a bruteforce search scales linearly with the number of data points, an increase in dimensionality can have a drastic impact on the amount of time it takes to perform a query.
[0100] In practice, brute-force searches should be avoided, and approximation should be used whenever possible. The approximate nearest-neighbor (ANN) search differs from the bruteforce method in that it does not search for the exact nearest-neighbor and instead searches for a neighbor within some radius, r, of the query. It should be noted here that the accuracy of ANN is highly dependent on the chosen value of r, thus care must be taken when choosing a value for r. Although both ANN and brute-force each attempt to minimize the distance between the query vector and the search result, ANN stops searching when a candidate is found within rof the query. ANN can be defined by the following equation: p* = min d(pi, Qi) < r piEP
[0101] Whereas the brute-force method is guaranteed to identify the exact nearest-neighbor, ANN returns the first candidate that meets the criteria in Equation 2. In most real-world applications approximation produces results comparable to brute-force while providing drastic improvements to computation time. By combining early stopping with indexing and a space-partitioning data structure, such as a k-d tree, one is able to drastically reduce the time complexity related to searching operations. The introduction of these optimizations reduces the time complexity from O(n) for brute-force search to O(log ri) for ANN. Unless a case can be made justifying the necessity of an exact solution, approximation should be used in lieu of brute-force.
[0102] Choosing the correct distance metric for a particular task is essential in order to produce accurate results. The most commonly used techniques for examining distance or similarity between two vectors are the L1 distance, the L2 distance, and cosine similarity.
[0103] The L1 distance is also known as the Manhattan distance because it measures the distance between two points as the number of discrete steps between them, similar to the navigating of Manhattan city blocks. The L1 distance between two vectors is defined as the summation of the distance between each pair of points, given by the equation below:
[0104] The L2 distance, also known as the Euclidean distance, is a distance metric that examines the distance between two points in the Euclidean space. Although the L1 and L2 distance formulas are similar, the L2 distance formula tends to scale the coefficients more evenly than the L1 distance formula. The definition of the L2 distance between two vectors is given by the following equation:
[0105] Cosine similarity is a method for measuring the similarity of two vectors by examining the cosine angle between the two vectors. As such, cosine similarity is not concerned with vector magnitude, merely the angle of the vector. The cosine similarity of two vectors is defined as 1 minus the dot product of the vectors divided by the product of their lengths, as shown in the following equation. It should be noted that the cosine similarity is identical to dot product similarity with an additional step to normalize the magnitudes of the vectors, hence dot product similarity should be used in lieu of cosine similarity in instances where vector magnitude needs to be taken into consideration.
[0106] In Figure 9 one sees two 1x5 dimensional vectors, pi and p2, together with a query vector q. The vector pi is equivalent to p2 mirrored across the origin and q was chosen to be equidistant to both pi and p2. In this contrived example, one can see how difficult it can be to determine the nearest neighbor. This example was designed to show the importance of choosing the correct distance metric. Calculating the L1 or L2 distance will result in the same distance for pi and p2. However, calculating cosine similarity or the dot product will result in different values for pi and p2, with the dot product retaining the magnitude information.
[0107] The EDC mostly encodes timbre and reverberation time, but does not account for the reflection pattern of the early reflections. The subset of best matches resulting for the selection process can be searched again by comparing another ’’fingerprint” which encodes the reflection pattern. A simple method to implement that is by using the cross-correlation of the squared signal, where the maximum value is a measure of high energy segments (reflections) at the same time. Any other procedures different from a cross-correlation processing can be used as well such as the described matching procedures comprising a Manhattan or L1 distance, an Euclidean or L2 distance, a cosine distance or a distance for measuring a proximity between vectors in a vector space or features in a feature space.
[0108] Fig. 13 illustrates a preferred implementation of the present invention in order to refine the database matching / selection process. Particularly, an initial database search with the test fingerprint is performed as discussed before with respect to Figs. 1-9, 11b, 12.
[0109] Particularly, an initial database search with the test fingerprint is performed as discussed before with respect to Figs. 1 to 9, 11b, and 12. As illustrated in item 34, an additional test fingerprint is calculated from the test acoustic room information. The initial database search performed in block 30 by the database processor 20 of Fig. 11a results in n best matching reference acoustic room information, i.e. , a group of at least two, for example room impulse responses found in the database in response to the test fingerprint derived from the test acoustic room information. The database processor 20 is configured to retrieve 30 a group of at least two reference acoustic room information comprising the reference acoustic room information indicated at the output of the database processor 20 in Fig. 11a and at least one additional acoustic room information. The test fingerprint processor 10 is configured to derive or calculate 34 an additional test fingerprint from the test acoustic room information, the additional test fingerprint being different from the first test fingerprint illustrated at the output of block 10 in Fig. 11a.
[0110] The database processor 20 is configured to determine an additional reference fingerprint for the reference acoustic room information and the at least one additional acoustic room information included in the group of n best matching reference acoustic room information items. The database processor 20 is configured to perform a matching operation between the additional test fingerprint generated by block 34 and the additional reference fingerprints generated by a block 32 which are two or more such as n additional reference fingerprints to determine the reference acoustic room information in the group having the additional reference fingerprint best matching with the additional test fingerprint.
[0111] The additional test fingerprint for the test acoustic room information on the one hand and the additional reference finger print for the group of n best matching reference acoustic room information items are calculated using the same calculation rule. Preferably, the blocks 34, 32 operate to use the full band information of the test acoustic room information / reference acoustic room information. This hybrid approach is useful, since in the first database operation, subband data has been taken and in the second stage incurring only a small number of the small n best matching reference acoustic room information, the full band information is used. The number of n best matching reference acoustic room information can be any number between n = 2 and n = 100, but is preferably below 20 or even below 10 and most preferably, even below 5 or equal to 5.
[0112] In one embodiment, a certain frame of the test acoustic room information excluding the direct portion and preferably also excluding the late reverberation portion, i.e., only comprising the early reflection portion is taken. In this frame, each sample of the preferably used room impulse response is squared and the squared samples represent the additional test fingerprint.
[0113] Alternatively, the full band samples of the portion or frame of the test acoustic room information can be used as well without squaring or a different operation from squaring such as using any power greater than 1 and preferably lower than 6 can be used as well. Using a power being an integer multiple of 2, for example, is preferred in order to obtain an absolute value. Other measures for obtaining an absolute value can be used as well.
[0114] Alternatively, the used information from the test acoustic room information is not the full band information, but is a sub-band information, but the sub-bands are preferably broader than the sub-bands as calculated from the filter bank 11 of Fig. 11 b or Fig. 1.
[0115] In order to refine the selection method, an optional second search mechanism is provided. In this method, the correlation between a defined time frame of the square (raised to the power of 2) of the base RIR (provided at 150) used as the input of the system and the square of each one of the n best ranked RIRs in the database 110 according to the main selection method defined in the previous sections is computed and ranked in block 50 according to the maximum value of the resulting correlation vector. The (e.g. RIR) selection stage, including the cross-correlation search stage 32, 34, 36 is described in Fig. 10a and 13.
[0116] There are a number of techniques that can be employed to improve the quality of queries. One simple technique is vector weighting. Although distance metrics assume that each dimension is equally useful, it is often the case that some dimensions are more important than others. To counteract this, one can implement vector weighting by applying a weight to each dimension by multiplying a particular dimension by its contribution to the solution.
[0117] In cases where multiple different vector representations exist, hybrid queries and multistage queries can be used to fine-tune the results of similarity searches. For example one set of acoustic parameters is used to make a fingerprint of the diffuse reflections and another set of acoustic parameters to make a fingerprint of the specular reflections and one wants to search for the most similar fingerprint to some query fingerprint. Using a multistage query, one would first search one set of fingerprints to reduce the size of the search space and search the other set of fingerprints for the final result. Multi-stage queries are commonly formed so that for each successive query the smallest remaining vector is used to remove the bulk of the remaining candidates, resulting in drastically reduced computation costs. Hybrid queries are similar to multi-stage queries in that they both involve multiple vectors with the difference being that hybrid queries combine the ranks of the successive queries into one overall rank, favoring candidates that rank highly across all queries. One caveat to be aware of is that increasing the dimensionality of vectors may result in less accurate results, this is an example of the curse of dimensionality. An increase in the dimensionality of the vectors used results in an increase in the size of the search space of the problem and increases the number of samples in the database needed to represent the problem. If the database does not contain enough samples, the accuracy of results will decrease. If the database does contain enough samples, the computation costs increase dramatically with the number of samples. To combat this it is advisable to use multiple smaller fingerprints containing related parameters rather than one large fingerprint.
[0118] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
[0119] Subsequently, a preferred implementation of the test fingerprint processor 10 and the database processor 20 is discussed particularly with respect to Fig. 15a and 15b. Both procedures as illustrated in Fig. 15a and Fig. 15b and also illustrated with respect to Fig. 11a to 13 and also illustrated, for example, in Fig. 1 and the other figures can be used within the input interface 100 of Fig. 14 in order to derive, from a low resolution single-channel acoustic data, which is input into the interface 100, a single-channel acoustic data such as an RIR or RTF with a higher resolution.
[0120] Figs. 16 to 20 illustrate preferred implementations for the implementation of the two-channel synthesizer 200 of Fig. 14 and the sound generator 300 of Fig. 14 so that, in the end, the procedure of Fig. 14 provides an audio signal processor for generating a two-channel audio signal that comprises, on the one hand, an apparatus illustrated in Fig. 11a with elements 10, 20 collectively implemented in the input interface 100 of Fig. 14, a subsequently connected two-channel synthesizer 200 for synthesizing two-channel acoustic data from the single-channel acoustic data using a listener position or listener rotation, and a sound generator 300 for generating the two-channel audio signal from an audio signal and the two- channel acoustic data. As illustrated in Fig. 15a, the apparatus is configured to acquire 150, as a raw representation related to the single-channel acoustic data, the test acoustic room information, and to derive 151 the single-channel data as the reference acoustic room information using additional data stored by the audio signal processor or accessible by the audio signal processor, for example, from the database included within the database processor 20. The database processor 20 comprises the database or has access to a spate database or is implemented in a mix mode, i.e., comprises a certain limited database included in the database processor and, additionally, has an access to an external database in order to increase the storage resources.
[0121] Aspects of the invention start from single-channel acoustic data describing an acoustic environment and result in an audio sound generation relying on a two-channel acoustic data for the specific setting consisting of the acoustic environment, the one or more sources and the listener.
[0122] In accordance with a first aspect of the audio signal processor of the present invention, the specific source characteristic and, particularly, a directivity information of the sound source is integrated into a two-channel synthesis for the purpose of synthesizing two-channel acoustic data from the single-channel acoustic data. This integration of sound source directivity information can be particularly performed in the processing of the direct sound (DS) part of the single-channel acoustic data describing the acoustic environment. However, the integration of the directivity information allowing a natural reproduction of a sound source having a non-omnidirectional directivity characteristic can also be integrated in the processing of the early reflections (ER) part of the single-channel acoustic data, or the directivity information can even be integrated into both, the direct sound processing and the early reflection processing in an efficient way.
[0123] In accordance with a second aspect of the audio signal processor of the present invention, the specific processing of the early reflections (ER) part of the single-channel acoustic data is enhanced. Particularly, the early reflection part is segmented into a plurality of segments where each segment comprises a certain reflection. Particularly, a plurality of image source positions representing the sources of reflection sound are determined, and these image source positions are associated to the segments using an inventive matching operation that relies on a time of sound arrival calculated for each image source to the listener position in an initial measurement. A matching is performed in order to associate the time of sound arrival for each image source to a certain segment, i.e., to a certain reflection in the segment. By this, an automated and high quality association of image source positions to the different early reflections is obtained. By means of an additional integration of the directivity information not only for the direct sound, but also for the individual image sources, a certain orientation of an image source can also be accounted for in order to arrive at a more natural sound reproduction. In accordance with a third aspect of the audio signal processor of the present invention, processing of the early reflection part of the single-channel acoustic data describing the acoustic environment is enhanced by calculating the two-channel acoustic data for the early reflection part not only using a specular part describing distinct earlier reflections, but also accounting for a diffuse part describing a diffuse influence in the early reflection part. It has been found that although the “second part” of the room impulse response shows prominent early reflections, it does not consist of only these. Instead, even this early reflections part has a significant diffuse part that has an even increasing influence in the course from the beginning of the early reflections part to the end of the early reflections part i.e., near the beginning of the late reverberation part of the room impulse response. Therefore, by calculating the two-channel acoustic data describing the acoustic environment using a diffuse contribution even in the early reflection part has resulted in better natural auralization of an artificial sound scene, for example, by headphones or by speakers fed with the two- channel audio data generated by the sound generator using the two-channel acoustic data for the early reflection part relying on not only the specular part by also relying on the diffuse part.
[0124] In accordance with a fourth aspect of the audio signal processor of the present invention, that is related to an improved calculation of the late reverberation (LR) part of the singlechannel acoustic data such as the BRIR or BRTF (binaural room transfer function) relies on the specific generation of the two-channel late reverberation part by means of combining magnitude-data derived from the single-channel acoustic data and a preferably binaural two-channel noise sequence. The generation of two channels from one channel is done by using the same magnitudes but different phase values.
[0125] Particularly, a preferably binaural noise sequence consisting of two channels is converted into a spectral domain with a short-time Fourier transform or any other time domain / frequency domain conversion algorithm. This results in two spectrograms. The late reverberation part or the combination of the early reverberation part and the late reverberation part of the single-channel acoustic data is converted into a spectral representation as well preferably using the same transform algorithm. The two-channel acoustic data for the environment are derived by relying on the same amplitudes that may also be e.g. low-pass filtered for the actual generation with the two phase spectra and these two resulting spectrograms are transformed into the time domain to obtain the processed late reverberation part and preferably also the diffuse part of the processed early reflection part as has been discussed before with respect to the fourth aspect. The specific procedure of diffuse signal calculation can be applied to the late reverberation part only or can be applied to the calculation of the diffuse part of the early reverberation part only or can be applied, as it is the case in the preferred embodiment of the present invention, to the calculation of both the early reflection part and the late reverberation part. Particularly, for the calculation of the combined early reflection and late reverberation part, a separation into these parts is not necessary at all, since the calculation of the binaural diffuse portion is done without any knowledge on any separation of an early reflection part and a late reverberation part so that any such separation between the early reflection part and the late reverberation part is not needed at all for this aspect of the present invention. This approach results in a specific saving of computational resources. A high audio quality is obtained which is even sufficient so that, specifically for the calculation of the late reverberation part, any changes depending on a listener position or a source position or orientation do not have to be accounted for enhancing the efficiency of the algorithm further. Any changes depending on a listener position or a source position or orientation do not have to be accounted for the calculation of the diffuse part in the early reflection part of the room impulse response, too.
[0126] In accordance with a fifth aspect of the audio signal processor of the present invention, the problem is addressed how to efficiently and flexibly obtain a high quality single-channel acoustic data such as a single-channel room impulse response that has sufficient quality for obtaining a high quality auralization. To this end, the input interface is configured to acquire a raw representation related to the single-channel acoustic data and the input interface is additionally configured for deriving the single channel acoustic data using the raw representation and additional data stored in the audio signal processor or accessible by the audio signal processor. By means of an initial measurement relying on natural sounds producible by a user such as a user clapping with his or her hands or stamping on the floor with his or her feet or even a speech signal can be used instead of a typically used sine sweep signal which is a highly unnatural signal and can, of course, not be generated by a listener at all.
[0127] The provision of the initial measurement can be performed by a low quality microphone as is, for example, included in a laptop or mobile phone or so and, based on this raw representation related to the single-channel acoustic data, a synthesis or, in general, a generation of high quality single-channel acoustic data can be done using a database matching process relying on test and reference fingerprints or a synthesis can be done with a single or several neural networks that rely on the acquired raw representation such as the initial measurement or even only geometric data on the acoustic environment and, probably, also an intended source position and an intended or initial listener position.
[0128] This procedure effectively addresses the problem to have a single channel room impulse response that is good enough to perform a useful calculation of a head related impulse response based on a certain listener and source position.
[0129] In accordance with a sixth aspect of the audio signal processor of the present invention, the processing tasks can be distributed to several different devices having different power supplies. This allows to do most of the tasks on a wearable device such as a headphone, an earbud, an in ear element or so, while the second device is a device that has a large battery such as a mobile phone, a smart watch, a tablet or a notebook computer or a stationary computer.
[0130] Particularly, it has been found that the computationally most expensive part is the calculation of the late reverberation and, to some extent, also the calculation of the earlier reflection part. However, it has been found that the update rate for these procedures can be lower compared to the update rate of the calculation of the direct sound. On the other hand, the calculation of the direct sound is computationally inexpensive, since this part is only a short portion in time, and therefore, only requires short filters that can be very efficiently processed.
[0131] Therefore, the processing task of calculating the direct sound part can easily be performed by a low-power device such as a wearable device while the more demanding tasks are performed by the separate second device. The incurred transmission latency is nonproblematic, since a lower update ratio is sufficient for the calculations that are computationally more offensive, i.e., the calculation of the early reflection part and, particularly, the calculation of the late reverberation part that has, depending on the certain acoustic environments, a considerable length in time when the room impulse response is considered. Particularly, in reverberant rooms such as churches, the late reverberation part can extend over several seconds of diffuse reverberation.
[0132] In its most basic embodiment, the audio signal processor consists of a single device, which contains all necessary sensors, components and sound transducers. The system might include the necessary components in a headphone or ear plug form factor and do all processing directly on the device. In other embodiments, the system works on distributed devices. The disclosed system consists of three principal, functional components, which work together to create a binaural signal in real time. The first component provides an omnidirectional RIR such as an RIR recorded from an omnidirectional loudspeaker or with an omnidirectional microphone or preferably with both omnidirectional elements, which has desired acoustic properties and contains the relevant acoustic cues of the environment. This especially includes the frequency dependent energy distribution of the reverb over time. In one form, the RIR provider holds qualitative in-situ measurements of the RIR. utilizing a loudspeaker and an omnidirectional microphone. Supplementary, the system can estimate (psycho-)acoustic parameters of low quality RIR measurements or surrounding noise and synthesize RIRs from these parameters or pick suitable higher quality RIRs from a database. This system also incorporates a machine learning approach, e.g. for supporting the parameter estimation. If necessary, multiple RIRs can be blended to improve on transition areas between different acoustic environments, e.g. coupled rooms.
[0133] The second component is a Binaural Synthesizer, which takes a RIR and adds binaural cues to it, turning the RIR into a BRIR. The Binaural Synthesizer further receives room geometric information as an input. In an embodiment, the room geometric information consists of a shoebox geometry, which approximates the users real environment, by fitting a rectangular room consisting of six surfaces. This gives estimations about acoustically reflective surfaces in the environment, especially floor, ceiling and walls close to the listener. While the simplification by the shoebox room geometry already yields good results, improvements can come from more accurate geometry models of the room. The RIR itself is processed in split segments motivated by basic research in the field psychoacoustics. The direct sound describes the first sound wave that directly reaches the listener. Here, influences are given by the according HRTFs and DTFs as well as the distance law for sound propagation. These cues can be applied straight forward. For the reflections in the room represented in the RIR, there is a transition from specularity to diffuseness. The given RIR is combined with phase information from a binaural noise sequence to yield the diffuse layer of the BRIR. The early reflection segment is split into blocks, that are assigned with an estimation for the ratio between specular and diffuse energy. The blocks are convolved with HRTFs and optionally DTFs to get the directional part, which is layered with a snippet from the diffuse part on the according indices. After combining the three segments, a BRIR is complete. The binaural synthesizer is connected to positional sensors, which are able to determine the user's head rotation and additionally its position relative to a reference system. These pose information (“pose” stands for listener position and listener orientation or source position and source orientation) are provided in real time by the position tracking system. The virtual source poses are provided by a preset, optionally changing over time as moving sound sources. The Binaural Synthesizer is connected to a system, which yields measured or synthesized HRTFs corresponding to directions of arrival. Similarly, a part of the system deploys directivity transfer functions (DTF) of a sound source, depending on the relative position. Like for HRTFs, the DTFs might be derived from measurements or a synthesis process. The synthesized BRIR is sent to the auralizer, where it is convolved with an Audio Signal in real time. For this, a state of the art blockwise real time convolution method might be used.
[0134] The resulting binaural audio signal is played back over headphones, but cross-talk cancelled loudspeakers might also be employed. In order to keep the illusion of an externalized sound source plausible, the BRI Rs need to be resynthesized regularly with current positional data. In some embodiments, the three segments might be calculated at different rates while maintaining the immersive experience. The described system represents a novelty in the field of binaural synthesis. It makes it possible to experience lifelike virtual, spatial sound.
[0135] Embodiments use psychoacoustic knowledge to both decrease the computational complexity of the system and to allow a distributed calculation of the binaural synthesis on devices, which are connected by transmission channels which add a greater delay to the signal processing than is otherwise acceptable.
[0136] BRI Rs combine multiple filter effects. They can be split at arbitrary points in time, resulting in any number of sub filters. They can be reassembled by either summing the individual parts, with respect to their individual delays, or by convolving the filter with a full or partial signal and summing the resulting signals with respect to their individual delays. The same basic segmentation and summation process is also valid, when parts of the auralized signal or all of it are not processed by means of convolving a BRIR with a signal, but instead are simulated directly, i.e. by using methods based on delay networks.
[0137] By using psycho acoustic domain knowledge about how different parts of the binaural filter are perceived differently, a rendering system can be designed in a way, that it calculates less important parts of the filter less often and distributes those calculations between devices.
[0138] In one form, the system consists of a single device, capable of synthesizing and auralizing a binaural signal in real time. It includes at least two loudspeakers, which are able to reproduce the sound for one ear each, i.e. all types of common headphones or crosstalk- canceled speakers.
[0139] The system includes one or more positional sensors, which are able to determine the head rotation of the user relative to a reference system. (This is commonly called three degree of freedom or 3DoF tracking.) In a different embodiment, the system includes one or more positional sensors instead, which are able to determine the head rotation of the user relative to a reference system, plus their position relative to the reference system. (This is commonly called six degree of freedom or 6DoF tracking.) The system is able to process the binaural filters or to directly simulate the auralized signal, by employing one or more adequate binaural synthesis algorithms. It does not depend on one specific method of auralization. Different embodiments of the system might use different binaural synthesis algorithms.
[0140] In this embodiment, the used binaural synthesis algorithm must be capable of calculating the filter for the direct sound path and the room reverb separately. Auralization of the direct sound path is typically achieved by blockwise convolution of a filter, which approximates the filter effects of the users head, ears and torso in relation to a sound source at a given position and distance (the HRTF), with the audio signal.
[0141] Processing of these filters needs to encode correct changes in the inter-aural-time- difference (ITD) and inter-aural-level-difference (ILD) as well as changes in the sound intensity and other cues. Human listeners are comparatively sensitive to even small changes in these values, which is why it is necessary to calculate these changes with good spatial- and time-resolution. These filters are however comparatively short and deriving them usually involves only few processing steps.
[0142] The room reverb simulates the filter effects on the sound which are caused by the environments geometry for sound which does not travel on a direct part from the sound source to the users ears. This includes reflection, refraction, absorption and resonance effects. Such a reverb filter is expected to be much longer than the short direct sound filter. A number of processes and algorithms and systems are capable of processing adequate binaural reverb, such as the image source algorithm, raytracing, parametric reverberators and a number of delay network based approaches.
[0143] Either the signal processor, or another processor is configured as an aggregator. In some embodiments, in which the employed binaural synthesis methods return a continuous stream of blockwise binaural audio signals, this aggregator simply sums up the blocks supplied by the direct- and reverberant processing paths and acts as a signal aggregator. This requires, that the blocks to be summed correspond to the same point in time or contain control data that identifies the time frame they correspond to. Alternatively, the aggregator can be configured to sum up the two partial filters, with respect to their time delays, as determined by the algorithms. It therefore reconstructs a full BRIR filter from the individual processors results and acts as a filter aggregator. The filter can be used to convolve audio signal blocks using state of the art real time (blockwise) convolution methods. The aggregator keeps a full BRIR filter in its memory at all times. The BRIR can therefore be partially updated at the individual rates of the individual processors, which process the partial filters. The resulting signal blocks contain the combined binaural signals for the direct sound path and the reverberation path. They are passed to a loudspeaker signal generator, to be played back over the system’s loudspeakers. The loudspeakers can be speakers in a wearable device or cross talk cancellation speakers or any speakers such as speakers placed with some kind of sound separation element in between. This enables auralization of binaural audio with a similar level of externalization and perceived quality to those of the individual algorithms, while lowering processing requirements significantly.
[0144] A further part or aspect of the solution receives the previously derived RIR as an input and synthesizes a BRIR from it. It uses further metadata, like available positional data of both the room, the listener and the sound source, for the synthesis process. The system tracks the users position, relative to the source to be auralized and the real room, using a tracking system, consisting of one or more sensors, like an IMU or an optical tracking device. It receives metadata about the virtual sound sources position and a (individual or general) HRTF set.
[0145] For processing, the system might split the received RIR into arbitrary time-segments, which can be processed in parallel with different algorithms and at different intervals. In one embodiment, the RIR is split into three parts, including direct sound, early reflections and late reverberation. The direct sound segment is truncated in a way, that it contains the part of the RIR that contains the sound directly transmitted from source to receiver, but does not contain the first reflection arriving at the receiver. The late reverberation segment might start at a point, after which no single, strong reflections are perceptible anymore. The segments are windowed appropriately, for instance using overlapping Tukey windows, so that they can be reconstructed later. The relative position of listener and source determines the incidence direction of the direct sound, which is used to select a fitting HRTF from the set, either directly or by interpolation, to convolve it per channel with the direct sound segment of the RIR.
[0146] For the full length of the two reverberant parts, a pseudo-diffuse RIR is calculated by modeling the frequency dependent energy envelope of the RIR onto the binaural white noise (a signal with uniformly distributed energy over all frequency bands, but the phase information of the perfectly diffuse field of a BRIR), while retaining the phase information of the high density reflection pattern. This can be done by separating frequency bands using a perfect reconstruction filter bank, determining bandwise, low-passed envelopes and multiplying the noise signal with it. Alternatively, RIR and binaural noise can be transformed to the time domain, for instance by using a STFT, before applying the magnitude of the RIR onto the noise while keeping the phase and transforming it back to the time domain. The hereby derived pseudo-diffuse part, windowed accordingly, is used by the system as the late reverberation of the BRIR.
[0147] The early reflections segment of the RIR is further windowed into sub-windows, which may or may not correspond to the location of single- or multiple early reflections. Similar to the direct sound, each of the detected sub-segments is assumed to have an incidence direction, if it corresponds to an early reflection. This direction of arrival is either derived from a room model of appropriate complexity, using an algorithm like the image-source algorithm, or chosen statistically. A HRTF is picked or interpolated based on that direction and convolved with the sub-segment. To overcome the sparseness of this approach, the system mixes the pseudo-diffuse part with the fully directional (“specular”) part, to simulate diffuseness from reflections arriving at a similar time, and / or non-linear parts of the RIR.
[0148] For that, a function to determine a coefficient for the diffuseness of each window is used, to linearly interpolate between diffuse- and specular parts for each sub-segment. An appropriate function might be formed out of the energy ratio of the low-passed average energy in a small window around the signal, to the ratio of the low- passed average energy in a larger window around the signal, therefore approximating the ratio of local energy in relation to the short term average energy, as a predictor for masking effects. The resulting sub-segments are windowed and reassembled. Depending on the used signals and HRTFs, further post processing like diffuse field or headphone equalization might be applied.
[0149] Fig. 14 illustrates an input interface 100 that can receive several inputs as will be described later on, and that provides a single-channel acoustic data describing an acoustic environment. The single-channel acoustic data can be a room impulse response or a room transfer function or any other description that describes an acoustic environment such as a room or an open room or a semi-open room. The acoustic environment can also be an environment out of room depending on the situation. Typically, the acoustic environment will comprise reflection objects such as room walls, furniture, etc. or absorption objects such as persons in a room or curtains in a room or any other “acoustic objects”.
[0150] The audio signal processor additionally comprises a two-channel synthesizer for synthesizing two-channel acoustic data from the single-channel acoustic data using a listener position or orientation as illustrated in Fig. 14. The result of the two-channel synthesizer 200 is a two-channel acoustic data such as a binaural room impulse response or a binaural room transfer function or any other two-channel impulse response or transfer function as the case may be. Other descriptions from the impulse response or the transfer function can also be applied as the acoustic data such as a certain parameterization, etc.
[0151] The two-channel acoustic data is input into a sound generator for generating the two- channel audio signal from an audio signal which is typically a mono signal also illustrated in Fig. 14 and the two-channel acoustic data received from the two-channel synthesizer 200 of Fig. 14. The input interface can also be termed in this specification as the RIR provider. The two-channel synthesizer is also termed to be a binaural synthesizer and the sound generator is also termed to be an auralizer in this specification. Nevertheless, both descriptions mean the same thing, i.e. , the RIR provider is generally an input interface, the binaural synthesizer is a general two-channel synthesizer and the sound generator is a general auralizer.
[0152] The two-channel synthesizer 200 is configured to separate the single-channel acoustic data into at least two parts that consist of a direct sound part, an early reflection part and a late reverberation part, and the two-channel synthesizer 200 is configured to individually process the at least two parts for generating two-channel acoustic data for each part. This is illustrated in Fig. 16. In block 210, the single-channel acoustic data is separated in at least two parts. Block 220 illustrates a direct sound processing. Block 230 illustrates an earl reflection processing and block 240 illustrates a late reverberation processing. All three two-channel acoustic data for each part are combined by the aggregation or combination of the two-channel acoustic data as illustrated in 250. It is to be noted that block 250 covers the two alternatives that can, in general, be performed. The first alternative is the individual parts of a BRIR are aggregated into a full BRIR and the full BRIR is applied to the audio signal by convolution. The convolution is performed by the sound generator 300 of Fig. 14 that also receives the audio signal.
[0153] The alternative embodiment comprises that each part is convolved with the audio signal separately so that three streams of binaural audio data are obtained and the binaural audio data are calculated by combining the three individual streams binaural audio 1 , binaural audio 2, and binaural audio 3. The processing with the audio signal and the aggregation of the audio signals is performed by the sound generator 300 as illustrated in Fig. 14.
[0154] In a preferred embodiment illustrated in Fig. 16, the direct sound processing relies on the source directivity, initial source or sink data from the initial measurement, current listener data and / or current source data. In this context, it is to be noted that current listener data referred to as listener position, listener orientation or both also termed to be a listener “pose” in the following text. The same is true for the source data. The source data can be a source position or a source rotation or both, the source position and the source rotation. Specifically, the source rotation can be advantageously accounted for even for non- omnidirectional sources using the source directivity information in accordance with the first or second aspect of the present invention.
[0155] The early reflection processing in block 230 relies on the listener position and / or orientation, and geometrical data on the acoustic environment and, typically, initial data such as the association of image sound sources to early reflections. The source directivity can be accounted for in the earlier reflection processing in block 230 as well. The late reverberation processing 240 relies on the two-channel noise data illustrated as two arrows in Fig. 16 in order to illustrate the transformation of the late reverberation part which is a single channel part into two output channels illustrated at the lower portion of block 240 in Fig. 16. Fig. 15a illustrates a preferred implementation related to the smart determination of the room impulse response from a raw representation related to the single-channel acoustic data. Particularly, the input interface 100 of the device illustrated in Fig. 14 is configured to acquire a raw representation related to the acoustic data as illustrated at 150. The input interface 100 is configured to derive the single-channel acoustic data using the raw representation obtained in block 150 and using additional data stored by the audio signal processor or accessible by the audio signal processor in order to obtain the single-channel acoustic data that is forwarded to the two-channel synthesizer 200.
[0156] Exemplarily, the input interface 100 comprises blocks 10 and 20 of Fig. 11a and is configured to acquire, as the raw representation, an initial measurement of raw singlechannel acoustic data to derive a test fingerprint of the raw single-channel acoustic data as illustrated in block 101 of Fig. 15b. Based on this test fingerprint, a pre-stored database 110 connected or included in the database processor 20 with an associated set of reference fingerprints is accessed, where each reference fingerprint is associated to a higher resolution single-channel acoustic data, where the high resolution single-channel acoustic data has a higher resolution than the initial measurement. From the pre-stored database 110, the high resolution single-channel acoustic data having a reference fingerprint best matching with the test fingerprint is retrieved as illustrated in block 113 of Fig. 15b.
[0157] Bock 101 corresponding to block 10 of Fig. 11a is configured to derive a test fingerprint as a set of at least one of the following parameters RT 60, EDC, DRR, and wherein the reference fingerprint additionally comprises at least one of the following parameters RT 60, EDC, DRR and particularly the subband-wise energy decay parameter as outlined with respect to Figs. 11a and 11 b.
[0158] A further implementation of the present invention is that the user generates a natural sound as illustrated in 160 of Fig. 15b. Such a natural sound is clapping, or speech or any transient sound producible a listener. This avoids the generation of an unpleasant measuring sound in a room such as sine sweep. Based on this sound, a (low resolution) RIR is recorded as the microphone signal, and is processed by any of the procedures shown in Figs. 1 to 15b in order to obtain, from this raw representation, the high resolution room impulse response for the purpose of further processing.
[0159] In accordance with Fig. 17, the two-channel synthesizer and, particularly, the early reflection processing block 230 is configured to segment the early reflection part into a plurality of segments as shown in block 231. Exemplarily a segmentation can be performed up to fifty segments or more. Naturally, less segments can also be used. In an embodiment, there exist blocks with 256 samples and an overlap of 128 samples. The number of segments is obtained by the length of the early reflection part (direct sound until the mixing time) having roughly 7700 samples. Dividing this number by an advance value of 128 per segment results in about 60 segments. But, the number can vary depending on the length of the early reflection part, the advance value and potential other parameters used.
[0160] As shown in block 232, a plurality of image source positions is determined preferably using a geometric model of the room such as a shoe box model. The image source positions represent source positions of reflection sound. An association of the image source positions to the segments is performed using a matching operation. In the matching operation, a time of sound arrival from each image source to the listener position is calculated as illustrated in block 233. Preferably, the initial listener position is used for this calculation, so that the initial listener position, i.e., the listener position when the RIR was provided by the input interface is input. The image source positions are associated to corresponding segments that match with the time of arrival for a specific image source in a best way as illustrated in block 234.
[0161] Therefore, the time of arrival for the sound from each image source position to the initial listener position is compared to the time index in a certain segment. Typically, the segments have a certain width and, therefore, for a segment, the time index in the middle of the segment is compared to the time of arrival. When a time of arrival for an image source position is equal to the time index associated with a segment such as the time index in the middle of the segment, this image source position is associated with this segment for the further calculation such as a calculation of a direction of arrival for this segment. Typically, the image source positions are calculated for the room model up to a certain order. Some first order image source positions being the first reflections are used. The second order reflections can also be constructed and refer to the physical effect that a reflection reaching the listener’s head travels on and is reflected at a second wall and reaches the listener again.
[0162] Depending on how complex the geometrical model is, a certain number of image sound sources are determined with respect to their position and are associated to corresponding segments. When, for example, fifty segments are used for segmenting the early reflection part, it is sufficient to determine the image source positions up to an order resulting in fifty sources. However, this can be quite complex and, in order to save computational resources, a preferred way of doing this is to only calculate image source positions up to a certain order resulting in less than fifty image source positions. The remaining image source positions can be selected in a random way as illustrated in block 235. Therefore, if it has been found that a certain segment results in a non-matching image source associated with this segment, either a random position is associated to this segment or, in the further calculation, a random direction of arrival and, therefore, a randomly selected HRIR is used for the processing of this segment.
[0163] The result of this procedure is illustrated in table 236 at the bottom of Fig. 17 illustrating that the first three segments are associated with source position 2, source position 1 , source position 4, respectively, and there also exists one or several segments typically in the end of the segments when the segments are counted from the direct sound / early reflection border to the early refl ection / reverbe rati on border that do not have a discrete image source position, but that have associated therewith a random source position or receives a random HRIR in the processing of this segment.
[0164] As outlined, the two-channel segment is configured to determine the plurality of image source positions using an initial source position and an initial sink position of an initial measurement and geometric data on the acoustic environment. Particularly, the image source method is preferred.
[0165] In this embodiment, the two-channel synthesizer is configured to determine, for the listener position and an image source position or orientation of an image sound source, directivity information of the image sound source, and to use the directivity information in the calculation 220 of the two-channel acoustic data for the earlier reflection sound part. Preferably, the directivity information for each image source is derived from the same set of directivity information determined for the direct sound part, or wherein an orientation of the image sound source is determined by an image source model, and the directivity information is in a specific embodiment determined and used for a predetermined subset of the segments in the early reflection part, which comprises less than ten segments and preferably only 2 segments. The remaining segments can be calculated without any directivity information of an image source.
[0166] In Fig. 18, the two-channel synthesizer is configured to calculate a specular part for segment n or, generally, for the early reflection part as discussed with respect to the second aspect and as shown in block 237. Additionally, the two-channel acoustic data for early reflection part are also calculated using a diffuse part as shown in block 238 that describes a diffuse influence in the early reflection part. Both blocks 237 and 238 receive a single-channel early reflection part for the segment n as provided by block 210 of Fig. 16. Both blocks 237, 238 output two channels of binaural data, and these two channels are combined in block 239 correspondingly so that a first channel and a second channel for segment n is obtained, and this two-channel data for the early reflection part not only represents the specular influence of the distinct early reflections as in prior art procedures, but also accounts for the diffuse portion that heavily contributes to natural and pleasing sound impression for the listener.
[0167] Particularly, the two-channel synthesizer 200 is configured to calculate the diffuse part using a combination of the early reflection part of the single-channel acoustic data and a two- channel noise sequence as input into block 238. Preferably, this two-channel noise sequence is a binaural noise sequence as measured when a certain noise signal is emitted by a speaker at a certain position with respect to an artificial head, and where a full HRTF is detected by a means of two microphones residing in the artificial head. Such binaural noise can be actually measured or can, alternatively, be synthesized or if this is, for some reason, not practical, even two different noise sequences can be used for the binauralization of the late reverberation part of the room impulse response.
[0168] Fig. 19a illustrates the subject-matter of the present that refers to the improved calculation of the diffuse part either for the early reflection part or the late reverberation part or only for the late reverberation part or for both parts using a magnitude spectrum of the early reflection part and / or the late reverberation part and using phase spectra for the two- channel (binaural) noise. The two-channel synthesizer 200 of Fig. 14 is configured to calculate a two-channel diffuse portion of the early reflection part or of the single-channel acoustic data without the direct sound part using a magnitude spectrum of the early reflection part or of the single-channel acoustic data without the direct sound part and a first channel noise phase spectrum for obtaining a first channel of the two-channel acoustic data and using a magnitude spectrum of the early reflection part or of the single-channel acoustic data without the direct sound part and a second channel noise phase spectrum.
[0169] Particularly, the first channel nose phase spectrum and the second channel noise phase spectrum are derived from a two-channel binaural noise sequence. This is illustrated by block 530 illustrating the calculation of the magnitude spectrum and block 532 in Fig. 19a illustrating the calculation of the phase spectra of the two-channel (binaural) noise.
[0170] The conversion of a single-channel data as derived by block 520 or as preferably derived subsequent to a smoothing of the magnitude spectrum in block 531 is transformed into a second-channel result by adding the phases of the first channel phase spectrum to the smoothed amplitudes of a spectrum in block 531 to obtain the first channel result and by adding the second channel phase spectrum of block 532 to the preferably smoothed magnitude spectrum of block 531 in the combiner 533 to obtain the second channel of the two-channel diffuse part for the late reverberation part or for the early reflection part plus the late reverberation part or, stated in other words, for the single-channel acoustic data without the direct sound part that is assumed to be non-diffuse and, therefore, does not receive and diffuse contribution. It is to be noted that the „adding“ of magnitude and phase is, in the specific mathematical sense, a multiplication as shown in block 444 of Fig. 19b, i.e., a multiplication of a magnitude spectrogram and a phase spectrogram: |RTF| * eA(angle(binauralNoise1)) and |RTF| * eA(angle(HbinauralNoise2)), where H stands for a transform into the spectral domain.
[0171] As illustrated in Fig. 19b, a mono RIR is provided in block 440 and absolute values of an STFT spectrogram consisting of a sequence of spectra is taken as shown in block 442. A binaural noise sequence 441 is provided as well and is subjected to a corresponding spectrogram processing by means of a time-to-frequency conversion and the phase angles of each channel of the binaural noise are taken as illustrated in block 443 and the phase angles are combined with the corresponding magnitudes as illustrated in block 444, preferably subsequent to the smoothing operation in block 531.
[0172] The smoothing operation in block 531 has the advantage that this smoothing along the frequency direction of the magnitude spectrum, naturally in each spectrum of the sequence of spectra covering e.g. the early reflection part and the late reverberation part avoids any peaks that might occur due to the inverse Fourier transform when a phase manipulation has been performed in the spectral domain as it is the case the present invention. On the other hand, the procedure of calculating spectrograms and simply “adding” the phases of the binaural sequences to the (smoothed) spectrogram is a computationally easy procedure that does not require a considerable amount of computational resources. Additionally, it has been found that this late reverberation processing has a pleasant sound for the listener which is particularly useful, since, due to its quality, the same late reverberation two-channel acoustic data for the acoustic environment can be used irrespective of whether the source position or orientation or the listener position or orientation changes. This situation results in a significant consequence in that the update ratio for the calculation of the late reverberation part can be made significantly smaller (typically by one or even two orders) which additionally reduces required computational resources and also allows to distribute the processing tasks to different elements as will be illustrated with respect to the sixth aspect of the present invention.
[0173] An overlapping blocks transform is applied to the late reverberation room impulse response or to both, the early reflection part and the late reverberation part. This results in a first spectrogram where, preferably, and as shown in block 531 , a low-pass filtering is performed in each magnitude spectrum over frequency. It is preferred to additionally perform a lowpass filtering over time, i.e., over two or more adjacent blocks and with respect to the same frequency bin, but in adjacent blocks, i.e., with time-adjacent frequency bins relating to the same frequency. A similar transform is performed for the time domain binaural noise sequence to obtain the second spectrogram and the third spectrogram, and the phases of the second and third spectrograms are added to the spectral domain and time domain lowpass filtered spectra.
[0174] The result is transformed into a Cartesian format and inversely transformed into the time domain. An overlap and add procedure is performed and, finally a truncation and windowing and an overlap with the earlier reflection part is performed which is only done for the diffuse signal for the late reverberation to obtain the two-channel acoustic data for the late reverberation part.
[0175] Fig. 20 illustrates a preferred embodiment, where the device illustrated in Fig. 14 is separated into a first device and a second device. Particularly, the two-channel synthesizer 200 is configured by two physically separate devices 901 , 902. The first device 901 of the two physically separate devices is configured to process the direct sound part as shown in block 916 and as illustrated in block 220 of Fig. 16. To this end, the processing requires the listener position or rotation. The second device 902 of the two physically separate devices is configured to process the at least one of the early reflection part and the late reverberation part. This block is illustrated at 923 and implements either one or both of the functionalities of block 230 and 240 of Fig. 16. Both devices are connected to each other via transmission interface 918 of the first device and 925 of the second device. This transmission interface is preferably a wireless interface and operates in accordance, for example, with the Bluetooth standard. One consequence of the separation of the two physically separated devices is that the first device 901 has an own power supply 917 and the second device 902 also has an own power supply 924.
[0176] Preferably, as illustrated in Fig. 20, the first device is configured to update the two-channel acoustic data for the direct sound part more often than the second device updates the two- channel audio data for the at least one of the early reflection part and the late reverberation part. In figures it is preferred to have a direct sound part update above 15 Hz, i.e., more than 15 updates per second, preferably more than 20 updates per second and even more preferably more than 50 updates per second. An update rate of the early reverberation part is preferably in a range between 5 Hz and 15 Hz, and a late reverberation part update is sufficient to be in the range between 0.5 Hz to 5 Hz. It appears that the parts that require a lower update rate are processed in the second device 902. It has been found that it is these parts that require a significant higher processing power due to the long filters which require, on the other hand, a lower update rate. The second device is implemented to be computationally and battery-power like significantly stronger and more powerful than the first device. The first device can be an earbud device, a headphone device, an in-ear device or any other wearable device that typically has a limited battery power. However, the second device can be a highly powered device such as a mobile phone, a smart watch, a laptop computer, a tablet or even a stationary computer connected to the power mains and also typically connected to a large area network such as the internet. Preferably, the first device not only comprises the processing block for the direct sound part 916 but also comprises a microphone for the recording of the acoustic measurement for the RIR provision and, additionally, the functionality for sound rendering as illustrated by the sound generator 300 of Fig. 14 and, additionally, the speakers when the device is a headphone device, for example. Alternatively, the speakers can also be separate when the speakers are provided with Bluetooth signals, for example, from the device 901 which has the communication interface rather than the actual speakers.
[0177] The reverberation processing is implemented in the second device 902 and the direct sound processing is implemented in the first device 901. Additionally, the functionalities of the input interface 100, the signal aggregator 310 and the signal generator 300 are also implemented in the first device 901 . Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed. Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier. Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier or a non-transitory storage medium. In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein. In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0178] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
Claims
Claims1. Apparatus for retrieving an acoustic room information from a database (110), comprising: a test fingerprint processor (10) for deriving a test fingerprint from a test acoustic room information; and a database processor (20) for retrieving a reference acoustic room information having a reference fingerprint matching with the test fingerprint, wherein the test fingerprint processor (10) comprises; a filter bank (11) configured for decomposing the test acoustic room information into a plurality of acoustic room information sub-bands; and a parameter calculator (12) configured for calculating a sub-band parameter for each acoustic room information sub-band of the plurality of acoustic room information subbands, the sub-band parameter indicating an energy change over time in the acoustic room information sub-band, wherein the test fingerprint is based on the sub-band parameters for the acoustic room information sub-bands.
2. Apparatus of claim 1 , wherein the parameter calculator (12) is configured to calculate an energy-related characteristic over time from the acoustic room information sub-band, and to parameterize the energy-related characteristic over time to obtain the sub-band parameter for the acoustic room information sub-band.
3. Apparatus of claim 2, wherein the parameter calculator (12) is configured to approximate the energy-related characteristic by calculating (15) one or more curve parameters by fitting a pre-defined curve to the energy-related characteristic over time to obtain one or more calculated parameters, wherein the one or more calculated parameters are used to represent the sub-band parameters for the acoustic room information sub-band.
4. Apparatus of claim 3, wherein the predetermined curve is a straight line and the one or more parameters is a slope of the straight line.
5. Apparatus of one of claims 2 - 4, wherein the parameter calculator (12) is configured to calculate (13) an energy decay curve from the acoustic room information subband and to convert (14) the energy decay curve for the acoustic room information sub-band into a logarithmic scale.
6. Apparatus of claim 5, wherein the parameter calculator (12) is configured to perform (15) a linear regression on the energy decay curve in the logarithmic scale.
7. Apparatus of one of the preceding claims, wherein the test fingerprint processor (10) is configured to weight (16) the sub-band parameters using a psychoacoustic characteristic, so that a parameter for a first acoustic room information sub-band, in which a listening perception is greater than in a second acoustic room information sub-band, has a higher influence on the test fingerprint compared to an influence of a parameter for the second acoustic room information sub-band on the test fingerprint.
8. Apparatus of claim 7, wherein the test fingerprint processor (10) is configured to use, for the weighting (16), a Fletcher-Munson-Curve or an equal loudness curve or an equal loudness contour in accordance with International Standard ISO226.
9. Apparatus of claim 7 or 8, wherein the test fingerprint processor (10) is configured to calculate (13, 14, 15), for each acoustic room information sub-band, the sub-band parameter, to weight (16), for each acoustic room information sub-band, the subband parameter with a weighting value for the acoustic room information sub-band to obtain a weighted sub-band parameter for each sub-band, and wherein the test fingerprint is based on the weighted sub-band parameters for the acoustic room information sub-bands.
10. Apparatus of claim 9, wherein the test fingerprint processor (10) is configured to calculate (15) a slope parameter for each acoustic room information sub-band, to select, from the psychoacoustic characteristic given as a logarithmic characteristic, a logarithmic value for a sub-band corresponding to the acoustic room information sub-band, andto divide (16) the slope parameter by the selected logarithmic value to obtain the weighted parameter for the acoustic room information sub-band.
11. Apparatus of one of the preceding claims, wherein the database processor (20) is configured to calculate (46), for a plurality of reference fingerprints for a plurality of reference acoustic room information sub-bands, a pair wise distance of the test fingerprint and the reference fingerprint and to select a reference acoustic room information associated with the reference fingerprint having a smallest distance to the reference fingerprint or a distance being smaller than a distance threshold, or wherein the database processor (20) is configured to perform a brute force nearest neighbor search or, alternatively an approximate nearest neighbor search using a predetermined query radius.
12. Apparatus of claim 11 , wherein the test fingerprint has a plurality of test fingerprint components, one test fingerprint component for each acoustic room information subband, wherein the reference fingerprint has a plurality of reference fingerprint components, one reference fingerprint component for each acoustic room information sub-band, wherein the test database processor (20) is configured to calculate (46) a component distance between a test fingerprint component for an acoustic room information sub-band and a reference fingerprint component for the acoustic room information sub-band, and to accumulate the component distances over the acoustic room information sub-bands to obtain the pair wise distance.
13. Apparatus of claim 12, wherein the database processor (20) is configured to calculate, as the pair wise distance, a Manhattan or L1 distance, an Euclidean or L2 distance, a cosine distance or a distance for measuring a proximity between vectors in a vector space.
14. Apparatus of one of the preceding claims, comprising a memory (110) for storing a plurality of reference acoustic room information items, wherein the database processor (20) is configured to store, in the memory, a test fingerprint or a referencefingerprint calculated in an earlier retrieving operation or being stored in the memory due to a memory initialization, wherein the database processor (20) is configured to check (40), for a test acoustic room information or a reference acoustic room information, whether an associated test fingerprint or an associated reference fingerprint is already stored in the memory, and to retrieve (42) a stored test fingerprint or a stored reference fingerprint, when a check result is positive or to calculate (44) the test fingerprint or the reference fingerprint, when the check result is negative.
15. Apparatus of one of the preceding claims, wherein the acoustic room information is a room impulse response, a binaural room impulse response, an impulse response, a binaural impulse response, a head related transfer function, a binaural head related transfer function, a room transfer function or a binaural room transfer function.
16. Apparatus of one of the preceding claims, wherein the database processor (20) is configured to retrieve a group of at least two reference acoustic room information comprising the reference acoustic room information and at least one additional acoustic room information, wherein the test fingerprint processor (10) is configured to derive (34) an additional test fingerprint from the test acoustic room information, the additional test fingerprint being different from the test fingerprint obtained by the test fingerprint, wherein the database processor (20) is configured to determine (32) an additional reference fingerprint for the acoustic room information and the at least one additional reference acoustic room information, and wherein the database processor (20) is configured to perform (36) a matching operation between the additional test fingerprint and the additional reference fingerprints to determine the reference acoustic room information in the group which is best matching with the additional test fingerprint,wherein a calculation rule for the calculation of the additional reference fingerprint is different from a calculation rule for calculating the test fingerprint, and wherein the calculation rule for calculating the additional test fingerprint is similar to the calculation of the additional reference fingerprint for the group of the at least two reference acoustic room information.
17. Apparatus of claim 16, wherein the test fingerprint processor (10) is configured to derive the additional test fingerprint from a representation of the test acoustic room information having a bandwidth being broader than a bandwidth defined by at least two acoustic room information subbands.
18. Apparatus of claim 16 or 17, wherein the test fingerprint processor (10) is configured to derive the additional test fingerprint from a full band representation of the test acoustic room information and / or a predefined time frame of the test acoustic room information excluding a direct sound part and / or excluding a late reverberation part and / or only comprising an early reflection part of the test acoustic room information.
19. Apparatus of one of claims 16 to 18, wherein the test fingerprint processor (10) is configured to derive a sample-by-sample squared representation of the test acoustic room information as the additional test fingerprint, and wherein the database processor (20) is configured to derive a sample-by-sample squared representation of each one of the reference acoustic room information items to obtain the at least two additional reference fingerprints.
20. Apparatus of one of claims 16 to 19, wherein the database processor (20) is configured to perform, in the matching operation (36), a cross correlation processing between the additional test fingerprint and the at least two additional reference fingerprints to select the reference acoustic room information having a cross- correlational result indicating a best matching situation.21 . Audio signal processor for generating a two-channel audio signal, comprising: an apparatus of one of claims 1 to 20, wherein the acoustic room information comprises single channel acoustic data;a two-channel synthesizer (200) for synthesizing two-channel acoustic data from the single-channel acoustic data using a listener position or rotation; and a sound generator (300) for generating the two-channel audio signal from an audio signal and the two-channel acoustic data, wherein the apparatus is configured to acquire (150), as a raw representation related to the single-channel acoustic data, the test acoustic room information, and to derive (151) the single-channel acoustic data as the reference acoustic room information.
22. Audio signal processor of claim 21 , wherein the test acoustic room information is an initial measurement of raw single-channel acoustic data, and wherein the reference acoustic room information comprises high-resolution single-channel acoustic data having a higher resolution than the initial measurement.
23. Audio signal processor of claim 21 or 22, wherein the database processor (20) is configured to select the single-channel acoustic data having the reference fingerprint that minimizes a distance to the test fingerprint.
24. Audio signal processor of one of the claims 21 to 23, wherein the apparatus is configured to use, for an initial measurement, to a natural sound producible by a listener.
25. Audio signal processor of claim 24, wherein the natural sound is clapping, or speech, or a transient sound producible by the listener.
26. Audio signal processor of one of the claims 21 to 25, wherein the apparatus comprises a speaker or a microphone embedded in a mobile device, and wherein the apparatus is configured to perform an initial measurement with the speaker or the microphone or only with the microphone embedded in the mobile device.
27. Audio signal processor of one of the claims 21 to 26, wherein the two-channel synthesizer (200) is configured to separate (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection partand a late reverberation part, and to individually process (220, 230, 240) the at least two parts for generating two-channel acoustic data for each part, to determine (222), for the listener position and a source position or orientation of a sound source, directivity information of the sound source, and to use the directivity information in the calculation (220) of the two-channel acoustic data for the direct sound part.
28. Audio signal processor of one of the claims 21 to 27, wherein the two-channel synthesizer (200) is configured to separate (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to individually process (220, 230, 240) the at least two parts for generating two-channel acoustic data for each part, wherein the two-channel synthesizer (200) is configured to segment (231) the early reflection part into a plurality of segments, to determine (232) a plurality of image source positions representing source positions of reflecting sound, to associate the image source positions to the segments using a matching operation, wherein the matching operation comprises calculating a time of sound arrival for each image source to the listener position and associating (234) the image source positions to corresponding segments that have time delays in the corresponding segments best matching with the time of sound arrival of the corresponding image source positions, and to calculate the two-channel acoustic data for the direct sound using the image source positions associated to the segments.
29. Audio signal processor of one of the claims 21 to 28, wherein the two-channel synthesizer (200) is configured to separate (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to individually process (220, 230, 240) the at least two parts for generating two-channel acoustic data for each part, and wherein the two-channel synthesizer (200) is configured to calculate (230) the two- channel acoustic data for the early reflection part using a specular part describingdistinct early reflections and a diffuse part describing a diffuse influence in the early reflection part.
30. Audio signal processor of one of the claims 21 to 29, wherein the two-channel synthesizer (200) is configured to separate (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to individually process (220, 230, 240) the at least two parts for generating two-channel acoustic data for each part, and wherein the two-channel synthesizer (200) is configured to calculate (230) the two- channel acoustic data for the early reflection part using a specular part describing distinct early reflections and a diffuse part describing a diffuse influence in the early reflection part.31 . Audio signal processor of one of the claims 21 to 30, wherein the two-channel synthesizer (200) is configured to separate (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to individually process (220, 230, 240) the at least two parts for generating two-channel acoustic data for each part, and wherein the two-channel synthesizer (200) is configured to calculate a two-channel diffuse portion of the early reflection part or of the single-channel acoustic data without the direct sound part or of the late reverberation part using a magnitude spectrum of the early reflection part or of the single-channel acoustic data without the direct sound part or of the late reverberation part and a first channel noise phase spectrum for obtaining a first channel of the two-channel acoustic data and using a magnitude spectrum of the early reflection part or of the single-channel acoustic data without the direct sound part or of the late reverberation part and a second channel noise phase spectrum.
32. Audio signal processor of one of the claims 21 to 31 , wherein the two-channel synthesizer (200) is configured to separate (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an earlyreflection part and a late reverberation part, and to individually process (220, 230, 240) the at least two parts for generating two-channel acoustic data for each part, and wherein the two-channel synthesizer comprises two physically separate device (901 , 902), wherein the first device (901) of the two physically separated devices is configured to process (220, 230) at least one of the direct sound part and the early reflection part, wherein the second device (903) of the two physically separated devices is configured to process (230, 240) at least one of the early reflection part and the late reverberation part, and wherein the first device (901) and the second device (902) are connected via a transmission interface (918, 925) and have separate power supplies (917, 924).
33. Method of retrieving an acoustic room information from a database, comprising: deriving a test fingerprint from a test acoustic room information; and retrieving a reference acoustic room information having a reference fingerprint matching with the test fingerprint, wherein the retrieving comprises; decomposing the test acoustic room information into a plurality of acoustic room information sub-bands; and calculating a sub-band parameter for each acoustic room information sub-band of the plurality of acoustic room information sub-bands, the sub-band parameter indicating an energy change over time in the acoustic room information sub-band, wherein the test fingerprint is based on the sub-band parameters for the acoustic room information sub-bands.
34. Method of generating a two-channel audio signal, comprising: a method of claim 33, wherein the acoustic room information comprises single channel acoustic data;synthesizing (200) two-channel acoustic data from the single-channel acoustic data using a listener position or rotation; and generating (300) the two-channel audio signal from an audio signal and the two- channel acoustic data, wherein the method comprises acquiring (150), as a raw representation related to the single-channel acoustic data, the test acoustic room information, and to deriving (151) the single-channel acoustic data as the reference acoustic room information.
35. Computer program for performing, when running on a computer or a processor, the method of claim 33 or claim 34.
Citation Information
Patent Citations
Systems and methods for modifying room characteristics for spatial audio rendering over headphones
US20200137508A1
Converter and method for converting an audio signal
WO2010057997A1
Augmented reality headphone environment rendering
WO2017136573A1
Devices and methods for binaural audio rendering
WO2023208333A1
Audio signal processor and related method and computer program for generating a two-channel audio signal using a smart determination of the single-channel acoustic data
WO2024089036A1